Enhanced prokaryotic gene expression regulators through nucleotide-level mapping
Patent Information
- Application Number
- PCT/US2024/055987
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-15
- Filing Date
- 2024-11-14
- Publication Date
- 2025-06-19
AI Technical Summary
Current prokaryotic gene expression systems face challenges with poor expression of gene products and inadequate regulation, necessitating improved tools for enhanced gene expression and regulation.
The development of an expression system comprising nucleic acid sequences with mutant pBAD, pRha, pBgal, pPTK, and pTETO1 promoter sequences operably linked to heterologous genes, along with constitutive promoters encoding AraC, RhaS, BgaR, AraR, and TetR proteins, which induce RNA polymerase binding in the presence of specific inducers.
This approach significantly enhances gene expression levels, achieving 5-fold to 10-fold greater expression compared to control systems, and provides improved regulation of gene expression, thereby addressing the limitations of existing systems.
Smart Images

Figure US2024055987_19062025_PF_FP_ABST
Abstract
Description
[0001] Atty. Dkt. No.: 136669-0122 ENHANCED PROKARYOTIC GENE EXPRESSION REGULATORS THROUGH NUCLEOTIDE-LEVEL MAPPING CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Patent Application No.63 / 599,160, filed November 15, 2023, the entire contents of which are incorporated herein by reference. GOVERNMENT SUPPORT This invention was made with government support under 1847226 awarded by the National Science Foundation. The government has certain rights in the invention. TECHNICAL FIELD The present technology relates generally to nucleic acid constructs, including promoter constructs, and methods of preparation thereof, that significantly improve prokaryotic gene expression and regulation thereof. BACKGROUND The following description of the background of the present technology is provided simply as an aid in understanding the present technology and is not admitted to describe or constitute prior art to the present technology. Prokaryotic expression systems are used to generate a variety of products. Efficiency can be hampered by poor expression of gene products or poor regulation of expression. Accordingly, there is an urgent need for improved tools for prokaryotic gene expression and regulation. SUMMARY OF THE PRESENT TECHNOLOGY In one aspect, the present disclosure provides an expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pBAD promoter sequence of any one of SEQ ID NOs: 1-14 that is operably linked to a heterologous gene and (b) a constitutive Pc promoter sequence that is operably linked to a gene sequence encoding an AraC protein, wherein the AraC protein is configured to induce an RNA polymerase to bind to the mutant pBAD promoter sequence in the presence of arabinose. 1 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 In another aspect, the present disclosure provides an expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pRha promoter sequence of any one of SEQ ID NOs: 16-25 that is operably linked to a heterologous gene and (b) a constitutive PRhaRS promoter sequence that is operably linked to a gene sequence encoding a RhaS protein, wherein the RhaS protein is configured to induce an RNA polymerase to bind to the mutant pRha promoter sequence in the presence of rhamnose. In a different aspect, the present disclosure provides an expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pBgal promoter sequence of any one of SEQ ID NOs: 27-46 that is operably linked to a heterologous gene and (b) a BgaR promoter sequence that is operably linked to a gene sequence encoding a BgaR protein, wherein the BgaR protein is configured to induce an RNA polymerase to bind to the mutant pBgal promoter sequence in the presence of lactose. In one aspect, the present disclosure provides an expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pPTK promoter sequence of any one of SEQ ID NOs: 48-67 that is operably linked to a heterologous gene and (b) an AraR promoter sequence that is operably linked to a gene sequence encoding an AraR protein, wherein the AraR protein is configured to induce an RNA polymerase to bind to the mutant pPTK promoter sequence in the presence of arabinose. In yet another aspect, the present disclosure provides an expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pTETO1 promoter sequence of any one of SEQ ID NOs: 69-87 that is operably linked to a heterologous gene and (b) a TetR promoter sequence that is operably linked to a gene sequence encoding a TetR protein, wherein the TetR protein is configured to induce an RNA polymerase to bind to the mutant pTETO1 promoter sequence in the presence of tetracycline or anhydrotetracyline (aTc). In any of the preceding embodiments, the expression system is integrated on a chromosome of a host cell. In any of the preceding embodiments, the expression system is an expression vector. In some embodiments the expression vector is a plasmid, a cosmid, a bacterial artificial chromosome (BAC) or a yeast artificial chromosomes (YAC). In any of the preceding embodiments, the heterologous gene encodes a protein, an enzyme, a structural polypeptide, a toxin, a fusion protein, an antibody agent, a drug, a cytokine, an enzyme inhibitor, a growth factor, a signaling protein, a bioluminescent protein, a fluorescent protein, a chemiluminescent protein, a catalytic RNA, or an inhibitory RNA. In some embodiments, the expression system expresses the heterologous gene to a greater degree than a control expression 2 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 system that comprises a wild type pBAD, pRha, pBgal, pPTK, or pTETO1 promoter operably linked to the heterologous gene. In some embodiments, the expression system expresses the heterologous gene at about a 5-fold to about a 10-fold greater degree than the control expression system. In one aspect, the present disclosure provides a prokaryotic host cell comprising the expression system of any one of the preceding embodiments. In some embodiments, the host cell is a gram-positive bacteria or a gram negative bacteria. In some embodiments, the host cell is E. coli or Clostridium. In another aspect, the present disclosure provides a nucleic acid construct comprising a mutant inducible promoter sequence operably linked to a heterologous gene, wherein the mutant promoter sequence is selected from the group consisting of SEQ ID NOs: 1-14, 16-25, 27-46, 48-67, and 69-87. In some embodiments, the nucleic acid construct is integrated on a chromosome of a host cell. In some embodiments, the nucleic acid construct is integrated in an expression vector. In specific embodiments, the expression vector is a plasmid, a cosmid, a bacterial artificial chromosome (BAC) or a yeast artificial chromosomes (YAC). In one aspect, the present disclosure provides a prokaryotic host cell comprising the nucleic acid construct of any one of the preceding embodiments. In some embodiments, the host cell is a gram-positive bacteria or a gram negative bacteria. In some embodiments, the host cell is E. coli or Clostridium. In a different aspect, the present disclosure provides a method for overexpressing a heterologous polypeptide or nucleic acid in a prokaryotic cell comprising contacting the prokaryotic host cell of any one of the preceding embodiments with an effective amount of an inducer molecule, wherein the heterologous polypeptide or nucleic acid is encoded by the heterologous gene of the expression system of any one of the preceding embodiments or the nucleic acid construct of the preceding embodiments. In some embodiments, the inducer molecule is selected from the group consisting of arabinose, rhamnose, lactose, tetracycline and anhydrotetracyline (aTc). In some embodiments, the method further comprises lysing the prokaryotic cell to isolate the heterologous polypeptide. In one aspect, the present disclosure provides a kit comprising a nucleic acid encoding the expression system of any one of the preceding embodiments or the nucleic acid construct of any one of the preceding embodiments and instructions for use thereof to express the heterologous gene in a host cell. 3 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 In another aspect, the present disclosure provides a kit comprising a host cell of any one of claims the preceding embodiments and instructions for use thereof to overexpress a heterologous polypeptide or nucleotide. In any of the preceding embodiments, the kit further comprises one or more inducer molecules. BRIEF DESCRIPTION OF THE DRAWINGS FIGs.1A-1B: Schematics and mechanisms of (FIG.1A) PBAD and (FIG.1B) PRha. Schematic and known mechanisms of E. coli PBADand PRhapromoters. The PBADpromoter is regulated by transcription factor AraC. In the absence of arabinose, AraC represses transcription by binding at araI1 (I1) and araO2 (O2) creating a DNA loop that prevents RNA polymerase and CRP from binding to PBAD, as well as reducing the expression of araC from the constitutive PC promoter. The presence of arabinose changes the AraC dimer conformation to the activating form and binds to araI1 (I1) and araI2 (I2) while recruiting the RNA polymerase (RNAP) to the - 35 and -10 boxes. Full activation of PBAD is achieved when CRP binds. The PRha promoter shares a similar structural arrangement of their binding sites as the PBADpromoter, however, differs in their repression mechanisms. The absence of L-rhamnose results in no binding of transcription factor RhaS. PRha is activated when transcription factor RhaS, expressed from PRhaRS, binds at rhaI1 (I1) and rhaI2 (I2) half-sites, which also recruits the RNA polymerase to the -35 and -10 boxes. Full activation is achieved when the co-activator CRP binds upstream of the rhaI half- sites. FIGs.2A-2B: PBADvalidations in E. coli. FIG.2A: Initial validations testing with 0% and 0.2% arabinose in E. coli NEB5α. Individual promoters were induced with 0% and 0.2% L- arabinose (w / v) and the median gfp expression was measured after t=5 hours. FIG.2B: Single mutant promoters with improved promoter expression strength were combined and assessed after 24 hours after induction with a range of L-arabinose [0%, 0.02%, 0.063%, 0.2%, 0.63%, 2%]. Statistical difference of the 0.2% L-arabinose induced samples between native and variants were determined by a two-tailed t-test. Error bars show the SD (n = 2); p-value summary: ****p < 0.0001, ***p < 0.001, *p < 0.05.m (n=2). FIGs.3A-3B: PRhavalidations in E. coli. FIG.3A: Single mutant promoters were induced with 0% and 0.2% L-rhamnose (w / v) and the median gfp expression was measured after 5 hours after induction by flow cytometry. FIG.3B: Single mutant promoters with improved promoter expression strength were combined and assessed 24 hours after induction with a range 4 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 of L-rhamnose [0%, 0.02%, 0.063%, 0.2%, 0.63%, 2%]. Statistical difference of the 0.2% L- rhamnose induced samples between native and variants were determined by a two-tailed t-test. Error bars show the SD (n=2); p-value summary: ****p < 0.0001, **p < 0.002, *p < 0.05. FIGs.4A-4B: Promoter library vector PBAD (FIG.4B) and PRha (FIG.4A) insert map. FIGs.5A-5D: Mechanisms and schematics of Clostridium inducible promoters. FIG. 5A: Lactose-inducible promoter, PBgaLis controlled by transcription factor BgaR. In the presence of lactose, BgaR binds to activate gfp expression. FIG.5B: Arabinose-inducible promoter, ARAi system is controlled by transcription factor AraR. AraR stays bound to its putative binding site, PPTK, to repress transcription. In the presence of arabinose, AraR releases from the PPTKpromoter to allow the RNA polymerase to activate transcription. FIG.5C (Anhydro)tetracycline-inducible promoter, PTETO1 is controlled by transcription factor TetR. TetR stays bound to its binding site, tetO1, in the PTETO1promoter to repress transcription. The presence of tetracycline or anhydrotetracyline (aTc) releases TetR from the binding to allow RNA polymerase to bind to activate gfp expression. FIG.5D: Schematics of additional modified promoter versions for PBgaLand ARAi system. In the PBgaLv2 promoter, PlacI promoter was inserted between the PBgaRand BgaR gene to control BgaR expression. Two additional versions were constructed for the ARAi system. In the second version of ARAi, Para promoter, the PPTK promoter was deleted. In the third version, the entire ParaR-Paradivergent promoter was replaced with PlacIto control araR gene expression. Promoter libraries for sort-seq were generated with error-prone PCR (epPCR) using PBgaLv1, PBgaLv2, PPTK, and PTETO1 versions of the native plasmid constructs. The epPCR regions for the four promoter libraries are indicated underneath the schematics of the chosen promoter versions. FIGs.6A-6C: PBgaL activity in E. coli NEB5α under lactose and IPTG. Lactose (left column) and IPTG inducers (right column) at (FIG.6A) 3 hours, (FIG.6B) 5 hours, and (FIG. 6C) 7 hours after induction. (n=2). As a negative control, the BgaR gene was deleted from the plasmid resulting in the ∆BgaR construct. FIGs.7A-7C: Initial evaluation of ARAi and additional promoter versions in E. coli NEB5α. Under arabinose inducer at (FIG.7A) 4 hours, (FIG.7B) 6 hours, and (FIG.7C) 22 hours post-induction. Two negative controls with the araR gene deleted from ARAi and Para constructs, denoted by ∆araR, were also tested. (n=1). FIGs.8A-8B: Evaluation of PPTK and ∆araR in both NEB5α and E. coli ∆araC KO strain from the Keio Collection, JW0063-1, at (FIG.8A) 5.5 hours and (FIG.8B) 24 hours after arabinose induction (0%, 0.02%, 0.06%, 0.2%, 0.63%, and 2% w / v). (n=2). 5 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 FIGs.9A-9C: PTETO1and ∆TetR activity in E. coli NEB5α under tetracycline (Tc) inducer at (FIG.9A) 1.8 hours, (FIG.9B) 3.5 hours, and (FIG.9C) 5 hours post-induction. (n=2). FIGs.10A-10C: PTETO1 activity in E. coli NEB5α under anhydrotetracycline (Tc) inducer at (FIG.10A) 6 hours, (FIG.10B) 8.5 hours, and (FIG.10C) 24 hours post-induction. (n=2). FIG.11: Individual PBgaLpromoter validations in E. coli NEB5α under 0 mM and 5.84 mM lactose. (n=2). The promoter variants are denoted as the native sequence(position)mutant sequence. * denotes promoter variants chosen for validations in C. acetobutylicum ATCC 824. FIGs.12A-12C: PBgaL validations in C. acetobutylicum under 0 mM, 2 mM, 6.3 mM, and 20 mM lactose at t=10 hours (or t=8 hours post-induction). FIG.12A YFAST fluorescence labeled with 5 µM TFLime in the positive gate, FIG.12B OD600, FIG.12C percent of cells in negative and positive gates set by pMTL85141. (n=2). Data measured at t=26 hours can be found in FIGs.16A-16C. FIG.13: PPTK validations in E. coli ∆araC JW00063-1 under 0% and 0.2% arabinose after 5 hours post-induction. (n=2). FIG.14: PTETO1validations in E. coli NEB5α under 0 ng / mL and 600 ng / mL aTc. (n=2). FIGs.15A-15C: Clostridium Promoter Plasmid Maps for the PPTK (FIG.15A), PBgaLv1 and PBgaLv2(FIG.15B), and PTETO1and PTETO2(FIG.15C) plasmids. The PTETO1plasmid contains an RBS and TetR sequence from E. coli amplified from the PTETO2 plasmid. *** denotes plasmids used for the rest of the study. FIGs.16A-16C: PBgaLvalidations in C. acetobutylicum at t=26 hours. YFAST fluorescence (FIG.16A), OD600(FIG.16B), percent cells in the positive and negatives (FIG. 16C). FIG.17: Native promoter sequences and modified promoters of the present disclosure. FIG.18: Exemplary pBAD plasmid sequence. FIG.19: Exemplary pBgal plasmid sequence. FIG.20: Exemplary pPTK plasmid sequence. FIG.21: Exemplary pRha plasmid sequence. FIG.22: Exemplary pTETO1 plasmid sequence. 6 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 DETAILED DESCRIPTION It is to be appreciated that certain aspects, modes, embodiments, variations and features of the present methods are described below in various levels of detail in order to provide a substantial understanding of the present technology. Disclosed herein are significant improvements in gene expression regulation over those that are commonly used and / or commercially available through the modified arabinose- (PBAD), rhamnose (PRha), lactose- (PBgaL), arabinose- (PPTK), and anhydrotetracycline- (PTETO1) inducible promoters or biosensors. The well-known and commonly used native promoter forms were engineered with enhancing mutations that were obtained by employing a high-throughput method to map functional regulatory sites and sequences at the nucleotide level. Furthermore, the expression vectors also include mutations that exhibit user-defined phenotypes and is particularly beneficial for the cost-effective and efficient production, modulation of any gene of interest, or circuit design. In practicing the present methods, many conventional techniques in molecular biology, protein biochemistry, cell biology, immunology, microbiology and recombinant DNA are used. See, e.g., Sambrook and Russell eds. (2001) Molecular Cloning: A Laboratory Manual, 3rd edition; the series Ausubel et al. eds. (2007) Current Protocols in Molecular Biology; the series Methods in Enzymology (Academic Press, Inc., N.Y.); MacPherson et al. (1991) PCR 1: A Practical Approach (IRL Press at Oxford University Press); MacPherson et al. (1995) PCR 2: A Practical Approach; Harlow and Lane eds. (1999) Antibodies, A Laboratory Manual; Freshney (2005) Culture of Animal Cells: A Manual of Basic Technique, 5th edition; Gait ed. (1984) Oligonucleotide Synthesis; U.S. Patent No.4,683,195; Hames and Higgins eds. (1984) Nucleic Acid Hybridization; Anderson (1999) Nucleic Acid Hybridization; Hames and Higgins eds. (1984) Transcription and Translation; Immobilized Cells and Enzymes (IRL Press (1986)); Perbal (1984) A Practical Guide to Molecular Cloning; Miller and Calos eds. (1987) Gene Transfer Vectors for Mammalian Cells (Cold Spring Harbor Laboratory); Makrides ed. (2003) Gene Transfer and Expression in Mammalian Cells; Mayer and Walker eds. (1987) Immunochemical Methods in Cell and Molecular Biology (Academic Press, London); and Herzenberg et al. eds (1996) Weir’s Handbook of Experimental Immunology. Methods to detect and measure levels of polypeptide gene expression products (i.e., gene translation level) are well-known in the art and include the use of polypeptide detection methods such as antibody detection and quantification techniques. (See also, Strachan & Read, Human Molecular Genetics, Second Edition. (John Wiley and Sons, Inc., NY, 1999)). 7 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Definitions Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this technology belongs. As used in this specification and the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the content clearly dictates otherwise. For example, reference to “a cell” includes a combination of two or more cells, and the like. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry, analytical chemistry and nucleic acid chemistry and hybridization described below are those well-known and commonly employed in the art. As used herein, the term “about” in reference to a number is generally taken to include numbers that fall within a range of 1%, 5%, or 10% in either direction (greater than or less than) of the number unless otherwise stated or otherwise evident from the context (except where such number would be less than 0% or exceed 100% of a possible value). As used herein, the terms “amplify” or “amplification” with respect to nucleic acid sequences, refer to methods that increase the representation of a population of nucleic acid sequences in a sample. Nucleic acid amplification methods are well known to the skilled artisan and include ligase chain reaction (LCR), ligase detection reaction (LDR), ligation followed by Q-replicase amplification, PCR, primer extension, strand displacement amplification (SDA), hyperbranched strand displacement amplification, multiple displacement amplification (MDA), nucleic acid strand-based amplification (NASBA), two-step multiplexed amplifications, rolling circle amplification (RCA), recombinase- polymerase amplification (RPA)(TwistDx, Cambridge, UK), transcription mediated amplification, signal mediated amplification of RNA technology, loop-mediated isothermal amplification of DNA, helicase-dependent amplification, single primer isothermal amplification, and self- sustained sequence replication (3SR), including multiplex versions or combinations thereof. Copies of a particular nucleic acid sequence generated in vitro in an amplification reaction are called “amplicons” or “amplification products.” The terms “complementary” or “complementarity” as used herein with reference to polynucleotides (i.e., a sequence of nucleotides such as an oligonucleotide or a target nucleic acid) refer to the base-pairing rules. The complement of a nucleic acid sequence as used herein refers to an oligonucleotide which, when aligned with the nucleic acid sequence such that the 5' end of one sequence is paired with the 3’ end of the other, is in “antiparallel association.” For example, the sequence “5'-A-G-T-3’” is complementary to the sequence “3’-T-C-A-5.” Certain 8 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 bases not commonly found in naturally-occurring nucleic acids may be included in the nucleic acids described herein. These include, for example, inosine, 7-deazaguanine, Locked Nucleic Acids (LNA), and Peptide Nucleic Acids (PNA). Complementarity need not be perfect; stable duplexes may contain mismatched base pairs, degenerative, or unmatched bases. Those skilled in the art of nucleic acid technology can determine duplex stability empirically considering a number of variables including, for example, the length of the oligonucleotide, base composition and sequence of the oligonucleotide, ionic strength and incidence of mismatched base pairs. A complement sequence can also be an RNA sequence complementary to the DNA sequence or its complement sequence, and can also be a cDNA. As used herein, “conjugation” refers to the temporary direct contact between two bacterial cells leading to an exchange of genetic material (DNA). This exchange is unidirectional, i.e. one bacterial cell is the donor of DNA and the other is the recipient. In this way, genes are transferred laterally amongst existing bacterial as opposed to vertical gene transfer in which genes are passed on to offspring. Conjugation is a convenient means for transferring genetic material to bacteria. As used herein, “expression” includes one or more of the following: transcription of the gene into precursor mRNA; splicing and other processing of the precursor mRNA to produce mature mRNA; mRNA stability; translation of the mature mRNA into protein (including codon usage and tRNA availability); and glycosylation and / or other modifications of the translation product, if required for proper expression and function. As used herein, an “expression control sequence” refers to polynucleotide sequences which are necessary to affect the expression of coding sequences to which they are operably linked. Expression control sequences are sequences which control the transcription, post- transcriptional events and translation of nucleic acid sequences. Expression control sequences include appropriate transcription initiation, termination, promoter and enhancer sequences; efficient RNA processing signals such as splicing and polyadenylation signals; sequences that stabilize cytoplasmic mRNA; sequences that enhance translation efficiency (e.g., ribosome binding sites); sequences that enhance protein stability; and when desired, sequences that enhance protein secretion. The nature of such control sequences differs depending upon the host organism; in prokaryotes, such control sequences generally include promoter, ribosomal binding site, and transcription termination sequence. The term “control sequences” is intended to encompass, at a minimum, any component whose presence is essential for expression, and can 9 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 also encompass an additional component whose presence is advantageous, for example, leader sequences. “Gene” as used herein refers to a DNA sequence that comprises regulatory and coding sequences necessary for the production of an RNA, which may have a non-coding function (e.g., a ribosomal or transfer RNA) or which may include a polypeptide or a polypeptide precursor. The RNA or polypeptide may be encoded by a full length coding sequence or by any portion of the coding sequence so long as the desired activity or function is retained. Although a sequence of the nucleic acids may be shown in the form of DNA, a person of ordinary skill in the art recognizes that the corresponding RNA sequence will have a similar sequence with the thymine being replaced by uracil, i.e., "T" is replaced with "U." As used herein, the term “genome” refers to the whole hereditary information of an organism that is encoded in the DNA (or RNA for certain viral species) including both coding and non-coding sequences. In various embodiments, the term may include the chromosomal DNA of an organism and / or DNA that is contained in an organelle such as, for example, the mitochondria or chloroplasts and / or extrachromosomal plasmid and / or artificial chromosome. As used herein, the term “group II intron” refers to a class of bacterial retrotransposons that insert site-specifically into DNA target sites by a mechanism termed “retrohoming” in which the excised intron RNA reverse splices into a DNA strand and is reverse transcribed by the intron-encoded protein (a reverse transcriptase). Retrohoming is mediated by a ribonucleoprotein particle that contains the intron-encoded protein and excised intron RNA, with target specificity determined largely by base pairing of the intron RNA to the DNA target sequence. This feature enabled the development of mobile group II introns into bacterial gene targeting vectors (“targetrons”) with programmable target specificity. The term “guide sequence” refers to the portion of a crRNA or guide RNA (gRNA) that is responsible for hybridizing with the target DNA. As used herein, a “heterologous nucleic acid sequence” is any nucleic acid sequence placed at a location where it does not normally occur. A heterologous nucleic acid sequence may comprise a sequence that does not naturally occur in a cell, or it may comprise only sequences naturally found in the cell, but placed at a non-normally occurring location in the cell. In some embodiments, the heterologous nucleic acid sequence is not an endogenous sequence. In certain embodiments, the heterologous nucleic acid sequence is an endogenous sequence that is derived from a different cell. In other embodiments, the heterologous nucleic acid sequence is 10 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 a sequence that occurs naturally in a cell but is then relocated to another site where it does not naturally occur, rendering it a heterologous sequence at that new site. “Homology” or “identity” or “similarity” refers to sequence similarity between two peptides or between two nucleic acid molecules. Homology can be determined by comparing a position in each sequence which may be aligned for purposes of comparison. When a position in the compared sequence is occupied by the same base or amino acid, then the molecules are homologous at that position. A degree of homology between sequences is a function of the number of matching or homologous positions shared by the sequences. A polynucleotide or polynucleotide region (or a polypeptide or polypeptide region) has a certain percentage (for example, at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98% or 99%) of “sequence identity” to another sequence means that, when aligned, that percentage of bases (or amino acids) are the same in comparing the two sequences. This alignment and the percent homology or sequence identity can be determined using software programs known in the art. In some embodiments, default parameters are used for alignment. One alignment program is BLAST, using default parameters. In particular, programs are BLASTN and BLASTP, using the following default parameters: Genetic code=standard; filter=none; strand=both; cutoff=60; expect=10; Matrix=BLOSUM62; Descriptions=50 sequences; sort by ═HIGH SCORE; Databases=non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+SwissProtein+SPupdate+PIR. Details of these programs can be found at the National Center for Biotechnology Information. Biologically equivalent polynucleotides are those having the specified percent homology and encoding a polypeptide having the same or similar biological activity. Two sequences are deemed “unrelated” or “non-homologous” if they share less than 40% identity, or less than 25% identity, with each other. As used herein, the phrase “homologous recombination” refers to the process in which nucleic acid molecules with similar nucleotide sequences associate and exchange nucleotide strands. A nucleotide sequence of a first nucleic acid molecule that is effective for engaging in homologous recombination at a predefined position of a second nucleic acid molecule can therefore have a nucleotide sequence that facilitates the exchange of nucleotide strands between the first nucleic acid molecule and a defined position of the second nucleic acid molecule. Thus, the first nucleic acid can generally have a nucleotide sequence that is sufficiently complementary to a portion of the second nucleic acid molecule to promote nucleotide base pairing. Homologous recombination requires homologous sequences in the two recombining partner nucleic acids but does not require any specific sequences. Homologous recombination can be used to introduce a heterologous nucleic acid and / or mutations into the host genome. Such 11 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 systems typically rely on sequence flanking the heterologous nucleic acid to be expressed that has enough homology with a target sequence within the host cell genome that recombination between the vector nucleic acid and the target nucleic acid takes place, causing the delivered nucleic acid to be integrated into the host genome. These systems and the methods necessary to promote homologous recombination are known to those of skill in the art. The term “hybridize” as used herein refers to a process where two substantially complementary nucleic acid strands (at least about 65% complementary over a stretch of at least 14 to 25 nucleotides, at least about 75%, or at least about 90% complementary) anneal to each other under appropriately stringent conditions to form a duplex or heteroduplex through formation of hydrogen bonds between complementary base pairs. Hybridizations are typically and preferably conducted with probe-length nucleic acid molecules, preferably 15-100 nucleotides in length, more preferably 18-50 nucleotides in length. Nucleic acid hybridization techniques are well known in the art. See, e.g., Sambrook, et al., 1989, Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Press, Plainview, N.Y. Hybridization and the strength of hybridization (i.e., the strength of the association between the nucleic acids) is influenced by such factors as the degree of complementarity between the nucleic acids, stringency of the conditions involved, and the thermal melting point (Tm) of the formed hybrid. Those skilled in the art understand how to estimate and adjust the stringency of hybridization conditions such that sequences having at least a desired level of complementarity will stably hybridize, while those having lower complementarity will not. For examples of hybridization conditions and parameters, see, e.g., Sambrook, et al., 1989, Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Press, Plainview, N.Y.; Ausubel, F. M. et al.1994, Current Protocols in Molecular Biology, John Wiley & Sons, Secaucus, N.J. In some embodiments, specific hybridization occurs under stringent hybridization conditions. An oligonucleotide or polynucleotide (e.g., a probe or a primer) that is specific for a target nucleic acid will “hybridize” to the target nucleic acid under suitable conditions. As used herein, the terms “individual”, “patient”, or “subject” are used interchangeably and refer to an individual organism, a vertebrate, a mammal, or a human. In a preferred embodiment, the individual, patient or subject is a human. As used herein, “microbiome” refers to the collective genetic content of the communities of microbes that live in and on the human body, both sustainably and transiently, including eukaryotes, fungi, archaea, bacteria, and viruses (including bacterial viruses (i.e., phage)), wherein “genetic content” includes genomic DNA, RNA such as micro RNA and ribosomal 12 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 RNA, the epigenome, plasmids, and all other types of genetic information. As used herein, the term “gut microbiome” refers to the collective genetic content of the communities of microbes present in the gastrointestinal tract (GIT). As used herein, “microbiota” refers to the collective microbes that live in and on the human body, both sustainably and transiently, including eukaryotes, fungi, archaea, bacteria, and viruses (including bacterial viruses (i.e., phage)). “Gut microbiota” as used herein refers to the totality of the microbes present in the GIT, including eukaryotes, fungi, archaea, bacteria, and viruses (including bacterial viruses (i.e., phage)). As used herein, “oligonucleotide” refers to a molecule that has a sequence of nucleic acid bases on a backbone comprised mainly of identical monomer units at defined intervals. The bases are arranged on the backbone in such a way that they can bind with a nucleic acid having a sequence of bases that are complementary to the bases of the oligonucleotide. The most common oligonucleotides have a backbone of sugar phosphate units. A distinction may be made between oligodeoxyribonucleotides that do not have a hydroxyl group at the 2' position and oligoribonucleotides that have a hydroxyl group at the 2' position. Oligonucleotides may also include derivatives, in which the hydrogen of the hydroxyl group is replaced with organic groups, e.g., an allyl group. Oligonucleotides of the method which function as primers or probes are generally at least about 10-15 nucleotides long and more preferably at least about 15 to 25 nucleotides long, although shorter or longer oligonucleotides may be used in the method. The exact size will depend on many factors, which in turn depend on the ultimate function or use of the oligonucleotide. The oligonucleotide may be generated in any manner, including, for example, chemical synthesis, DNA replication, restriction endonuclease digestion of plasmids or phage DNA, reverse transcription, PCR, or a combination thereof. The oligonucleotide may be modified e.g., by addition of a methyl group, a biotin or digoxigenin moiety, a fluorescent tag or by using radioactive nucleotides. As used herein, “operably linked” means that expression control sequences are positioned relative to a nucleic acid of interest to initiate, regulate or otherwise control transcription of the nucleic acid of interest. In some embodiments, transcription of a polynucleotide operably linked to an expression control element (e.g., a promoter) is controlled, regulated, or influenced by the expression control element. As used herein, the term “polynucleotide” or “nucleic acid” means any RNA or DNA, which may be unmodified or modified RNA or DNA. Polynucleotides include, without limitation, single- and double-stranded DNA, DNA that is a mixture of single- and double- 13 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 stranded regions, single- and double-stranded RNA, RNA that is mixture of single- and double- stranded regions, and hybrid molecules comprising DNA and RNA that may be single-stranded or, more typically, double-stranded or a mixture of single- and double-stranded regions. In addition, polynucleotide refers to triple-stranded regions comprising RNA or DNA or both RNA and DNA. The term polynucleotide also includes DNAs or RNAs containing one or more modified bases and DNAs or RNAs with backbones modified for stability or for other reasons. As used herein, the term “primer” refers to an oligonucleotide, which is capable of acting as a point of initiation of nucleic acid sequence synthesis when placed under conditions in which synthesis of a primer extension product which is complementary to a target nucleic acid strand is induced, i.e., in the presence of different nucleotide triphosphates and a polymerase in an appropriate buffer (“buffer” includes pH, ionic strength, cofactors etc.) and at a suitable temperature. One or more of the nucleotides of the primer can be modified for instance by addition of a methyl group, a biotin or digoxigenin moiety, a fluorescent tag or by using radioactive nucleotides. A primer sequence need not reflect the exact sequence of the template. For example, a non-complementary nucleotide fragment may be attached to the 5′ end of the primer, with the remainder of the primer sequence being substantially complementary to the strand. The term primer as used herein includes all forms of primers that may be synthesized including peptide nucleic acid primers, locked nucleic acid primers, phosphorothioate modified primers, labeled primers, and the like. The term “forward primer” as used herein means a primer that anneals to the anti-sense strand of dsDNA. A “reverse primer” anneals to the sense-strand of dsDNA. As used herein, “primer pair” refers to a forward and reverse primer pair (i.e., a left and right primer pair) that can be used together to amplify a given region of a nucleic acid of interest. The term “promoter” as used herein refers to any sequence that regulates the expression of a coding sequence, such as a gene. Promoters may be constitutive, inducible, repressible, or tissue-specific, for example. A “promoter” is a control sequence that is a region of a polynucleotide sequence at which initiation and rate of transcription are controlled. It may contain genetic elements at which regulatory proteins and molecules may bind such as RNA polymerase and other transcription factors. As used herein, the term “recombinant” when used with reference, e.g., to a cell, or nucleic acid, protein, or vector, indicates that the cell, nucleic acid, protein or vector, has been modified by the introduction of a heterologous nucleic acid or protein or the alteration of a native nucleic acid or protein, or that the material is derived from a cell so modified. Thus, for 14 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 example, recombinant cells express genes that are not found within the native (non-recombinant) form of the cell or express native genes that are otherwise abnormally expressed, under expressed or not expressed at all. As used herein, an endogenous nucleic acid sequence in the cell of an organism (or the encoded protein product of that sequence) is deemed “recombinant” herein if a heterologous sequence is placed adjacent to the endogenous nucleic acid sequence, such that the expression of this endogenous nucleic acid sequence is altered. In this context, a heterologous sequence is a sequence that is not naturally adjacent to the endogenous nucleic acid sequence, whether or not the heterologous sequence is itself endogenous to the organism (originating from the same organism or progeny thereof) or exogenous (originating from a different organism or progeny thereof). By way of example, a promoter sequence can be substituted (e.g., by homologous recombination) for the native promoter of a gene in the cell of an organism, such that this gene has an altered expression pattern. This gene would be “recombinant” because it is separated from at least some of the sequences that naturally flank it. A nucleic acid is also considered “recombinant” if it contains any modifications that do not naturally occur in the corresponding nucleic acid in a cell. For instance, an endogenous coding sequence is considered “recombinant” if it contains an insertion, deletion or a point mutation introduced artificially, e.g., by human intervention. A “recombinant nucleic acid” also includes a nucleic acid integrated into a host cell chromosome at a heterologous site and a nucleic acid construct present as an episome. As used herein, the term “replication origins”, “origins of replications” or “rep origins” refers to a unique DNA sequence of a replicon at which DNA replication is initiated and proceeds bidirectionally or unidirectionally. It contains the sites where the first separation of the complementary strands occurs, a primer RNA is synthesized, and the switch from primer RNA to DNA synthesis takes place. As used herein, a “reporter gene” refers to a polynucleotide sequence encoding a gene product (e.g., polypeptide) that can generate, under appropriate conditions, a detectable signal that allows detection of the presence and / or quantity of the gene product. Reporter genes are often used as an indication of whether a certain gene has been introduced into or expressed in the host cell or organism. Examples of commonly used reporters include: antibiotic resistance genes, fluorescent proteins, auxotrophic selection modules, β-galactosidase (encoded by the bacterial gene lacZ), luciferase (from lightning bugs), chloramphenicol acetyltransferase (CAT; from bacteria), GUS (β-glucuronidase; commonly used in plants) and green fluorescent protein (GFP; from jelly fish). Reporters or selection modules can be selectable or screenable. 15 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 As used herein, “selection marker” refers to a gene that confers a trait suitable for artificial selection. Typically host cells expressing the selectable selection marker is protected from a selective agent that is toxic or inhibitory to cell growth. Examples of commonly used selective markers include antibiotic resistance genes. A screenable selection marker (e.g., gfp, lacZ) generally allows researchers to distinguish between wanted cells (expressing the selection module) and unwanted cells (not expressing the selection module or expressing at insufficient level). The term “stringent hybridization conditions” as used herein refers to hybridization conditions at least as stringent as the following: hybridization in 50% formamide, 5xSSC, 50 mM NaH2PO4, pH 6.8, 0.5% SDS, 0.1 mg / mL sonicated salmon sperm DNA, and 5x Denhart's solution at 42oC. overnight; washing with 2x SSC, 0.1% SDS at 45oC; and washing with 0.2x SSC, 0.1% SDS at 45oC. In another example, stringent hybridization conditions should not allow for hybridization of two nucleic acids which differ over a stretch of 20 contiguous nucleotides by more than two bases. As used herein, “16S ribosomal RNA” or “16S rRNA”, is a component of the prokaryotic ribosome 30S subunit. The 16S rRNA gene is the DNA sequence corresponding to rRNA encoding bacteria, which exists in the genome of all bacteria. 16S rRNA is highly conserved and specific, and the gene sequence is long enough (about 1,500 base pairs) for informatics purposes. 16S rRNA sequences are used for phylogenetic reconstruction as they are generally highly conserved, but contain specific hypervariable regions that harbor sufficient nucleotide diversity to differentiate genera and species of most bacteria. As used herein, a "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. One type of vector is a "plasmid," which generally refers to a circular double stranded DNA loop into which additional DNA segments may be ligated, but also includes linear double- stranded molecules such as those resulting from amplification by the polymerase chain reaction (PCR) or from treatment of a circular plasmid with a restriction enzyme. Other vectors include cosmids, bacterial artificial chromosomes (BAC) and yeast artificial chromosomes (YAC). Another type of vector is a viral vector, wherein additional DNA segments may be ligated into the viral genome. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., vectors having an origin of replication which functions in the host cell). Other vectors can be integrated into the genome of a host cell upon introduction into the host cell, and are thereby replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of 16 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 genes to which they are operatively linked. Such vectors are referred to herein as "recombinant expression vectors" (or simply "expression vectors"). Promoter Systems pBAD and pRha The E. coli arabinose- (PBAD) and rhamnose-inducible promoters (PRha) are well-studied and widely used promoter systems. Both promoters provide tight expression while exhibiting a wide dynamic range in the presence of their respective inducers, L-arabinose and L-rhamnose, making them attractive for diverse biotechnological applications [1-6]. Moreover, they have also been adapted to respond to various effector ligands [7-9]. Transcription factors (TF) AraC and RhaS are primary regulators of PBAD and PRha, respectively, with both promoters involving cyclic AMP receptor protein (CRP) to enhance expression, including their own regulator expression from their respective divergent operons, PCand PRhaRS[10, 11]. Although PBADand PRhapromoters share a similar structural arrangement of their binding sites on the promoter
[0012] , they are unique in their mechanisms (FIGs.1A-1B). PBAD becomes transcriptionally active when AraC binds to araI1 and araI2 half-sites
[0012] (FIG.1A). In the absence of arabinose, AraC occupies two sites that are separated by 210 bp, araO2 and araI1 half-sites, leading to a DNA loop structure that represses transcription
[0012] . In the presence of arabinose, AraC undergoes conformational change, releasing its binding to araO2. The DNA loop opens with the assistance of CRP, which binds upstream of the araI half-sites, allowing AraC to bind to both araI1 and araI2, thereby recruiting RNA polymerase
[0012] . As the AraC concentration increases, expression is downregulated by binding to araO1 located in the Pc promoter. While AraC plays a dual role as activator and repressor in regulating PBAD, activation of PRha is controlled by the binding of rhamnose-bound RhaS to rhaI1 and rhaI2 half-sites. While RhaS can activate transcription independently, full activation is achieved when the co-activator CRP binds upstream of the rhaI half-sites (FIG.1B). Expression of RhaS from PRhaRSis regulated by RhaR which also responds to rhamnose. Together, the presence of rhamnose triggers RhaR to bind to PRhaRS, co-activated by CRP, and expresses both rhaR and rhaS. Subsequently, rhamnose bound RhaS binds to PRha, co-activated by its own CRP located upstream, to activate transcription. Despite the precise control offered by PBAD and PRha promoters, they may not always be sufficient for certain applications, such as achieving maximum protein expression [13-15]. Instead, these promoters have been utilized to provide titratable expression of the T7 RNA polymerase or LacI protein, which binds to the lacO within the T7 promoter system derived from 17 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 the T7 bacteriophage, thus reducing leakiness in commercial strains (BL21-AI from Thermofisher and KRX Competent Cells from Promega) [2, 3]. Expression levels of PBADand PRha are orders in magnitude lower than those of T7 promoter systems
[0014] . While various improvements have been made to enhance the dynamic range, sensitivity, and other characteristics, these improvements have mostly focused on altering other components of the system without modifying the promoters themselves, such as varying TF expression by using different promoters [16, 17], optimizing RBS [18-20], and altering plasmid copy number [5, 21]. Adjusting expression levels of the arabinose transporter gene, araE, can tune the expression strength of PBAD[22, 23]. Similarly, tunability of PRhahas been achieved in a RhaT-mediated rhamnose transport and catabolism double knockout E. coli strain [6]. On the protein level, increase in the sensitivity of AraC to arabinose was achieved to reduce its response to IPTG, allowing simultaneous use PBADand lactose-inducible promoter (PLac) at the protein level
[0024] . Previous research focusing directly on the characterization and improvements of PBAD and PRha promoters has mainly concentrated on the parts level, isolating short sequence motifs [25-32]. Notably, reduction of basal activity in PBADhas been achieved through the assembly of hybrid promoters that combine the araI half-sites with repressor bindings sites, such as lacO and tetO, inserted within non-native -35 and -10 boxes. This arrangement repressed basal activity through the binding of LacI or TetR repressor but activated with the addition of lactose and arabinose [27, 29]. However, such approaches have resulted in only moderate improvements induced expression strength. The comprehensive understanding of the regulation and characterization of PBADand PRhapromoters has been achieved through decades of study involving a laborious techniques and experiments [10, 11, 31, 32]. While these studies have been essential for the utilization of PBAD and PRha in synthetic biology applications, a more comprehensive assessment using high- throughput methodologies is needed to thoroughly investigate the known binding sites, including the surrounding regions that were most often considered non-functional
[0033] . Massively parallel reporter assays, such as sort-seq, in combination with statistical methods, can systemically interrogate a library of pooled sequences based on the expressed phenotype, thereby identifying functional sites in the sequence space under a specific condition [34-36]. Clostridium Promoters Clostridium spp. represents one of the many biotechnological platforms positioned to address the global climate crisis through the sustainable production of valuable chemicals. Their inherent capability to tolerate and convert diverse substrates into valuable chemicals through 18 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 their numerous and diverse pathways has rekindled interest in leveraging Clostridium spp. as cell factories [63, 64]. Furthermore, recombinant proteins produced by Clostridium strains are considered endotoxin-free making them ideal expression platforms for the production of therapeutic enzymes and proteins
[0065] . In addition, their ability to germinate and thrive exclusively in hypoxic and necrotic areas of tumors has made them attractive hosts as therapeutic vectors for cancer therapy
[0066] . However, despite the significant achievements made in fermentation processes over the past century
[0067] and their expanding applications in potential medical fields, gene expression tools in Clostridium remain relatively limited compared to other industrially relevant aerotolerant microorganisms [68, 69]. Promoters are one of many essential synthetic biology tools that can accelerate various applications utilizing Clostridium as a platform for gene function studies, expression modulation, pathway engineering, and drug delivery. Although constitutive and inducible promoters are available for Clostridium strains, they remain insufficiently characterized. While efforts have been made to improve constitutive promoters such as the thiolase (thl) promoter [70, 71], inducible promoters lack the same level of development. Massively parallel reporter assays (MPRAs) have significantly expanded the repertoire and performance of promoters and regulatory genetic parts in well-studied organisms like E. coli and emerging organisms [72, 73]. However, implementing high-throughput workflows in non- model anaerobic microorganisms like those within the Clostridium genus are lagging, resulting in the frequent use of suboptimal inducible promoters in their native forms, which remain relatively uncharacterized
[0074] . Efforts to enhance promoter properties in solventogenic clostridia have been limited [70, 71, 75]. Challenges arise concerning the availability and use of oxygen-independent reporters. While attempts have been made to construct promoter libraries using flavin-binding fluorescent proteins (FbFPs) like iLOV, phiLOV, creiLOV, BsFbFP, PpFbFP, and EcFbFP)
[0076] issues related to cell autofluorescence have prompted a return to the use of low-throughput enzymatic reporters such as gusA [71, 75]. Recent developments with the fluorescence-activating and absorption-shifting tag (FAST) protein show promise in various Clostridium species and other anaerobes [77, 78], including anaerobic thermophiles
[0079] . These advancements potentially enable high-throughput workflows for studying crucial gene functions, pathway engineering, and biosensor development for high-throughput screening
[0078] . The fluorescence of the expressed FAST protein is detected upon binding with fluorogenic ligands, that are commercially available from Twinkle Bioscience S.A.S.). Similarly, HaloTag and SNAP-tag, another set of promising 19 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 anaerobic fluorescent reporters with a variety of multicolor fluorogens, have been used in combination with FAST to differentiate two different species in co-culture systems
[0080] . However, the development of genetic tools using MPRAs requires large libraries and efficient DNA transformation into the target organism for effective screening of desirable traits. Due to the highly active restriction and modification (RM) systems acting as the organism’s “immune system,” transformation efficiency remains a limiting factor [64, 81, 82] in introducing large-scale DNA libraries into Clostridium. Although existing methylation methods are available to protect foreign plasmid DNA from degradation through methyltransferases [64, 82], the efficiency is still low, requiring labor-intensive transformations to ensure a large enough library size. In this study, sort-seq was employed in E. coli to investigate three frequently used Clostridium inducible promoters, spanning from well-characterized to the less characterized, with the aim of identifying binding sites within Clostridium promoters. The objective was to identify mutations that can alter both the expression strength and basal activity of these promoters: the lactose-inducible promoter PBgaLand its activator BgaR from Clostridium perfringens [77, 83-86] (FIG.5A), the arabinose-inducible promoter PPTK(from the ARAi system
[0087] ) regulated by AraR repressor sourced from Clostridium acetobutylicum (FIG.5B), and (anhydro)tetracycline-inducible promoter PTETO1and its repressor TetR (FIG.5C)
[0088] . Despite their frequent use, no attempts have been made to improve the dynamic range properties of PBgaL and PPTK promoters. While the gram-positive version of the PTETO1 promoter has undergone improvements in previous studies, primarily at the parts level, involving the insertion of one versus two tetO operators in a strong promoter, PxylAfrom Bacillus subtilis, and modifying the -10 box of the promoter controlling tetR expression, PTetR, to increase TetR expression for higher induction of the Pxyl / tet promoter, which is renamed in this study as PTETO1
[0089] . The identified sequences are validated in E. coli and optionally in Clostridium acetobutylicum. This effort is designed to enhance the performance of existing inducible promoters. Expression Systems for Inducible Expression of Heterologous Genes and Methods of Use Thereof In one aspect, the present disclosure provides a bacterial expression system nucleic acid sequence comprising (a) a mutant inducible promoter sequence operably linked to a heterologous nucleic acid and (b) a promoter operably linked to a gene encoding a repressor protein, wherein the repressor protein is configured to induce an RNA polymerase to bind to bind to the mutant inducible promoter in the presence of an inducing agent. Examples of mutant inducible 20 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 promoters include pBAD, pRha, pBgal, pPTK, and pTETO1 promoters. Examples of genes encoding a repressor protein which represses the inducible promoter includes the AraC gene, the RhaS gene, the BgaR gene, the AraR gene, and the TetR gene. In some embodiments, the inducing agent is arabinose, rhamnose, lactose, tetracycline or anhydrotetracycline. In one embodiment, the mutant inducible promoter is pBAD, the gene encoding a repressor protein which represses the inducible promoter is the AraC gene, and the inducing agent is arabinose. In one embodiment, the mutant inducible promoter is pRha, the gene encoding a repressor protein which represses the inducible promoter is the RhaS gene, and the inducing agent is rhamnose. In one embodiment, the mutant inducible promoter is pBgal, the gene encoding a repressor protein which represses the inducible promoter is the BgaR gene, and the inducing agent is lactose. In one embodiment, the mutant inducible promoter is pPTK, the gene encoding a repressor protein which represses the inducible promoter is the AraR gene, and the inducing agent is arabinose. In one embodiment, the mutant inducible promoter is pTETO1, the gene encoding a repressor protein which represses the inducible promoter is the TetR gene, and the inducing agent is tetracycline or anhydrotetracycline. In some embodiments, the mutant inducible promoter sequence comprises any one of the mutant promoters in FIG.17. In one aspect, the present disclosure provides an expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pBAD promoter sequence of any one of SEQ ID NOs: 1-14 that is operably linked to a heterologous gene and (b) a constitutive Pc promoter sequence that is operably linked to a gene sequence encoding an AraC protein, wherein the AraC protein is configured to induce an RNA polymerase to bind to the mutant pBAD promoter sequence in the presence of arabinose. In another aspect, the present disclosure provides an expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pRha promoter sequence of any one of SEQ ID NOs: 16-25 that is operably linked to a heterologous gene and (b) a constitutive PRhaRS promoter sequence that is operably linked to a gene sequence encoding a RhaS protein, wherein the RhaS protein is configured to induce an RNA polymerase to bind to the mutant pRha promoter sequence in the presence of rhamnose. In a different aspect, the present disclosure provides an expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pBgal promoter sequence of any one of SEQ ID NOs: 27-46 that is operably linked to a heterologous gene and (b) a BgaR promoter sequence that is operably linked to a gene sequence encoding a BgaR 21 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 protein, wherein the BgaR protein is configured to induce an RNA polymerase to bind to the mutant pBgal promoter sequence in the presence of lactose. In yet another aspect, the present disclosure provides an expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pPTK promoter sequence of any one of SEQ ID NOs: 48-67 that is operably linked to a heterologous gene and (b) an AraR promoter sequence that is operably linked to a gene sequence encoding an AraR protein, wherein the AraR protein is configured to induce an RNA polymerase to bind to the mutant pPTK promoter sequence in the absence of arabinose. In one aspect, the present disclosure provides an expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pTETO1 promoter sequence of any one of SEQ ID NOs: 69-87 that is operably linked to a heterologous gene and (b) a TetR promoter sequence that is operably linked to a gene sequence encoding a TetR protein, wherein the TetR protein is configured to induce an RNA polymerase to bind to the mutant pTETO1 promoter sequence in the absence of tetracycline or anhydrotetracyline (aTc). The expression systems of the present disclosure may be incorporated into an expression vector, such as, but not limited to, a plasmid, cosmid, bacterial artificial chromosome (BAC), or a yeast artificial chromosome (YAC). Additionally or alternatively, the expression system of the present disclosure may be integrated on a chromosome of a host cell. One of ordinary skill in the art would be able to select an appropriate delivery system or chromosome integration site to ensure expression system activity in a host cell. In some embodiments, the heterologous gene encodes a protein, an enzyme, a structural polypeptide, a toxin, a fusion protein, an antibody agent, a drug, a cytokine, an enzyme inhibitor, a growth factor, a signaling protein, a bioluminescent protein, a fluorescent protein, a chemiluminescent protein, a catalytic RNA, or an inhibitory RNA. In some embodiments, the heterologous gene encodes a protein. In some embodiments, the heterologous gene encodes an enzyme. In some embodiments, the heterologous gene encodes a structural polypeptide. In some embodiments, the heterologous gene encodes a toxin. In some embodiments, the heterologous gene encodes a fusion protein. In some embodiments, the heterologous gene encodes an antibody agent. In some embodiments, the heterologous gene encodes a drug. In some embodiments, the heterologous gene encodes a cytokine. In some embodiments, the heterologous gene encodes an enzyme inhibitor. In some embodiments, the heterologous gene encodes a growth factor. In some embodiments, the heterologous gene encodes a signaling protein. In some embodiments, the heterologous gene encodes a bioluminescent protein. In 22 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 some embodiments, the heterologous gene encodes a fluorescent protein. In some embodiments, the heterologous gene encodes a chemiluminescent protein. In some embodiments, the heterologous gene encodes a catalytic RNA. In some embodiments, the heterologous gene encodes an inhibitory RNA. In any of the preceding embodiments, the heterologous gene product may comprise a component that assists with purification of the gene product (e.g. an amino acid or nucleotide tag). In some embodiments, the expression systems of the present disclosure have increased expression upon induction and / or improved regulation of expression as compared to the same system with a wild type version of the mutant inducible promoter. In some embodiments, expression of the heterologous gene in the expression systems of the present technology is increased upon induction by about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 100%, about 150%, about 200%, about 250%, about 300%, about 350%, about 400%, about 450%, about 500%, about 550%, about 600%, about 650%, about 700%, about 750%, about 800%, about 850%, about 900%, about 950%, or about 1000%, as compared to a control expression system comprising a wild type inducible promoter operably linked to the heterologous nucleic acid. In some embodiments, expression of the heterologous gene in the expression systems of the present technology is increased upon induction by about 20% to about 700% as compared to the control expression system. In some embodiments, expression of the heterologous gene in the expression systems of the present technology is increased upon induction by about 1-fold, about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 6-fold, about 7-fold, about 8-fold, about 9-fold, or about 10-fold as compared to a control expression system comprising a wild type inducible promoter operably linked to the heterologous nucleic acid. Improved regulation of expression can be indicated, for example, by an increase in the dynamic range of the system (ratio of induced expression versus uninduced expression) as compared to a control expression system comprising a wild type inducible promoter operably linked to the heterologous nucleic acid. In some embodiments, the dynamic range of the expression systems of the present technology are increased by about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 100%, about 150%, about 200%, about 250%, about 300%, about 350%, about 400%, about 450%, or about 500%, as compared to a control expression system comprising a wild type inducible promoter operably linked to the heterologous nucleic acid. In some embodiments, the expression systems 23 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 of the present technology allow for tunable expression of a gene product, wherein induction of the mutant promoter allows for precise expression of the heterologous gene. In one aspect, the present disclosure provides a prokaryotic host cell comprising the expression system of any one of the preceding embodiments. In some embodiments, the host cell is a gram positive bacteria or a gram negative bacteria. In some embodiments, the host cell is a gram positive bacteria. In some embodiments, the host cell is a gram negative bacteria. In some embodiments the host cell is E. coli or Clostridium. In some embodiments, the host cell is E. coli. In some embodiments, the host cell is Clostridium. One of skill in the art will appreciate that the expression systems of the present disclosure may be used in any appropriate host cell. In another aspect, the present disclosure provides a nucleic acid construct comprising a mutant inducible promoter sequence operably linked to a heterologous gene, wherein the mutant promoter sequence is selected from the group consisting of SEQ ID NOs: 1-14, 16-25, 27-46, 48-67, and 69-87. In some embodiments, the nucleic acid construct is integrated on a chromosome of a host cell. Ine some embodiments, the nucleic acid construct is integrated in an expression vector. In specific embodiments, the expression vector is a plasmid, a cosmid, a bacterial artificial chromosome (BAC) or a yeast artificial chromosomes (YAC). In some embodiments, the heterologous gene encodes a protein, an enzyme, a structural polypeptide, a toxin, a fusion protein, an antibody agent, a drug, a cytokine, an enzyme inhibitor, a growth factor, a signaling protein, a bioluminescent protein, a fluorescent protein, a chemiluminescent protein, a catalytic RNA, or an inhibitory RNA. In some embodiments, the nucleic acid constructs of the present disclosure have increased expression upon induction and / or improved regulation of expression as compared to the same system with a wild type version of the mutant inducible promoter. In some embodiments, expression of the heterologous gene in the nucleic acid constructs of the present technology is increased upon induction by about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 100%, about 150%, about 200%, about 250%, about 300%, about 350%, about 400%, about 450%, about 500%, about 550%, about 600%, about 650%, about 700%, about 750%, about 800%, about 850%, about 900%, about 950%, or about 1000%, as compared to a control expression system comprising a wild type inducible promoter operably linked to the heterologous nucleic acid. In some embodiments, expression of the heterologous gene in the nucleic acid constructs of the present technology is increased upon induction by about 20% to about 700% as compared to the control expression system. In some 24 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 embodiments, expression of the heterologous gene in the nucleic acid constructs of the present technology is increased upon induction by about 1-fold, about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 6-fold, about 7-fold, about 8-fold, about 9-fold, or about 10-fold as compared to a control expression system comprising a wild type inducible promoter operably linked to the heterologous nucleic acid. Improved regulation of expression can be indicated, for example, by an increase in the dynamic range of the system (ratio of induced expression versus uninduced expression) as compared to a control expression system comprising a wild type inducible promoter operably linked to the heterologous nucleic acid. In some embodiments, the dynamic range of the nucleic acid constructs of the present technology are increased by about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 100%, about 150%, about 200%, about 250%, about 300%, about 350%, about 400%, about 450%, or about 500%, as compared to a control expression system comprising a wild type inducible promoter operably linked to the heterologous nucleic acid. In some embodiments, the nucleic acid constructs of the present technology allow for tunable expression of a gene product, wherein induction of the mutant promoter allows for precise expression of the heterologous gene. In some embodiments, the heterologous gene encodes a protein, an enzyme, a structural polypeptide, a toxin, a fusion protein, an antibody agent, a drug, a cytokine, an enzyme inhibitor, a growth factor, a signaling protein, a bioluminescent protein, a fluorescent protein, a chemiluminescent protein, a catalytic RNA, or an inhibitory RNA. In some embodiments, the heterologous gene encodes a protein. In some embodiments, the heterologous gene encodes an enzyme. In some embodiments, the heterologous gene encodes a structural polypeptide. In some embodiments, the heterologous gene encodes a toxin. In some embodiments, the heterologous gene encodes a fusion protein. In some embodiments, the heterologous gene encodes an antibody agent. In some embodiments, the heterologous gene encodes a drug. In some embodiments, the heterologous gene encodes a cytokine. In some embodiments, the heterologous gene encodes an enzyme inhibitor. In some embodiments, the heterologous gene encodes a growth factor. In some embodiments, the heterologous gene encodes a signaling protein. In some embodiments, the heterologous gene encodes a bioluminescent protein. In some embodiments, the heterologous gene encodes a fluorescent protein. In some embodiments, the heterologous gene encodes a chemiluminescent protein. In some embodiments, the heterologous gene encodes a catalytic RNA. In some embodiments, the heterologous gene encodes an inhibitory RNA. In any of the preceding embodiments, the heterologous gene 25 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 product may comprise a component that assists with purification of the gene product (e.g. an amino acid or nucleotide tag). In another aspect, the present disclosure includes a prokaryotic host cell comprising the nucleic acid construct of any one of the preceding embodiments. In some embodiments, the host cell is a gram-positive bacteria or a gram negative bacteria. In some embodiments, the host cell is a gram positive bacteria. In some embodiments, the host cell is a gram negative bacteria. In some embodiments the host cell is E. coli or Clostridium. In some embodiments, the host cell is E. coli. In some embodiments, the host cell is Clostridium. One of skill in the art will appreciate that the nucleic acid constructs of the present disclosure may be used in any appropriate host cell. In one aspect, the present disclosure provides a method for overexpressing a heterologous polypeptide in a prokaryotic cell comprising contacting the prokaryotic host cell of any one of the preceding embodiments with an effective amount of an inducer molecule, wherein the heterologous polypeptide is encoded by the heterologous gene of the expression system of any one of the preceding embodiments or the nucleic acid construct of any one of the preceding embodiments. In some embodiments, the inducer molecule is selected from among arabinose, rhamnose, lactose, tetracycline and anhydrotetracyline (aTc). In some embodiments, the method further compress lysing the prokaryotic cell to isolate the heterologous polypeptide. Kits of the Present Technology Also provided herein are kits comprising any and all embodiments of the expression systems and nucleic acid constructs described herein and instructions for using the expression systems and nucleic acid constructs to express the heterologous gene in a host cell. In one aspect, a kit comprising a nucleic acid encoding the expression system of any one of the preceding embodiments or the nucleic acid construct of any one of the preceding embodiments and instructions for use thereof to express the heterologous gene in a host cell. In another aspect, the present disclosure provides a kit comprising a host cell of any one of the preceding embodiments and instructions for use thereof to overexpress a heterologous polypeptide. In any of the preceding embodiments, the kit further comprises one or more inducer molecules, optionally wherein the inducer molecules are selected from the group consisting of arabinose, rhamnose, lactose, tetracycline and anhydrotetracyline (aTc). In some embodiments, the kits comprise buffers, preservatives, prokaryotic cells, and / or reagents for the transformation of a prokaryotic cell. 26 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 A kit may further contain a means for measuring the expression of one of the heterologous genes, or products produced therefrom, described herein. In some embodiments, the kit comprises reagents for the purification and / or measurement of a heterologous polypeptide or nucleotide encoded by one of the heterologous genes described herein. The kit may also comprise instructions for use, software for automated analysis, containers, packages such as packaging intended for commercial sale and the like. The kits of the present technology can also include other necessary reagents to perform any of the NGS techniques disclosed herein. For example, the kit may further comprise one or more of: adapter sequences, barcode sequences, reaction tubes, ligases, ligase buffers, wash buffers and / or reagents, hybridization buffers and / or reagents, labeling buffers and / or reagents, and detection means. The buffers and / or reagents are usually optimized for the particular amplification / detection technique for which the kit is intended. Protocols for using these buffers and reagents for performing different steps of the procedure may also be included in the kit. EXAMPLES The present technology is further illustrated by the following Examples, which should not be construed as limiting in any way. Example 1: Characterization and Improvement of the PBADand Promoters Materials and Methods Chemicals and Reagents. Q5® High-Fidelity 2X Master Mix (cat. #M0492), Taq 2X Master Mix (cat. #M0270), NEBuilder® HiFi DNA Assembly Master Mix (cat. #E2621), DpnI (cat. #R0176), T4 Polynucleotide Kinase (cat. # M0201), and Instant Sticky-end Ligase Master Mix (cat. #M0370) were purchased from New England Biolabs (Ipswich, MA). GeneMorph II Random Mutagenesis Kit (cat. #200550) were purchased from Agilent (Santa Clara, CA). QIAGEN Plasmid Miniprep Kit (catalog #27106), QIAquick PCR Purification Kit (cat. #28106), and QIAquick gel purification kit (cat. #28706) were purchased from Qiagen. Sanger sequencing and NGS were outsourced to Azenta Life Sciences, Inc. (Research Triangle Park, NC and South Plainfield, NJ). Primers for plasmid construction, library generation, and sanger and NGS sequencing used in this study were synthesized by Integrated DNA Technologies, Inc. (Coralville, IA) and are reproduced in the tables below. Table 1: Library generation primers. Primer binding regions are in upper case letters. Overhangs in hifi assembly primers are in lower case letters. Library Primers Forward Primer (5'-->3') Reverse Primer (5'-->3') 27 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PBADLibrary Vector NK404 TTCTCCATACCCGT NK405 TGAAAAGTATGG TTTTTTG (SEQ ID NO: 94) PBADepPCR Library NK406 cgcttcagccatacttttcaTA NK407 aaaaaaacgggtatggag Generation CTCCCGCCATTCAG aaACAGTAGAGA AG (SEQ ID NO: 96) GTTGCGATAAAA AG (SEQ ID NO: 97) PRhaLibrary Vector NK366 GTAATGAACAATTC NK367 TCCTGAAAATTC TTAAGAAGG (SEQ ACGCTG (SEQ ID ID NO: 98) NO: 99) PRhaepPCR Library NK368 tacagcgtgaattttcaggaAA NK369 tcttaagaattgttcattacG Generation TGCGGTGAGCATCA ACCAGTCTAAAA CATC (SEQ ID NO: AGCGC (SEQ ID 100) NO: 101) Table 2: PBAD& PRhaNGS Primers. Primer binding regions are in upper case letters. Primers with overhangs containing the Illumina partial adapters with 6 nucleotide barcode sequence are in lower case letters. Barcodes are underlined. Library & NGS Forward Primer (5'-->3') Reverse Primer (5'-->3') Primers PBADAmplicon NK437 acactctttccctacacgacgctcttccgatctcgat NK441 gactggagttcagacgtgtg primers with gtTTTCATACTCCCGCCATTCAG ctcttccgatctAGAAA Illumina partial (SEQ ID NO: 102) CAGTAGAGAGTT adapters NK438 acactctttccctacacgacgctcttccgatcttgac GCGATAAAAAG caTTTCATACTCCCGCCATTCAG (SEQ ID NO: 106) (SEQ ID NO: 103) NK439 acactctttccctacacgacgctcttccgatctacag tgTTTCATACTCCCGCCATTCAG (SEQ ID NO: 104) NK440 acactctttccctacacgacgctcttccgatctgcca atTTTCATACTCCCGCCATTCAG (SEQ ID NO: 105) PRhaAmplicon NK413 acactctttccctacacgacgctcttccgatctcgat NK417 gactggagttcagacgtgtg primers with gtAATGCGGTGAGCATCACATC ctcttccgatctTGTGCC Illumina partial (SEQ ID NO: 107) CATTAACATCAC adapters NK414 acactctttccctacacgacgctcttccgatcttgac CATCTAATTC caAATGCGGTGAGCATCACATC (SEQ ID NO: 111) (SEQ ID NO: 108) NK415 acactctttccctacacgacgctcttccgatctacag tgAATGCGGTGAGCATCACATC (SEQ ID NO: 109) 28 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 NK416 acactctttccctacacgacgctcttccgatctgcca atAATGCGGTGAGCATCACATC (SEQ ID NO: 110) Table 3: Bacterial Strains & Plasmids used in this study. Bacterial Strains Relevant characteristics Source Reference Escherichia coli NEB5αNew England Biolabs Escherichia coli DH5α-EF– φ80lacZΔM15 Δ(lacZYA-argF)U169 recA1ThermoFisher endA1 hsdR17(rK–, mK+) gal–phoA supE44 λ–thi-1 Scientific gyrA96 relA1 Plasmids Relevant characteristics Purpose Source Reference sfGFP-pBAD (#54519) PBADPromoter source PBADPromoter
[0059] source pJeM1 (#135088) PRhaPromoter source PRhaPromoter
[0060] source WT∆araCPBmoR-bmoR, PBMO-gfp, colE1 ori,Plasmid [36, 61] ampRbackbone BASIC_7_J23101- J23101-mCherry source for Vector backbone
[0062] RBS34-mCherry-B0015 vector backbone (#68141) PBAD ∆mCherryPc-araC, PBAD-gfp in pNK25Vector backbone This study vector backbone PBADPc-araC, PBAD-gfp, J23101- PBADnative This study mCherry in pNK25 vector backbone PBADValidation This study mCherry in pNK25 vector backbone PBAD∆araC Pc-∆araC, PBAD-gfp, J23101- native plasmid This study mCherry in pNK25 vector w / o TF backbone PRhanative This study mCherry in vector backbone PRha∆rhaS PRhaRS-∆rhaS, PRha-gfp, J23101- native plasmid This study mCherry in vector backbone w / o TF PRha-G41T PRhaRS-rhaS, PRha-gfp with G41T PRhaValidation This study PRha-A45T PRhaRS-rhaS, PRha-gfp with A45T PRhaValidation This study 29 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PBAD-P39-48(combo) Pc-araC, PBAD-gfp, with PBADValidation This study AxxxTCCATA39- 48GxxxATGGAT PBAD-A275T+G279A Pc-araC, PBAD-gfp, with PBADValidation This study A275T+G279A PBAD-T47A+A275T Pc-araC, PBAD-gfp, with PBADValidation This study T47A+A275T PBAD-T47A+G279A Pc-araC, PBAD-gfp, with PBADValidation This study T47A+G279A PRha-T91G+G94A PRhaValidation This study PRha-A45T+T91G PRhaRS-rhaS, PRha-gfp with PRhaValidation This study A45T+T91G PRha-A45T+G94A Validation This study 30 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Table 4: Primers used to construct plasmids in this study. Primers used to construct plasmids in this study. Mutations and overhangs for hifi assembly are in lower case letters. Plasmid Forward Primer (5'-->3') Reverse Primer (5'-->3') Plasmid Template PBADNK259 ATGAGTAAAGGAGA NK260 GGATCTGAAGCTT PBMO(∆mCherry) AGAACTTTTCAC GGGCC (SEQ ID
[0036] (SEQ ID NO: 112) NO: 113) NK257 cgggcccaagcttcagatccTT NK258 agttcttctcctttactcatAT sfGFP- ATGACAACTTGACG GTATATCTCCTTCT pBAD GC (SEQ ID NO: 114) TAAAGTTAAAC
[0059] (SEQ ID NO: 115) PBADNK267 AAAGTGCCACCTAG NK268 TCGGGGAAATGTG PBADTTCACCG (SEQ ID CGCGG (SEQ ID (∆mCherry) NO: 116) NO: 117) NK265 ttccgcgcacatttccccgaTC NK266 ggtgaactaggtggcactttC BASIC_7 TAGAAAGATCGATA TCGAGTTTTTCAG
[0062] GGTC (SEQ ID NO: CAAG (SEQ ID NO: 118) 119) PBAD(∆araO2) NK384 ATTGCATCAGACAT NK385 TTCTCTGAATGGC PBADTGCC (SEQ ID NO: GGGAG (SEQ ID 120) NO: 121) PBADNK386 ATTCAGAGAAGAAA NK387 ATGGCTGAAGCGC PBAD GCC (SEQ ID NO: 124) CTTGG (SEQ ID NO: 125) PRhaNK259 ATGAGTAAAGGAGA NK260 GGATCTGAAGCTT PBADAGAACTTTTCAC GGGCC (SEQ ID (SEQ ID NO: 126) NO: 127) NK276 cgggcccaagcttcagatccTT NK277 agttcttctcctttactcatAT pJeM1 ATTGCAGAAAGCCA GTATATCTCCTTCT
[0060] TC (SEQ ID NO: 128) TAAGAATTG (SEQ ID NO: 129) PRha∆rhaSNK351 ACTGGCCTCCTGATNK352 TAAGGATCTGAAG PRhaGTCG (SEQ ID NO: CTTGGGC (SEQ ID 130) NO: 131) PRha- NK463 AGCAAATTGTtAACA NK464 GAATTGTGGTGAT PRhaG41T TCATCACG (SEQ ID GTGATG (SEQ ID NO: 132) NO: 133) 31 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PRha- NK465 AATTGTGAACtTCAT NK466 TGCTGAATTGTGG PRhaA45T CACGTTC (SEQ ID TGATG (SEQ ID NO: NO: 134) 135) PRha- NK467 AATTGTGAACcTCAT NK468 TGCTGAATTGTGG PRha C64T TGCCAATG (SEQ ID ACAATTTG (SEQ ID NO: 138) NO: 139) PRha- NK471 TCCCTGGTTGaCAAT NK472 AAGATGAACGTGA PRhaC72A GGCCCA (SEQ ID NO: TGATGTTCACAAT 140) TTGC (SEQ ID NO: 141) PRha- NK473 AATGGCCCATaTTCC NK474 GGCAACCAGGGAA PRhaT84A TGTCAG (SEQ ID NO: AGATG (SEQ ID 142) NO: 143) PRha- NK475 CATTTTCCTGgCAGT NK476 GGCCATTGGCAAC PRhaT91G AACGAGAAGG (SEQ CAGGG (SEQ ID ID NO: 144) NO: 145) PRha- NK477 ATTTTCCTGTtAGTA NK478 GGGCCATTGGCAA PRhaC92T ACGAGAAGGTCGC CCAGG (SEQ ID (SEQ ID NO: 146) NO: 147) PRha- NK479 TTTCCTGTCAaTAAC NK480 ATGGGCCATTGGC PRhaG94A GAGAAGGTC (SEQ AACCA (SEQ ID ID NO: 148) NO: 149) PRha- NK481 CCTGTCAGTAtCGAG NK482 AAAATGGGCCATT PRha T43A ATTGCATCAG (SEQ CGGGA (SEQ ID ID NO: 152) NO: 153) PBAD- NK567 ACCAATTGTCgATAT NK568 TTCTTCTCTGAATG PBADC45G TGCATCAG (SEQ ID GCGG (SEQ ID NO: NO: 154) 155) PBAD- NK569 CCAATTGTCCgTATT NK570 TTTCTTCTCTGAAT PBADA46G GCATCAG (SEQ ID GGCG (SEQ ID NO: NO: 156) 157) PBAD- NK571 CAATTGTCCAaATTG NK572 GTTTCTTCTCTGAA PBADT47A CATCAG (SEQ ID NO: TGGC (SEQ ID NO: 158) 159) 32 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PBAD- NK575 GCGTCACACTaTGCT NK576 CGTGCAAATAATC PBADT232A ATGCCA (SEQ ID NO: AATGTGGAC (SEQ 160) ID NO: 161) PBAD- NK577 ACTTTGCTATaCCAT NK578 GTGACGCCGTGCA PBADG239A AGCATTTTTATCC AATAA (SEQ ID A242T ATTTTTATCCATAAG GCAAA (SEQ ID (SEQ ID NO: 164) NO: 165) PBAD- NK581 TAAGATTAGCaGATA NK582 TGGATAAAAATGC PBADG268A CTACCTGAC (SEQ ID TATGGC (SEQ ID NO: 166) NO: 167) PBAD- NK583 GATTAGCGGAaACT NK584 TTATGGATAAAAA PBADT271A ACCTGAC (SEQ ID TGCTATGG (SEQ ID NO: 168) NO: 169) PBAD- NK585 TTAGCGGATAtTACC NK586 TCTTATGGATAAA PBADC273T TGACGC (SEQ ID NO: AATGCTATGG 170) (SEQ ID NO: 171) PBAD- NK587 AGCGGATACTtCCTG NK588 AATCTTATGGATA PBADA275T ACGCTT (SEQ ID NO: AAAATGCTATGG 172) (SEQ ID NO: 173) PBAD- NK589 GATACTACCTaACGC NK590 CGCTAATCTTATG PBADG279A TTTTTATC (SEQ ID GATAAAAATG NO: 174) (SEQ ID NO: 175) PBAD- NK591 ATACTACCTGgCGCT NK592 CCGCTAATCTTAT PBADA280G TTTTATC (SEQ ID GGATAAAAATG G282A TTATCGC (SEQ ID ATGGATAAAAATG NO: 178) (SEQ ID NO: 179) PBAD- NK595 tggatTTGCATCAGAC NK596 tCAAcTGGTTTCTT PBADP39-48(combo) ATTGCCG (SEQ ID CTCTGAATGG NO: 180) (SEQ ID NO: 181) PBAD- NK660 AGCGGATACTtCCTa NK588 AATCTTATGGATA PBADA275T+G279A ACGCTTTTTATC AAAATGCTATGG (SEQ ID NO: 183) (SEQ ID NO: 182) PBAD- NK587 AGCGGATACTtCCTG NK588 AATCTTATGGATA PBAD- T47A+A275T ACGCTT (SEQ ID NO: AAAATGCTATGG T47A 184) (SEQ ID NO: 185) 33 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PBAD- NK589 GATACTACCTaACGC NK590 CGCTAATCTTATG PBAD- T47A+G279A TTTTTATC (SEQ ID GATAAAAATG T47A NO: 186) (SEQ ID NO: 187) PRha- NK661 CATTTTCCTGgCAaT NK476 GGCCATTGGCAAC PRhaT91G+G94A AACGAGAAGGTC CAGGG (SEQ ID (SEQ ID NO: 188) NO: 189) PRha- NK479 TTTCCTGTCAaTAAC NK480 ATGGGCCATTGGC PRha- A45T+G94A GAGAAGGTC (SEQ AACCA (SEQ ID A45T ID NO: 192) NO: 193) Equipment. All DNA concentration and purity and OD600 were measured on the DeNovix DS-11+ Spectrophotometer (DeNovix Inc., Wilmington, DE) using microvolume and cuvette absorbance modes, respectively. Optical densities at an absorbance of 600λ (OD600) of bacterial cultures in 45 mm x 10 mm x 10 mm (H x W x D) cuvettes (Greiner, cat. #613101) were measured on DeNovix DS-11+ Spectrophotometer (DeNovix Inc., Wilmington, DE). Uninoculated LB medium was used as a blank prior to measurement. All cultures were grown in Thermo Scientific™ MaxQ™ 6000 Incubated Stackable Shakers (ThermoFisher Scientific), shaking at 250 rpm, at 37°C. Biological Resources. Bacterial strains and plasmids used in the study are listed in Table 3. NEB® 5-alpha Competent E. coli (High Efficiency) (cat. #C2987) and NEB® 5-alpha Competent E. coli (Subcloning Efficiency) (cat. #C2988J) were purchased from New England Biolabs (Ipswich, MA). Electrocompetent ElectroMAX™ DH5α-E Competent Cells were purchased from ThermoFisher Scientific (cat. #11319019). All bacterial strains were grown in liquid or solid Luria broth, pH 7.0 (LB medium) which includes: 10 g / L tryptone (VWR, cat. #J859-500G), 5 g / L yeast extract (VWR, cat. #J850-5KG), 10 g / L sodium chloride (VWR, Cat. #0241-5KG), 15 g / L agar for solid (Fisher, cat. #BP1423-500) and supplemented with 100 µg / mL ampicillin (Amp100) (VWR, cat. #0339-25G). For volumes less than 10 mL, cultures were either grown in 5 mL culture tubes (VWR, cat. #60818-500), 15 mL conical tube (VWR, cat. #89039-664), 50 mL conical tubes (VWR, cat. #89039-658), or 2 mL deep-well 96-well plates (VWR, cat. #89237-526) sealed with breathable rayon film for culture plates (VWR, Cat. #60941-086). 34 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Plasmid construction. Plasmid construction and testing were completed in chemically competent NEB5α cells and grown on liquid and solid LB medium Amp100. Plasmid libraries were constructed and tested in electrocompetent ElectroMAX™ DH5α-E cells. Construction of native promoter plasmids. The pNK93 plasmid was first constructed by amplifying the vector backbone (3176 bp) of pNK25 plasmid
[0036] and AraC-PBAD(1286 bp) fragment from pJeM1 plasmid (Addgene, Cat. # 135088) were amplified with Q5® High- Fidelity 2X Master Mix. Once fragment sizes were verified on 0.8% agarose (VWR, Cat. #0710- 500G), PCR products were purified, digested with DpnI to remove the parental template, and then assembled using NEBuilder® HiFi DNA Assembly Master Mix. The ligated mixture was transformed into 25 µL NEB5α using the manufacturer’s instructions. After 1 hour recovery at 37°C (250 rpm), 100 µL of the transformed cells were plated on LB Amp100 and incubated at 37°C for ~18 hours. Following transformation, 3 individual colonies were screened using colony PCR with Taq 2X Master Mix to verify the desired insert in the correct orientation was present. Positive clones were grown in 5 mL LB Amp100 in 15 mL conical tubes (placed at an angle) overnight at 37°C shaking at 250 rpm for ~18 hours and miniprepped with QIAGEN Plasmid Miniprep Kit. Isolated plasmids were verified with Sanger sequencing. AraC-PBAD and RhaS-PRha plasmids, carrying the arabinose- and rhamnose-inducible promoters, respectively, were constructed similarly to pNK93 by HiFi assembly of two PCR fragments. J23101-RBS0015-mCherry (1019 bp) from BASIC_7_J23101-RBS34-mCherry- B0015 plasmid (Addgene, cat. #68141) was inserted into the backbone of pNK93 (4418 bp), resulting in plasmid PBAD. The PBAD plasmid was then used as the vector backbone of PRha plasmid, with the RhaS-PRhaBADinsert (1179 bp) amplified from sfGFP-pBAD plasmid (Addgene, cat. #54519). Construction of ∆TF controls and validation mutants. Point mutations were introduced via PCR using Q5® High-Fidelity 2X Master Mix with primers containing the desired mutations listed in Table 4. Deletion of the transcription factor gene in the native promoter plasmids were removed via PCR with non-mutagenized primers flanking the transcription factor gene. Fragment sizes were verified on 0.8% agarose, purified, and digested with DpnI to remove the parental template. Blunt ends of the fragments were phosphorylated with T4 Polynucleotide Kinase and then ligated with Instant Sticky-end Ligase Master Mix. The ligated mixture was transformed into 25 µL chemically competent NEB5α, screened using colony PCR, miniprepped, and verified with Sanger sequencing. 35 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Promoter analysis in E. coli via flow cytometry. Native promoter and control plasmids were transformed into NEB5α (subcloning efficiency). Individual colonies of each construct were grown in 3 mL LB Amp100 in 5 mL culture tubes overnight at 37°C, shaking at 250 rpm. Overnight cultures were diluted to an OD600of ~0.05 in pre-warmed 3 mL LB Amp100 in 15 mL conical tubes (placed at an angle) and grown at 37°C. When the OD600 reached ~0.1-0.2, 375 µL of cultures were transferred to 5-mL culture tubes with 125 µL LB Amp100 or 4X L-rhamnose or L-arabinose and incubated at 37°C shaking at 250 rpm. GFP fluorescence was measured at various time points on the Attune NxT (blue solid-state laser (488 nm excitation), an optical filter at 530 / 30 nm for GFP fluorescence, and 488 / 10 nm optical filter for side scatter (SSC)). For each time point, 10 μL of each sample were transferred to 500 μL PBS, pH 7.4 in 5 mL polystyrene tubes. The median fluorescence intensity of the FITC-A fluorescence of 10,000 events per sample was measured at a flow rate of 12.5 µL / min. Promoter library generation and assembly. Promoter libraries were generated by random mutagenesis using GeneMorph II Random Mutagenesis Kit (Table 1). The first round of error-prone PCR (epPCR) was amplified off the native plasmid as the template. The resulting PCR product was purified and used as the starting template for the next round of epPCR. A total of five to seven rounds of epPCR were completed to achieve high mutation rates. The last four rounds of purified epPCR products were gel purified using QIAquick gel purification kit using the manufacturer’s recommendations with minor modifications in the protocol (see
[0036] ) to remove contamination of the plasmid template prior to assembly with the NEBuilder® HiFi DNA Assembly Master Mix into the library vector. The library vector with the gfp gene underwent two rounds of PCR; the first was amplified from the purified native PBADand PRhaplasmid templates, and the second from the PCR product from the first round of PCR after purification and DpnI digestion. The vector from the second round of PCR was purified and then used as the library vector to minimize the native plasmid used as the PCR template in the transformed library. Similarly, HiFi assembly of the library vector only was prepared and transformed into 25 µL electrocompetent DH5α to calculate the library background of vectors without the insert within the library. Small-scale library transformation and selection. The last four rounds of epPCR library were assessed prior to generating the final library at a larger scale. Prior to transformation, the HiFi assembly library mix were desalted using a modified method described in
[0052] . Instead of using a micropipette tip, a 200 µL PCR tube was used to create a conical- shaped well in the agarose and glucose mix. After 90 minutes on ice, the desalted library was 36 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 transferred to a clean tube. Then 1-2.5 µL of the hifi assembly library was electroporated into 25 µL DH5α per transformation in 1 mm gap cuvettes (VWR, cat. #76102-576), pulsed at 1800V in BioRad MicroPulser. Following electroporation, 250 μL pre-warmed SOC media was added directly into the cuvette, and then transferred into 5-mL culture tube for recovery at 37°C shaking at 250 rpm for 1 hour. Prior to antibiotic selection, 20 μL of the recovered library was plated on solid LB Amp100 to estimate the library size. The rest of the transformed library was transferred to a 15-mL conical tube and 2725 μL of liquid LB Amp100 was added for antibiotic counterselection. Cell death and viability were monitored by sampling 20 μL of the library, diluted in 500 μL of PBS, pH 7.4, over time on the flow cytometer based on the %gated of the FSC and SSC dot plot. After about 3 hours of antibiotic selection, the %gated of the culture plateaued and immediately placed on ice to stop the growth. The selected library culture was mixed with sterile 80% glycerol (final glycerol concentration of 20%) directly into the 15-mL conical tube and then stored at -80°C. The colonies of the plated libraries from the day prior were counted to calculate the transformation efficiency, as well as the vector only background, to ensure that the library size was large enough but contained <1% vector only background in the entire library population. Small-scale library analysis via flow cytometry. Assessment of the promoter library median GFP distribution was done as described under ‘Promoter analysis in E. coli via flow cytometry’ under the ‘Methods & Materials’ section with few differences described here: Overnight cultures of the controls and the different epPCR round libraries were grown overnight at 30°C, shaking at 250 rpm; controls were grown from individual colonies in 3 mL LB Amp100 from transformations plated within 2 days; libraries were thawed on ice (~15 minutes) from the - 80°C stored glycerol stocks prepped using the protocol described under ‘Small-scale library transformation and selection’; and transferred to 50 mL conical tubes containing 15 mL LB Amp100 for growth. The plasmids isolated from overnight cultures of the tested libraries were harvested by centrifugation at 6,800xg for 10 minutes. The supernatant was discarded, and the plasmid libraries were isolated to be prepped for NGS, as described in ‘NGS Sample Preparation’ in the ‘Materials and Methods’ section, to assess the library diversity and mutation frequency of the last four rounds of the epPCR library. Large-scale library transformation and selection. Once the library round was determined for sort-seq, a total of eight transformations per library were selected, where 20 μL of the desalted HiFi assembly library was combined with 200 µL electrocompetent DH5α cells. After 37 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 gently mixing, 27.5 μL of the DNA-cell mixture were transferred into pre-chilled 1 mm gap electroporation cuvettes. After 10 minutes of incubation, the cuvette was placed in the pulser. Immediately after the DNA-cell mixture was pulsed at 1800V, the cells were recovered in 250 µL of prewarmed SOC media per transformation directly in the cuvette. Using a sterile disposable ultra-fine tip transfer pipet (VWR, cat. #414004-019), the cells were transferred into a 5 mL-culture tube per transformation. After recovery for 1 hour at 37°C shaking at 250 rpm, all recovered cells were pooled into a sterile 50-mL conical tube. To calculate the approximate size of the library, 20 μL of the recovered cells were serially diluted and plated on solid LB medium Amp100 before the antibiotic selection. To start the library selection, 14 mL of LB Amp100 was added to tube with the recovered library and then monitored cell death and viability of the library as described in “Small-scale library transformation and selection” of the methods. Library sorting. Seed culture and induction. Libraries were started from four 1.6-mL frozen glycerol stocks in 200 mL of LB Amp100 in a 1 L baffled flask at 30°C for ~15 hours at 150 rpm. As controls, the native promoter plasmids were also grown overnight in 10 mL LB Amp100 in 50 mL conical tubes. Overnight cultures were diluted in pre-warmed 100 mL LB Amp100 to OD600of ~0.1 to keep the cells in the log phase and grown in 500 mL baffled flasks for ~30 min at 37°C, 250 rpm. The controls were grown similarly but in a 250 mL baffled flask with pre-warmed 50 mL LB Amp100. Once the OD600 reached ~0.2, each of the inoculated cultures were split into two baffled flasks (50 mL each for libraries in 250 mL baffled flasks; 25 mL each for controls in 125 mL baffled flasks) in which one was induced with the final concentration of 0.2% (w / v) inducer and the other with no inducer. The induced and uninduced cultures were grown at 37°C for ~5-6 hours, shaking at 250 rpm. Post-induction prep for sorting. Samples were removed from the incubator and then immediately placed on ice. A sample amount of the cultures (both induced and uninduced) were collected for initial analysis on the flow cytometer and OD600readings prior to centrifugation. Both induced and uninduced libraries were transferred to pre-chilled 50-mL conical tubes where each condition were aliquoted 40 mL and 10 mL. A total of four tubes were centrifuged at 4°C, 3000xg, for 10 minutes. Supernatant were carefully discarded. The tubes with the 40 mL library cultures were saved for future use, while tubes with 10 mL library cultures were resuspended in 10 mL of pre-chilled PBS, pH 7.4 to be used for sorting. Finally, the resuspended libraries were diluted to OD600 of 0.02 – 0.025 (1:200 – 1:400) in ice cold PBS (5-mL tube) to achieve a sorting rate of <8000 events / sec on the iSort™ Automated Cell Sorter (blue solid-state laser at 38 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 488 nm and 165 mW, optical filters 525 / 50BP for GFP and 488 / 10 SSC, 85 μm ceramic nozzle, fixed sample flow rate of 23 μL / minute). The diluted libraries were placed on ice prior to sorting. Library sorting and post-sort analysis. To set the sorting gates, the uninduced and induced native promoter were used as guides. Because the iSort™ Automated Cell Sorter is a two-way sorter, low (B1) and high (B4) gates were first sorted for the induced library, while the induced middle two gates (B2 and B3) were sorted together. Then the B1 and B4 gates of the uninduced libraries were paired for the third sort. Cells were sorted based on the GFP fluorescence until all bins reached a minimum of 250,000 cells. The total number of cells sorted per bin, percent target of the library, sorting time, and the event rate per second for both libraries were recorded (data not shown). To ensure proper sorting (purity and cross-over), each sorted bin was sampled on the Attune NxT flow cytometer based on the mCherry fluorescence gate, using the constitutively expressed mCherry located on the plasmid library vector. The plasmid library bins, including the unsorted library, were isolated after recovering in 10 mL LB Amp100 medium overnight (where the PBS:LB Amp100 media ratio was at a maximum of ~1:3). The extracted plasmid libraries were subsequently used to generate PCR amplicons for NGS. NGS sample preparation. Plasmid libraries of the recovered sorted and unsorted bins were isolated using Qiagen QIAprep Spin Miniprep Kit according to the manufacturer’s recommendations. The region of interest was amplified with Q5® High-Fidelity 2X Master Mix for ten cycles using assigned primers (Table 2) with the annealing region, 6-nt unique barcode sequence, and partial Illumina adapters overhangs from 1000 ng of the extracted plasmids per 50 µL PCR reaction. Six PCR reactions tubes per bin were combined and prepped for gel electrophoresis. The desired PCR amplicon size was excised from 0.7% (w / v) agarose gel in 0.5X TAE buffer with a clean razor blade and gel purified using the Qiagen QIAquick Gel Extraction Kit with few modifications in the manufacturer’s protocol as described in
[0036] . Additional impurities were removed using the Qiagen QIAquick PCR Purification Kit. In some cases, library amplicons were pooled into one tube with equal molar mixture and were ran on a 0.8% agarose gel in 0.5X TBE to ensure purity and size (~100 ng loaded per lane) (data not shown). The final concentrations of the amplicons were normalized to 20 ng / μL in 10 mM Tris- Cl, pH 8.0 with A260 / 280 at 1.8-2.0 prior to outsourcing to AZENTA Life Sciences using 2x250 bp paired-end Amplicon EZ Sequencing service. NGS pre-processing workflow. Galaxy web-based platform (on the public server at usegalaxy.org and usegalaxy.eu)
[0053] was used to preprocess the raw library FASTQ data prior to generating the nucleotide counts per position, mutation frequency per position plots, 39 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 information footprints, and enrichment analysis. In brief, paired-end reads were first merged using Paired-End read merger (PEAR)
[0054] . Then, the ‘Barcode Splitter’ tool
[0055] was used to obtain merged reads with intact 5’ fixed ends and to separate multiplexed bins, if needed. Prior to mapping, the barcodes and the fixed 5’ and 3’ ends were trimmed using the ‘Trim sequences’ tools
[0055] . Reads were mapped against the reference sequence of native promoter using the BBMap tool
[0056] . ‘Trim sequences’ was used again trim off any overhangs resulted from the mapping. Reads were converted from BAM to FASTA file format with Samtools fastx tool from the SAMtools software package
[0057] and then unique reads were obtained using the ‘Collapse’ tool
[0055] . To calculate the nucleotide counts per position, the aligned reads underwent a series of formatting changes using the ‘Replace’ tool
[0058] and ‘Convert delimiters to TAB’ in Galaxy prior to exporting the reads as a CSV file. This file was imported into Microsoft Excel for further processing of the data. During the conversion from BBMap to FASTA file formats, some of the read alignments were misaligned if deletions were found in the 5’ end. Indications of misaligned sequences were based on the calculated hamming distance between each read and the native sequence using the following syntax in Microsoft Excel: =COUNT(RANGE1)-SUMPRODUCT(-- (RANGE1=RANGE2)) Hypermutated sequences were filtered from the read set using the calculated hamming distance. Reads with a hamming distance greater than >20 between each read and native sequence was removed. Calculating Mutual information and Enrichment Analysis. Information footprints and the enrichment analysis were calculated as described in [35, 36]. In brief, the nucleotide occurrence of each base at each position were counted and normalized to the total number of reads in each library for the input library and the sorted high and low expression bin libraries. The mutual information between the expression bins and the bases at each nucleotide position were calculated using: where ^^^^^^^^is the base at the ^^^^th position, ^^^^ is the expression activity bin, and ^^^^(^^^^^^^^,^^^^) is the jointfrequency distribution and ^^^^(^^^^^^^^)and ^^^^(^^^^)is the marginal frequency distributions. To calculatethe enrichment of each base at each position, the ratio of the number of nucleotide occurrence in each of the expression bins were normalized to the counts of the unsorted bins. 40 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Dose-response promoter validations. Validation plasmids in NEB5α were grown overnight at 37°C, shaking at 250 rpm in 3 mL LB Amp100 in 5 mL culture tubes. Overnight cultures were diluted to an OD600 of ~0.05 in pre-warmed 10 mL LB Amp100 in 50 mL conical tubes and grown at 37°C. Once the OD600reached ~0.1-0.2, 450 µL of cultures were transferred to 2 mL deep-well plates, and subsequently induced with 150 µL LB Amp100 or 4X rhamnose or arabinose. Plates were sealed with breathable rayon film to prevent evaporation during incubation at 37°C. Endpoint measurements of the GFP fluorescence were measured on the flow cytometer. For flow cytometer measurements, 10 μL of each sample were transferred to 500 μL PBS, pH 7.4 in 5 mL polystyrene tubes. Generation of GFP fluorescence, OD600, information footprint, and enrichment plots. GFP distribution histograms from flow cytometry experiments were generated with the Attune NxT acquisition software. GFP fluorescence and OD600 measurements were averaged and plotted with the standard deviations in GraphPad Prism 10.0.2 (232). All replicates represent biological samples grown from individual colonies. The information footprints and enrichment analysis calculations were completed in Microsoft Excel and then transferred to GraphPad Prism 10.0.2 (232) to generate the plots. Results PBAD and PRha Library Generation To identify nucleotides sites that alter gene expression when mutated, promoter libraries were constructed by mutating the upstream regions of the RNA polymerase (RNAP) binding sites, including known binding sites of AraC and RhaS (FIGs.1A-1B). The mutagenized region for PBAD, 253 bp in length, includes the araO2, araO1, cAMP receptor protein (CRP), and araI1 and araI2 half-sites, while the 97 bp mutagenized region for PRha includes the CRP and the RhaS binding sites rha1 and rha2. In both promoters, the transcription factor (TF) binding sites partially overlap the -35 box, thus they were included in the mutagenized regions. PBADand PRhapromoters underwent four and five rounds of error-prone PCR (epPCR), respectively, and cloned into the vector to drive the expression of the gfp reporter gene in E. coli DH5α (FIGs.4A-4B). The vector also carries a constitutively expressed mCherry reporter gene to isolate viable cells harboring the library plasmids after sorting. The final library sizes and average mutation rate per position of PBAD and PRha were approximately 661,000 with 5.7% and 583,000 with 6.4%, respectively, based on the transformation efficiencies (data not shown). 41 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Promoter libraries were sorted based on GFP fluorescence under induced (0.2% arabinose or rhamnose) and uninduced (0% arabinose or rhamnose) conditions. Induced libraries were sorted into four expression bins (B1 to B4, ranging from low to high expression) while uninduced libraries were sorted into two expression bins (low, B1 and high, B4). The promoters of the sorted and unsorted (B0) bins were deep sequenced via NGS and pre-processed to obtain unique reads across all bins for PBADand PRha. The importance of mutations at each nucleotide position on bin location is quantified by mutual information and relative enrichment [35, 36]. Sort-seq identified canonical binding sites and mutations that alter promoter response and strength. The first goal was to identify all sites that affect promoter regulation. Using information footprints and enrichment analysis generated for each promoter, all the well-annotated regulatory sites were identified for each promoter at the nucleotide level. Mutual information provides a strong prediction of site-specific effect while the enrichment analyses are used to determine the specific nucleotide residues. Since mutual information combines data from all four bins to indicate the importance of each location, it compresses the information at each site. Therefore, enrichment analyses that focus on individual bins were used to identify which sequence was the most important. PBAD. Under induced (0.2% arabinose) conditions, mutual information values were highest in the araI2 and araI1 binding sites, and were moderate in the CRP and araO2 sites (data not shown). The mutual information values in the araO1 site were not higher than background. These results are consistent with the described mechanism of PBADpositive regulation where arabinose-bound AraC binds to araI1 and araI2, recruiting RNA polymerase, and co-activator CRP binds to its binding site (FIG.1A). Importantly, differential importance was observed within each binding site, indicating nucleotide positions that are more or less likely to drive the DNA-protein interaction strength. The araI1 sequence, and to a lesser extent the CRP site, is highly conserved, as indicated by the depletion of mutations in the high gfp bin (B4). Certain mutations in araI2 were enriched in the high gfp bin (B4), suggesting that this region might be of interest to tune gfp expression. The araO2 was highly enriched with mutations in the TSS proximal region when highly expressing gfp, aligning with previous research that destroying araO2 prevents binding of AraC for repression [12, 37]. When uninduced, a similar mutual information and enrichment profile was observed in the araO2 region as in the induced condition; any mutation contributes to increased GFP activity in P43 – P48. While it has been reported negative regulation in the 42 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 absence of arabinose occurs when AraC is bound to araO2 and araI1 in the absence of arabinose
[0012] , an accumulation of mutations in the araI1 site was not observed (data not shown). Within the 0% data set, certain sites showed relatively high mutual information values, displaying a mix of congruence and divergence from the existing literature. The sites that aligned with the literature were within the araO2 and CRP binding sites. Also, the requirement of CRP for full activation of PBADare demonstrated by the depleted mutations in the CRP binding sites (P216, P227, P231, and P233) in the 0.2% B4 enrichment analysis. Conversely, sites that deviated from the existing literature were found between the CRP binding site and araI1 (P238, P239, P241), within araI1 (P247), and in araI2 (P269, P279). The mutual information at araI1 was low for the TCCATA motif site in the 0% data (P254-P259). Among the identified sites, P238 (MI: 0.86), P241 (MI: 1.38), P247 (MI: 0.10), and P269 (MI: 0.64) had 0.2% mutual information below 1.4 millibits, while the 0% mutual information ranged from 2.9 to 27.5 millibits. Notably, P239 had the highest mutual information at 54.5 millibits in the 0% data with a mutual information of 4.6 millibits in the 0.2% data. Upstream of the araI1 binding site, the data reveal overlapping priorities: aligning with the consensus CRP binding motif and a proto-promoter site (P213 – ATTTGC – 17 bp gap – TATGCC – P241). Mutations enriched in the 0% B4 heat maps form a sequence matching the canonical σ70with notable enrichments of A213T, T215G, C218A, G239A, and C241T. Taken together, these mutations show a potential promoter of P213 – TTGTGA – 17 bp gap – TATACT – P241 (mutations bolded) and the same time more closely aligning with the CRP binding site consensus sequence (P 214 – TGTGA – 6 bp gap – TCACA – P229). A similar pattern can be found in the induced condition data, but with much smaller mutual information values due to the promoter being in the ‘on’ state. Enriched mutations at a location with a significant mutual information value that did not align with consensus sequences were found within and near araI2 (G268A, T271A, C273T, A275T, G279A, A280G, G282A). PRha. Similarly, regulatory elements of PRha, rhaI1, rhaI2, and CRP, were identified under induced conditions (data not shown). The two half-sites, rhaI1 and rhaI2, as well as the CRP binding site, were in alignment with the literature
[0032] . Like PBAD, the most proximal RhaS binding site, rhaI2, contributed the most to gfp expression, followed by rhaI1, and then CRP. In the presence of rhamnose, the outer and inner regions of the rhaI half-sites increased in information content, specifically at P57-P62 and P69-P74 in rhaI1 and P89-P93 and P102-P105 in rhaI2. These regions were previously predicted to be where RhaS monomers of each dimer interacted with the DNA at each rhaI half-site [32, 38]. Unlike PBAD, the CRP binding site of 43 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PRha was especially important at P41 in the 0% information footprints, where mutation at this site could lead to leaky expression. Although the -35 box was only included in the analysis because it shared overlapping sites with rhaI2, it was interesting to see that the two non-overlapping sites of the -35 box were not as prominent in contributing to the gfp expression, whereas it mattered more in PBAD. Additional significant sites were found that extended the annotated rhaI half-sites. While AraC binds to direct repeat sequences that are separated by 4 bp to activate transcription, RhaS binds to rhaI1 and rhaI2 inverted repeat sequences separated by 16 bp [32, 38]. Interestingly, the data suggested that both half-sites extend by one nucleotide (P74 and P89), each resulting in an 18 bp half-site and shortening the space in between from 16 to 14 bp. Validation of inducible promoter variants. Having identified the regulatory sites for both PBADand PRha, the next step was to validate the predicted sequences that outperformed the native sequence based on the information footprint and enrichment analysis. Sequences were selected based on the enrichment analysis of the induced (0.2%) bins, while being cognizant of the mutants predicted to raise basal expression in the uninduced (0%) data. Individual promoter variants and their respective native promoters were evaluated in E. coli NEB5α by flow cytometry under uninduced (0%) and induced conditions (0.2% - L-arabinose for PBAD; L-rhamnose for PRha) (t=5 hours, n=2). PBAD. Mutations in araO2: T43A (position 43 mutated from T to A), C45G, A46G, T47A; as well as a combination construct with araO2 site mutations (araO2 Combo): A39G, T43A, C44T, C45G, A46G, T47A, and A48T, were assessed. These mutations led to a 2.5x to 3.1x increase in GFP expression under the 0.2% condition while also increasing the uninduced expression by 1.7x to 1.9x (FIG.2A). Although there was an increase in basal expression, the dynamic range of these mutants exhibited 1.4x to 1.8x improvement over the native promoter. Mutations in the araI1 half-site were avoided due to its predicted negative effect on promoter activation. Instead, two mutations located in between CRP and araI1 were chosen: G239A and A242T. Their dynamic ranges remained unchanged. G239A showed a slight increase in fluorescence levels under both 0% and 0.2% arabinose despite having the most significant mutual information in the 0% information footprint. Five sites in the araI2 and the overlap -35 box region were chosen for validations: G268A, T271A, C273T, A275T, and G279A. These mutations demonstrated improved induced GFP fluorescence, with G268A, C273T, and T271A showing 2.3x, 1.9x, and 1.5x increase, 44 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 respectively. P275 and P279 exhibited a 3.2x increase in basal fluorescence, with induced GFP fluorescence of 3.5x and 2.6x increase, respectively. G282A, located in the -35 box, was also included. Furthermore, both T232A and A280G were chosen to demonstrate moderate levels of GFP expression. Enrichment analysis of B4 show that the native sequences of P232 (located in CRP binding site) and P280 (araI2 / -35 box) were important in full induction of the promoter. It was hypothesized that mutations at P232 and P280 would result in decreased expressed and included to validate mutations in both directions. Both mutants performed as predicted, demonstrating moderate levels of GFP expression. The response of promoters exhibiting a significant GFP fluorescence output and increase in induction fold compared to the native promoter was assessed in response to variable amounts of the inducer concentration (0% to 2% arabinose) using half-log dilutions (FIGs.2A-2B). This included all four single nucleotide in araO2, araO2 combined promoter, and three in araI2 (G268A, A275T, and G279A). Three double mutant promoter variants were also created with combinations of A275T, G279A, and T47A. Although these variants had elevated basal expression in FIG.2A, it was next assessed whether the combination of the two variants would result in increased induction GFP fluorescence while maintaining the same level of basal expression as to when variants were tested individually. Lastly, the araO2 motif deleted from PBAD(∆araO2) to compare differences between the single nucleotide promoters variants in araO2. Dose-response validations are shown in FIG.2B were measured at 24 hours post- induction (n=2). All single variants located in araO2, including the ∆araO2, showed similar responses to various concentrations of arabinose, with the araO2 combo promoter variant performing at the lower maximal fluorescence in comparison. The three double mutant promoters increased in their maximum induction fluorescence by 16.2x – 19.1x over the native promoter and sensitivity at 0.2% arabinose, while the native and the rest of the promoter constructs plateaued at 0.63% arabinose. Basal expression of the double mutant promoters also increased but the dynamic ranges of A275T+T47A and G279A+T47A improved by ≥2x over the native promoter. Promoter variant G279A+T47A had the highest basal expression. PRha. Initial validations for PRhamutants included eight constructs with a variety of promoter performance ranges, including those that exhibit improved induction range (FIG.3A). For the best performing mutants, T91G and G94A, were chosen due to significant mutual information values and high enrichment in the induced B4 and the lack of such features in the uninduced condition, indicating little effect on basal expression. Unlike PBAD, there were fewer 45 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 sites in RhaS binding regions in PRha that contributed to basal activity, where A97T stood out. A97T was included to demonstrate high basal expression levels since it was highly enriched in B4 of both the induced and uninduced condition. The most unusual mutation, C92T, showed enrichment in the high gfp expression B4 in the uninduced condition while the same mutation was enriched in the induced B1 and B2 and depleted in the higher bins. T84A, just upstream of the rhaI2 site, was also included where the mutation appeared to be tolerated, but not preferred. Promoter variants within the CRP binding site (G41T, A45T, and A45C) were also constructed for analysis. Mutation G41T was chosen due to its unique enrichment pattern where the mutation is severely depleted in the induced and uninduced B1, highly enriched in the induced B3 and uninduced B4, but not in the induced B4. This indicates that this mutation leads to moderate expression levels with higher basal expression. Also, promoter variants were chosen with mutations at P45 located in between the CRP binding site consensus sequences. Even though the mutual information at P45 (MI: 2.3 millibits) in the 0% information footprint was lower compared to the other chosen sites, the rationale for including this mutant was primarily because PRhahad a better induction fold with its native CRP sequences despite having lower affinity to CRP than the E. coli CRP consensus sequences in a different study
[0020] . The E. coli CRP consensus sequence and the PRha CRP sequence differs by one base at the underlined base (TGTGAxxxxxxTCACA). Furthermore, of the six non- consensus regions in P43-P47, P45 had the highest information content and contained two mutants enriched in the highest gfp bin but depleted in the lowest gfp bin: A45T and A45C. The results shown in FIG.3A observed G94A, followed by T91G, to exhibit improved induced maximum fluorescence signal and maintained low basal expression compared to the native promoter. Three of the highest information content sites in the 0% condition predicted to have high basal activity but different ranges of inducibility were also chosen: C92T (low inducibility), G41G (moderate inducibility), and AP97T (high inducibility). Unlike PBAD, validation of PRha variants behaved as predicted by the information footprints and enrichment analysis (FIG.3A). High basal expression was observed for C92T and A97T, with the former being overall weak and the latter strong in both 0 and 0.2% rhamnose induction. Mutant promoters with T91G and G94A had low levels of basal expression with slightly but significantly increased expression in the 0.2% induced condition. The T84A mutation shows only a slight reduction in induced expression and no change when uninduced. In the CRP binding region, G41T resulted in leaky expression and lower gfp expression compared to the native sequence when induced as predicted 46 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 by enrichment in B2 and B3 (0.2%). The variants at position 45 yielded slightly higher GFP output. The best performing single mutations (A45T, T91G, and G94A) and all two-mutation combinations were evaluated under rhamnose inductions ranging from 0% to 2% w / v (FIG.3B). All constructs retained the low basal expression of the native sequence. The single mutation constructs yielded consistent results at the 0.2% induction condition from initial analysis (FIG. 3A). Of the single mutants, the G94A mutation yielded significantly higher expression at the 2% induction condition, resulting in a ~40% increase in fluorescence. All three double mutant promoter variants outperformed the single variants and the native promoter in their fluorescence output by 2.4x when induced with 2% rhamnose over the native promoter while maintaining low basal levels improving their dynamic range. In summary, PRhapromoter variants, in particular, A45T+G94A, T91G+G94A, and A45T+T91G, exhibited 2x to 2.4x improvements in dynamic range over the native promoter, while maintaining low basal levels. For the PBAD combinations, 275+279, 275+47, and 279+47 increased the promoter strength nearly 7x. Improved sensitivity was also observed in these variants, with maximum expression elicited at 0.2% arabinose compared to 0.63% arabinose in other constructs and the native sequence. Increase in promoter strength and sensitivity makes the constructs of the present disclosure ideal for various applications as an alternative to the T7 promoter, which is known to be 2x to 10x stronger than the native PBAD
[0014] , for max production
[0013] . Although increased basal expression was observed, basal expression can be controlled using glucose in the medium for glucose-responsive catabolite repression. Alternatively, promoters with single mutations located in araO2 exhibited 2x increase in basal expression and 3x increase in induction. Moreover, some PBAD promoter variants exhibited moderate increases in dynamic range without increasing the basal levels, such as G268A and C273T. Accordingly, the constructs and promoters of the present technology are useful in constructs, compositions, and methods for the expression of gene products in prokaryotic organisms. Example 2: Characterization and Improvement of Clostridium promoters Materials and Methods Chemicals and reagents. Q5® High-Fidelity 2X Master Mix (Cat. #M0492), Taq 2X Master Mix (Cat. #M0270), NEBuilder® HiFi DNA Assembly Master Mix (Cat. #E2621), DpnI (Cat. #R0176), T4 Polynucleotide Kinase (Cat. # M0201), and Instant Sticky-end Ligase Master Mix (Cat. #M0370) were purchased from New England Biolabs (Ipswich, MA). GeneMorph II Random Mutagenesis Kit (Cat. #200550) were purchased from Agilent Technologies (Santa 47 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Clara, CA). QIAGEN Plasmid Miniprep Kit (Catalog #27106), QIAquick PCR Purification Kit (Catalog #28106), QIAquick gel purification kit (Catalog # 28706), and DNeasy® Blood & Tissue Kit were purchased from Qiagen (Catalog #69506) were purchased from Qiagen. Sanger sequencing and NGS were outsourced to Azenta Life Sciences, Inc. (Research Triangle Park, NC and South Plainfield, NJ). Primers for plasmid construction, library generation, and sanger and NGS sequencing used in this study are listed below and were synthesized by Integrated DNA Technologies, Inc. (Coralville, IA).TFLime was purchased from Twinkle Bioscience S.A.S. (Paris, France). Table 5: Library generation primers. Primer binding regions are in upper case letters. Overhangs in hifi assembly primers are in lower case letters Library & NGS Primers Forward Primer (5'-->3') Reverse Primer (5'-->3') PBgaLLibrary Vector NK36 AATTATGTTTCTTA NK36 TGTTTAATATAAG 2 AATATACAATCATG 3 ACTACTATAAAAT (SEQ ID NO: 194) TGG (SEQ ID NO: 195) PBgaLepPCR Library NK36 tagtagtcttatattaaacaTTT NK36 tgtatatttaagaaacataattC Generation 4 ACATGAGAGCTTTG 5 CATATAAATCATT C (SEQ ID NO: 196) TTTCAAAATAGTT TTTAC (SEQ ID NO: 197) PPTKLibrary Vector NK40 TAATTAATGATTAT NK40 ATAATATCTCTTA 0 AAATTTAGGAGGA 1 AATTGGTATTATG ATATTATG (SEQ ID ATC (SEQ ID NO: NO: 198) 199) PPTKepPCR Library NK40 accaatttaagagatattatAA NK40 cctaaatttataatcattaattaT Generation 2 AAGAGACTTCAAG 3 GAATTCATCTAGT ATAAATTG (SEQ ID GTGATTC (SEQ ID NO: 200) NO: 201) PTETO1Library Vector NK38 AAAGGAGAAAATT NK38 ACTAGTTTGACAA 0 TTATGAGTAAAG 1 ATAACTCTATC (SEQ ID NO: 202) (SEQ ID NO: 203) PTETO1epPCR Library NK38 gagttatttgtcaaactagtTTT NK38 ctcataaaattttctcctttACT Generation 2 TTATTTCGATGCCC 3 GCAGGAGCTCAGA TGG (SEQ ID NO: 204) TC (SEQ ID NO: 205) *reverse-complement the original fwd / rev primers to keep primer length 60 bp 48 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Table 6. Clostridium promoters NGS Primers. Primer binding regions are in upper case letters. Primers with overhangs containing the Illumina partial adapters with 6 nucleotide barcode sequence are in lower case letters. Barcodes are underlined. NGS Primers Forward Primer (5'-->3') Reverse Primer (5'-->3') PBgaLAmplicon primers with 3 CGCTCTTCCGATCTCGATG GTGTGCTCTTCCGATC Illumina partial Tttctacctcctaacctataaaattagcc TCAGATCataattccatataaat adapters (SEQ ID NO: 206) catttttcaaaatagtttttac (SEQ NK43 ACACTCTTTCCCTACACGA ID NO: 210) 4 CGCTCTTCCGATCTTGACC Attctacctcctaacctataaaattagcc (SEQ ID NO: 207) NK43 ACACTCTTTCCCTACACGA 5 CGCTCTTCCGATCTACAGT Gttctacctcctaacctataaaattagcc (SEQ ID NO: 208) NK43 ACACTCTTTCCCTACACGA 6 CGCTCTTCCGATCTGCCAA Tttctacctcctaacctataaaattagcc (SEQ ID NO: 209) PPTKAmplicon NK44 ACACTCTTTCCCTACACG NK452 GACTGGAGTTCAGAC primers with 8 ACGCTCTTCCGATCTCGA GTGTGCTCTTCCGATC Illumina partial TGTTGAATTCATCTAGTG TAAAAGAGACTTCAA 9 ACGCTCTTCCGATCTTGA CCATGAATTCATCTAGTG TGATTC (SEQ ID NO: 212) NK45 ACACTCTTTCCCTACACG 0 ACGCTCTTCCGATCTACA GTGTGAATTCATCTAGTG TGATTC (SEQ ID NO: 213) NK45 ACACTCTTTCCCTACACG 1 ACGCTCTTCCGATCTGCC AATTGAATTCATCTAGTG TGATTC (SEQ ID NO: 214) 49 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PTETO1Amplicon NK41 ACACTCTTTCCCTACACG NK422 GACTGGAGTTCAGAC primers with 8 ACGCTCTTCCGATCTCGA GTGTGCTCTTCCGATC Illumina partial TGTtttttatttcgatgccctgg (SEQ Tcatcaccttcaccctctc (SEQ adapters ID NO: 216) ID NO: 220) NK41 ACACTCTTTCCCTACACG 9 ACGCTCTTCCGATCTTGA CCAtttttatttcgatgccctgg (SEQ ID NO: 217) NK42 ACACTCTTTCCCTACACG 0 ACGCTCTTCCGATCTACA GTGtttttatttcgatgccctgg (SEQ 1 ACGCTCTTCCGATCTGCC AATtttttatttcgatgccctgg (SEQ ID NO: 219) Table 7 Bacterial Strains & Plasmids used in this study. Bacterial Strains Relevant characteristics Source Reference Escherichia coli NEB5a F– φ80lacZΔM15 Δ(lacZYA- New England argF)U169 recA1 endA1 hsdR17(rK–, Biolabs mK+) phoA supE44 λ–thi-1 gyrA96 relA1 Escherichia coli DH5a-E F– φ80lacZΔM15 Δ(lacZYA-argF)U169 recA1 ThermoFisher endA1 hsdR17(rK–, mK+) gal–phoA supE44 λ– Scientific thi-1 gyrA96 relA1 Escherichia coli ER1821 F-endA1 glnV44 thi-1 relA1? e14-(mcrA-) New England rfbD1? spoT1? Δ(mcrC-mrr)114::IS10 Biolabs E. coli K-12 Keio F-, Δ(araD-araB)567, ΔaraC771::kan, Baba et al.2006; E. collection JW0063-1 ΔlacZ4787(::rrnB-3), λ-, rph-1, Δ(rhaD-rhaB)568, coli Genetic Stock hsdR514 Center (CGSC) at Yale C. acetobutylicum ATCC Wild-type strain American Type 824 Culture Collection Plasmids Relevant characteristics Purpose Source Reference sfGFP-pBAD PBAD Promoter souPBAD Promoter (#54519)rcesourcePédelacq et al. 2006BASIC_7_J23101- RBS34-mCherry- J23101-mCherry source for vector backVector backbone Storch et al. 2015B0015 (#68141)bonepNK25 bmoR, PBMO-gfp,colE1 ori, ampR Vector backbone This studyPBAD DmCherryPc-araC, PBAD- in pNK25vector backbone Vector backbone This studyPBADPc-araC, PBAD-gfp, J23101-mCherry in vector backbone Plasmid backbone This studypRPF185PTETO1 Promotersource Fagan et al. 201150 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122p95thlsupFASTthlsup promoter; Ampr MLSrColE1 Ori repL FAST Reporter source Streett et al. 2019pKO_mazFThr MLSr; ccdB repL; ori; bgaRand PbgaL upstream of mazFPromoter source Al-Hinai et al.2012 PBgaR-bgaR, PBgaL-gfp, PBgaLv1 J23101-mCherry in vector native promoter This study backbone Plac-bgaR, PBgaR PBgaL-gfp, PBgaLv2 J23101-mCherry in vector modificaThis studybackbonetionsPBgaR-DbgaR, PBgaL-gfp, plasmid w / o PBgaLDbgaR J23101-mCherry in vector This studybackbonePPTKv1 ParaR-araR, PPTK-gfp, J23101-mCherry in vector backbonenative This studyPPTKv3Plac-araR, PPTK-gfp, J23101-native w / mCherry in vector backbone modificationsThis studyParaR-DaraR, PPTK-gfp, PPTKv1DaraR J23101-mCherry in vector native plasmid w / o TFThis studybackbone PtetR-tetR, PTETO1-gfp, PTETO2 J23101-mCherry in vector aThis studybackbonend PtetR-tetRPtetR-tetR, PTETO1-gfp, PTETO1 J23101-mCherry in vector native promoter This study backbone PtetR-DtetR, PTETO1-gfp, PTETO1DtetR J23101-mCherry in vector native w / o This studybackbonePBgaL-A(144)CPBgaR-bgaR, PBgaL-gfp withPBgaL Validation in A144C NEB5a PBgaL-A(153)GPBgaR-bgaR, PBgaL-gfp withPBgaL Validation in A153G NEB5aThis studyPBgaL-T(158)GPBgaR-bgaR, PBgaL-gfp withPBgaL Validation in A158G NEB5a PBgaL-T(159)APBgaR-bgaR, PBgaL-gfp withPBgaL Validation in T159A NEB5aThis studyPBgaL-T(162)APBgaR-bgaR, PBgaL-gfp withPBgaL Validation in T162A NEB5a PBgaL-A(164)GPBgaR-bgaR, PBgaL-gfp withPBgaL Validation in A164G NEB5aThis studyPBgaL-T(166)CPBgaR-bgaR, PBgaL-gfp withPBgaL Validation in T166C NEB5a PBgaL-C(170)GPBgaR-bgaR, PBgaL-gfp withPBgaL Validation in C170G NEB5a PBgaL-C(170)TPBgaR-bgaR, PBgaL-gfp withPBgaL Validation in C170T NEB5aThis studyPBgaL-T(171)CPBgaR-bgaR, PBgaL-gfp withPBgaL Validation in T171C NEB5a PBgaL-T(171)APBgaR-bgaR, PBgaL-gfp withPBgaL Validation in T171A NEB5aThis studyPBgaL-APBgaR-bgaR, PBgaL-gfp withPBgaL Validation in A172C NEB5aThis studyPBgaL-T(179)CPBgaR-bgaR, PBgaL-gfp withPBgaL Validation in T179C NEB5astudy51 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PBgaL-A(193)TPBgaR-bgaR, PBgaL-gfp withPBgaL Validation in A193T NEB5aThis study TTGACA PBgaL-GATT(176- PBgaR-bgaR, PBgaL-gfp with PBgaL Validation in 179)CCGC GATT(176-179)CCGCNEB5aThis studyPTETO1-A(21)GPTETR-tetR, PTETO1-gfp withPTETO1 Validation A21G in NEB5aThis studyPTETO1-A(30)TPTETR-tetR, PTETO1-gfp withPTETO1 Validation A30T in NEB5aThis studyPTETO1-T(53)GPTETR-tetR, PTETO1-gfp withPTETO1 Validation A53G in NEB5aThis studyPTETO1-T(55)GPTETR-tetR, PTETO1-gfp withPTETO1 Validation T55G in NEB5aThis studyPTETO1-A(56)GPTETR-tetR, PTETO1-gfp withPTETO1 Validation A56G in NEB5aThis studyPTETO1-T(57)GPTETR-tetR, PTETO1-gfp withPTETO1 Validation T57G in NEB5aThis studyPTETO1-C(58)GPTETR-tetR, PTETO1-gfp withPTETO1 Validation C58G in NEB5aThis studyPTETO1-T(61)APTETR-tetR, PTETO1-gfp withPTETO1 Validation T61A in NEB5aThis studyPTETO1-G(62)CPTETR-tetR, PTETO1-gfp withPTETO1 Validation G62C in NEB5aThis studyPTETO1-A(63)CPTETR-tetR, PTETO1-gfp withPTETO1 Validation A63C in NEB5aThis studyPTETO1-T(64)CPTETR-tetR, PTETO1-gfp withPTETO1 Validation T64C in NEB5aThis study - 102)CgatcgGaTAgttTacain NEB5a52 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 102)CgatcgGaTAgtt Taca PPTK-A(207)GPlac-araR, PPTK-gfp withPPTK Validation in A207G Keio DaraCThis studyPPTK-T(165)CPlac-araR, PPTK-gfp withPPTK Validation in T165C Keio DaraCThis studyPPTK-T(156)GPlac-araR, PPTK-gfp withPPTK Validation in T156G Keio DaraCThis studyPPTK-A(155)GPlac-araR, PPTK-gfp withPPTK Validation in A155G Keio DaraCThis studyPPTK-A(141)CPlac-araR, PPTK-gfp withPPTK Validation in A141C Keio DaraCThis studyPPTK-A(133)TPlac-araR, PPTK-gfp withPPTK Validation in A133T Keio DaraCThis studyPPTK-A(122)TPlac-araR, PPTK-gfp withPPTK Validation in A122T Keio DaraCThis study GTGTCGTA GTGTCGTA PPTK- ATTTATACGTAC Plac-araR, PPTK-gfp with Validation in AAAT(99- ATTTATACGTACAAAT(99- )GTGTCGTAA 114)GTGTCGTAAACCAGAGKeio study114DaraCACCAGAG pAN3Kmr; B. subtilis phage Φ3T IEscherichia coli methyltransferase gene ER1821Al-Hinai et al. 2012E. coli - Clostridium shuttle Clostridium-E. coli pMTL85141 vector (Cmr; ColE1 ori; pIM13 shuttle vector; Heap et al.2009 ori) negative control 53 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PBgaR-bgaR, PBgaLnative promoter inPBgaLv1-yfast,pMTL85141 shuttle vectorC. acetobutylicum This study ATCC824 native promoter inPTETO1 pMTL85141 vectorC. acetobutylicum This study ATCC824 Plac-araR, PPTK-native promoter inPPTKv3yfast,pMTL85141 shuttle vectorC. acetobutylicum This study ATCC824 PTETR-tetR, PTETO1-yfast PTETO1 Validation pMTL85141- PTETO1-A(21)G with A21G, pMTL85141 shuttle in C. acetobutylicum This study vector ATCC824 PTETR-tetR, PTETO1-yfast PTETO1 Validation pMTL85141- PTETO1-A(63)C with A63C, pMTL85141 shuttle in C. acetobutylicum This study vector ATCC824 pMTL85141- PTETR-tetR, PTETO1-yfast PTETO1 Validation with T86C, pMTL85141 shuttle in C. acetobutylicum This study vector ATCC824 pMTL85141- PTETR-tetR, PTETO1-yfast PTETO1 Validation PTETO1-G(94)T with G94T, pMTL85141 shuttle in C. acetobutylicum This study vector ATCC824 pMTL85141- PTETR-tetR, PTETO1-yfast PTETO1 Validation PTETO1-C(95)A with C95A, pMTL85141 shuttle in C. acetobutylicum This study vector ATCC824 pMTL85141- PTETO1- PTETR-tetR, PTETO1-yfast PTETO1 Validation actctatcattgatagag(5 with actctatcattgatagag(52- 2- 68)cGcGGGGatACCCGgag, in C. acetobutylicum This study 68)cGcGGGGatAC pMTL85141 shuttle vector CCGgag pMTL85141-PPTK- Plac-araR, PPTK-yfast with PPTKv3 Validation A207G, pMTL85141 shuttle in C. acetobutylicum This study A(207)G vector ATCC824 pMTL85141-PPTK- Plac-araR, PPTK-yfast with PPTKv3 Validation A133G, pMTL85141 shuttle in C. acetobutylicum This study A(133)T vector ATCC824 pMTL85141-PPTK- Plac-araR, PPTK-yfast with PPTKv3 Validation AAAT(99- ATTTATACGTACAAAT 114)GTGTCGTAAACCAGAG, in C. acetobutylicum This study 114)GTGTCGTAA pMTL85141 shuttle ATCC824 ACCAGAG vector pMTL85141- PbgaR-bgaR, PBgaL-yfast with PBgaL Validation in PBgaL-A(153)G A153G, pMTL85141 shuttle C. acetobutylicum This study vector ATCC824 PbgaR-bgaR, PBgaL-yfast with PBgaL Validation in pMTL85141- PBgaL-T(159)A T159A, pMTL85141 shuttle C. acetobutylicum This study vector ATCC824 54 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 pMTL85141- PbgaR-bgaR, PBgaL-yfast with PBgaL Validation in T162A, pMTL85141 shuttle C. acetobutylicum This study PBgaL-T(162)A vector ATCC824 pMTL85141- PbgaR-bgaR, PBgaL-yfast with PBgaL Validation in PBgaL-A(164)G A164G, pMTL85141 shuttle C. acetobutylicum This study vector ATCC824 pMTL85141- PbgaR-bgaR, PBgaL-yfast with PBgaL Validation in PBgaL-T(166)C A166C, pMTL85141 shuttle C. acetobutylicum This study vector ATCC824 PbgaR-bgaR, PBgaL-yfast with PBgaL Validation in pMTL85141- PBgaL-C(170)T C170T, pMTL85141 shuttle C. acetobutylicum This study vector ATCC824 PbgaR-bgaR, PBgaL-yfast with PBgaL Validation in pMTL85141- PBgaL-T(171)C T171C, pMTL85141 shuttle C. acetobutylicum This study vector ATCC824 pMTL85141- PbgaR-bgaR, PBgaL-yfast with PBgaL Validation in A193T, pMTL85141 shuttle C. acetobutylicum This study vector ATCC824 PbgaR-bgaR, PBgaL-yfast with PBgaL Validation in TTCTAA(168-173)TTGCCA, C. acetobutylicum This study pMTL85141 shuttle vector ATCC824 PbgaR-bgaR, PBgaL-yfast with PBgaL Validation in TTCTAA(168-173)TTGACA, C. acetobutylicum This study pMTL85141 shuttle vector ATCC824 PbgaR-bgaR, PBgaL-yfast with PBgaL Validation in GATT(176-179)CCGC, C. acetobutylicum This study pMTL85141 shuttle vector ATCC824 pMTL85141- PBgaR-DbgaR, PBgaL-yfast, PBgaL Validation in PBgaLv1DbgaR pMTL85141 shuttle vector C. acetobutylicum This study ATCC824 pMTL85141- PtetR-DtetR, PTETO1-yfast, PTETO1 Validation PTETO1DtetR pMTL85141 shuttle vector in C. PPTKv3 Validation pMTL85141- Plac-DaraR, PPTK-yfast, PPTKv3DaraR pMTL85141 shuttle vector in C. acetobutylicum This study ATCC824 Table 8. Primers used to construct plasmids in this study. Primers used to construct plasmids in this study. Mutations and overhangs for hifi assembly are in lower case letters. Plasmid Forward Primer (5'-->3') Reverse Primer (5'-->3') Templ ate PBADNK259 ATGAGTAAAGGAGAAG NK260 GGATCTGAAGCTTG pNK25 (DmCherry) AACTTTTCAC (SEQ ID GGCC (SEQ ID NO: NO: 221) 222) NK257 cgggcccaagcttcagatccTTAT NK258 agttcttctcctttactcatATGT AraC- GACAACTTGACGGC ATATCTCCTTCTTAA pBAD (SEQ ID NO: 223) AGTTAAAC (SEQ ID NO: 224) PBADNK267 AAAGTGCCACCTAGTTC NK268 TCGGGGAAATGTGC PBADACCG (SEQ ID NO: 225) GCGG (SEQ ID NO: (DmCh 226) erry) 55 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 NK265 ttccgcgcacatttccccgaTCTAG NK266 ggtgaactaggtggcactttCT BASIC AAAGATCGATAGGTC CGAGTTTTTCAGCA _7 (SEQ ID NO: 227) AG (SEQ ID NO: 228) PBgaLNK259 ATGAGTAAAGGAGAAG NK260 GGATCTGAAGCTTG PBADAACTTTTCAC (SEQ ID GGCC (SEQ ID NO: NO: 229) 230) NK337 cgggcccaagcttcagatccttatatac NK338 agttcttctcctttactcatTTTA pKO_ ttggtttatttacttgattatttctgt CCCTCCCAATACATT mazF (SEQ ID NO: 231) TAAAATAA (SEQ ID NO: 232) PBgaL(PlacIq:bg NK372 TTCTACCTCCTAACCTA NK373 ATGCAAATATTGTG PBgaLv1 aR) TAAAATTAG (SEQ ID GAAAAAG (SEQ ID NO: 233) NO: 234) NK374 tttttccacaatatttgcatATTCAC NK375 ttataggttaggaggtagaaCG pGEX- CACCCTGAATTG (SEQ CCTGATGCGGTATTT iLOV ID NO: 235) TC (SEQ ID NO: 236) PPTKv1 NK259 ATGAGTAAAGGAGAAG NK260 GGATCTGAAGCTTG PBADAACTTTTCAC (SEQ ID GGCC (SEQ ID NO: NO: 237) 238) NK341 cgggcccaagcttcagatccTTAT NK342 tgctcaactgTATGATCTT C. TTCAATTTTAAGGTTGA CCATAACTTAACTT acetob ATC (SEQ ID NO: 239) AC (SEQ ID NO: 240) utylicu m ATCC 824 genom e NK346 gaagatcataCAGTTGAGCA NK347 agttcttctcctttactcatAATA C. AGTTTATG (SEQ ID NO: TTCCTCCTAAATTTA acetob 241) TAATCATTAATTATG utylicu (SEQ ID NO: 242) m ATCC 824 genom e PPTKv3 NK376 CAGTTGAGCAAGTTTAT NK377 ATGAAGCACAAATA PBgaLv1 G (SEQ ID NO: 243) TGAAG (SEQ ID NO: 244) NK378 tcttcatatttgtgcttcatATTCAC NK379 gtcataaacttgctcaactgCG pGEX- CACCCTGAATTG (SEQ CCTGATGCGGTATTT iLOV ID NO: 245) TC (SEQ ID NO: 246) PTETO2NK259 ATGAGTAAAGGAGAAG NK260 GGATCTGAAGCTTG PBADAACTTTTCAC (SEQ ID GGCC (SEQ ID NO: NO: 247) 248) NK282 cgggcccaagcttcagatccTTAA NK283 agttcttctcctttactcatATGT pBBR2 GACCCACTTTCACATTT ATATCTCCTTCTTAA K AAG (SEQ ID NO: 249) AAGATC (SEQ ID NO: 250) PTETO1NK286 ATGAGTAAAGGAGAAG NK287 AAAAATTAGGAATT PTETO2AAC (SEQ ID NO: 251) AATGATGTCTAG (SEQ ID NO: 252) NK284 catcattaattcctaatttttGTTGAC NK285 agttcttctcctttactcatAAA pRPF1 ATTATATCATTGATAGA ATTTTCTCCTTTACT 85 56 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PBgaLNK350 TTCTACCTCCTAACCTA NK349 TAAGGATCTGAAGC PBgaLv1 (DbgaL) TAAAATTAG (SEQ ID TTGG (SEQ ID NO: NO: 255) 256) PPTKv1 NK355 TTTTTTCTCCTATAAAA NK349 TAAGGATCTGAAGC PPTKv1 (DaraR) TAATTGTTTTTTATAG TTGG (SEQ ID NO: (SEQ ID NO: 257) 258) PPTKv2 NK355 TTTTTTCTCCTATAAAA NK349 TAAGGATCTGAAGC PPTKv2 (DaraR) TAATTGTTTTTTATAG TTGG (SEQ ID NO: (SEQ ID NO: 259) 260) PTETO2NK356 CATTAATTCCTAATTTT NK349 TAAGGATCTGAAGC PTETO2(DtetR) TGTTGAC (SEQ ID NO: TTGG (SEQ ID NO: 261) 262) PTETO1NK356 CATTAATTCCTAATTTT NK349 TAAGGATCTGAAGC PTETO1(DtetR) TGTTGAC (SEQ ID NO: TTGG (SEQ ID NO: 263) 264) pMTL85141- NK213 agatctttttttaacaaaacTTTTTA NK214 acgacggccagtgccacataT p95thlsPthl-yfast ACAAAAAGTATTGAAA CATACCCTCTTAACupFAST TTTGG (SEQ ID NO: 265) GAAAAC (SEQ ID NO: 266) NK215 TATGTGGCACTGGCCGT NK216 GTTTTGTTAAAAAA pMTL C (SEQ ID NO: 267) AGATCTGCAGGAGC 85141 (SEQ ID NO: 268) pMTL85141- NK453 ATGGAACACGTAGCAT NK454 GTTTTGTTAAAAAA PBgaLv1 PBgaLv1 TTG (SEQ ID NO: 269) AGATCTGC (SEQ ID NO: 270) NK455 agatctttttttaacaaaacTTATAT NK456 ccaaatgctacgtgttccatTTT pMTL ACTTGGTTTATTTACTT ACCCTCCCAATACA 85141 GATTATTTC (SEQ ID TTTAAAATAATTATG NO: 271) (SEQ ID NO: 272) pMTL85141- NK453 ATGGAACACGTAGCAT NK454 GTTTTGTTAAAAAA PBgaLv2 PBgaLv2 TTG (SEQ ID NO: 273) AGATCTGC (SEQ ID NO: 274) NK455 agatctttttttaacaaaacTTATAT NK456 ccaaatgctacgtgttccatTTT pMTL ACTTGGTTTATTTACTT ACCCTCCCAATACA 85141 GATTATTTC (SEQ ID TTTAAAATAATTATG NO: 275) (SEQ ID NO: 276) pMTL85141- NK453 ATGGAACACGTAGCAT NK454 GTTTTGTTAAAAAA PTETO1PTETO1TTG (SEQ ID NO: 277) AGATCTGC (SEQ ID NO: 278) NK457 agatctttttttaacaaaacTTAAG NK458 ccaaatgctacgtgttccatAA pMTL ACCCACTTTCACATTTA AATTTTCTCCTTTAC 85141 AG (SEQ ID NO: 279) TGC (SEQ ID NO: 280) pMTL85141- NK453 ATGGAACACGTAGCAT NK454 GTTTTGTTAAAAAA PPTKv3 PPTKv3 TTG (SEQ ID NO: 281) AGATCTGC (SEQ ID NO: 282) NK459 agatctttttttaacaaaacTTATTT NK460 ccaaatgctacgtgttccatAA pMTL CAATTTTAAGGTTGAAT TATTCCTCCTAAATT 85141 C (SEQ ID NO: 283) TATAATCATTAATTA TG (SEQ ID NO: 284) PBgaL- NK485 AAACATTATAcCATATT NK486 CTTATAATCTTATAA PBgaLv1 A(144)C TTAGAACTTTTTAAC ATTTTAATAACTAAT (SEQ ID NO: 285) ATATAAAG (SEQ ID NO: 286) PBgaL- NK487 AACATATTTTgGAACTT NK488 ATAATGTTTCTTATA PBgaLv1 A(153)G TTTAACTATTC (SEQ ID ATCTTATAAATTTTA NO: 287) 57 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 ATAAC (SEQ ID NO: 288) PBgaL- NK489 ATTTTAGAACgTTTTAA NK490 ATGTTATAATGTTTC PBgaLv1 T(158)G CTATTCTAAAAGATTAA TTATAATCTTATAAA TTTAC (SEQ ID NO: 289) TTTTAATAAC (SEQ ID NO: 290) PBgaL- NK491 TTTTAGAACTaTTTAAC NK492 TATGTTATAATGTTT PBgaLv1 T(159)A TATTCTAAAAGATTAAT CTTATAATCTTATAA TTAC (SEQ ID NO: 291) ATTTTAATAAC (SEQ ID NO: 292) PBgaL- NK493 TAGAACTTTTaAACTAT NK494 AAATATGTTATAAT PBgaLv1 T(162)A TCTAAAAGATTAATTTA GTTTCTTATAATCTT C (SEQ ID NO: 293) ATAAATTTTAATAA C (SEQ ID NO: 294) PBgaL- NK495 GAACTTTTTAgCTATTC NK496 TAAAATATGTTATA PBgaLv1 A(164)G TAAAAGATTAATTTAC ATGTTTCTTATAATC (SEQ ID NO: 295) (SEQ ID NO: 296) PBgaL- NK497 ACTTTTTAACcATTCTA NK498 TCTAAAATATGTTAT PBgaLv1 T(166)C AAAGATTAATTTACATA AATGTTTCTTATAAT TTAAC (SEQ ID NO: 297) C (SEQ ID NO: 298) PBgaL- NK499 TTTAACTATTgTAAAAG NK500 AAGTTCTAAAATAT PBgaLv1 C(170)G ATTAATTTACATATTAA GTTATAATGTTTC C (SEQ ID NO: 299) (SEQ ID NO: 300) PBgaL- NK501 TTTAACTATTtTAAAAG NK502 AAGTTCTAAAATAT PBgaLv1 C(170)T ATTAATTTACATATTAA GTTATAATGTTTC C (SEQ ID NO: 301) (SEQ ID NO: 302) PBgaL- NK503 TTAACTATTCcAAAAGA NK504 AAAGTTCTAAAATA PBgaLv1 T(171)C TTAATTTACATATTAAC TGTTATAATGTTTC (SEQ ID NO: 303) (SEQ ID NO: 304) PBgaL- NK505 TTAACTATTCaAAAAGA NK506 AAAGTTCTAAAATA PBgaLv1 T(171)A TTAATTTACATATTAAC TGTTATAATGTTTC (SEQ ID NO: 305) (SEQ ID NO: 306) PBgaL- NK507 TAACTATTCTcAAAGAT NK508 AAAAGTTCTAAAAT PBgaLv1 A(172)C TAATTTACATATTAAC ATGTTATAATGTTTC (SEQ ID NO: 307) (SEQ ID NO: 308) PBgaL- NK509 TCTAAAAGATcAATTTA NK510 ATAGTTAAAAAGTT PBgaLv1 T(179)C CATATTAACATTTAATT CTAAAATATGTTAT ATG (SEQ ID NO: 309) AATG (SEQ ID NO: 310) PBgaL- NK511 TTACATATTAtCATTTAA NK512 ATTAATCTTTTAGAA PBgaLv1 A(193)T TTATGGGTAAAAAC TAGTTAAAAAGTTC (SEQ ID NO: 311) (SEQ ID NO: 312) PBgaL- NK513 ATATTAACATaTAATTA NK514 GTAAATTAATCTTTT PBgaLv1 T(197)A TGGGTAAAAACTATTTT AGAATAGTTAAAAA G (SEQ ID NO: 313) G (SEQ ID NO: 314) PBgaL- NK517 ACATTTAATTgTGGGTA NK518 TAATATGTAAATTA PBgaLv1 A(203)G AAAACTATTTTG (SEQ ATCTTTTAGAATAGT ID NO: 315) TAAAAAG (SEQ ID NO: 316) PBgaL- NK519 TTTTTAACTAttgccaAAG NK520 GTTCTAAAATATGTT PBgaLv1 TTCTAA(16 ATTAATTTACATATTAA ATAATGTTTCTTATA 8- CATTTAATTATG (SEQ ATC (SEQ ID NO: 318) 173)TTGCC ID NO: 317) A 58 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PBgaL- NK521 TTTTTAACTAttgacaAAG NK522 GTTCTAAAATATGTT PBgaLv1 TTCTAA(16 ATTAATTTACATATTAA ATAATGTTTCTTATA 8- CATTTAATTATG (SEQ ATC (SEQ ID NO: 320) 173)TTGAC ID NO: 319) A PBgaL- NK525 TATTCTAAAAccgcAATT NK526 GTTAAAAAGTTCTA PBgaLv1 GATT(176- TACATATTAACATTTAA AAATATGTTATAAT 179)CCGC TTATGGGTAAAAAC G (SEQ ID NO: 322) (SEQ ID NO: 321) PTETO1- NK527 GATGCCCTGGgCTTCAT NK528 GAAATAAAAAACTA PTETO1A(21)G GAAA (SEQ ID NO: 323) GTTTGACAAATAAC TC (SEQ ID NO: 324) PTETO1- NK529 GACTTCATGAtAAACTA NK530 CAGGGCATCGAAAT PTETO1A(30)T AAAAAAATATTGAC AAAAAAC (SEQ ID (SEQ ID NO: 325) NO: 326) PTETO1- NK531 ATATTGACACgCTATCA NK532 TTTTTTTAGTTTTTC PTETO1T(53)G TTGATAG (SEQ ID NO: ATGAAGTCC (SEQ ID 327) NO: 328) PTETO1- NK533 ATTGACACTCgATCATT NK534 ATTTTTTTTAGTTTT PTETO1T(55)G GATAGAG (SEQ ID NO: TCATGAAGTC (SEQ 329) ID NO: 330) PTETO1- NK535 TTGACACTCTgTCATTG NK536 TATTTTTTTTAGTTT PTETO1A(56)G ATAGAG (SEQ ID NO: TTCATGAAGTC (SEQ 331) ID NO: 332) PTETO1- NK537 TGACACTCTAgCATTGA NK538 ATATTTTTTTTAGTT PTETO1T(57)G TAGAG (SEQ ID NO: 333) TTTCATGAAGTC (SEQ ID NO: 334) PTETO1- NK539 GACACTCTATgATTGAT NK540 AATATTTTTTTTAGT PTETO1C(58)G AGAGTATAATTAAAAT TTTTCATGAAGTC AAG (SEQ ID NO: 335) (SEQ ID NO: 336) PTETO1- NK541 ACTCTATCATaGATAGA NK542 GTCAATATTTTTTTT PTETO1T(61)A GTATAATTAAAATAAG AGTTTTTCATG (SEQ (SEQ ID NO: 337) ID NO: 338) PTETO1- NK543 CTCTATCATTcATAGAG NK544 TGTCAATATTTTTTT PTETO1G(62)C TATAATTAAAATAAG TAGTTTTTCATG (SEQ ID NO: 339) (SEQ ID NO: 340) PTETO1- NK545 TCTATCATTGcTAGAGT NK546 GTGTCAATATTTTTT PTETO1A(63)C ATAATTAAAATAAG TTAGTTTTTC (SEQ (SEQ ID NO: 341) ID NO: 342) PTETO1- NK547 CTATCATTGAcAGAGTA NK548 AGTGTCAATATTTTT PTETO1T(64)C TAATTAAAATAAG (SEQ TTTAGTTTTTC (SEQ ID NO: 343) ID NO: 344) PTETO1- NK549 TATCATTGATgGAGTAT NK550 GAGTGTCAATATTTT PTETO1A(65)G AATTAAAATAAGC (SEQ TTTTAGTTTTTC ID NO: 345) (SEQ ID NO: 346) PTETO1- NK551 AAAATAAGCTcGATCGT NK552 AATTATACTCTATCA PTETO1T(86)C AGCG (SEQ ID NO: 347) ATGATAGAG (SEQ ID NO: 348) PTETO1- NK553 AGCTTGATCGgAGCGTT NK554 TATTTTAATTATACT PTETO1T(92)G AACA (SEQ ID NO: 349) CTATCAATGATAGA GTG (SEQ ID NO: 350) PTETO1- NK555 CTTGATCGTAtCGTTAA NK556 CTTATTTTAATTATA PTETO1G(94)T CAGATC (SEQ ID NO: CTCTATCAATGATA 59 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PTETO1- NK557 TTGATCGTAGaGTTAAC NK558 GCTTATTTTAATTAT PTETO1C(95)A AGATC (SEQ ID NO: 353) ACTCTATCAATG (SEQ ID NO: 354) PTETO1- NK559 TCGTAGCGTTtACAGAT NK560 TCAAGCTTATTTTAA PTETO1A(99)T CTGAG (SEQ ID NO: 355) TTATACTCTATC (SEQ ID NO: 356) PTETO1- NK561 acccggagTATAATTAAAA NK562 atccccgcgTGTCAATAT PTETO1actctatcattgat TAAGCTTGATCGTAG TTTTTTTAGTTTTTC agag(52- (SEQ ID NO: 357) ATG (SEQ ID NO: 358) 68)cGcGGG GatACCCGg ag PTETO1- NK563 agtttacaGATCTGAGCTCC NK564 atccgatcgAGCTTATTT PTETO1tgatcgtagcgtt TGCAGT (SEQ ID NO: TAATTATACTCTATC aaca(86- 359) AATGATAG (SEQ ID 102)CgatcgG NO: 360) aTAgttTaca PPTK- NK601 AAAAATTAGTgGAAAA NK602 AATTAAATGTCTTA PPTKv3 A(207)G TCTAGAAATATACC AAACAAATATTATA (SEQ ID NO: 361) TG (SEQ ID NO: 362) PPTK-T(165)C NK603 ATACGTATAGcACATAT NK604 AGTAAAATATTTGT PPTKv3 AATATTTGTTTTAAG ACGCATATG (SEQ ID (SEQ ID NO: 363) NO: 364) PPTK- NK605 TATTTTACTAgACGTAT NK606 TTTGTACGCATATGT PPTKv3 T(156)G AGTACATATAATATTTG CAATG (SEQ ID NO: (SEQ ID NO: 365) 366) PPTK- NK607 ATATTTTACTgTACGTA NK608 TTGTACGCATATGTC PPTKv3 A(155)G TAGTACATATAATATTT AATG (SEQ ID NO: G (SEQ ID NO: 367) 368) PPTK- NK609 ACATATGCGTcCAAATA NK610 CAATGACTTTTTAGT PPTKv3 A(141)C TTTTACTATAC (SEQ ID AATTTGTAC (SEQ ID NO: 369) NO: 370) PPTK- NK611 AGTCATTGACtTATGCG NK612 TTTTAGTAATTTGTA PPTKv3 A(133)T TACAAATATTTTAC CGTATAAATTTAATT (SEQ ID NO: 371) ATATTTTG (SEQ ID NO: 372) PPTK- NK613 AATTACTAAAtAGTCAT NK614 TGTACGTATAAATTT PPTKv3 A(122)T TGACATATGC (SEQ ID AATTATATTTTG NO: 373) (SEQ ID NO: 374) PPTK- NK615 TACGTACAAAgTACTAA NK616 TAAATTTAATTATAT PPTKv3 T(114)G AAAGTC (SEQ ID NO: TTTGACAACTATG 375) (SEQ ID NO: 376) PPTK- NK617 TATACGTACAgATTACT NK618 AATTTAATTATATTT PPTKv3 A(112)G AAAAAGTCATTG (SEQ TGACAACTATGAAG ID NO: 377) (SEQ ID NO: 378) PPTK- NK619 ATTTATACGTcCAAATT NK620 TTAATTATATTTTGA PPTKv3 A(109)C ACTAAAAAGTC (SEQ ID CAACTATGAAG NO: 379) (SEQ ID NO: 380) PPTK- NK621 AAATTTATACaTACAAA NK622 AATTATATTTTGACA PPTKv3 G(107)A TTACTAAAAAGTCATTG ACTATGAAGTTAAT (SEQ ID NO: 381) TAG (SEQ ID NO: 382) PPTK- NK623 ATTAAATTTAgACGTAC NK624 TATATTTTGACAACT PPTKv3 T(104)G AAATTACTAAAAAG ATGAAGTTAATTAG 60 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 PPTK- NK625 AATTAAATTTgTACGTA NK626 ATATTTTGACAACTA PPTKv3 A(103)G CAAATTACTAAAAAG TGAAGTTAATTAG (SEQ ID NO: 385) (SEQ ID NO: 386) PPTK- NK643 AAATATACCAggGAATC NK628 CTAGATTTTCTACTA PPTKv3 TA(228- ACACTAGATGAATTC ATTTTTAATTAAATG 229)GG (SEQ ID NO: 387) (SEQ ID NO: 388) PPTK- NK644 AGAAATATACggTAGAA NK630 AGATTTTCTACTAAT PPTKv3 CA(226- TCACACTAGATGAATTC TTTTAATTAAATG 227)GG (SEQ ID NO: 389) (SEQ ID NO: 390) PPTK- NK645 ACAAATTACTgtgttGTCA NK632 ACGTATAAATTTAA PPTKv3 AAAAA(119 TTGACATATGCG (SEQ TTATATTTTGACAAC -123)GTGTT ID NO: 391) (SEQ ID NO: 392) PPTK- NK646 caacATATTTTACTATAC NK647 actccTATGTCAATGA PPTKv3 TGCGTACA GTATAGTACATATAATA CTTTTTAGTAATTTG A(136- TTTG (SEQ ID NO: 393) (SEQ ID NO: 394) 144)GGAGT CAAC PPTK- NK648 aacCAAATTACTAAAAA NK649 gtcgAAATTTAATTAT PPTKv3 ATACGTA( GTCATTGAC (SEQ ID ATTTTGACAACTATG 103- NO: 395) (SEQ ID NO: 396) 109)CGACA AC PPTK- NK650 agagTACTAAAAAGTCAT NK651 ggttGTATAAATTTAA PPTKv3 GTACAAAT TGACATATG (SEQ ID TTATATTTTGACAAC (107- NO: 397) TATG (SEQ ID NO: 114)AACCA 398) GAG PPTK- NK652 cgtaGTACAAATTACTAA NK653 acacTTAATTATATTT PPTKv3 ATTTATAC AAAGTCATTG (SEQ ID TGACAACTATGAAG (99- NO: 399) (SEQ ID NO: 400) 106)GTGTC GTA PPTK- NK654 aaccagagTACTAAAAAGT NK655 tacgacacTTAATTATAT PPTKv3 ATTTATAC CATTGACATATG (SEQ TTTGACAACTATGA GTACAAAT ID NO: 401) AG (SEQ ID NO: 402) (99- 114)GTGTC GTAAACCA GAG pMTL85141- NK527 GATGCCCTGGgCTTCAT NK528 GAAATAAAAAACTA pMTL PTETO1- GAAA (SEQ ID NO: 403) GTTTGACAAATAAC 85141- A(21)G TC (SEQ ID NO: 404) PTETO1pMTL85141- NK545 TCTATCATTGcTAGAGT NK546 GTGTCAATATTTTTT pMTL PTETO1- ATAATTAAAATAAG TTAGTTTTTC (SEQ 85141- A(63)C (SEQ ID NO: 405) ID NO: 406) PTETO1pMTL85141- NK551 AAAATAAGCTcGATCGT NK552 AATTATACTCTATCA pMTL PTETO1- AGCG (SEQ ID NO: 407) ATGATAGAG (SEQ ID 85141- T(86)C NO: 408) PTETO1pMTL85141- NK555 CTTGATCGTAtCGTTAA NK556 CTTATTTTAATTATA pMTL PTETO1- CAGATC (SEQ ID NO: CTCTATCAATGATA 85141- G(94)T 409) G (SEQ ID NO: 410) PTETO1pMTL85141- NK557 TTGATCGTAGaGTTAAC NK558 GCTTATTTTAATTAT pMTL PTETO1- AGATC (SEQ ID NO: 411) ACTCTATCAATG 85141- 61 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 pMTL85141- NK561 acccggagTATAATTAAAA NK562 atccccgcgTGTCAATAT pMTL PTETO1- TAAGCTTGATCGTAG TTTTTTTAGTTTTTC 85141- actctatcattgat (SEQ ID NO: 413) ATG (SEQ ID NO: 414) PTETO1agag(52- 68)cGcGGG GatACCCGg ag pMTL85141- NK601 AAAAATTAGTgGAAAA NK602 AATTAAATGTCTTA pMTL PPTK- TCTAGAAATATACC AAACAAATATTATA 85141- A(207)G (SEQ ID NO: 415) TG (SEQ ID NO: 416) PPTKv3 pMTL85141- NK611 AGTCATTGACtTATGCG NK612 TTTTAGTAATTTGTA pMTL PPTK- TACAAATATTTTAC CGTATAAATTTAATT 85141- A(133)T (SEQ ID NO: 417) ATATTTTG (SEQ ID PPTKv3 NO: 418) pMTL85141- NK615 TACGTACAAAgTACTAA NK616 TAAATTTAATTATAT pMTL PPTK- AAAGTC (SEQ ID NO: TTTGACAACTATG 85141- T(114)G 419) (SEQ ID NO: 420) PPTKv3 pMTL85141- NK625 AATTAAATTTgTACGTA NK626 ATATTTTGACAACTA pMTL PPTK- CAAATTACTAAAAAG TGAAGTTAATTAG 85141- A(103)G (SEQ ID NO: 421) (SEQ ID NO: 422) PPTKv3 pMTL85141- NK643 AAATATACCAggGAATC NK628 CTAGATTTTCTACTA pMTL PPTK- ACACTAGATGAATTC ATTTTTAATTAAATG 85141- TA(228- (SEQ ID NO: 423) (SEQ ID NO: 424) PPTKv3 229)GG pMTL85141- NK654 aaccagagTACTAAAAAGT NK655 tacgacacTTAATTATAT pMTL PPTK- CATTGACATATG (SEQ TTTGACAACTATGA 85141- ATTTATAC ID NO: 425) AG (SEQ ID NO: 426) PPTKv3 GTACAAAT (99- 114)GTGTC GTAAACCA GAG pMTL85141- NK487 AACATATTTTgGAACTT NK488 ATAATGTTTCTTATA pMTL PBgaL- TTTAACTATTC (SEQ ID ATCTTATAAATTTTA 85141- A(153)G NO: 427) ATAAC (SEQ ID NO: PBgaLv1 428) pMTL85141- NK491 TTTTAGAACTaTTTAAC NK492 TATGTTATAATGTTT pMTL PBgaL- TATTCTAAAAGATTAAT CTTATAATCTTATAA 85141- T(159)A TTAC (SEQ ID NO: 429) ATTTTAATAAC (SEQ PBgaLv1 ID NO: 430) pMTL85141- NK493 TAGAACTTTTaAACTAT NK494 AAATATGTTATAAT pMTL PBgaL- TCTAAAAGATTAATTTA GTTTCTTATAATCTT 85141- T(162)A C (SEQ ID NO: 431) ATAAATTTTAATAA PBgaLv1 C (SEQ ID NO: 432) pMTL85141- NK495 GAACTTTTTAgCTATTC NK496 TAAAATATGTTATA pMTL PBgaL- TAAAAGATTAATTTAC ATGTTTCTTATAATC 85141- A(164)G (SEQ ID NO: 433) (SEQ ID NO: 434) PBgaLv1 pMTL85141- NK497 ACTTTTTAACcATTCTA NK498 TCTAAAATATGTTAT pMTL PBgaL- AAAGATTAATTTACATA AATGTTTCTTATAAT 85141- T(166)C TTAAC (SEQ ID NO: 435) C (SEQ ID NO: 436) PBgaLv1 pMTL85141- NK501 TTTAACTATTtTAAAAG NK502 AAGTTCTAAAATAT pMTL PBgaL- ATTAATTTACATATTAA GTTATAATGTTTC 85141- C(170)T C (SEQ ID NO: 437) (SEQ ID NO: 438) PBgaLv1 pMTL85141- NK503 TTAACTATTCcAAAAGA NK504 AAAGTTCTAAAATA pMTL PBgaL- TTAATTTACATATTAAC TGTTATAATGTTTC 85141- T(171)C (SEQ ID NO: 439) (SEQ ID NO: 440) PBgaLv1 62 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Biological Resources. Bacterial strains and plasmids used in the study are listed in Table 7. NEB® 5-alpha Competent E. coli (High Efficiency) (Cat. #C2987) and NEB® 5-alpha Competent E. coli (Subcloning Efficiency) (Cat. #C2988J) were purchased from New England Biolabs (Ipswich, MA) while electrocompetent ElectroMAX™ DH5α-E Competent Cells were purchased from ThermoFisher (Cat. # 11319019). C. acetobutylicum ATCC 824 were purchased from American Type Culture Collection (Manassas, VA). E. coli K-12 BW25113 JW0063-1 (∆araC) from the Keio Collection were purchased from the E. coli stock center (New Haven, CT)
[0101] . Microorganisms and growth media. All E. coli bacterial strains in this study were grown in liquid or solid Luria broth, pH 7.0 (LB medium) which includes: 10 g / L tryptone (VWR, cat. #J859-500G), 5 g / L ultra-pure yeast extract*** (VWR, cat. #J850-5KG), 10 g / L sodium chloride (VWR, Cat. #0241-5KG), and 15 g / L agar for solid (Fisher, cat. #BP1423-500) in ultra-pure water. The LB medium was autoclaved at 121°C for 15-30 minutes. For cultures with plasmids, LB medium was supplemented with 100 µg / mL ampicillin (Amp100), 35 µg / mL Kanamycin (Kan35), or 35 µg / mL Chloramphenicol (Cm35). (***The VWR brand ultra-pure yeast extract results in lower background fluorescence on the flow cytometer compared to the 63 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 BD Bacto™ yeast extract (data not shown). Thus, cells sampled for flow cytometry experiments were not washed with PBS, pH 7.4.) Wild-type C. acetobutylicum ATCC 824 cultures were grown on liquid and solid 2xYTG, pH 6.5 medium or liquid Clostridium Growth Medium (CGM), pH 6.8, supplemented with 15 µg / mL thiamphenicol (Tp15) for C. acetobutylicum strains carrying recombinant plasmids. Components for making 2xYTG, pH 6.5 include: 16 g / L tryptone (VWR, cat. #J859-500G), 10 g / L ultra-pure yeast extract (VWR, cat. #J850-5KG), 4 g / L sodium chloride (NaCl, VWR, Cat. #0241-5KG), 5 g / L D-glucose anhydrous (C6H12O6, VWR, cat. #0188-5KG) and 15 g / L agar for solid (Fisher, cat. #BP1423-500) in ultra-pure water. Once dissolved, the pH was adjusted to pH 6.5 with hydrochloride acid (HCl). The LB medium was autoclaved at 121°C for 15-30 minutes. To make CGM, pH 6.8, first the following components were dissolved in 750 mL ultra-pure water: 0.713 g / L MgSO4∙7H2O, 2 g / L C4H8N2O3, 5 g / L ultra-pure yeast extract, 2 g / L (NH4)2SO4, and 2.46 g / L CH3COONa. In addition, 1 mL of 0.1 g / mL MnSO4∙H2O and 0.1 g / mL FeSO4∙7H2O was added to the solution. Once dissolved, the pH was adjusted to 6.8 with HCl prior to autoclaving. A glucose solution with 80 g C6H12O6 in 250 mL ultra-pure water was autoclaved in a separate flask. Both autoclaved solutions were combined under a sterile environment. Once cooled, 10 mL of sterile 100X potassium phosphate solution and 100 µL of sterile 4-aminobenzoic acid (PABA) were added. The sterile 100X potassium phosphate solution was prepared by dissolving 7.5 g KH2PO4, 7.5 g K2HPO4, and 10 g NaCl in 100 mL ultra-pure water, and subsequently filter-sterilized with 0.2 µm filter. To prepare the PABA solution, 0.2 g of PABA was dissolved in 50 mL of ultra-pure water, and subsequently filter-sterilized with 0.2 µm syringe filter. The leftover PABA solution was stored at room temperature and protected from light. All liquid and solid growth medium for wild-type C. acetobutylicum ATCC 824 were deoxygenated for at least two days in the anaerobic chamber (Coy Labs, Grass Lake, MI), pressurized to ~20 psig with 7% / 10% mixture of H2 / CO2(Airgas, Catalog #X03NI83C30036A8) and ultra-high purity grade nitrogen (Airgas, Catalog #NI UHP300). For volumes less than 10 mL, cultures were either grown in 5 mL culture tubes (VWR, cat. #60818-500), 15 mL conical tube (VWR, cat. #89039-664), 50 mL conical tubes (VWR, cat. #89039-658), or 5 mL deep-well 48-well plates (VWR, cat. #43001-0062), or 10 mL deep-well 24-well plate (VWR, cat. #43001- 0066) sealed with breathable rayon film for culture plates (VWR, Cat. #60941-086). Bacterial Growth Measurements. Optical density (OD600) of bacterial cultures was measured on DeNovix DS-11+ Spectrophotometer (DeNovix Inc., Wilmington, DE) for E. coli and C. acetobutylicum. OD600 measurements for preparing electrocompetent C. acetobutylicum 64 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 cells were done on the Biochrom WPA CO8000 Cell density meter located inside the anaerobic chamber. For both spectrophotometers, samples were blanked using the uninoculated medium in 45 mm x 10 mm x 10 mm (H x W x D) cuvettes (Greiner, cat. #613101). All cultures were diluted to OD600 less than 0.6 in fresh medium prior to measurement. Genome extraction. Genomic DNA isolation from C. acetobutylicum was performed according to the manufacturer’s recommendations. In brief, C. acetobutylicum cells were grown in 20 mL culture in 2xYTG, pH 6.5 until an OD600=0.5-0.7 at 37°C. Then cells were harvested by centrifugation for 10 min at 5000xg. After the supernatant was discarded, the cell pellet was resuspended in 180 µL enzymatic lysis buffer (20 mM Tris, pH 8.0; 2 mM EDTA; 1.2% triton X-100; 200 µg / mL lysozyme added the day of use) and incubated at 37°C for 45 minutes. Then 25 µL Proteinase K and 200 µL Buffer AL (without ethanol) were added, vortexed, and then incubated at 56°C for 30 min in a heat block. The lysed cells were thoroughly mixed with 200 µL ethanol (96-100%) to the sample. The entire mixture (including any precipitate) was transferred into the Dneasy Mini spin column placed in a 2 mL collection tube. The flow-through was discarded after centrifugation at 6000xg. The Dneasy Mini spin column with the captured genomic DNA were placed in a new 2 mL collection tube, then 500 µL Buffer AW1 was added. The flow-through was discarded via centrifugation for 1 minute at 6000xg. Afterwards, the spin column was placed in a new 2 mL collection tube where 500 µL of Buffer AW2 was added into the column and centrifuged for 3 minutes at 20,000xg (14,000 rpm) to dry the Dneasy membrane. The flow-through and the collection tube were discarded. Finally, the spin column was placed in a new clean 1.5 mL microcentrifuge tube where the purified genomic DNA was incubated for 1 minute at room temperature in 100 µL Buffer AE, and subsequently eluted by centrifugation for 1 minute at 6000xg. The purified C. acetobutylicum was then used as the PCR template to amplify the regions to construct the ARAi promoter system. Plasmid construction. Construction of native promoter plasmids (FIGs.15A-15C). The backbone of pNK94, which includes gfp reporter gene and J23101-mCherry to isolate cell populations for flow analysis, was used as the backbone of the native promoter plasmids, PBgaL, ARAi, and PTETO1. Both the vector backbone and the BgaR-PBgaL (pKO_mazF
[0085] ), ARAi (PCR from C. acetobutylicum ATCC 824 based on plasmid from
[0087] ), and PTETO1promoter (pRPF185
[0088] , Addgene, cat. #106367) expression cassettes were amplified with Q5® High-Fidelity 2X Master Mix according to the manufacturer’s instructions. For the PTETO1 plasmid, the E. coli RBS and the tetR gene from pBBR2k-GFPuv (
[0102] , Addgene, #106383) were amplified. Once fragment sizes were confirmed visually on 0.8% TBE agarose gel, PCR products were purified 65 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 and digested with DpnI to remove any residual parental plasmid DNA template. The vector and promoter cassette insert were assembled using NEBuilder® HiFi DNA Assembly Master Mix according to the manufacturer’s recommendations, and subsequently 2 µL of the ligated samples were chemically transformed into 25 µL NEB5α using the manufacturer’s instructions. For each transformation, 100 µL of the transformed cells were plated on LB Amp100 and incubated at 37°C for ~18 hours following 1 hour recovery at 37°C (250 rpm). Three individual colonies were screened using colony PCR with Taq 2X Master Mix to verify the desired insert in the correct orientation was present. Positive clones that were visually verified in 0.8% agarose gel were grown in 5 mL LB Amp100 in 15 mL conical tubes (placed at an angle) overnight at 37°C shaking at 250 rpm for ~18 hours. Cultures were miniprepped using QIAGEN Plasmid Miniprep Kit to isolated plasmids to verify with Sanger sequencing. Construction of ∆TF controls and validation mutants. Introduction of point mutations and deletions of the transcription factor genes were incorporated or removed via PCR using Q5® High-Fidelity 2X Master Mix with primers containing the desired mutations listed in Table 8. PCR product sizes were first verified on 0.8% agarose, purified, and then digested with DpnI to remove the parental template. Blunt ended fragments were phosphorylated with T4 Polynucleotide Kinase and then ligated with Instant Sticky-end Ligase Master Mix. Then 2 µL of the ligation was transformed into 25 µL NEB5α, screened using colony PCR, miniprepped, and verified with Sanger sequencing. Promoter analysis in E. coli via flow cytometry. The native promoter and control plasmids were transformed into NEB5α (subcloning efficiency) at least 2 days prior to the promoter analysis. Cultures of each construct from individual colonies were grown overnight at 37°C in 3 mL LB Amp100 in 5 mL culture tubes, shaking at 250 rpm. The next day the cultures were diluted to an OD600of ~0.05 in pre-warmed 3 mL LB Amp100 in 15 mL conical tubes (placed at an angle) and grown at 37°C. When the OD600 reached ~0.1-0.2, 375 µL of cultures were transferred to 5-mL culture tubes with 125 µL LB Amp100 or 4X D-Lactose Monohydrate (VWR, cat. #470300-964), L-arabinose (Acros organics, cat. #365181000), or anhydrotetracycline hydrochloride (ThermoFisher Scientific, cat. #J66688.MA). After induction, cultures were incubated at 37°C, shaking at 250 rpm. Measurements of the GFP fluorescence were measured at various time points on the Attune NxT (blue solid-state laser (488 nm excitation), an optical filter at 530 / 30 nm for GFP fluorescence, and 488 / 10 nm optical filter for side scatter (SSC)). Approximately, 10 μL of each sample were transferred to 500 μL PBS, 66 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 pH 7.4 in 5 mL polystyrene tubes, where the median fluorescence intensity of the FITC-A fluorescence of 10,000 events per sample was measured at a flow rate of 12.5 µL / min. Promoter library generation and assembly. Promoter libraries were generated by random mutagenesis using GeneMorph II Random Mutagenesis Kit. The first round of error- prone PCR (epPCR) was amplified off the native plasmid as the template. The resulting PCR product was purified and used as the starting template for the next round of epPCR. A total of six to seven rounds of epPCR were completed to achieve high mutation rates. The last four rounds of purified epPCR products were gel purified using QIAquick gel purification kit using the manufacturer’s recommendations with minor modifications in the protocol (see
[0036] ) to remove contamination of the plasmid template prior to assembly with the NEBuilder® HiFi DNA Assembly Master Mix into the library vector. The library vector with the gfp gene underwent two rounds of PCR; the first was amplified from the purified native PBgaLv1, PBgaLv2, PPTK, and PTETO1plasmid templates, and the second from the PCR product from the first round of PCR after purification and DpnI digestion. The vector from the second round of PCR was purified and then used as the library vector to minimize the native plasmid used as the PCR template in the transformed library. Similarly, HiFi assembly of the library vector only was prepared and transformed into 25 µL electrocompetent DH5α to calculate the library background of vectors without the insert within the library. Small-scale library transformation and selection. The last four rounds of epPCR library were assessed prior to generating the final library at a larger scale. Prior to transformation, the HiFi assembly library mix were desalted using a modified method described in
[0052] . Instead of using a micropipet tip, a 200 µL PCR tube was used to create a conical-shaped well in the agarose and glucose mix. After 90 minutes on ice, the desalted library was transferred to a clean tube. Then 1-2.5 µL of the hifi assembly library was electroporated into 25 µL DH5α per transformation in 1 mm gap cuvettes (VWR, cat. #76102-576), pulsed at 1800V in BioRad MicroPulser. Following electroporation, 250 μL pre-warmed SOC media was added directly into the cuvette, and then transferred into 5-mL culture tubes for recovery at 37°C shaking at 250 rpm for 1 hour. About 20 μL of the recovered library was plated on solid LB Amp100 to estimate the library size. The rest of the transformed library was transferred to a 15-mL conical tube and 2725 μL of liquid LB Amp100 was added for antibiotic counterselection. To monitor cell death and viability by sampling 20 μL of the library, diluted in 500 μL of PBS, pH 7.4, over time on the flow cytometer based on the %gated of the FSC and SSC dot plot. After about 3 hours of antibiotic selection, the cultures were placed on ice to stop the growth. The selected library 67 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 culture was mixed with sterile 80% glycerol (final glycerol concentration of 20%) directly into the 15-mL conical tube and then stored at -80°C. The colonies of the plated libraries from the day prior were counted to calculate the transformation efficiency, as well as the vector only background, to ensure that the library size was large enough but contained <1% vector only background in the entire library population. Small-scale library analysis via flow cytometry. Assessment of the promoter library GFP distribution was done as described under ‘Promoter analysis in E. coli via flow cytometry’ under the ‘Methods & Materials’ section with few differences described here: Overnight cultures of the controls and the different epPCR round libraries were grown overnight at 30°C, shaking at 250 rpm; controls were grown from individual colonies in 3 mL LB Amp100 from transformations plated within 2 days; libraries were thawed on ice (~15 minutes) from the -80°C stored glycerol stocks prepped using the protocol described under ‘Small-scale library transformation and selection’; and transferred to 50 mL conical tubes containing 15 mL LB Amp100 for growth. The plasmids isolated from overnight cultures of the tested libraries were harvested by centrifugation at 6,800xg for 10 minutes. The supernatant was discarded, and the plasmid libraries were isolated to be prepped for NGS, as described in ‘NGS Sample Preparation’ in the ‘Materials and Methods’ section, to assess the library diversity and mutation frequency of the last four rounds of the epPCR library. Large-scale library transformation and selection. Once the library round was determined for sort-seq, a total of eight transformations per library were selected, where 20 μL of the desalted HiFi assembly library was combined with 200 µL electrocompetent DH5α cells. After gently mixing, 27.5 μL of the DNA-cell mixture were transferred into pre-chilled 1 mm gap electroporation cuvettes (VWR, cat. #76102-576). After 10 minutes of incubation, the cuvette was placed in the pulser. Immediately after the DNA-cell mixture was pulsed at 1800V, the cells were recovered in 250 µL of prewarmed SOC media per transformation directly in the cuvette. Using an ultra-fine tip pipet bulb (VWR, cat. #414004-019), the cells were transferred into a 5 mL-culture tube per transformation. After recovery for 1 hour at 37°C shaking at 250 rpm, all recovered cells were pooled into a sterile 50-mL conical tube. To calculate the approximate size of the library, 20 μL of the recovered cells were serially diluted and plated on solid LB medium Amp100 before the antibiotic selection. To start the library selection, 14 mL of LB Amp100 was added to tube with the recovered library and then monitored cell death and viability of the library as described in “Small-scale library transformation and selection” of the methods. 68 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Library sorting. Seed culture and induction. Libraries were started four 1.6-mL frozen glycerol stocks in 200 mL of LB Amp100 in a 1 L baffled flask at 30°C for ~15 hours at 150 rpm. As controls, the native promoter plasmids were also grown overnight in 10 mL LB Amp100 in 50 mL conical tubes. Overnight cultures were diluted in pre-warmed 100 mL LB Amp100 to OD600 of 0.1 to keep the cells in the log phase and grown in 500 mL baffled flasks for ~30 min at 37°C, 250 rpm. The controls were grown similarly but in a 250 mL baffled flask with pre- warmed 50 mL LB Amp100. Once the OD600reached ~0.2, each of the inoculated cultures were split into two baffled flasks (50 mL each for libraries in 250 mL baffled flasks; 25 mL each for controls in 125 mL baffled flasks) in which one was induced with the final concentration of 0.2% (w / v) inducer and the other with no inducer. The induced and uninduced cultures were grown at 37C for ~5-6 hours, shaking at 250 rpm. Post-induction prep for sorting. Samples were removed from the incubator and then immediately placed on ice. A sample amount of the cultures (both induced and uninduced) were collected for initial analysis on the flow cytometer and OD600readings prior to centrifugation. Both induced and uninduced libraries were transferred to pre-chilled 50-mL conical tubes where each condition were aliquoted 40 mL and 10 mL. A total of four tubes were centrifuged at 4°C, 3000xg, for 10 minutes. Supernatant were carefully discarded. The tubes with the 40 mL library cultures were saved for future use, while tubes with 10 mL library cultures were resuspended in 10 mL of pre-chilled PBS, pH 7.4 to be used for sorting. Finally, the resuspended libraries were diluted to OD600 of 0.02 – 0.025 (1:200 – 1:400) in ice cold PBS (5-mL tube) to achieve a sorting rate of <8000 events / sec on the iSort™ Automated Cell Sorter (blue solid-state laser (488 nm, 165 mW), optical filters 525 / 50BP for GFP and 488 / 10 SSC, 85 μm ceramic nozzle, fixed sample flow rate of 23 μL / minute). The diluted libraries were placed on ice prior to sorting. Library sorting and post-sort analysis. To set the sorting gates, the uninduced and induced native promoter were used as guides. Because the iSort™ Automated Cell Sorter is a two-way sorter, low (B1) and high (B4) gates were first sorted for the induced library, while the induced middle two gates (B2 and B3) were sorted together. Then the B1 and B4 gates of the uninduced libraries were paired for the third sort. A fourth sort for the uninduced PPTKlibrary was performed pairing the two middle gates (B2 and B3). Cells were sorted based on the GFP fluorescence until all bins reached a minimum of 250,000 cells. The total number of cells sorted per bin, percent target of the library, sorting time, and the event rate per second for both libraries were recorded (data not shown). To ensure proper sorting (purity and cross-over), each sorted bin was sampled on the Attune NxT flow cytometer based on the mCherry fluorescence gate, using the constitutively expressed mCherry located on the plasmid library vector. The plasmid 69 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 library bins, including the unsorted library, were isolated after recovering in 10 mL LB Amp100 medium overnight (where the PBS:LB Amp100 media ratio was at a maximum of ~1:3). The extracted plasmid libraries were subsequently used to generate PCR amplicons for NGS. NGS sample preparation. Plasmid libraries of the recovered sorted and unsorted bins were isolated using Qiagen QIAprep Spin Miniprep Kit according to the manufacturer’s recommendations. The region of interest was amplified with Q5® High-Fidelity 2X Master Mix for ten cycles using assigned primers with the annealing region, 6-nt unique barcode sequence, and partial Illumina adapters overhangs from 1000 ng of the extracted plasmids per 50 µL PCR reaction. Due to the AT-rich content of the PBgaLlibraries, a modified PCR protocol was used to generate the PCR library amplicons for NGS. Six PCR reactions tubes per bin were combined and prepped for gel electrophoresis. The desired PCR amplicon size was excised from 0.7% (w / v) agarose gel in 0.5X TAE buffer with a clean razor blade and gel purified using the Qiagen QIAquick Gel Extraction Kit with few modifications in the manufacturer’s protocol as described in
[0036] . Additional impurities were removed using the Qiagen QIAquick PCR Purification Kit. In some cases, library amplicons were pooled into one tube with equal molar mixture and were ran on a 0.8% agarose gel in 0.5X TBE to ensure purity and size (~100 ng loaded per lane). The final concentrations of the amplicons were normalized to 20 ng / μL in 10 mM Tris-Cl, pH 8.0 with A260 / 280 at 1.8-2.0 prior to outsourcing to AZENTA Life Sciences using 2x250 bp paired-end Amplicon EZ Sequencing service. NGS pre-processing workflow. Galaxy web-based platform (on the public server at usegalaxy.org and usegalaxy.eu)
[0053] was used to preprocess the raw library FASTQ data prior to generating the nucleotide counts per position, mutation frequency per position plots, information footprints, and enrichment analyses. In brief, paired-end reads were first merged using Paired-End read merger (PEAR)
[0054] . Then, the ‘Barcode Splitter’ tool
[0055] was used to obtain merged reads with intact 5’ fixed ends and to separate multiplexed bins, if needed. Prior to mapping, the barcodes and the fixed 5’ and 3’ ends were trimmed using the ‘Trim sequences’ tools
[0055] . Reads were mapped against the reference sequence of native promoter using the BBMap tool
[0056] . ‘Trim sequences’ was used again trim off any overhangs resulted from the mapping. Reads were converted from BAM to FASTA file format with Samtools fastx tool from the SAMtools software package
[0057] and then unique reads were obtained using the ‘Collapse’ tool
[0055] . To calculate the nucleotide counts per position, the aligned reads underwent a series of formatting changes using the ‘Replace’ tool
[0058] and ‘Convert delimiters to TAB’ in Galaxy prior to exporting the reads as a CSV file. This file was imported into Microsoft Excel for to 70 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 calculate the number of nucleotides, mutation frequency, mutual information, and the relative sequence enrichment at each position. Calculating Mutual information for Generating Information Footprints and Enrichment Analyses. Information footprints and the enrichment analyses were calculated as described in [35, 36]. In brief, the nucleotide occurrence of each base at each position were counted and normalized to the total number of reads in each library for the input library and the sorted high and low expression bin libraries. The mutual information between the expression bins and the bases at each nucleotide position were calculated using: where ^^^^^^^^is the base at the ^^^^th position, ^^^^ is the expression activity bin, and ^^^^(^^^^^^^^,^^^^)is the joint frequency distribution and ^^^^(^^^^^^^^) and ^^^^(^^^^) is the marginal frequency distributions. To calculate the enrichment of each base at each position, the ratio of the number of nucleotide occurrence in each of the expression bins were normalized to the counts of the unsorted bins. Preparation of chemically competent E. coli ER1821 and JW0063-1 ∆araC cells from the Keio collection. Prior preparing chemically competent cells, Transformation Buffer (“TB”) was made by dissolving 3 g PIPES (10 mM), 2.2 g CaCl22H2O (15 mM), and 18.6 g in KCl (250 mM) in 1L ultra-pure H2O. The pH was “TB” solution was adjusted to 6.7-6.8 with 5 N KOH. Then 10.9 g MnCl2(55 mM) was dissolved to complete the “TB” solution prior to filtration with a 0.22 µm filter. Cells were grown in Super Optimal Broth (SOB). To make SOB media, 20 g tryptone, 5 g yeast extract, 2 mL 5 M NaCl, 1.25 mL 2 M KCl were dissolved in 990 mL of ultra-pure H2O and then autoclaved at 121°C for 15-30 minutes. Once the SOB was cooled to room temperature, 10 mL of filter-sterilized magnesium solution (1 M MgSO47H2O and 1 M MgCl26H2O) was added to the autoclaved SOB. The E. coli ER1821 strain, used for methylating plasmids for Clostridium acetobutylicum transformations, was streaked on a LB agar plate with no antibiotics. A single colony was cultivated in 10 mL SOB in a 50 mL conical tube (loose caps) for ~6-8 hours at 37°C shaking at 250 rpm. Subsequently, 0.5 mL and 0.1 mL of the culture was diluted into two 200 mL of SOB in a 1 L baffled flask and were grown at 20°C-22°C overnight (~18 hours), shaking at 100 rpm, until the OD600reached between 0.4 and 0.6. The centrifuge was set to 0°C the day prior to harvesting the grown cultures. Once the cultures reached OD600=~0.55, they were placed on ice for ~10 minutes and harvested at 2,500xg for 10 minutes in four 50 mL conical tubes. When the tubes were out of the 0°C centrifuge, all cells were kept on ice. Once the supernatant was removed, the cells were gently resuspended in 71 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 ~20 mL of cold “TB” per tube. Cells were centrifuged at 2,500xg for 10 minutes and the supernatant was discarded. Once the cells were resuspended in 4 mL of chilled “TB” in each tube (for a total of 16 mL), 280 µL DMSO was added and incubated on ice for 10 minutes.50 µL competent cells were aliquoted on chilled 1.5 mL microcentrifuge tubes and frozen at -80°C until use. The same protocol was used to make chemically competent Keio JW0063-1 ∆araC strain but were grown in LB agar plates and SOB supplemented with Kan35. The prepped chemical competent ER1821 and Keio JW0063-1 (∆araC) cells were tested for its transformation efficiency. In vivo methylation of E. coli – C. acetobutylicum shuttle plasmids. Target plasmids and pAN3 were simultaneously transformed into chemically competent E. coli ER1821. First, the target and the pAN3 plasmids were each diluted to 2 ng / mL and equally mixed. Then 2 µL of the target and the pAN3 plasmid mixture was added to 50 µL of chemically competent E. coli ER1821 cells. The DNA-cell mixture was gently mixed. After 30 minutes of incubation on ice, a 30 second heat shock at 42°C was applied and immediately placed on ice for ~2 minutes. Transformed E. coli ER1821 cells were recovered in 450 μL of pre-warmed SOC media and incubated for 1 hour at 37°C while shaking at 250 rpm. Then 100 μL of the transformation were plated on solid LB medium supplemented with both 35 μg / mL kanamycin and 35 μg / mL chloramphenicol to select cells carrying both plasmids. Single colonies were cultured overnight in 10 mL LB medium with 35 μg / mL kanamycin and 35 μg / mL chloramphenicol at 37°C, shaking at 250 rpm. Methylated plasmids, along with the methylating plasmids were extracted from 5 mL of the overnight cultures using the QIAGEN Plasmid Miniprep Kit. The pAN3 plasmid carrying the φ3TI methyltransferase gene from Bacillus subtilis phage φ3tI protects target plasmids from Cac824I restriction activity in C. acetobutylicum
[0082] . Due to the incompatible origin of replication, pAN3 does not propagate in C. acetobutylicum. Preparation of electrocompetent C. acetobutylicum ATCC 824 cells. Six to ten single colonies from freshly streaked C. acetobutylicum plates (~3-4 days old) were grown in 10 mL pre-warmed 2xYTG, pH 6.5 at 37°C until the OD600 reached mid-exponential phase (OD600 = 0.4 – 0.8). Then 10% of the cultures were inoculated into 40 mL 2xYTG, pH 6.5 in 50 mL conical tube and grown at room temperature until the cultures reached an OD600 of 0.5. Cells were harvested at 3,500xg for 10 minutes at 0°C, and washed twice in 10 mL ice-cold electroporation buffer (270 mM sucrose, 5 mM NaH2PO4, pH 6.5). After the final wash, cells were resuspended in 2 mL electroporation buffer, enough for four transformations. For transformations greater than four, additional number of pre-aliquoted 50 mL conical tubes with 40 mL 2xYTG, pH 6.5 were 72 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 inoculated instead of scaling up in one large vessel. This was to prevent repeated stress on the cell and save time when transferring to 50 mL conical tubes to be centrifuged. Transformation of C. acetobutylicum ATCC 824 cells. Electroporation of target plasmids were carried out with 500 µL of electrocompetent cells mixed with 5 µL of methylated plasmid (0.5 – 5 µg) in 4 mm cuvettes (VWR, cat. #76102-580) at 2000 V, 25 µF, and infinite resistance. Immediately after electroporation, 1 mL of pre-warmed 2xYTG, pH 6.5 medium was added directly into the electroporation cuvette, and subsequently transferred to 9 mL of pre- warmed 2xYTG, pH 6.5 medium. Recovery cultures were incubated at 37°C for 4 hours of outgrowth before being pelleted and resuspended in the residual media after removing the supernatant. The entire cell pellet was plated on pre-warmed solid 2xYTG, pH 6.5 supplemented with 15 µg / mL thiamphenicol and incubated at 37°C for 48 hours. To make freezer stocks of the C. acetobutylicum transformants, individual transformants were grown overnight in 10 mL CGM, pH 6.8 supplemented with 15 μg / mL thiamphenicol until mid-exponential phase (OD600=0.5-0.9) and centrifuged. Once the supernatant was removed, the cells were resuspended in 2xYTG, pH 6.5 and 60% glycerol in 2xYTG, pH 6.5 at a 3:1 ratio (for a final 15% glycerol concentration) and subsequently transferred to 2 mL cryovials to be stored at -80°C. Promoter analysis in C. acetobutylicum ATCC 824 via flow cytometry. Frozen freezer stocks of C. acetobutylicum harboring the transformed plasmids were streaked on solid 2xYTG pH 6.5 supplemented with 15 µg / mL thiamphenicol (Tp15). After 24 hours of incubation, six individual colonies per construct were resuspended in 4 mL CGM pH 6.8 supplemented with Tp15 in a 48-well plate. The 48-well plate was sealed with a breathable membrane and grown overnight at 37°C. For all OD600readings, a P1200 multichannel pipette was used to gently mix the cultures prior to transferring to cuvettes to ensure accurate readings. Cultures in the mid- exponential phase (OD600= 0.6 – 0.8) were inoculated into 9 or 12 mL pre-warmed CGM, pH 6.8 supplemented with Tp15 to an OD600=0.05 in either 24-well plate or 15-mL conical tubes, respectively. After approximately 2-3 hours, 2.25 mL cultures were aliquoted into 4 separate wells in a 48-well plate (every other well in a column) with the serological pipette and then induced with or without the 750 µL of 4X concentrated inducer stock in CGM, pH 6.8 with Tp15 with the P1200 multichannel pipette when the OD600 reached 0.2 – 0.3. Cultures that were growing faster than the rest of the cultures were back diluted to OD600=0.2 prior to induction. Samples were gently mixed using the same tips to transfer the inducers to the cultures. The 48- well plates were sealed with a permeable membrane and then incubated at 37°C. After 10 hours post-inoculation, 750 µL of cultures were transferred to 1.5 mL microfuge tubes with caps with 73 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 the P1200 multi-channel pipette (while ensuring proper mixing) to take end-point measurements on for fluorescence, OD600, and the UPLC. From the collected 750 µL of cultures, 8.5 µL (for t=10 hours measurement) or 4.25 µL (for t=26 hours measurement) were transferred to 250 µL of oxygen-free PBS, pH 7.4 with 5 µMTFLime in 1.5 mL microfuge tubes and incubated for at least 15 seconds prior to reading the fluorescence on the flow cytometer measured at the flow rate of 12.5 µL / min (blue solid-state laser (488 nm excitation), an optical filter at 530 / 30 nm for GFP fluorescence, and 488 / 10 nm optical filter for side scatter (SSC)). For OD600, 100 µL were transferred to 10 mm cuvettes with 900 µL of CGM. The leftover collected samples were centrifuged at 3,000xg for 5 min at 0°C and about 500 µL of the supernatant were syringe- filtered (hydrophobic PTFE, 13 mm, 0.22 µm, VWR, cat. #76479-010) into a 96-well plate (Waters) for metabolite detection on the UPLC. Measurements were taken at 10 hours and 26 hours. Regarding the YFAST fluorescence measurements on the Attune NxT flow cytometer, allTFLime stained samples increased in background fluorescence reducing the signal-to-noise ratio at higher flow rates. Therefore, the YFAST fluorescence was measured at the flow rate of 12.5 µL / min to reduce background fluorescence. Gating strategy for promoter analysis in C. acetobutylicum ATCC 824 via flow cytometry. Due to the high autofluorescence of C. acetobutylicum cells, the empty pMTL85141 vector strain stained withTFLime was used to set up the negative and positive gates. First, the whole cell population was gated on the FSC (x-axis) vs. SSC (y-axis) dot plot. Then the negative and positive gates of the whole cell population were set on the FITC-A fluorescence (x-axis) vs. %cells gated (y-axis) histogram, while making sure that the percent of pMTL85141 carrying cells in positive gate was <1% of the gated population. For obtaining the final fluorescence, the median fluorescence intensity in the positive gate was used to plot the data. Preparation ofTFLime stocks.TFLime stocks were prepared according to the manufacturer’s recommendations. First, the lyophilized 250 nmolTFLime stocks were centrifuged 1 minute in the microcentrifuge at the instrument’s max speed (17,000xg). Then 50 µL DMSO was added directly to the tube and vortexed to make a 1000X stock solution. The 5 mM stockTFLime were aliquoted based on the number of samples to run at each time point to prevent multiple freeze / thaw cycles. The aliquoted stocks were stored at -20°C protected from light until use (<6 months in DMSO). To prepare theTFLime labeling solution, the 5 mM stock solution was diluted 1:1000 in filtered oxygen-free PBS, pH 7.4 prior to sample collection for each time point, and protect from light until adding samples for YFAST measurements on the flow cytometer. 74 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Generation of GFP fluorescence, OD600, information footprint, and enrichment plots. GFP distribution histograms from flow cytometry experiments were generated with the Attune NxT acquisition software. GFP fluorescence and OD600 measurements were averaged and plotted with the standard deviations in GraphPad Prism 10.0.2 (232). All replicates represent biological samples grown from individual colonies. The information footprints and enrichment analysis calculations were completed in Microsoft Excel and then transferred to GraphPad Prism 10.0.2 (232) to generate the plots. Calculations of the information footprints and enrichment maps. The information footprints and enrichment analysis calculations were completed in Microsoft Excel and then transferred to GraphPad Prism 10.0.2 (232) to generate the plots. Creation of figures and figure layouts of the generated plots were executed on Adobe Illustrator. Results Initial assessment of Clostridial promoter activities in E. coli: To conduct Sort-Seq experiments in E. coli, the three inducible promoters, PBgaL, ARAi, and PTETO1, were assessed within the E. coli system (FIGs.5A-5E). For each native promoter constructs, plasmids with deleted transcription factor genes were also constructed, ΔbgaR, ΔaraR, and ΔtetR. In addition to the native promoters, different variations of PBgaLand ARAi promoters were also included. For PBgaL, a second version, designated as PBgaLv2, was created by placing the bgaR gene under the control of the constitutive PLacIpromoter while keeping the native PBgaRin the construct (FIGs. 5A and 5E). This modification was carried out due to the fact that BgaR is a putative AraC transcription factor
[0083] . Also, PBgaR and PBgaL are divergent promoters, sharing a structural similarity with the arabinose-inducible promoter, PBAD. Similarly, two additional versions of the ARAi system were devised with the objective of reducing its length (FIGs.5B and 5E). The first version encompasses the native ARAi system (680 bp), featuring the araR gene regulated by its native PAraR promoter and the activating promoter PPTK. In the second version (313 bp), the PPTKpromoter was deleted from ARAi, resulting in the native PArapromoter controlling the expression of gfp. In the third iteration, referred to as the PPTK, the entire PAraR and PAra of the ARAi system was eliminated and replaced by PLacI to regulate araR, reducing the size to 505 bp. Lastly, only the native PTETO1promoter was considered since it has been the most characterized out of the three inducible promoters (FIG.5C). PBgaL and BgaR activity in E. coli: The lactose-inducible promoter was first assessed under various lactose concentrations (0 mM, 0.58 mM, 1.85 mM, 5.84 mM, 18.48 mM, and 58.4 mM) over time. The promoter activity under IPTG (0 mM, 0.1 mM, 0.32 mM, 1 mM, and 3.16 mM), an analog of allolactose, was also evaluated since lactose is converted to allolactose when 75 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 upon entering the cells. Performance of PBgaLv1 and PBgaLv2 was comparable, with a slightly higher increase in GFP activation observed in PBgaLv2. In both versions, however, showed leakiness under both lactose and IPTG induction in E. coli (FIGs.6A-6C). Over time, the basal activity (0 mM) continued to increase, narrowing the dynamic range of PBgaL. However, the absence of GFP fluorescence in the absence of BgaR confirms its role in activating PBgaLin E. coli, as demonstrated by the loss of GFP activity in the ∆bgaR plasmid. Interestingly, while PBgaLactivities under lactose induction reached and maintained its GFP fluorescence saturation starting at 0.58 mM, the effects of IPTG resulted in decreased GFP fluorescence as the IPTG concentration increased. This effect is likely attributed to the effect of IPTG toxicity on cells at higher concentrations or differences in binding kinetics and membrane transport of IPTG and allolactose. ARAi (PPTK and AraR) activity in E. coli: The three versions of the arabinose-inducible promoter, ARAi, were evaluated under various arabinose concentrations (0%, 0.13%, 0.25%, 0.5%, 1%, and 2% w / v) over time (FIGs.7A-7C). To determine the functional version in E. coli, a single replicate for each variant was carried out since the primary interest was in identifying which version exhibited functionality. AraR serves as a repressor in the absence of arabinose, and the addition of arabinose transcription is activated. Among the three versions assessed, only ARAi (v1) and pNK98v3 demonstrated functionality, with pNK98v3 showing a notably higher increase in GFP activation. Therefore, pNK98v3, identified as PPTK, was used in future experiments. Upon observing the activity depicted in the uninduced and the induced of ARAi∆araR (FIGs.7A-7C), it was evident that there was some endogenous activity in NEB5α. Subsequently, PPTK and ∆araR were thoroughly tested (n=2) in both NEB5α and E. coli ∆araC KO strain from the Keio Collection, JW0063-1 (FIGs.8A-8B). Endogenous activity of ∆araR plasmid was observed within NEB5α, whereas GFP fluorescence remained consistently maintained in the JW0063-1 strain across arabinose concentrations. Furthermore, the activation of PPTK in the JW0063-1 strain were notable higher than in NEB5α. PTETO1 and TetR activity in E. coli: Due to its broad use in bacteria and eukaryotic cells, only one version of PTETO1 was constructed and tested with two different analogs of the inducer, tetracycline (Tc) and its antibiotic derivative, anhydrotetracyline (aTc). PTETO1and the ∆tetR control plasmids were first tested with Tc (0 ng / mL, 200 ng / mL, 355.7 ng / mL, 632.5 ng / mL, 1124.7 ng / mL, and 2000 ng / mL) in the NEB5α strain over time (FIGs.9A-9C). There were high levels of endogenous activity in NEB5α strain. GFP fluorescence was lower in the PTETO1 plasmid. This is likely due to autoregulation of the promoter that controls tetR expression, PTetR. 76 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Using Tc as the inducer for PTETO1 did lead to increased variability among replicates and decrease in GFP fluorescence due to cell death. Therefore, the PTETO1promoter was retested under anhydrotetracycline (aTc) (FIGs.10A-10C). Although cell death was observed in the high inducer concentration ranges (FIGs.10A-10C), changing the inducer to aTc resulted in much higher GFP fluorescence and less variability among the replicates at a longer time of incubation at t=6 h. Therefore, aTc was used as the inducer in future experiments. Sort-Seq of the promoter libraries and Individual Promoter Validations: Given the functionality of the native promoters in E. coli, sort-seq experiments were performed. First, promoter libraries were generated by multiple rounds of error-prone PCR (epPCR) and subsequently cloned upstream of a green fluorescent protein (gfp) reporter gene. This reporter gene was sourced from a plasmid vector
[0036] that also contained an mCherry fluorescent protein reporter driven by the constitutive promoter J23101
[0062] . The GFP fluorescence, activated by the promoter libraries, was used to sort the cells using fluorescence-activated cell sorting (FACS) into two to four activity bins per condition. Additionally, mCherry fluorescence was utilized to assess the purity and viability of the sorted bins post each sort. To determine the distribution of GFP and the averaged mutation frequency, a small subset of the final four rounds of epPCR- generated promoter libraries was evaluated under both induced and uninduced conditions. This assessment was carried out using flow cytometry and next-generation sequencing (NGS), respectively. The epPCR round characterized by the broadest spectrum of promoter activities (based on the GFP fluorescence) and an averaged mutation frequency of approximately 5% was selected for sort-seq. Detailed information on the experimental procedures for each promoter is provided described in their respective sections and Extended Experimental Protocol. Lastly, mutual information-based information footprints and enrichment heat maps, generated at each position for every promoter, served as a guide to identify functional regulatory elements contributing to increased or decreased promoter activity under induced and uninduced conditions. A small subset of these sequences was selected for validation first in E. coli and then in C. acetobutylicum. PBgaLAssessment PBgaLLibrary Generation and Characterization. Since the native PBgaLpromoters are uncharacterized, the entire PBgaR and PBgaL regions were mutagenized. To investigate if the constitutive PBgaRplays a role in regulating PBgaLbeyond its role in expressing bgaR, PBgaLv2was included, in addition to PBgaLv1. 77 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 The PBgaR and PBgaL divergent promoters underwent one set of seven rounds of epPCR and cloned into two vectors for PBgaLv1 and PBgaLv2 to drive the expression of the gfp reporter gene. The last four epPCR rounds were transformed into DH5α E. coli cells to evaluate promoter activities for both libraries at a small-scale prior to scaling up for sort-seq. Each library was grown overnight from previously prepared frozen stocks which were subsequently back diluted to an OD600of 0.1. After about 30 minutes of growth, the libraries were induced with 0 mM, 0.58 mM, 1.04 mM, 1.85 mM, 3.29 mM, and 5.84 mM lactose and measured for GFP fluorescence activity after five to six hours post-induction on the flow cytometer. The leftover portion of the overnight library cultures were collected to be prepped for assessing the library diversity by NGS. Due to the relatively lower GC content of the PBgaL promoter region (~13% GC compared to 20% GC and 26% GC in PPTK of ARAi and PTETO1, respectively), there were some challenges in amplifying epPCR rounds 5 to 7. Although epPCR round 4 also presented difficulties in amplification, modifications to the extension temperature in the PCR protocol significantly improved the yield of the desired PCR product size. Fortunately, epPCR round 4 generated the widest GFP distribution in the small-scale library experiments, and hence was chosen for the large-scale sort-seq experiment. The final library size for PBgaLepPCR round 4 was approximately 430,105 with a vector background of 0.04% for PBgaLv1 library and while PBgaLv2 had a vector background of 0.03% of the library size of 392,187, based on the transformation efficiencies. The average mutation rate per position was estimated to be 4% for both versions of the PBgaL libraries. PBgaL Library Sorting: Both PBgaLv1 and PBgaLv2 promoter libraries were sorted into four expression bins (B1 to B4, ranging from low to high expression) and two bins (low, B1 and high, B4) based on the GFP fluorescence when induced with 5.84 mM and 0 mM lactose, respectively, after 5-6 hours post-induction. All sorted bins for each condition and the unsorted (B0) bins were sequenced by NGS. To calculate the importance and relative enrichment of mutations at each nucleotide position on bin location, raw reads were pre-processed to obtain unique reads across all bins [35, 36]. Sort-seq identifies putative RNA polymerase and BgaR binding sites of the PBgaLpromoter: To elucidate the regulatory sequences of PBgaL, information footprints and enrichment analyses were employed based on the sequenced bins for both versions of the library. In the mutual information analysis of PBgaLv1 and PBgaLv2 libraries under induced conditions, it was evident that all sites within the PBgaRpromoter exhibited mutual information values that were 78 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 close to 0 bits. This suggests that the PBgaR promoter does not contribute to gfp expression. The mutual information of PBgaRin the uninduced PBgaLv2library was also low, however, there appeared to be some noise in this region compared to the induced information footprints. Whether this is due to the biological effects of the high AT-content in this region or noise from experimental workflows, it appears that the PBgaRpromoter does not contribute to the maintenance of the low basal activity. Furthermore, there were no major differences in the data between the two versions of the library. The PBgaLv2footprints displaying the highest mutual information were clustered around positions P168-P173 and P191-196, with a 17-bp separation, hinting at the possibility of these regions serving as the σ70RNA polymerase binding site. The full activation of PBgaL favored the native sequence, as indicated by the rare sequences in B3 and B4 and the enriched sequences in B1 and B2.While the native sequence identified through sort-seq as the -35 box (P168-P173, TTCTAA) and -10 box (P191-P196, TAACAT) did not precisely match the typical σ70binding consensus sequence, TTGACA-17 bp-TATAAT, the enrichment analysis indicated a preference for T168, T / A169, G / T170, A171, C / A172, and A / T173for the -35 box and T191, A192, T / A / C193, A / T194, A195, and T196sequences for the -10 boxes. The identification of the binding sequence for BgaR through the information footprints was not straightforward especially since both the induced and uninduced information footprints had a similar profile. Deletion of BgaR resulted in no activation in the presence of lactose indicating that BgaR should not bind in the absence of lactose (FIGs.6A-6C). On the other hand, it made sense that both the induced and uninduced information footprints would share similar profiles since the native PBgaL is extremely leaky in E. coli. While the exact binding locations of BgaR were not definitive, there were indications of binding activity within the region spanning P134-P174, which also overlaps with the putative -35 box (P168-P173), as evidenced by the information footprints and corresponding enrichment analysis. Further analysis of P146-P178 revealed the presence of a palindromic sequence. DNA structural analysis on mFold
[0090] also uncovered an additional secondary structure upstream of this site, specifically in P128-P132, which is complementary to P139-P144. The mutual information between P128-P132 and P139-P144 was considered significant. PBgaLValidations: The primary objective was to identify sequences that could improve the performance of PBgaL, specifically on increasing the dynamic range while minimizing basal activity. Using information footprints and sequence enrichment analyses, 16 single mutants and 3 combinations of mutants were chose for the initial validation in E. coli prior to testing in 79 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Clostridium. Among the chosen mutants included one upstream at A(144)C and six within the putative BgaR binding site (palindrome) at A(153)G, A(158)G, T(159)A, T(162)A, A(164)G, and T(166)C. Within the identified in the putative -35 box, chosen sequences included C(170)G, C(170)T, T(171)A, and A(172)C, as well as combinations like TTCTAA(168-173)TTGCCA and TTCTAA(168-173)TTGACA at the putative -35 box. A mutant in the putative spacer, T(179)C and CATT(176-179)CCGC, was also selected along with A(193)T in the -10 box which was shown to increase GFP expression. No mutant combinations were constructed for the -10 box as the native sequence was critical for GFP expression. Additionally, mutations downstream of the putative -10 box, including T(197)A and A(203)G , were explored. T(197)C was also included, however, it was not included in the validation set due to unusual sanger results of the final cloned plasmid. Chosen validation sequences were cloned into the same plasmid vector that was used in the sort-seq experiments to be tested first in E. coli as shown in FIG.11. Overall, none of the chosen mutations reduced the leaky expression of PBgaL, however, all variants increased the GFP output compared to the native promoter. P159, P162, P164, P170, P193, and CATT(176- 179)CCGC did retain about the same level of dynamic range as the native sequence, which are all located within the palindromic sequence. Mutating the putative -35 box at P168-P173 from TTCTAA to TTGCCA and TTGACA increased the GFP output by ~4.8-fold compared to the native sequence, however, the dynamic range between induced and uninduced was abolished. Validations in C. acetobutylicum: Despite the increased leakiness in E. coli, 11 of 19 promoter mutants that demonstrated the highest gfp expression were tested in C. acetobutylicum. Previous studies have indicated that while the native PBgaLin E. coli exhibited severe basal expression, this same level of leakiness did not transfer to C. acetobutylicum
[0085] . To validate the promoter mutants in the intended host, the identified mutations were incorporated into the native PBgaLpromoter to drive the expression of the anaerobic fluorescent reporter yfast gene in the E. coli – Clostridium shuttle vector, pMTL85141
[0091] . Clostridium promoter validations were conducted in 48-well plates. For each promoter construct, two out of the six colonies grown were inoculated with an OD600range of 0.4-0.8 in CGM, pH 6.8 and allowed to grow for approximately 2-3 hours. When the OD600reached ~0.2, the cells were induced with various lactose concentrations (0 mM, 2 mM, 6.3 mM, and 20 mM lactose). Cultures were collected for fluorescence, OD600, and analyte detection measurements at approximately 10 hours and 26 hours after inoculation (or ~8 hours and 24 hours after induction, respectively). 80 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 To evaluate the mutant promoter performance, the 10-hour time point was focused on (FIGs.12A-12C). The promoter activities were tighter within the exponential phase (t=~10 h), whereas the fluorescence (FIG.12A) and OD600 (FIG.12B) were not consistent within replicates (n=2) during the stationary phase (t=26 h) (FIGs.16A-16B). While decreased cell growth was observe as the lactose concentration increased at t=10 h (FIG.12B), which was also previously reported in C. saccharoperbutylacetonicum
[0065] , it was difficult to accurately measure the OD600 at the t=26-hour time point due to the presence of cell clumps primarily within the 0 mM lactose samples (FIG.16B). Unlike in E. coli, the selected promoter mutants showed responsiveness to lactose with increasing concentrations in C. acetobutylicum at t=10 hours (FIG.12A). Among the best performers, which exhibited 2-fold or greater increase in dynamic range and sensitivity, were observed in promoter mutants T(159)A, T(162)A, T(166)C, T(171)C, A(193)T, and TTCTAA(168-173)TTGCCA where the fold-change fluorescence over the native promoter was at the highest at 2 mM lactose. The induction fold over the native promoter for these promoters at higher lactose concentrations was lower. This could be attributed to cell toxicity of lactose as shown by the decreased cell density as the lactose concentration increased at t=10 h (FIG.12B). Regardless, promoter variant TTCTAA(168-173)TTGCCA exhibited 2.9-fold higher in induction over the native sequence at 2 mM lactose. However, promoter variant TTCTAA(168-173)TTGCCA had a slight increase (1.3-fold) in basal expression over the native sequence at 0 mM lactose where 50.4% of cells measured were found within the fluorescent positive cell gate (FIG.12C, see FIG.16C for t = 26 h). Other promoter variants T(171)C and T(166)C had about 14.7% and 10.6% of the total cell population, respectively, were found within the fluorescent positive cell gate in the absence of lactose (0 mM), whereas the native PBgaL promoter had 1.5% of the total cell population in the fluorescent positive cell gate. Although none of the promoter constructs reduced basal levels below the native promoter, promoter variants T(159)A, T(162)A, and A(193)T exhibited comparable basal activity and with <1%, 2.93%, and <1% of fluorescent cells in the positive cell gate, respectively. No significant changes were observed for A(164)G, C(170)T, and CATT(176-179)CCGC. However, A(153)G and TTCTAA(168-173)TTGAC exhibited a reduction in dynamic range. Notably, TTCTAA(168-173)TTGACA had the highest increase in basal activity among the 11 promoter mutants tested, which is evident in the percentage of positively fluorescent cells under 0 mM lactose (FIG.12C). PPTK Assessment 81 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 Library Generation and Characterization: In a previous study, the binding motif in the PPTKfor the transcription factor AraR was determined using Electrophoretic Mobility Shift Assay (EMSA) employing a 180 bp DNA fragment of PPTK and purified recombinant AraR protein
[0092] . Accordingly, the 206 bp region which included the same region that was investigated in the EMSA analysis for mutagenesis was targeted. The PPTKpromoter underwent six rounds of epPCR where the last four rounds of epPCR were cloned into the vector with the gfp reporter gene in E. coli DH5α. The last four epPCR rounds of PPTK library were evaluated before being scaled up for sort-seq analysis. Given that the PPTK promoter exhibited endogenous activity in response to E. coli AraC transcription factor in NEB5α, the library was evaluated in the JW0063-1 strain. To ensure efficient transformation of the libraries, the library was first electroporated into DH5α where the transformed plasmid libraries were isolated after an overnight culture were transformed into JW0063-1. The vector lacking the library insert was utilized to determine the library background when electroporated into DH5α cells. Transformation efficiencies and library background data were recorded (data now shown). Under 0% arabinose conditions, the GFP distribution in all four epPCR rounds of the PPTK library exhibited a wider range compared to the native promoter (data not shown). However, when arabinose was present (0.02%, 0.063%, 0.2%, 0.63%, and 2%), the libraries displayed a bimodal histogram profile. This profile suggested that the far-left peak contained promoters with little to low GFP expression, while the right peak comprised high gfp-expressing promoters. As the number of epPCR rounds increased, both peaks in each round shifted towards the left. Additionally, the abundance of high gfp-expressing promoters decreased while the peak containing low gfp-expressing promoters increased. This observation indicates that as the mutation rate within the promoter library increased, more inactive promoters were present in the higher epPCR rounds. In the end, PPTKepPCR round 4 was selected for the large-scale sort-seq experiment. This was based on several factors, including the broadest GFP distribution observed in the presence of arabinose. Notably, this epPCR round also featured a third, smaller peak positioned between the two largest peaks. The number of unique reads in each transformed epPCR library, which were used to calculate the averaged mutation frequency of each round, was significantly reduced due to two series of transformations (data not shown). To maintain the diversity and prevent skewing of individual variants in the large-scale library pool, a greater amount of DNA was transformed into DH5α and cultivated at a lower temperature (30°C), distinct from the conditions used for the PBgaL and PTETO1 libraries. Extracted plasmid libraries from DH5α were subsequently 82 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 transformed into JW0063-1. The large-scale PPTKlibrary transformations in DH5α yielded in approximately 2,254,560 library size, with a 1.49% library vector background, determined by the transformation efficiencies. The final library size transformed into JW0063-1 strain was 7,560,000 with an averaged mutation frequency of 4.5%. Library Sorting: PPTK promoter libraries were sorted based on the GFP fluorescence after 5 to 6 hours of induction with 0.2% and 0% arabinose. The libraries under both conditions were sorted into four expression bins (B1 to B4, ranging from low to high expression). All 8 of the sorted bins and the unsorted (B0) bins were deep sequenced by NGS. To obtain unique reads across all bins for mutual information and the relative enrichment analysis, the raw NGS data files were pre-processed on the Galaxy Platform. Sort-seq identifies putative RNA polymerase and AraR binding sites of the PPTK promoter: The identification of RNA polymerase and AraR binding sites in PPTK were clearly identified in the uninduced information footprints and enrichment analyses. Considering prior data and other relevant studies [87, 92], AraR is required to repress transcription. The mutual information peaks observed at P128-P133 and P151-P156 strongly suggest the presence of the - 35 and -10 boxes, while the mutual information within the P98-P115 region points to the location of the AraR binding site. This observation is further supported in the enrichment analyses which emphasize that by maintaining the native sequences at the -35 and -10 boxes, where the preferred sequences for the identified -35 and -10 boxes are TTGACA and TANNAT, respectively, and introducing mutations in P98-P115 combined would destroy the AraR binding regions resulting in increased basal activity as shown in B3 and B4 Furthermore, P98-P115 is completely unimportant in the induced mutual information data. Additionally, apart from the established AraR binding site, another region was identified between the -35 and -10 boxes, spanning P134-P144, that also shows similar enrichment in mutant sequences in B3 and B4 of the uninduced data (data now shown). Notably, P98-P115 (established AraR binding site) and P134-P144 were not significant in the induced mutual information data and did not demonstrate enrichment or depletion (data not shown). The AraR binding site may extend to this region, regulating the transcription of PPTK. Beyond the two major crucial sites (RNA polymerase and AraR binding sites), other noteworthy sites were observed under uninduced conditions. While no additional important sites were identified to contribute to promoter activity upstream of the AraR binding site under uninduced conditions, there were several crucial sites located downstream of the RNA polymerase binding site, particularly at P207, P210, and P226-P229. Specific mutations in B3 83 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 and B4 at P207, P210, and P226-P227 were enriched, while sequence enrichment at P228-P229 within B1 indicate decreased basal activity. In the induced mutual information data, five positions at P119-P123 located in between the -35 box and the AraR binding site, that consisted entirely of native ‘A’ nucleotides, and another five positions at P143-P147 were recognized as important, while not observed in the uninduced mutual information data. P143-P147 is positioned between the -35 and -10 boxes and features the native sequence AAATA, where mutations were enriched in B2 to B4. With the exception of P144, these two regions exhibit low mutual information in the uninduced data compared to the induced data. It is plausible that introducing mutations in the P119-P123 and P143-P147 regions may be altering the DNA structure, potentially resulting in a destroyed AraR binding site structure, or exposing the -35 box for RNA polymerase to bind and activate transcription. This may occur despite AraR being bound, as AraR operates through a simple repression mechanism by binding to block RNA polymerase from binding, without other direct effect on activation. PPTK Validations in E. coli: A total of 20 PPTK promoter mutants were selected to be validated in JW0063-1 strain under 0% and 0.2% (v / w) arabinose. After 5 hours of induction, the GFP fluorescence was measured by flow cytometry (FIG.13). A common issue that comes across engineering inducible repressors is that modifications to the repressor binding site to decrease basal expression often leads to reduced inducibility and vice versa
[0035] . In addition to finding mutations that increase GFP expression in the presence of arabinose, it was also investigated whether the basal expression can be reduced without the expense of reducing the GFP expression. First, six single mutations located within the AraR binding site were tested: A(103)G, T(104)G, G(107)A, A(109)C, A(112)G, T(114)G. Although the last 5 promoter mutants were found enriched in B4 of the uninduced condition, the information footprint had a distinct pattern. The AraR binding site is a palindromic sequence with a 7-bp stem and a 4-bp loop in between. Sort-seq identified different levels of importance within the binding site where P104, P107, and P109 had the highest mutual information while the P100, P102, P111, and P113 had the lowest mutual information within the binding site. A(103)G was chosen for validation was because it was the only site that contained a depleted mutation in B4 of the uninduced data. Results are shown in FIG.13. Interestingly, A(103)G did not alter the promoter activity. While all mutations, except A(103)G, in the hairpin were expected to increase basal activity, the basal expression exhibited by G(107)A was found significant. Promoter mutants T(114)G and A(112)G exhibited an increase in their inducibility with moderate increase in basal expression. The native sequences at positions P104 and P109 are complementary. Promoters with the single mutations at these two sites, T(104)G and A(109)C, performed similarly to one another where 84 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 the level of basal and induced expressions were equal. Furthermore, combining the mutations in the AraR binding site resulted in constitutive expression as expected, regardless of combining mutants on the left half (ATTTATAC(99-106)GTGTCGTA) or the right half (ATACGTA(103- 109)CGACAAC) of the palindromic sequence. Interestingly, combining the mutations on the left and right half of the binding site (ATTTATACGTACAAAT(99- 114)GTGTCGTAAACCAGAG) did not elevate the GFP fluorescence at the same level as ∆araR. Mutations located outside of the established AraR binding site were also validated. Mutations in the poly A region that were found to be important in the 0.2% information footprints and enriched were included: A(122)T and AAAAA(119-123)GTGTT. While an increase in GFP expression in the presence of arabinose was expected, the maximum level GFP expression significantly decreased especially in the combination mutation promoter, AAAAA(119-123)GTGTT. Another odd surprising mutant was A(133)T, which is part of the - 35 box. At P133, this site was considered important in both the 0% and 0.2% information footprint. However, there were conflicting enrichment data at this site. While any mutation at P133 was completely depleted in B3 and B4 of the 0% enrichment data, the 0.2% enrichment data predicted nucleotide A to be enriched within B2 and B4 but not in B3. Therefore, A(133)T was included in the validations, which resulted in lower inducibility. Two mutations in the -10 box performed as predicted. The inducibility of promoter mutants A(155)G and A(156)G were nearly abolished in the presence of arabinose. For mutations within the spacer region, A(141)C and TgCgtACaA(136-144)GgAgtCAaC, resulted in increased basal activity as expected, however, A(141)C increased in gfp expression while lowered in the combination mutant promoter. Lastly, mutations located downstream of the RNA polymerase binding site that were identified as important in the 0% condition were included. While T(165)C and TA(228-229)GG significantly decreased in their inducibility, A(207)G and CA(226-227)GG increased in expression. PTETO1Assessment PTETO1 Library Generation and Characterization: The PTETO1 promoter library included the known tetO1 sequence within an 82 bp region. Four of the seven epPCR rounds were cloned into the vector to drive gfp expression to be assessed under 0 ng / mL, 200 ng / mL, 355.7 ng / mL, 632.5 ng / mL, 1124.7 ng / mL, and 2000 ng / mL tetracycline (Tc) and anhydrotetracycline (aTc). Transformation efficiencies and averaged mutation frequencies of the small-scale libraries for PTETO1were recorded. The PTETO1libraries exhibited a narrow GFP fluorescence distribution at 85 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 all inducer concentrations which could be due to the autoregulation of the PTetR promoter by excess TetR. Furthermore, inducing the libraries with Tc further narrowed the GFP distribution and increased in lower gfp-expressing cells. For sort-seq, epPCR round 5 was transformed into DH5α, resulting in a library size of 251,797 with a vector background of 1.77% of the library. The average mutation frequency for the final library is 5.1%. Library Sorting and Sort-seq identifies putative RNA polymerase and TetR binding sites of the PTETO1promoter: Similarly, the PTETO1library was sorted into four and two gfp expressing bins under induced 1124.7 ng / mL aTc and uninduced 0 ng / mL aTc, respectively. The mechanism of action for TetR on PTETO1is similar to how AraR regulates PPTK. Unlike AraR and PPTK, the mutual information of the palindromic TetR binding site in the PTETO1promoter was found to be significant in both the induced and uninduced conditions. Additionally, the information footprint pattern of tetO1 exhibited distinct features between 0.2% and 0% at P51-69, with the exception of P67. Interestingly, the mutual information at P67 was lower in the presence of aTc than in the 0% data. Furthermore, mutual information at the -35 and -10 boxes were visible in the presence of aTc but was diminished in its absence, where the -35 box contained far more conserved sequences compared to the -10 box. Like PPTK, a poly-A region was identified in PTETO1positioned upstream of the -35 box at positions P37-41, although the poly-A region spanned from P36-43. In addition, sort-seq unveiled previously unreported important regions under both induced and uninduced conditions. Located upstream of the RNA polymerase binding sites, isolated footprints with mutual information were observed at P21, P23, and P30 in the induced data, with P30 also detected in the uninduced condition. Downstream of the -10 box in the induced data, an unknown 18 bp region with mutual information spanning P82-P100 emerged. Within this region, nine native sequences at P82-P91, except for P86, were found to be conserved or essential for achieving high promoter activation. The other nine nucleotides, including P86, consisted of enriched sequences associated with increased the GFP expression in B4. Previous efforts aimed at reducing basal expression included the insertion of a second tetO operator within the unidentified region. The aTc-inducible promoter version employed in this study has undergone multiple iterations to optimize its use in gram-positive bacteria. This involved placing the tetO operator between the -35 and -10 boxes of a strong B. subtilis PxylA constitutive promoter and introducing a poly A sequence upstream of the -35 box (which is required in B. subtilis
[0093] )
[0094] . Additionally, the expression of tetR was made constitutive
[0095] . As a result, this version of the 86 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 promoter was reported to have strong inducibility albeit exhibited some level of basal expression. In another study
[0096] , further refinements were made by inserting a second tetO operator downstream leading to no detectable basal expression but a reduction in inducibility. In this study, the insertion of the second tetO corresponds to nucleotide positions P85 and P86
[0089] . PTETO1Validations in E. coli: A total of 19 PTETO1promoter variants were constructed and tested with 0 ng / mL and 600 ng / mL aTc in NEB5α E. coli after 6 hours of induction before testing in C. acetobutylicum (FIG.14). For mutations expected to increase GFP expression, the following single mutations were chosen: A(21)G, T(86)C, T(92)G, G(94)A, C(95)A, and A(99)T. In addition to these, mutants T(86)C, T(92)G, G(94)A, C(95)A, and A(99)T were combined to create another promoter variant: tgatcgtagcgttaaca(86-102)CgatcgGaTAgttTaca. Eleven more constructs that were anticipated to increase GFP expression but were also expected to raise basal activity were also included in the validation process: A(30)T, T(53)G, T(55)G, A(56)G, T(57)G, C(58)G, T(61)A, G(62)C, A(63)C, T(64)C, A(65)G. These latter mutation group was chosen based on their high mutual information scores within the tetO region, and they were also the most enriched mutations in B4 of the induced data at their respective positions. There were also combined in the PTETO1construct, ctctatcattgatagag(52- 68)cGcGGGGatACCCGgag, which had a calculated ∆G of -6.64 kcal / mol, compared to the native -3.47 kcal / mol, on mfold
[0090] . Of the single ‘up’ mutations, all except A(99)T and T(92)G, increased in GFP expression without raising basal activity (FIG.14). Among this group, the best performing promoter variants were G(94)A and C(95)A, which increased their induced expression by 1.6-fold and 1.5- fold, respectively, compared to the native PTETO1. T(86)C also had a slight increase by 1.2-fold over the native sequence while maintaining the same level of basal expression. However, combining the mutants T(86)C, T(92)G, G(94)A, C(95)A, and A(99)T led to a significant reduction in expression. Among the mutations predicted to increase both GFP expression and basal expression, four performed as predicted: A(30)T, C(58)G, G(62)C, and the combination mutant, COMBO(52-68). A(63)C resulted in a slight increase in GFP expression without significantly raising the basal expression, while the promoter constructs with mutations located on the left side of tetO, namely T(53)G, T(55)G, A(56)G, T(57)G, had minor effects on both the induced and uninduced expression. While no increase in basal expression was observed for T(61)A and T(64)C, the induced expression decreased, whereas A(65)G increased only in the basal activity. 87 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 In summary, three commonly employed Clostridium inducible promoters: PBgaL (lactose), PPTK(arabinose), and PTETO1(anhydrotetracycline) were engineered to increase expression levels and responsiveness to induction. Accordingly, the constructs and promoters of the present technology are useful in constructs, compositions, and methods for the expression of gene products in prokaryotic organisms. EQUIVALENTS The present technology is not to be limited in terms of the particular embodiments described in this application, which are intended as single illustrations of individual aspects of the present technology. Many modifications and variations of this present technology can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the present technology, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the present technology. It is to be understood that this present technology is not limited to particular methods, reagents, compounds compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting. In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group. As will be understood by one skilled in the art, for any and all purposes, particularly in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like, include the number recited and refer to ranges which can be subsequently broken down into subranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 cells refers to groups having 1, 2, or 3 cells. Similarly, a group having 1-5 cells refers to groups having 1, 2, 3, 4, or 5 cells, and so forth. 88 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 All patents, patent applications, provisional applications, and publications referred to or cited herein are incorporated by reference in their entirety, including all figures and tables, to the extent they are not inconsistent with the explicit teachings of this specification. REFERENCES 1. Guzman, L.-M., et al., Tight regulation, modulation, and high-level expression by vectors containing the arabinose PBAD promoter. Journal of bacteriology, 1995.177(14): p.4121-4130. 2. Chao, Y.P., C.J. Chiang, and W.B. Hung, Stringent Regulation and High‐Level Expression of Heterologous Genes in Escherichiacoli Using T7 System Controllable by the araBAD Promoter. Biotechnology progress, 2002.18(2): p.394-400. 3. Kim, S.K., et al., Tunable control of an Escherichia coli expression system for the overproduction of membrane proteins by titrated expression of a mutant lac repressor. ACS Synthetic Biology, 2017.6(9): p.1766-1773. 4. Lim, H.-K., et al., Production characteristics of interferon-α using an L-arabinose promoter system in a high-cell-density culture. Applied microbiology and biotechnology, 2000. 53: p.201-208. 5. Giacalone, M.J., et al., Toxic protein expression in Escherichia coli using a rhamnose- based tightly regulated and tunable promoter system. Biotechniques, 2006.40(3): p.355-364. 6. Hjelm, A., et al., Tailoring Escherichia coli for the l-Rhamnose PBAD promoter-based production of membrane and secretory proteins. ACS Synthetic Biology, 2017.6(6): p.985-994. 7. Romano, E., et al., Engineering AraC to make it responsive to light instead of arabinose. Nature Chemical Biology, 2021.17(7): p.817-827. 8. Tang, S.-Y., H. Fazelinia, and P.C. Cirino, AraC regulatory protein mutants with altered effector specificity. Journal of the American Chemical Society, 2008.130(15): p.5267-5271. 9. Tang, S.Y. and P.C. Cirino, Design and application of a mevalonate‐responsive regulatory protein. Angewandte Chemie International Edition, 2011.50(5): p.1084-1086. 10. Schleif, R., A Career's Work, the l-Arabinose Operon: How It Functions and How We Learned It. EcoSal Plus, 2022.10(1): p. eESP-0012-2021. 11. Egan, S.M. and R.F. Schleif, A regulatory cascade in the induction of rhaBAD. Journal of molecular biology, 1993.234(1): p.87-98. 12. Schleif, R., AraC protein, regulation of the l-arabinose operon in Escherichia coli, and the light switch mechanism of AraC action. FEMS microbiology reviews, 2010.34(5): p.779- 796. 13. Goulding, C.W. and L.J. Perry, Protein production in Escherichia coli for structural studies by X-ray crystallography. Journal of structural biology, 2003.142(1): p.133-143. 14. Francis, D.M. and R. Page, Strategies to optimize protein expression in E. coli. Current protocols in protein science, 2010.61(1): p.5.24.1-5.24.29. 15. Balzer, S., et al., A comparative analysis of the properties of regulated promoter systems commonly used for recombinant gene expression in Escherichia coli. Microbial cell factories, 2013.12: p.1-14. 16. Meyer, A.J., et al., Escherichia coli “Marionette” strains with 12 highly optimized small- molecule sensors. Nature chemical biology, 2019.15(2): p.196-204. 17. Kelly, C.n.L., et al., Synthetic chemical inducers and genetic decoupling enable orthogonal control of the rhaBAD promoter. ACS synthetic biology, 2016.5(10): p.1136-1145. 18. Ding, N., et al., Programmable cross-ribosome-binding sites to fine-tune the dynamic range of transcription factor-based biosensor. Nucleic Acids Research, 2020.48(18): p.10602- 10613. 19. Shilling, P.J., et al., Signal amplification of araC pBAD using a standardized translation initiation region. Synthetic Biology, 2022.7(1): p. ysac009. 89 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 20. Wegerer, A., T. Sun, and J. Altenbuchner, Optimization of an E. coli L-rhamnose- inducible expression vector: test of various genetic module combinations. BMC biotechnology, 2008.8(1): p.1-12. 21. Rogers, J.K., et al., Synthetic biosensors for precise gene control and real-time monitoring of metabolites. Nucleic acids research, 2015.43(15): p.7648-7660. 22. Khlebnikov, A., et al., Regulatable arabinose-inducible gene expression system with consistent control in all cells of a culture. Journal of Bacteriology, 2000.182(24): p.7029-7034. 23. Khlebnikov, A., T. Skaug, and J.D. Keasling, Modulation of gene expression from the arabinose-inducible araBAD promoter. Journal of Industrial Microbiology and Biotechnology, 2002.29(1): p.34-37. 24. Lee, S.K., et al., Directed evolution of AraC for improved compatibility of arabinose-and lactose-inducible promoters. Applied and environmental microbiology, 2007.73(18): p.5711- 5715. 25. Lagator, M., et al., Epistatic interactions in the arabinose cis-regulatory element. Molecular Biology and Evolution, 2016.33(3): p.761-769. 26. Wickstrum, J.R., et al., Transcription activation by the DNA-binding domain of the AraC family protein RhaS in the absence of its effector-binding domain. Journal of bacteriology, 2007. 189(14): p.4984-4993. 27. Chen, Y., et al., Tuning the dynamic range of bacterial promoters regulated by ligand- inducible transcription factors. Nature Communications, 2018.9(1): p.64. 28. Lutz, R. and H. Bujard, Independent and tight regulation of transcriptional units in Escherichia coli via the LacR / O, the TetR / O and AraC / I1-I2 regulatory elements. Nucleic acids research, 1997.25(6): p.1203-1210. 29. Lutz, R., et al., Dissecting the functional program of Escherichia coli promoters: the combined mode of action of Lac repressor and AraC activator. Nucleic Acids Research, 2001. 29(18): p.3873-3881. 30. Tamsir, A., J.J. Tabor, and C.A. Voigt, Robust multicellular computing using genetically encoded NOR gates and chemical ‘wires’. Nature, 2011.469(7329): p.212-215. 31. Niland, P., R. Hühne, and B. Müller-Hill, How AraC interacts specifically with its target DNAs. Journal of molecular biology, 1996.264(4): p.667-674. 32. Egan, S.M. and R.F. Schleif, DNA-dependent renaturation of an insoluble DNA binding protein: identification of the RhaS binding site at rhaBAD. Journal of molecular biology, 1994. 243(5): p.821-829. 33. Mejía-Almonte, C., et al., Redefining fundamental concepts of transcription initiation in bacteria. Nature Reviews Genetics, 2020.21(11): p.699-714. 34. Kinney, J.B. and D.M. McCandlish, Massively parallel assays and quantitative sequence–function relationships. Annual review of genomics and human genetics, 2019.20: p. 99-127. 35. Rohlhill, J., N.R. Sandoval, and E.T. Papoutsakis, Sort-seq approach to engineering a formaldehyde-inducible promoter for dynamically regulated Escherichia coli growth on methanol. ACS synthetic biology, 2017.6(8): p.1584-1595. 36. Kim, N.M., et al., Elucidation of Sequence–Function Relationships for an Improved Biobutanol In Vivo Biosensor in E. coli. Frontiers in bioengineering and biotechnology, 2022. 10: p.821152. 37. Dunn, T.M., et al., An operator at-280 base pairs that is required for repression of araBAD operon promoter: addition of DNA helical turns between the operator and promoter cyclically hinders repression. Proceedings of the National Academy of Sciences, 1984.81(16): p.5017-5020. 38. Bhende, P.M. and S.M. Egan, Amino acid-DNA contacts by RhaS: an AraC family transcription activator. Journal of bacteriology, 1999.181(17): p.5185-5192. 39. Carra, J.H. and R.F. Schleif, Variation of half‐site organization and DNA looping by AraC protein. The EMBO journal, 1993.12(1): p.35-44. 90 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 40. Fernandez-Lopez, R., et al., Structural basis of direct and inverted DNA sequence repeat recognition by helix–turn–helix transcription factors. Nucleic Acids Research, 2022.50(20): p. 11938-11947. 41. Brodsky, S., et al., Intrinsically disordered regions direct transcription factor in vivo binding specificity. Molecular cell, 2020.79(3): p.459-471. e4. 42. Martin, K., L. Huo, and R.F. Schleif, The DNA loop model for ara repression: AraC protein occupies the proposed loop sites in vivo and repression-negative mutations lie in these same sites. Proceedings of the National Academy of Sciences, 1986.83(11): p.3654-3658. 43. Reeder, T. and R. Schleif, AraC protein can activate transcription from only one position and when pointed in only one direction. Journal of molecular biology, 1993.231(2): p.205-218. 44. Seabold, R.R. and R.F. Schleif, Apo-AraC actively seeks to loop. Journal of molecular biology, 1998.278(3): p.529-538. 45. Zhang, X., T. Reeder, and R. Schleif, Transcription Activation Parameters atara pBAD. Journal of molecular biology, 1996.258(1): p.14-24. 46. Shahein, A., et al., Systematic analysis of low-affinity transcription factor binding site clusters in vitro and in vivo establishes their functional relevance. Nature Communications, 2022.13(1): p.5273. 47. Dhiman, A. and R. Schleif, Recognition of overlapping nucleotides by AraC and the sigma subunit of RNA polymerase. Journal of Bacteriology, 2000.182(18): p.5076-5081. 48. Decker, K.B. and D.M. Hinton, Transcription regulation at the core: similarities among bacterial, archaeal, and eukaryotic RNA polymerases. Annual review of microbiology, 2013.67: p.113-139. 49. Holcroft, C.C. and S.M. Egan, Roles of cyclic AMP receptor protein and the carboxyl- terminal domain of the α subunit in transcription activation of the Escherichia coli rhaBAD operon. Journal of Bacteriology, 2000.182(12): p.3529-3535. 50. Wickstrum, J.R. and S.M. Egan, Amino acid contacts between sigma 70 domain 4 and the transcription activators RhaS and RhaR. Journal of bacteriology, 2004.186(18): p.6277- 6285. 51. Via, P., et al., Transcriptional regulation of the Escherichia coli rhaT gene. Microbiology, 1996.142(7): p.1833-1840. 52. Atrazhev, A.M. and J.F. Elliott, Simplified desalting of ligation reactions immediately prior to electroporation into E. coli. Biotechniques, 1996.21(6): p.1024. 53. The Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2022 update. Nucleic Acids Research, 2022.50(W1): p. W345-W351. 54. Zhang, J., et al., PEAR: a fast and accurate Illumina Paired-End reAd mergeR. Bioinformatics, 2014.30(5): p.614-620. 55. Gordon, A. and G.J. Hannon, Fastx-toolkit. FASTQ / A short-reads preprocessing tools (unpublished) http: / / hannonlab. cshl. edu / fastx_toolkit, 2010.5. 56. Bushnell, B., J. Rood, and E. Singer, BBMerge–accurate paired shotgun read merging via overlap. PloS one, 2017.12(10): p. e0185056. 57. Li, H., et al., The sequence alignment / map format and SAMtools. bioinformatics, 2009. 25(16): p.2078-2079. 58. Gruening, B.A., Galaxy Wrapper.2014, Github: https: / / github.com / bgruening / galaxytools. 59. Pédelacq, J.-D., et al., Engineering and characterization of a superfolder green fluorescent protein. Nature biotechnology, 2006.24(1): p.79-88. 60. Jeske, M. and J. Altenbuchner, The Escherichia coli rhamnose promoter rhaP BAD is in Pseudomonas putida KT2440 independent of Crp–cAMP activation. Applied microbiology and biotechnology, 2010.85: p.1923-1933. 61. Dietrich, J.A., et al., Transcription factor-based screens and synthetic selections for microbial small-molecule biosynthesis. ACS synthetic biology, 2013.2(1): p.47-58. 91 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 62. Storch, M., et al., BASIC: a new biopart assembly standard for idempotent cloning provides accurate, single-tier DNA assembly for synthetic biology. ACS synthetic biology, 2015. 4(7): p.781-787. 63. Tracy, B.P., et al., Clostridia: the importance of their exceptional substrate and metabolite diversity for biofuel and biorefinery applications. Current opinion in biotechnology, 2012.23(3): p.364-381. 64. Charubin, K., et al., Engineering Clostridium organisms as microbial cell-factories: challenges & opportunities. Metabolic engineering, 2018.50: p.173-191. 65. Ma, Y., et al., Development of an Efficient Recombinant Protein Expression System in Clostridium saccharoperbutylacetonicum Based on the Bacteriophage T7 System. ACS Synthetic Biology, 2023. 66. Heap, J.T., et al., Spores of Clostridium engineered for clinical efficacy and safety cause regression and cure of tumors in vivo. Oncotarget, 2014.5(7): p.1761. 67. Moon, H.G., et al., One hundred years of clostridial butanol fermentation. FEMS microbiology letters, 2016.363(3): p. fnw001. 68. Gyulev, I.S., et al., Part by part: synthetic biology parts used in solventogenic Clostridia. ACS Synthetic Biology, 2018.7(2): p.311-327. 69. Joseph, R.C., et al., Metabolic Engineering and the Synthetic Biology Toolbox for Clostridium. Metabolic Engineering: Concepts and Applications, 2021.13: p.611-651. 70. Yang, G., et al., Rapid generation of universal synthetic promoters for controlled gene expression in both gas-fermenting and saccharolytic Clostridium species. ACS Synthetic Biology, 2017.6(9): p.1672-1678. 71. Mordaka, P.M. and J.T. Heap, Stringency of synthetic promoter sequences in Clostridium revealed and circumvented by tuning promoter library mutation rates. ACS synthetic biology, 2018.7(2): p.672-681. 72. Kim, N.M., R.W. Sinnott, and N.R. Sandoval, Transcription factor-based biosensors and inducible systems in non-model bacteria: current progress and future directions. Current opinion in biotechnology, 2020.64: p.39-46. 73. Riley, L.A. and A.M. Guss, Approaches to genetic tool development for rapid domestication of non-model microorganisms. Biotechnology for Biofuels, 2021.14: p.1-17. 74. Fackler, N., et al., Annual Review of Chemical and Biomolecular Engineering. Annu. Rev. Chem. Biomol. Eng., 2021. 75. Zhang, Y., et al., Heterologous Gene Regulation in Clostridia: Rationally Designed Gene Regulation for Industrial and Medical Applications. ACS Synthetic Biology, 2022.11(11): p. 3817-3828. 76. Mukherjee, A., et al., Characterization of Flavin-Based Fluorescent Proteins: An Emerging Class of Fluorescent Reporters. PLOS ONE, 2013.8(5): p. e64753. 77. Flaiz, M., et al., Establishment of green-and red-fluorescent reporter proteins based on the fluorescence-activating and absorption-shifting tag for use in acetogenic and solventogenic anaerobes. ACS Synthetic Biology, 2022.11(2): p.953-967. 78. Streett, H.E., K.M. Kalis, and E.T. Papoutsakis, A strongly fluorescing anaerobic reporter and protein-tagging system for Clostridium organisms based on the fluorescence- activating and absorption-shifting tag protein (FAST). Applied and environmental microbiology, 2019.85(14): p. e00622-19. 79. Hocq, R., et al., A fluorescent reporter system for anaerobic thermophiles. Frontiers in Bioengineering and Biotechnology, 2023.11. 80. Charubin, K., H. Streett, and E.T. Papoutsakis, Development of strong anaerobic fluorescent reporters for Clostridium acetobutylicum and Clostridium ljungdahlii using HaloTag and SNAP-tag proteins. Applied and Environmental Microbiology, 2020.86(20): p. e01271-20. 81. Pyne, M.E., et al., Technical guide for genetic advancement of underdeveloped and intractable Clostridium. Biotechnology advances, 2014.32(3): p.623-641. 92 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 82. Mermelstein, L. and E. Papoutsakis, In vivo methylation in Escherichia coli by the Bacillus subtilis phage phi 3T I methyltransferase to protect plasmids from restriction upon transformation of Clostridium acetobutylicum ATCC 824. Applied and environmental microbiology, 1993.59(4): p.1077-1081. 83. Hartman, A.H., H. Liu, and S.B. Melville, Construction and characterization of a lactose-inducible promoter system for controlled gene expression in Clostridium perfringens. Applied and environmental microbiology, 2011.77(2): p.471-478. 84. Banerjee, A., et al., Lactose-inducible system for metabolic engineering of Clostridium ljungdahlii. Applied and environmental microbiology, 2014.80(8): p.2410-2416. 85. Al-Hinai, M.A., A.G. Fast, and E.T. Papoutsakis, Novel system for efficient isolation of Clostridium double-crossover allelic exchange mutants enabling markerless chromosomal gene deletions and DNA integration. Applied and environmental microbiology, 2012.78(22): p.8112- 8121. 86. Wang, Y., et al., Bacterial genome editing with CRISPR-Cas9: deletion, integration, single nucleotide modification, and desirable “clean” mutant selection in Clostridium beijerinckii as an example. ACS synthetic biology, 2016.5(7): p.721-732. 87. Zhang, J., et al., A novel arabinose-inducible genetic operation system developed for Clostridium cellulolyticum. Biotechnology for biofuels, 2015.8(1): p.1-13. 88. Fagan, R.P. and N.F. Fairweather, Clostridium difficile has two parallel and essential Sec secretion systems. Journal of Biological Chemistry, 2011.286(31): p.27483-27493. 89. Bertram, R., B. Neumann, and C.F. Schuster, Status quo of tet regulation in bacteria. Microbial Biotechnology, 2022.15(4): p.1101-1119. 90. Zuker, M., Mfold web server for nucleic acid folding and hybridization prediction. Nucleic acids research, 2003.31(13): p.3406-3415. 91. Heap, J.T., et al., A modular system for Clostridium shuttle plasmids. Journal of microbiological methods, 2009.78(1): p.79-85. 92. Zhang, L., et al., Ribulokinase and transcriptional regulation of arabinose metabolism in Clostridium acetobutylicum. Journal of bacteriology, 2012.194(5): p.1055-1064. 93. Moran, C.P., et al., Nucleotide sequences that signal the initiation of transcription and translation in Bacillus subtilis. Molecular and General Genetics MGG, 1982.186: p.339-346. 94. Geissendörfer, M. and W. Hillen, Regulated expression of heterologous genes in Bacillus subtilis using the Tn 10 encoded tet regulatory elements. Applied microbiology and biotechnology, 1990.33: p.657-663. 95. Corrigan, R.M. and T.J. Foster, An improved tetracycline-inducible expression vector for Staphylococcus aureus. Plasmid, 2009.61(2): p.126-129. 96. Helle, L., et al., Vectors for improved Tet repressor-dependent gradual gene induction or silencing in Staphylococcus aureus. Microbiology, 2011.157(12): p.3314-3323. 97. Gomes, A.L., et al., Genome and sequence determinants governing the expression of horizontally acquired DNA in bacteria. The ISME Journal, 2020.14(9): p.2347-2357. 98. Kobayashi, M., K. Nagata, and A. Ishihama, Promoter selectivity of Escherichia coli RNA polymerase: effect of base substitutions in the promoter− 35 region on promoter strength. Nucleic acids research, 1990.18(24): p.7367-7372. 99. Urtecho, G., et al., Systematic dissection of sequence elements controlling σ70 promoters using a genomically encoded multiplexed reporter assay in Escherichia coli. Biochemistry, 2018.58(11): p.1539-1551. 100. Woolston, B.M., et al., Rediverting carbon flux in Clostridium ljungdahlii using CRISPR interference (CRISPRi). Metabolic engineering, 2018.48: p.243-253. 101. Baba, T., et al., Construction of Escherichia coli K‐12 in‐frame, single‐gene knockout mutants: the Keio collection. Molecular systems biology, 2006.2(1): p.2006.0008. 102. Cook, T.B., et al., Genetic tools for reliable gene expression and recombineering in Pseudomonas putida. Journal of Industrial Microbiology and Biotechnology, 2018.45(7): p. 517-527. 93 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 103. Dietrich, J.A. and J.D. Keasling, Transcription factor-based biosensor.2013, Lawrence Berkeley National Lab.(LBNL), Berkeley, CA (United States). 104. Hocq, R., et al., σ54 (σL) plays a central role in carbon metabolism in the industrially relevant Clostridium beijerinckii. Scientific Reports, 2019.9(1): p.7228. 105. Yu, H., et al., Engineering transcription factor BmoR for screening butanol overproducers. Metabolic engineering, 2019.56: p.28-38. 106. Yu, H., et al., Establishment of BmoR-based biosensor to screen isobutanol overproducer. Microbial cell factories, 2019.18(1): p.1-11. 107. Wu, T., et al., Engineering transcription factor BmoR mutants for constructing multifunctional alcohol biosensors. bioRxiv, 2021. 108. Gruber, T.M. and L.L. Huang, Mutant arabinose promoter for inducible gene expression. 2011, Google Patents. 109. Keasling, J.D. and S.K. Lee, Inducible expression vectors and methods of use thereof. 2012, Google Patents. 110. McClain, S., M. Valasek, and A. Gruber, Coordinated coexpression of thrombin.2019, Google Patents. 111. Jia, X. and J.A. Claypool, Controlled lysis of bacteria.2011, Google Patents. 112. Sabbadini, R.A., N. Berkley, and M.W. Surber, Rhamnose-inducible expression constructs and methods.2011, Google Patents. 113. Sabbadini, R.A., N. Berkley, and M.W. Surber, Eubacterial minicells and their use as vectors for nucleic acid delivery and expression.2007, Google Patents. 114. Brass, J., et al., Rhamnose Promoter Expression System.2008, Google Patents. 115. Caron, K. and S.C. Trowell, Highly sensitive and selective biosensor for a disaccharide based on an AraC-like transcriptional regulator transduced with bioluminescence resonance energy transfer. Analytical chemistry, 2018.90(21): p.12986-12993. 116. Karim, A.S., et al., Modular cell-free expression plasmids to accelerate biological design in cells. Synthetic Biology, 2020.5(1): p. ysaa019. 117. Minton, N.P. and Y. Zhang, Conditional vectors and uses thereof.2017, Google Patents. 118. T.C.N.I.P.A. (CNIPA), Editor.2014, Qingdao Institute of Bioenergy and Bioprocess Technology of CAS: China. 119. Fackler, N., et al., Transcriptional control of Clostridium autoethanogenum using CRISPRi. Synthetic Biology, 2021.6(1): p. ysab008. 120. Gossen, M. and H. Bujard, Tight control of gene expression in eucaryotic cells by tetracycline-responsive promoters.1995, Google Patents. 121. Bujard, H. and R. Loew, Tetracycline inducible transcription control sequence.2015, Google Patents. 122. Bujard, H. and M. Gossen, Tetracycline-inducible transcriptional inhibitor fusion proteins.2001, Google Patents. 123. Bujard, H., et al., Tetracycline regulated transcriptional modulators with altered DNA binding specificities.1996, Google Patents. 124. Dobrovolsky, V.N. and R.H. Heflich, On the use of the T‐REx™ tetracycline‐inducible gene expression system in vivo. Biotechnology and bioengineering, 2007.98(3): p.719-723. 125. Riggs, P.D., Overview of protein expression vectors for E. coli. Current Protocols Essential Laboratory Techniques, 2018.17(1): p. e23. 126. Ferreira, R.d.G., A.R. Azzoni, and S. Freitas, Techno-economic analysis of the industrial production of a low-cost enzyme using E. coli: the case of recombinant β-glucosidase. Biotechnology for biofuels, 2018.11: p.1-13. 127. Liu, X., et al., Expression of recombinant protein using Corynebacterium glutamicum: progress, challenges and applications. Critical reviews in biotechnology, 2016.36(4): p.652- 664. 94 4871-1271-3210.1 Atty. Dkt. No.: 136669-0122 128. Terpe, K., Overview of bacterial expression systems for heterologous protein production: from molecular and biochemical fundamentals to commercial systems. Applied microbiology and biotechnology, 2006.72: p.211-222. 129. Donahue Jr, R.A. and R.L. Bebee, BL21-SI™ competent cells for protein expression in E. coli. Protein Expr. Purif, 1999.7: p.289. 130. Elvin, C.M., et al., Modified bacteriophage lambda promoter vectors for overproduction of proteins in Escherichia coli. Gene, 1990.87(1): p.123-126. 131. Skerra, A., Use of the tetracycline promoter for the tightly regulated production of a murine antibody fragment in Escherichia coli. Gene, 1994.151(1-2): p.131-135. 132. Jajesniak, P. and T.S. Wong, From genetic circuits to industrial-scale biomanufacturing: bacterial promoters as a cornerstone of biotechnology. AIMS Bioengineering, 2015.2(3): p. 277-296. 133. Puetz, J. and F.M. Wurm, Recombinant proteins for industrial versus pharmaceutical purposes: a review of process and pricing. Processes, 2019.7(8): p.476. 134. Schofield, D.M., et al., Promoter engineering to optimize recombinant periplasmic Fab′ fragment production in Escherichia coli. Biotechnology Progress, 2016.32(4): p.840-847. 135. Yuan, S., et al., New expression system to increase the yield of phloroglucinol. Biotechnology & Biotechnological Equipment, 2020.34(1): p.405-412. 136. Walsh, G., Biopharmaceutical benchmarks 2014. Nature biotechnology, 2014.32(10): p. 992-1000. 137. Walsh, G., Biopharmaceutical benchmarks 2018. Nature biotechnology, 2018.36(12): p. 1136-1145. 138. Walsh, G., Biopharmaceutical benchmarks 2010. Nature biotechnology, 2010.28(9): p. 917-924. 139. Walsh, G. and E. Walsh, Biopharmaceutical benchmarks 2022. Nature Biotechnology, 2022.40(12): p.1722-1760. 140. d’Oelsnitz, S., et al., Using fungible biosensors to evolve improved alkaloid biosyntheses. Nature chemical biology, 2022.18(9): p.981-989. 141. Snoek, T., et al., Evolution-guided engineering of small-molecule biosensors. Nucleic acids research, 2020.48(1): p. e3-e3. 142. Della Corte, D., et al., Engineering and application of a biosensor with focused ligand specificity. Nature communications, 2020.11(1): p.4851. 143. Dai, X., et al., Inducible CRISPR genome-editing tool: classifications and future trends. Critical reviews in biotechnology, 2018.38(4): p.573-586. 144. Liang, R. and J. Liu, Scarless and sequential gene modification in Pseudomonas using PCR product flanked by short homology regions. BMC microbiology, 2010.10(1): p.1-9. 145. Mowday, A.M., et al., Advancing clostridia to clinical trial: past lessons and recent progress. Cancers, 2016.8(7): p.63. 146. Minton, N.P., Clostridia in cancer therapy. Nature Reviews Microbiology, 2003.1(3): p. 237-242. 147. Minton, N.P. and J.T. Heap, Treatment for cancer.2019, Google Patents. 95 4871-1271-3210.1
Claims
Atty. Dkt. No.: 136669-0122 CLAIMS 1. An expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pBAD promoter sequence of any one of SEQ ID NOs: 1-14 that is operably linked to a heterologous gene and (b) a constitutive Pc promoter sequence that is operably linked to a gene sequence encoding an AraC protein, wherein the AraC protein is configured to induce an RNA polymerase to bind to the mutant pBAD promoter sequence in the presence of arabinose. 2 An expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pRha promoter sequence of any one of SEQ ID NOs: 16-25 that is operably linked to a heterologous gene and (b) a constitutive PRhaRS promoter sequence that is operably linked to a gene sequence encoding a RhaS protein, wherein the RhaS protein is configured to induce an RNA polymerase to bind to the mutant pRha promoter sequence in the presence of rhamnose. 3 An expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pBgal promoter sequence of any one of SEQ ID NOs: 27-46 that is operably linked to a heterologous gene and (b) a BgaR promoter sequence that is operably linked to a gene sequence encoding a BgaR protein, wherein the BgaR protein is configured to induce an RNA polymerase to bind to the mutant pBgal promoter sequence in the presence of lactose. 4 An expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pPTK promoter sequence of any one of SEQ ID NOs: 48-67 that is operably linked to a heterologous gene and (b) an AraR promoter sequence that is operably linked to a gene sequence encoding an AraR protein, wherein the AraR protein is configured to induce an RNA polymerase to bind to the mutant pPTK promoter sequence in the presence of arabinose. 5 An expression system comprising a nucleic acid sequence, wherein the nucleic acid sequence includes (a) a mutant pTETO1 promoter sequence of any one of SEQ ID NOs: 69-87 that is operably linked to a heterologous gene and (b) a TetR promoter sequence that is operably linked to a gene sequence encoding a TetR protein, wherein the TetR protein is configured to induce an RNA polymerase to bind to the mutant pTETO1 promoter sequence in the presence of tetracycline or anhydrotetracyline (aTc). 6 The expression system of any one of claims 1-5, wherein the expression system is integrated on a chromosome of a host cell. 7 The expression system of any one of claims 1-5, wherein the expression system is an expression vector. 8 The expression system of claim 7, wherein the expression vector is a plasmid, a cosmid, a bacterial artificial chromosome (BAC) or a yeast artificial chromosomes (YAC). 96 4871-1271-3210.1Atty. Dkt. No.: 136669-0122 9. The expression system of any one of claims 1-8, wherein the heterologous gene encodes a protein, an enzyme, a structural polypeptide, a toxin, a fusion protein, an antibody agent, a drug, a cytokine, an enzyme inhibitor, a growth factor, a signaling protein, a bioluminescent protein, a fluorescent protein, a chemiluminescent protein, a catalytic RNA, or an inhibitory RNA.
10. The expression system of any one of claims 1-9, wherein the expression system expresses the heterologous gene to a greater degree than a control expression system that comprises a wild type pBAD, pRha, pBgal, pPTK, or pTETO1 promoter operably linked to the heterologous gene.
11. The expression system of claim 10, wherein the expression system expresses the heterologous gene at about a 5-fold to about a 10-fold greater degree than the control expression system.
12. A prokaryotic host cell comprising the expression system of any one of claims 1-11.
13. The prokaryotic host cell of claim 12, wherein the host cell is a gram-positive bacteria or a gram negative bacteria.
14. The prokaryotic host cell of claim 13, wherein the host cell is E. coli or Clostridium.
15. A nucleic acid construct comprising a mutant inducible promoter sequence operably linked to a heterologous gene, wherein the mutant promoter sequence is selected from the group consisting of SEQ ID NOs: 1-14, 16-25, 27-46, 48-67, and 69-87.
16. The nucleic acid construct of claim 15, wherein the nucleic acid construct is integrated on a chromosome of a host cell.
17. The nucleic acid construct of claim 15, wherein the nucleic acid construct is integrated in an expression vector.
18. The nucleic acid construct of claim 17, wherein the expression vector is a plasmid, a cosmid, a bacterial artificial chromosome (BAC) or a yeast artificial chromosomes (YAC).
19. A prokaryotic host cell comprising the nucleic acid construct of any one of claims 15-18.
20. The prokaryotic host cell of claim 19, wherein the host cell is a gram-positive bacteria or a gram negative bacteria.
21. The prokaryotic host cell of claim 20, wherein the host cell is E. coli or Clostridium.
22. A method for overexpressing a heterologous polypeptide or nucleic acid in a prokaryotic cell comprising contacting the prokaryotic host cell of any one of claims 12-14 or 19-21 with an effective amount of an inducer molecule, wherein the heterologous polypeptide or nucleic acid is encoded by the heterologous gene of the expression system of any one of claims 1-11 or the nucleic acid construct of any one of claims 15-18.
23. The method of claim 22, wherein the inducer molecule is selected from the group consisting of arabinose, rhamnose, lactose, tetracycline and anhydrotetracyline (aTc).
24. The method of claim 22 or 23, further comprising lysing the prokaryotic cell to isolate the heterologous polypeptide. 97 4871-1271-3210.1Atty. Dkt. No.: 136669-0122 25. A kit comprising a nucleic acid encoding the expression system of any one of claims 1-11 or the nucleic acid construct of any one of claims 15-18 and instructions for use thereof to express the heterologous gene in a host cell.
26. A kit comprising a host cell of any one of claims 12-14 or 19-21 and instructions for use thereof to overexpress a heterologous polypeptide or nucleotide.
27. The kit of claims 25 or 26, further comprising one or more inducer molecules. 98 4871-1271-3210.1
Citation Information
Patent Citations
Recombinant vector for deleting specific regions of chromosome and method for deleting specific chromosomal regions of chromosome in the microorganism using the same
US20090305421A1
Autotrophic hydrogen bacteria and uses thereof
US20170298395A1
Mutant arabinose promoter for inducible gene expression
US7998702B2