Genetically engineered bacteria as a chassis host for high yield terpenoid production
Optimized E. coli strains with polycistronic operons and modified RBS enhance terpenoid production by addressing toxicity and yield limitations, achieving high-capacity terpenoid synthesis with cost-effective inducer use.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BOARD OF TRUSTEES OPERATING MICHIGAN STATE UNIV
- Filing Date
- 2025-04-08
- Publication Date
- 2026-05-07
Smart Images

Figure US2025023675_07052026_PF_FP_ABST
Abstract
Description
Genetically Engineered Bacteria as a Chassis Host for High Yield Terpenoid ProductionCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This patent application claims priority to U. S. Provisional Patent Application Serial No. 63 / 631,246, filed on April 8, 2024, the contents of which is specifically incorporated herein by reference in its entirety.FEDERAL FUNDING
[0002] This invention was made with government support under grant numbers DESCOO 18409 awarded by the Department of Energy Great Lakes Bioenergy Research Center Center Cooperative Agreement. The government has certain rights in the invention.INCORPORATION OF SEQUENCE LISTING
[0003] The instant application contains a sequence listing, which has been submitted in XML file format by electronic submission and is hereby incorporated by reference in its entirety. The XML file, created on April 4, 2025, is named P14839WOOO.xml and is 243,792 bytes in size.BACKGROUND
[0004] Global demand for the development of environmentally favorable microbial processes to produce chemicals, fuels, and pharmaceuticals is increasing dramatically owing to environmental concerns, depletion of fossil fuels, and increasing petrochemical prices (Xu et al. 2014, Yao et al. 2014). Achieving sustainable bioproduction (a central facet of Synthetic Biology) requires the iterative design, construction and testing cycle of microbial cell factories capable of converting renewable feedstock into bioproducts at maximum yield and productivity (Lim et al. 2013). Terpenoids, represent the largest and most complex class of specialized metabolites with over 86,000 reported natural compounds found in plants and many living organisms (Dictionary of Natural Products 31.1. Terpenoids are useful in applications spanning medicine, nutrition, agriculture, flavors, fragrances, colorants and cosmetics (Wang et al. 2018). All terpenoids are biosynthesized from the two C5 isoprenoid building blocks isopentenyl diphosphate (IPP) and its isomer dimethylallyl diphosphate (DMAPP). Thecondensation of these precursors generates larger isoprenoid molecules including geranyl diphosphate (GPP, CIO), famesyl diphosphate (FPP, Cl 5), and geranylgeranyl diphosphate (GGPP, C20) which are first cyclized by terpene synthases to form the diverse terpene scaffolds and subsequently functionalized by cytochromes P450 into terpenoids (Chatzivasileiou et al. 2019).SUMMARY
[0005] Described herein are compositions and methods for generating a chassis for high-capacity production of different classes of high-value terpenoids. The chassis is a microorganism host that is modified with one or more expression systems that comprise a first operon comprising a first promoter operably linked to a nucleic acid segment encoding a modified ribosomal binding site, a Hmg-CoA reductase (HmgR) polypeptide, a 3 -hydroxy-3 -methylglutaryl CoA synthase (HmgS) polypeptide, and a P-ketothiolase (PhaA) polypeptide; and a second operon comprising a second promoter operably linked to a nucleic acid segment encoding a modified ribosomal binding site, a mevalonate kinase (MvKl) polypeptide, a phospho-mevalonate kinase (MvK2) polypeptide, and a diphosphomevalonate decarboxylase (MvD) polypeptide.
[0006] Described herein are methods and expression systems that can provide IPP- and DMAPP Cs terpene precursors and products of downstream terpenoid pathways including viridiflorol and lycopene. For example expression systems are described herein that include at least one expression cassette having a heterologous promoter operably linked to a nucleic acid segment encoding an enzyme with at least 90% sequence identity to SEQ ID NO: 18, 20, 21, 23, 25, 26, 28, 30, 31, 33, 34, 36, 38, 39, 41, 42, 44, 45, 47, 48, 50, 51, 53, 54, 56, 58, or a combination thereof. Also described herein are host cells that include such expression systems.
[0007] Methods of synthesizing a terpenoids are also described herein. For example, such methods of synthesizing a terpenoids can include incubating a host cell that has such expression system and isolating the terpenoids.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Fig. 1 IPP / DMAPP / Isoprene biosynthesis pathways: 2-C-methyl-D-erythritol 4-phosphate (MEP) and Mevalonate pathway (MYA).
[0009] Fig. 2A-B RBSHUS sequence. RBS for gene expression in polycistronic operon.A) RBS sequences used in this study. B) Effect of RBS-variants on EGFP expression. EGFP is in the second order in a polycistronic operon.
[0010] Fig. 3A-C Modification of the RBS of the tetR-repressor. A) Toxicity examination of doxycycline (DOC) on E. coli. B) The effect of different concentrations of the inducer AHT on EGFP expression after modification of the RBS of / e / ’-repressor. C) The effect of different concentrations of the inducer DOC on EGFP expression in response to modifications of the RBS of Zc / A’-repressor.
[0011] Fig. 4A-C The upper mevalonate pathway (UMVP). A) The enzymes catalyze the production of mevalonate: two molecules of Acetyl-CoA are condensed together by thiolase (PhaA, 2.3.1.9) to afford acetoacetyl-CoA. A third molecule of Acetyl-CoA is further condensed via hydroxymethylglutaryl-CoA synthase (HmgS, 2.3.3.10). The last step is the reduction of hydroxymethylglutaryl-CoA by a reductase (HmgR, 1.1.1.34) consuming two molecules of NAD(P)H to produce mevalonate. B) The UMVP genes and their order in a polycistronic operon: thiolase encoding gene (PhaA) from A. eutropha H16, hydroxymethylglutaryl-CoA synthase encoding gene (HmgS) from E. faecalis. and the reductase encoding gene (HmgR) from S. aureus. C) The production of mevalonate using E. coli DH (MCI 061) and BL21 (C41) strains in different media.
[0012] Fig. 5 The original LMVP (native pathway) and the archaea pathway (modified pathway) for IPP / DMAPP formation.
[0013] Fig. 6 Strategy used to construct the LMVP system. Five different constructs architectures were designed by varying the source of MvK2 and MvD. Sc; is S. cerevisiae, Sa;. S'. aureus.
[0014] Fig. 7A-7B The operons used for lycopene production. A) The 12 steps catalyzed by an 11 -enzyme cascade for lycopene production from acetyl-CoA. B) Separation of the lycopene pathway into three different operons: Upper- MVA (UMVP) under 7b / C-promoter; lower-MVA (LMVP) under TetP or AtzcP-promoter; lycopene operon (Lyco) under TetP-promoter. dT, rrndTl terminator.
[0015] Fig. 8A-8C Lycopene production using different modules. A) Upper-MVP (UMVP) under 7b / C-promoter, lower-MVP (LMVP) under LrzcT’-promoter, lycopene operon (Lyco) under 7c / / J-promoter. The inducer is anhydrotetracycline (AHT, 200 ng / ml) and IPTG(1 mM). B) Upper-MVP (UMVP) under 7 / C-promoter, lower-MVP (LMVP) and lycopene operon (Lyco) under TeZ -promoter. The inducer is anhydrotetracycline (AHT, 200 ng / ml). C) The color of the lycopene producing strains with different modules. S, comb 1-4 refers to the differently assembled LMVP (see Fig. 4).
[0016] Fig. 9A-9B Effect of using different LMVP-modules on OD and residual mevalonate in culture. A) Residual mevalonate during lycopene fermentation using different LMVP-modules. B) Growth of the cells carrying different LMVP-modules undertow different promoters TetP or LacP.
[0017] Fig. 10A-10D Effect of modification of lycopene-operons on biomass and lycopene production. A) the lycopene production without inducing agent (-AHT / or IPTG). B) Lycopene production after adding (+AHT / IPTG). Side panel, cell culture of lycopene producing strains with / or without the inducer. C) growth of the cells carrying different lycopene-operons. D) Scaled shake flask production of lycopene.
[0018] Fig. 11A-11B Effect of the type of inducers and media on biomass and lycopene production. A) Lycopene production with different concentrations of AHT or DOC in different media (M9 or LB). B) Effect of the inducer types and concentration on the growth of the cells in different media.
[0019] Fig. 12A-12B Full lycopene pathway in one plasmid. A) Construction of one-plasmid for the full pathway expression with 3 -different origin of replication. B) Lycopene production in different E. coli strains (DH and C41 strains).
[0020] Fig. 13A-13B DAP-auxotrophic generating E. coli MCI 061. A) pACTR-dapA plasmid map that express dapA genes on a level sufficient to facilitate auxotrophic growth. B) Growth behavior of dap A- E. coli strain (Hus 10).
[0021] Fig. 14A-14B Viridiflorol and viridiflorin production in E. coli. A) the two VidS-operons constructed in this work. B) Viridiflorol (VidOl) and viridiflorin (Vidin) yield using the two VidS-operons.
[0022] Fig. 15A-15C VidS truncation s. A) VidS-sequence analysis identifying cysteine residues. B) Prediction of the disulfide bonds in VidS. C) Native and truncated versions of VidS.
[0023] Fig. 16 Viridiflorol and viridiflorin yield from different versions constructed by engineering the viridiflorol synthase.
[0024] Fig. 17 Map of plasmid pESG-MVA with a nucleic acid sequence of SEQ ID NO:5 for mevalonate production in E. coll.
[0025] Fig. 18 Map of plasmid pPCTR-Combl-Tl. UMVP with a nucleic acid sequence of SEQ ID NO:6 with Lac-promotor used for LMVP and UMVP in E. coli.
[0026] Fig. 19 Map of plasmid pPCTR-Comb3-Tl. UMVP with a nucleic acid sequence of SEQ ID NOV.
[0027] Fig. 20 Map of plasmid pACTR-Combl-Tl. UMVP with a nucleic acid sequence of SEQ ID NO: 8 with Z / P-promoter used for LMVP and UMVP in E. coli.
[0028] Fig. 21 Map of plasmid pACTR-Comb3-Tl. UMVP with a nucleic acid sequence of SEQ ID NO:9 with Z ZP-promoter used for LMVP and UMVP in E. coli.
[0029] Fig. 22 Map of plasmid pESG-crtEIB-IspA with a nucleic acid sequence of SEQ ID NO: 10 with Lycopene-operon.
[0030] Fig. 23 Map of plasmid pEF-BR42-LMVP-UMVP-Lyco with a nucleic acid sequence of SEQ ID NO: 11 with the full lycopene pathway.
[0031] Fig.24 Map of plasmid pACTR-LMVP-UMVP-Lyco with a nucleic acid sequence of SEQ ID NO: 12 with the full lycopene pathway.
[0032] Fig. 25 Map of plasmid pACTR(dapA)-LMVP-UMVP with a nucleic acid sequence of SEQ ID NO: 13 with dapA for auxotrophic growth of E. coli.
[0033] Fig. 26 Map of plasmid pACTR(dapA)-LMVP-UMVP-Lyco with a nucleic acid sequence of SEQ ID NO: 14 with dapA for auxotrophic growth of E. coli.
[0034] Fig. 27 Map of plasmid pESG-VidS-IspA with a nucleic acid sequence of SEQ ID NO: 15 for viridiflorol production in E. coli.
[0035] Fig. 28 Map of plasmid pESG-VidS (del.91)-IspA with a nucleic acid sequence of SEQ ID NO: 16 for viridiflorol production in E. coli.
[0036] Fig. 29 Map of plasmid pESG-VidS(del.91).fus. IspA with a nucleic acid sequence of SEQ ID NO: 17 for viridiflorol production in E. coli.
[0037] Figs.30A-30C (A) Schematic of the LMVP and lycopene operons under the TolCl-promoter. B) Plasmid map of pEE-BR42-Tl. Comb3-Tl. MVA which was constructed with the native genes codon usage and C. necator Hl 6 codon usage for the LMVP. C) Images of lycopene production by C. necator Hl 6 when mevalonate was supplied.
[0038] Figs.31 A-31C Images of plated (A) and centrifuged culture (B) of C. necator YL16 expressing the full lycopene pathway split into 3 operons (mevalonate pathway “UMVA”, intermediate pathway “LMVA”, and lycopene pathway “Lyc”). C) Image of centrifuged culture of C. necator Hl 6 expressing the split lycopene operon when the culture is grown in three different media (Luria Broth “LB”, minimal media “M9”, and minimal media +yeast Extract “M9Y”). Cell growth for 24 hr was good in both M9 and M9Y.DETAILED DESCRIPTION
[0039] Described herein are compositions and methods for generating genetically modified host cells that demonstrate versatility in high-capacity production of different classes of high-value terpenoids. The terpenoids produced can range from simple Cs (Hemiterpene), Cio (Monoterpene), Cis (Sesquiterpene), C20 (Diterpene), C25 (Sesterterpene), C30 (Triterpene) and C40 (Tetraterpene, carotenoid natural colorants) to larger molecules including polyisoprene (rubber). For example, the terpenoids produced can include custom carotenoids, lycopene, viridiflorol or a combination thereof.
[0040] All naturally produced terpenoids are derived from isopentenyl diphosphate (IPP) and its isomer dimethylallyl diphosphate (DMAPP). In nature, two different routes of IPP biosynthesis occur: the mevalonate pathway (MV A) and the deoxyxylulose 5-phosphate (DXP / MEP) pathway. In this work, high production of the terpenoid-precursors in E. coli is achieved through MVA pathway engineering using candidate genes from Staphylococcus aureus and Saccharomyces cerevisiae.
[0041] Isoprenoid Precursors IPP and DMAPP
[0042] The isoprenoid precursors IPP and DMAPP are naturally produced in the cell through one of two pathways (Fig.l): the Mevalonate pathway (MVA) (Yang et al. 2012, Chotani 2013, Zheng et al. 2013, Lv et al. 2014, Ramos et al. 2014, Vickers and Sabri 2015) and / or the Deoxyxylulose 5-phosphate pathway (DXP) / 2-C-methyl-D-erythritol 4-phosphate (MEP) (Zhao et al. 2011, Chotani 2013, Liu et al. 2014, Vickers and Sabri 2015). The MVA is mainly found in animals, plants, and many gram-positive bacteria, while the MEP exists in most bacteria and all plants (Primak et al. 2011, Wang et al. 2019). Calculations of net energy as well as carbon consumption (Fig.l) is presented in detail in our previous work (Huibin et al. 2017).
[0043] Escherichia coli and yeast were the first microorganisms engineered for terpenoid production, and significant progress has been achieved in these microorganisms (Vickers and Sabri 2015, Yao et al. 2018). However, the yield was far lower than the theoretical yield (Yang et al. 2012, Yang et al. 2012, Zheng et al. 2013, Campbell 2014, Liu et al. 2014, Vickers and Sabri 2015). Several strategies were individually employed to improve the yield in engineered E. coll such as tuning gene expression level by various promoters as well as Ribosome Binding Site (RBS)-optimization (Kim et al. 2016, Li et al. 2019) by suppression of unwanted byproducts from acetyl-CoA via deletion genes encoding shunt pathways (Kim et al. 2016, Liu et al. 2019) by screening of key enzymes (Primak et al. 2011, Kazieva et al. 2017, Yao et al.2018, Li et al. 2019, Liu et al. 2019); by recycling the redox cofactors and enhanced NADPH production (Guo et al. 2019, Liu et al. 2019); or by introduction of an orthogonal pathway (Chatzivasileiou et al. 2019). After these improvements, the isoprene yields reached 92 mg / 1 (Chatzivasileiou et al. 2019, Liu et al. 2019), 587 mg / 1 (Liu et al. 2019), 665 mg / 1 (Guo et al.2019), or 695 mg / 1 (Li et al. 2019). The major challenges with E. coll for terpenoid production are high toxicity of the intermediate metabolites (IPP, DMAPP), growth inhibition after induction especially when grown on glucose, and yields far from the theoretical yield. These challenges require selection of modified candidate genes and balancing of the expression level to increase the yield. When both MVA and DXP -pathways were combined for lycopene production in E. coli (Xu et al. 2018), the engineered strain accumulated around 60 mg / 1 lycopene in flask cultivation and with further screening for the fermentation media, the highest achieved lycopene production by fermentation was 925 mg / 1 (67 mg / g CDW after 40hr). Table 1 shows further strategies and production titer (reviewed in (Zhang and Hong 2020)).
[0044] Table 1: Strategies used to improve terpenoids production in E. coli {Zhang and Hong 2020).Product Strategy and features Culture conditions Titer Total Non-natural route to isoprenoid biosynthesis (isoprenoid alcohol Shake flasks 0.6 g / L monoterpenoids pathway / IPA)Pinene Adaptive laboratory evolution for improving pinene tolerance; E. Shake flasks 166.5 mg / L coli co-culture system; whole-cell biocatalysisLimonene Cell-free enzyme systems Glass vials 12.5 g / L Geranyl acetate Two-phase system; convert monoterpenoid geraniol to its acetyl Fed-batch 4.8 g / L ester to avoid geraniol toxicity fermentationViridiflorol Promoters and RBSs engineering Fed-batch 25.7 g / L fermentationFPP-resistant mevalonate kinase 1; lower MVA pathway; Fed-batch 8.5 g / L a-Bisabolol Optimization of inducer concentration, aeration and enzymatic fermentationcofactorLongifolene Codon optimization of longifolene synthase, investigate into Fed-batch 382 mg / L different FPP synthases fermentationOphiobolin F Ophiobolin synthase with SUMO tag; phylogenetics based Shake flasks 150 mg / L mutationCarotenoids Scaffold-free enzyme assemblies (IDI and CrtE) Fed-batch 276 mg / L fermentationAstaxanthin Promoters and RBSs engineering; multidimensional heuristic Fed-batch 320 mg / L process (MHP) fermentationCRISPR-mediated morphology and oxidative stress engineering Shake flasks 12 mg / g DCWZeaxanthin Dynamic control of MVA path way by IPP / FPP- responsive Fed-batch 722 mg / L promoter. fermentationCRISPRi-guided balancing of MVA pathway Shake flasks 71 mg / L T / vconeneOptimization of the lycopene biosynthetic; Overexpressing the Shake flasks 448 mg / g MEP pathway DCW
[0045] Compositions and methods described herein provide genetically engineered host cells that produce one or more terpenoids, wherein the host cells comprise one or more expression system that include one or more polycistronic operons encoding a ribosomal binding site (RBS), one or more promoters, an inducer, a selection system, genes to establish the pathway(s) for synthesizing one or more terpenoids, or a combination thereof.
[0046] For example, some of the methods described herein involves (a) incubating a population of host cells comprising an expression system that includes one or more polycistronic operons encoding a ribosomal binding site (RBS), one or more promoters, an inducer, and one or more genes for synthesizing one or more terpenoids; (b) isolating terpenoids from the population of host cells. The expression system can also include a selection system.
[0047] In some cases, methods of producing terpenes and / or terpenoids described herein can include, for example, (a) incubating a population of host cells comprising an expression system that includes one or more an expression cassettes (or expression vectors) comprising an upper-MVA (UMVP) operon and a lower-MVA (LMVP) operon, wherein the UMVP operon includes a heterologous promoter operably linked to a nucleic acid segment encoding Hmg-CoA reductase (HmgR), a nucleic acid segment encoding 3 -Hydroxy-3 -methylglutaryl CoA synthase (HmgS), and a nucleic acid segment encoding P-ketothiolase (PhaA) enzyme, and the LMVP operon includes a heterologous promoter operably linked to a nucleic acid segment encoding mevalonate kinase (MvaKl), phospho-mevalonate kinase (MvK2), diphosphomevalonate decarboxylase (MvD), and isopentenyl diphosphate isomerase (IDI).
[0048] Host Cells
[0049] The expression systems can be introduced into a variety of host cells to produce a chassis. Host cells that may be modified by these methods (e.g., by incorporation of nucleic acids and expression systems) include any microbial cells with metabolic and cellular properties suitable for use as chassis hosts for high-capacity production of terpenes and / or terpenoids. The host cells used herein are Escherichia coli (E. colt) expression strain BL21 and cloning strain DH5-alpha and Cupriavidus necator Hl 6. However, other microorganisms that can be modified to be such chassis hosts are contemplated. For example, microorganisms such as Saccharomyces cerevisiae, Corynebacterium glutamicum, Pichia pastoris, Pseudomonas putida, Yarrowia lipolytica, Clostridium ljungdahlii, Clostridium autoethanogenum, Clostridium kluyveri, Cupriavidus metallidurans; Pseudomonas fluorescens, Pseudomonas oleavorans; Delftia acidovorans, Bacillus subtilis, Lactobacillus delbrueckii, Lactococcus lactis, Aspergillus niger, Candida tropicalis, Candida albicans, Candida cloacae, Candida guillermondii, Candida intermedia, Candida maltosa, Candida parapsilosis, Candida zeylenoides, Issathenkia orientalis, Debaryomyces hansenii, Arxula adenoinivorans, Kluyveromyces lactis, or Exophiala, Mucor, Trichoderma, Cladosporium, Phanerochaete, Cladophialophora, Paecilomyces, Scedosporium, and / or Ophiostoma may be used as chassis hosts. In embodiments, the chassis host may include plant systems. For examples, the chassis hosts can be plants such as Nicotiana benthamiana, Nicotiana tabacum, Nicotiana rustica, Nicotiana excelsior, and / or Nicotiana excelsiana.
[0050] Polycistronic operons
[0051] The expression systems can include one or more expression cassette(s) (or expression vector(s)) having one or more polycistronic operons for expression of genes encoding enzymes for synthesizing the one or more terpenes and / or terpenoids. The polycistronic operon(s) can provide co-transcription of multiple genes in the same mRNA molecule from the same upstream promoter. The expression of the genes from polycistronic operon(s) can ensure production of each enzyme encoded therein and provide simultaneous inducibility of the gene expressed with an inducer encoded in the operon(s).
[0052] Ribosomal Binding Site (RBS)
[0053] Expression of genes from a polycistronic-operon requires an efficient ribosomal binding site (RBS) to maintain the level of expression of each individual enzyme encoded therein. The number of the nucleotides in the RBS as well their sequences and secondary structures can affect the expression level of the downstream gene (Tietze and Laie 2021). Read-through by the ribozyme can generate non-functional transcripts. Described herein is a new RBS that shifts the Open Reading Frame (ORF) of a downstream gene and ensures efficient translational termination. The 3’ end of the preceding ORF was modified via introduction of three nucleotides (GGG) coding for the amino acid glycine and changing the stop codon from ochre (TAA) to umber (TGA) (Fig. 2A). An additional ochre stopcodon (TAA) was introduced before the start codon of the downstream gene. This yielded the synthetic RBSHUS of only 11 nucleotides with a mere 5 nucleotides between the genes. The nucleic acid sequence of RBSHUS is:
[0054] GGGTGAGGCTA (SEQ ID NO: 1 )
[0055] The nucleic acid sequence of RBS3a is:
[0056] CGAGGGCAAAAA (SEQ ID NO:2)
[0057] RBSHUS was compared with the functionally established RBS3a for E. colt at 12-nuclotides and genes in the same reading frame using the green fluorescence protein (EGFP) as a reporter. RBSHUS showed almost 2-fold improvement in the expression level compared with RBS3a (Fig. 2B). The improved RBSHUS was used throughout this study in polycistronic operons.
[0058] Promoter
[0059] The Tet -promoter is inducible by anhydrotetracycline (AHT, Supelco $233 / 100 mg), through inhibition of its repressor tetR. The compatibility of the expression systems described herein was determined with another more cost-efficient induction agent doxycycline (DOC, Sigma Aldrich $140 / 5 g). As DOC is a broad-spectrum antibiotic synthetically derived from oxytetracycline, the limit of working concentration was determined. An E. coll strain carrying a pESG-TetP-EGFP plasmid with different DOC concentrations showed a significant growth inhibition via measurement of optical density of the E. coll culture at ODeoo with concentrations above 50 ng / ml (Fig. 3A). The native plasmid requires high concentration of AHT to bind with the tetR-repressor and induce the EGFP-expression. Significant expression of EGFP could be observed only with AHT-concentration above 50 ng / ml, with the highest expression level at 200 ng / ml of AHT (Fig. 3B). With DOC, efficient EGFP expression could achieved with a high concentration of 100 ng / ml (Fig.3C), which is beyond the limit of growth inhibition. Therefore, regulating the expression level of the tetR-repressor gene via modification of its RBS was developed.
[0060] Fifteen (15) different modifications of the RBS were constructed, and two modified RBS were identified that lead to increased sensitivity of the system to DOC, plausibly by decreasing expression of the repressor. Those RBS are RBS. Ml “ TA A TTGAA TTAA TG” (SEQ ID NO:3) and RBS. M2 “ TGAGAATTAATG” (SEQ ID NO:4). These modifications increased sensitivity for AHT and EGFP expression was achieved with concentrations as low as 12.5-50 ng / ml AHT. RBS. M2 showed a 5-fold increase in EGFP expression compared with the native RBS or RBS. Ml (Fig. 3B). With DOC, RBS. M2 achieved a concentration of 25 ng / ml with a 3.5-fold increase in EGFP expression in comparison with both other RBS (Fig. 3C). This supports full compatibility DOC with our system and likely with other commonly used A. coli strains and represents an approximately 100-fold reduction in cost of the inducing agent. Consequently, we used the RBS. M2 to control the expression level of te R-repressor gene in the following expression systems.
[0061] High mevalonate production in E. coli'.
[0062] The initial steps in formation of terpene building blocks in E. coli are the conversion of the central intermediate metabolite acetyl-CoA into mevalonate (Fig.4A, three-step upper mevalonate pathway (UMVP)). The key enzyme, and rate limiting step is the Hmg-CoAreductase (HmgR, EC: 1.1.1.34). There are two types of HmgR depending on the c-o-factors: NADH-dependent that is usually associated with mevalonate degradation (reverse reaction), and NADPH-dependent involved in formation of mevalonate (Ma et al. 2011, Miziorko 201 / ). Several microbial sources of HmgR (e.g., E.faecalis, S. cerevisiae, P. mevalonii, S. aureus, B. petrii, and D. acidovorans) were used to improve mevalonate production (Ma el al. 2011, Yang et al. 2012). Among them, HmgR from S. aureus has high specific activity and mevalonate yields (Ma et al. 2011). To build an efficient recombinant and complete UMVP, 5. aureus HmgR was combined with the microbial model enzyme Enterococcus faecalis 3-Hydroxy-3 -methylglutaryl CoA synthase (HmgS), and the -ketothiolase (PhaA) condensing two acetyl-CoA molecules to acetoacetyl-CoA from Ralslonia eutropha Hl 6 (Fig. 4B). The system was tested for mevalonate production in E. coli expression strain BL21 and the cloning strain DH5-alpha
[0063] Comparable production capacities were found for both strains (Fig. 4C). The impact of the culture media was also investigated. In mineral media (M9) supplemented with 1% glycerol, mevalonate accumulated to almost 800 mg / 1 after 24hr post induction in shake flask. The yield is further increased in LB-media and reached around 1.5 g / 1. The highest observed mevalonate production with approximately 2.5 g / 1 was achieved in LB-media with 1% glycerol as a carbon source. With the theoretical yield of mevalonate from glycerol at 0.536 g / g (0.33 mol / mol), 5.36 g / L, our highest production level reached 50% of the theoretical yield. This indicates that we have built a highly efficient synthetic UMVP.
[0064] A nucleic acid sequence for PhaA, acetyl-CoA C -acetyltransferase [EC:2.3.1.9], from Cupria\'idus necator H16: H16 A1438 is provided below7(SEQ ID NO: 18):1 atgactgacg ttgtcatcgt atccgccgcc cgcaccgcgg tcggcaagtt tggcggctcg 61 ctggccaaga tcccggcacc ggaactgggt gccgtggtca tcaaggccgc gctggagcgc 121 gccggcgtca agccggagca ggtgagcgaa gtcatcatgg gccaggtgct gaccgccggt 181 tcgggccaga accccgcacg ccaggccgcg atcaaggccg gcctgccggc gatggtgccg 241 gccatgacca tcaacaaggt gtgcggctcg ggcctgaagg ccgtgatgct ggccgccaac 301 gcgatcatgg cgggcgacgc cgagatcgtg gtggccggcg gccaggaaaa catgagcgcc 361 gccccgcacg tgctgccggg ctcgcgcgat ggtttccgca tgggcgatgc caagctggtc 421 gacaccatga tcgtcgacgg cctgtgggac gtgtacaacc agtaccacat gggcatcacc 481 gccgagaacg tggccaagga atacggcatc acacgcgagg cgcaggatga gttcgccgtc 541 ggctcgcaga acaaggccga agccgcgcag aaggccggca agtttgacga agagatcgtc 601 ccggtgctga tcccgcagcg caagggcgac ccggtggcct tcaagaccga cgagttcgtg 661 cgccagggcg ccacgctgga cagcatgtcc ggcctcaagc ccgccttcga caaggccggc 721 acggtgaccg cggccaacgc ctcgggcctg aacgacggcg ccgccgcggt ggtggtgatg 781 tcggcggcca aggccaagga actgggcctg accccgctgg ccacgatcaa gagctatgcc 841 aacgccggtg tcgatcccaa ggtgatgggc atgggcccgg tgccggcctc caagcgcgcc 901 ctgtcgcgcg ccgagtggac cccgcaagac ctggacctga tggagatcaa cgaggccttt961 gccgcgcagg cgctggcggt gcaccagcag atgggctggg acacctccaa ggtcaatgtg 1021 aacggcggcg ccatcgccat cggccacccg atcggcgcgt cgggctgccg tatcctggtg 1081 acgctgctgc acgagatgaa gcgccgtgac gcgaagaagg gcctggcctc gctgtgcatc 1141 ggcggcggca tgggcgtggc gctggcagtc gagcgcaaat aa
[0065] An amino acid sequence for PhaA encoded by nucleic acid sequence of SEQ ID NO: 18 is (SEQ ID NO: 19):MTDVVIVSAARTAVGKFGGSLAKIPAPELGAVVIKAALERAGVKPEQVSE VIMGQVLTAGSGQNPARQAAIKAGLPAMVPAMTINKVCGSGLKAVMLAA NAIMAGDAEIVVAGGQENMSAAPHVLPGSRDGFRMGDAKLVDTMIVDGL WDVYNQYHMGITAENVAKEYGITREAQDEFAVGSQNKAEAAQKAGKFD EEIVPVLIPQRKGDPVAFKTDEFVRQGATLDSMSGLKPAFDKAGTVTAAN ASGLNDGAAAVVVMSAAKAKELGLTPLATIKSYANAGVDPKVMGMGPV P ASKRAL SRAEWTPQDLDLMEINEAF AAQ AL AVHQQMGWDT SKVNVNG GAIAIGHPIGASGCRILVTLLHEMKRRDAKKGLASLCIGGGMGVALAVER
[0066] A nucleic acid sequence for HmgS, hydroxymethylglutaryl-CoA synthase [EC:2.3.3.10], from Enterococcus faecalis D32, AFO44063is provided below (SEQ ID NO:20):1 atgacaattg ggattgataa aattagtttt tttgtgcccc cttattatat tgatatgacg 61 gcactggctg aagccagaaa tgtagaccct ggaaaatttc atattggtat tgggcaagac 121 caaatggcgg tgaacccaat cagccaagat attgtgacat ttgcagccaa tgccgcagaa 181 gcgatcttga ccaaagaaga taaagaggcc attgatatgg tgattgtcgg gactgagtcc 241 agtatcgatg agtcaaaagc ggccgcagtt gtcttacatc gtttaatggg gattcaacct 301 ttcgctcgct ctttcgaaat caaggaagct tgttacggag caacagcagg cttacagtta 361 gctaagaatc acgtagcctt acatccagat aaaaaagtct tggttgtagc agcagatatt 421 gcaaaatatg gattaaattc tggcggtgag cctacacaag gagctggggc ggttgcaatg 481 ttagttgcta gtgaaccgcg catcttggct ttaaaagagg ataatgtgat gctgacgcaa 541 gatatctatg acttttggcg tccaacaggc catccgtatc ctatggtcga tggtcctttg 601 tcaaacgaaa cctacatcca atcttttgcc caagtctggg atgaacataa aaaaagaacc 661 ggtcttgatt ttgcagatta tgatgcttta gcgttccata ttccttacac aaaaatgggc 721 aaaaaagcct tattagcaaa aatctccgac caaactgaag cagaacagga acgaatttta 781 gcccgttatg aagaaagcat catctatagt cgtcgcgtag gaaacttgta tacgggttca 841 ctttatctgg gactcatttc ccttttagaa aatgcaacga ctttaaccgc aggcaatcaa 901 attgggttat tcagttatgg ttctggtgct gtcgctgaat ttttcactgg tgaattagta 961 gctggttatc aaaatcattt acaaaaagaa actcatttag cactgctaga taatcggaca 1021 gaactttcta tcgctgaata tgaagccatg tttgcagaaa ctttagacac agatattgat 1081 caaacgttag aagatgaatt aaaatatagt atttctgcta ttaataatac cgttcgctct 1141 tatcgaaact ag
[0067] A nucleic acid sequence of SEQ ID NO:20 that is codon optimized for C. necator H 16 is provided below (SEQ ID NO:21):1 atgaccatcg gcatcgacaa gatctcgttc ttcgtgccgc cgtactacat cgacatgacc 61 gccctggccg aagcccgcaa cgtggacccg ggcaagttcc acatcggcat cggccaggac 121 cagatggccg tgaacccgat ctcgcaggac atcgtgacct tcgccgccaa cgccgccgaa 181 gccatcctga ccaaggaaga caaggaagcc atcgacatgg tgatcgtggg caccgaatcg 241 tcgatcgacg aatcgaaggc cgccgccgtg gtgctgcacc gcctgatggg catccagccg 301 ttcgcccgct cgttcgaaat caaggaagcc tgctacggcg ccaccgccgg cctgcagctg 361 gccaagaacc acgtggccct gcacccggac aagaaggtgc tggtggtggc cgccgacatc 421 gccaagtacg gcctgaactc gggcggcgaa ccgacccagg gcgccggcgc cgtggccatg 481 ctggtggcct cggaaccgcg catcctggcc ctgaaggaag acaacgtgat gctgacccag 541 gacatctacg acttctggcg cccgaccggc cacccgtacc cgatggtgga cggcccgctg 601 tcgaatgaaa cctacatcca gtcgttcgcc caggtgtggg acgaacacaa gaagcgcacc 661 ggcctggact tcgccgacta cgacgccctg gccttccaca tcccgtacac caagatgggc 721 aagaaggccc tgctggccaa gatctcggac cagaccgaag ccgaacagga acgcatcctg 781 gcccgctacg aagaatcgat catctactcg cgccgcgtgg gcaacctgta caccggctcg 841 ctgtacctgg gcctgatctc gctgctggaa aacgccacca ccctgaccgc cggcaaccag 901 atcggcctgt tctcgtacgg ctcgggcgcc gtggccgaat tcttcaccgg cgaactggtg 961 gccggctacc agaaccacct gcagaaggaa acccacctgg ccctgctgga caaccgcacc 1021 gaactgtcga tcgccgaata cgaagccatg ttcgccgaaa ccctggacac cgacatcgac 1081 cagaccctgg aagacgaact gaagtactcg atctcggcca tcaacaacac cgtgcgctcg 1141 taccgcaact ag
[0068] An amino acid sequence for HmgS encoded by nucleic acid sequence of SEQ ID NO:21 is provided below (SEQ ID NO: 22):MTIGIDKISFFVVPYYlDMTALAEARNVDPGKFHIGIGQDQMAVNPISQDIVTFAA NAAEAILTKEDKEAIDMVIVGTESSIDESKAAAVVLHRLMGIQPFARSFEIKEACY GATAGLQLAKNHVALHPDKKVLVVAADIAKYGLNSGGEPTQGAGAVAMLVAS EPRIL ALKEDN LTQDIYDF WRPTGHP YPMATIGPL SNET YIQ SF AQ VWDEHKK RTGLDF DYDALAFELIPYTKMGKKALLAKISDQTEAEQERILARYEESIIYSRRVG NLYTGSLYLGLISLLENATTLTAGNQIGLFSYGSGAVAEFFTGELVAGYQNHLQK ETHIIALLDNRTELSIAEYEAMFAETLDTDIDQTLEDELKYSIISAINNTVRSYRN
[0069] A nucleic acid sequence for HmgR, hydroxymethylglutaryl-CoA reductase [EC: 1.1.1.34], from Staphylococcus aureus subsp. Aureus that is codon optimized for C. necator HI6 is provided below (SEQ ID NO:23):1 atgcagtcgc tggacaagaa cttccgccac ctgtcgcgcc agcagaagct gcagcagctg 61 gtggacaagc agtggctgtc ggaagaacag ttcaacatcc tgctgaacca cccgctgatc 121 gacgaagaag tggccaactc gctgatcgaa aacgtgatcg cccagggcgc cctgccggtg 181 ggcctgctgc cgaacatcat cgtggacgac aaggcctacg tggtgccgat gatggtggaa 241 gaaccgtcgg tggtggccgc cgcctcgtac ggcgccaagc tggtgaacca gaccggcggc 301 ttcaagaccg tgtcgtcgga acgcatcatg atcggccaga tcgtgttcga cggcgtggac 361 gacaccgaaa agctgtcggc cgacatcaag gccctggaaa agcagatcca caagatcgcc 421 gacgaagcct acccgtcgat caaggcccgc ggcggcggct accagcgcat cgccatcgac 481 accttcccgg aacagcagct gctgtcgctg aaggtgttcg tggacaccaa ggacgccatg 541 ggcgccaaca tgctgaacac catcctggaa gccatcaccg ccttcctgaa gaacgaattc 601 ccgcagtcgg acatcctgat gtcgatcctg tcgaaccacg ccaccgcctc ggtggtgaag 661 gtgcagggcg aaatcgacgt gaaggacctg gcccgcggcg aacgcaccgg cgaagaagtg 721 gccaagcgca tggaacgcgc ctcggtgctg gcccaggtgg acatccaccg cgccgccacc781 cacaacaagg gcgtgatgaa cggcatccac gccgtggtgc tggccaccgg caacgacacc 841 cgcggcgccg aagcctcggc ccacgcctac gcctcgcgcg acggccagta ccgcggcatc 901 gccacctggg gctacgacca ggaacgccag cgcctgatcg gcaccatcga agtgccgatg 961 accctggcca tcgtgggcgg cggcaccaag gtgctgccga tcgccaaggc ctcgctggaa 1021 ctgctgaacg tggactcggc ccaggaactg ggccacgtgg tggccgccgt gggcctggcc 1081 cagaacttcg ccgcctgccg cgccctggtg tcggaaggca tccagcaggg ccacatgtcg 1141 ctgcagtaca agtcgctggc catcgtggtg ggcgccaagg gcgacgaaat cgcccaggtg 1201 gccgaagccc tgaagcagga accgcgcgcc aacacccagg tggccgaacg catcctgcag 1261 gacctgcgct cgcagcagta g
[0070] An amino acid sequence for HmgR encoded by nucleic acid sequence of SEQ ID NO:23 is provided below (SEQ ID NO: 24):MQSLDKNFRHLSRQQKLQQLVDKQWLSEEQFNILLNHPLIDEEVANSLIENVIAQ GALPVGLLPNIIVDDKAYVVPMMEEPSVVAAASYGAKLVNQTGGFKTVSSERI MIGQIVFDGVDD TEKLS AD IK ALEKQI HK LAD ELAY PSIK RGGGYQRIAID TFPEQ QLLSLKVFVDTKDAMGANMLNTILEAITAFLKNEFPQSDILMSILSNHATASVVK VQGEIDVKDLARGERTGEEVAKRMERASVLAQVDIHRAATHNKGVMNGIHAVV LATGNDTRGAEASAHAYASRDGQYRGIATWGYDQERQRLIGTIEVPMTLAIVGG GTKVLIAKASLELLNVDSAQELGHVVAAVGLAQNFAACRALVSEGIQQGHMSL QYKSLAIVVGAKGDEIAQVAEALKQEPRANTQVAERILQDLRSQQ
[0071] A nucleic acid sequence for MvaKl, mevalonate kinase [EC:2.7.1.36], from Staphylococcus aureus subsp. aureus is provided below (SEQ ID NO:25):1 atgacaagaa aaggatatgg ggaatcgaca ggtaagatta ttttaatagg agaacatgct 61 gttacatttg gagagcctgc tattgcagta ccgtttaacg caggtaaaat caaagtttta 121 atagaagcct tagagagcgg gaactattcg tctattaaaa gcgatgttta cgatggtatg 181 ttatatgatg cgcctgacca tcttaagtct ttggtgaacc gttttgtaga attaaataat 241 attacagagc cgctagcagt aacgatccaa acgaatttac caccatcacg tgggttagga 301 tcgagtgcag ctgtcgcggt tgcttttgtt cgtgcgagtt atgatttttt agggaaatca 361 ttaacgaaag aagaactcat tgaaaaggct aattgggcag agcaaattgc acatggtaaa 421 ccaagtggta ttgatacgca aacgattgta tcaggcaaac cagtttggtt ccaaaaaggt 481 catgctgaaa cgttgaaaac gttaagttta gacggctata tggttgttat tgatactggt 541 gtgaaaggtt caacaagaca agcagtagaa gatgttcata aactttgtga ggaccctcag 601 tacatgtcac atgtaaaaca tatcggtaag ttagttttac gtgcgagtga tgtgattgaa 661 catcataaat ttgaagcctt agcggatatt tttaatgaat gtcatgcgga tttaaaggcg 721 ttgacagtta gtcatgataa aatagaacaa ttaatgaaaa ttggtaaaga aaatggtgcg 781 attgctggaa aacttactgg tgctggtcgt ggtggaagta tgttattgct tgccaaagat 841 ttaccaacag cgaaaaatat tgtaaaagct gtagaaaaag ctggtgcagc acatacatgg 901 attgagaatt tagga
[0072] A nucleic acid sequence of SEQ ID NO:25 that is codon optimized for C. necator H16 is provided below (SEQ ID NO:26):1 atgacccgca agggctatgg cgaatcgacc ggcaagatca tcctgatcgg cgaacatgcc 61 gtgacctttg gcgagccggc catcgccgtg ccgtttaacg ccggcaagat caaggtgctg 121 atcgaagccc tggagagcgg caactattcg tcgatcaaga gcgatgtgta cgatggcatg 181 ctgtatgatg cgccggacca tctgaagtcg ctggtgaacc gctttgtgga actgaataat 241 atcaccgagc cgctggccgt gacgatccag acgaatctgc cgccgtcgcg cggcctgggc 301 tcgtcggccg ccgtcgcggt ggcctttgtg cgcgcgtcgt atgattttct gggcaagtcg 361 ctgacgaagg aagaactgat cgaaaaggcc aattgggccg agcagatcgc ccatggcaag 421 ccgtcgggca tcgatacgca gacgatcgtg tcgggcaagc cggtgtggtt ccagaagggc 481 catgccgaaa cgctgaagac gctgtcgctg gacggctata tggtggtgat cgataccggc 541 gtgaagggct cgacccgcca ggccgtggaa gatgtgcata agctgtgcga ggacccgcag 601 tacatgtcgc atgtgaagca tatcggcaag ctggtgctgc gcgcgtcgga tgtgatcgaa 661 catcataagt ttgaagccct ggcggatatc tttaatgaat gccatgcgga tctgaaggcg 721 ctgaccgtgt cgcatgataa gatcgaacag ctgatgaaga tcggcaagga aaatggcgcg 781 atcgccggca agctgaccgg cgccggccgc ggcggctcga tgctgctgct ggccaaggat 841 ctgccgaccg cgaagaatat cgtgaaggcc gtggaaaagg ccggcgccgc ccatacctgg 901 atcgagaatc tgggc
[0073] An amino acid sequence for MvaKI encoded by nucleic acid sequence of SEQ ID NO:26 is provided below (SEQ ID NO:27):MTRKGYGESTGKIILIGEHAVTFGEPAIAVPFNAGKIKVLIEALESGNYSSIKSDVYDGMLYDAPDHLKSLVNRFVELNNITEPLAVTIQTNLPPSRGLGSSAAVAVAFVRASYDFLGKSLTKEELIEKANWAEQIAHGKPSGIDTQTIVSGKPVWFQKGHAETLKTLSLDGYMVVIDTGVKGSTRQAVEDVHKLCEDPQYMSHVTCHIGKLVLRASDVIEHEIKFEALADIFNECHADLKALTVSHDKIEQLMKIGKENGAIAGKLTGAGRGGSMLLLAKLPTAKNIKVAVEKAGAAHTWIENLG
[0074] A nucleic acid sequence for MvaKI, mevalonate kinase [NP_013935.1], from Saccharomyces cerevisiae subsp. cerevisiae is provided below (SEQ ID NO:65):ATGTCATTACCGTTCTTAACTTCTGCACCGGGAAAggttattattttggtGAACACTCT GCTGTGTACAACAAGCCTGCCGTCGCTGCTAGTGTGTCTGCGTTGAGAACCTA CCTGCTAATAAGCGAGTCATCTGCACCAGATACTATTGAATTGGACTTCCCGG ACATTAGCI'TTAATC / XTAAGTGGTCCATCA. ATG TTTCAATGCCATCAGCGAG GATCAAGTAAACTCCCAAAAATTGGCCAAGGCTCAACAAGCCACCGATGGCT TGTCTCAGGAACTCGTTAGTCTTTTGGATCCGTTGTTAGCTCAACTATCCGAA TCCTTCCACTACCATGCAGCGTTTTGTTTCCTGTATATGTTTGTTTGCCTATGC CCCCATGCCAAGAATATtaagttttctttaaagTCTACTTTACCCATCGGTGCTGGGTTG GGCTCAAGCGCCTCTATTTCTGTATCACTGGCCTTAGCTATGGCCTACTTGGG GGGGTTAATAGGATCTAATGACTTGGAAAAGCTGTCAGAAAACGATAAGCAT ATAGTGAATCAATGGGCCTTCATAGGTGAAAAGTGTATTCACGGTACCCCTTCAGGAATAGATAACGCTGTGGCCACTTATGGTAATGCCCTgctattgaaaaagacTCAC ATAATGGAACAATAAACACAAACAATTTTAAGTTCTTAGATGATTTCCCAGC CATTCCAATGATCCTAACCTATACTAGAATTCCAAGGTCTACAAAAGATCTTG TTGCTCGCGTTCGTGTGTTGGTCACCGAGAAATTTCCTGAAGTTATGAAGCCA ATTCTAGATGCCATGGGTGAATGTGCCCTACAAGGCTTAGAGATCATGACTA AGT TAAG FAAATGT AA AGGC ACCGA FGACGAGGCTGT AG A C FAA FAA I GA ACTGT ATG A AC A ACT ATTGG AATTG AT A AG A ATA A ATCA TGG AC TGCTTGTC T CAATCGGTGTTTCTCATCCTGGATTAGAACTTATTAAAAATCTGAGCGATGAT TTGAGAATTGGCTCCACAAAACTTACCGGTGCTGGTGGCGGCGGTTGCTCTTT GACTTTGTTACGAAGAGACATTACTCA GAGCAAATTGACAgcttcaaaaagaaattgc aAGATGATTTTAGTTACGAGACATTTGAAACAGACTTGGGTGGGACTGGCTGC TGTTTGTTAAGcgcaaaaaatttgaataaagaTCTTAAAATCAAATCCCTAGTATTCcaatattt gaaaataaaactacC CAAAGCAACAAATTGACGATCTATTATTGCCAGGAAACACG A ATTTACCATGGACTTCATAA
[0075] An amino acid sequence for MvaKl encoded by nucleic acid sequence of SEQ ID NO:65 is provided below (SEQ ID NO:66) (Uniprot: P07277):MSLPFLTSAPGKVIIFGEHSAVYNKPAVAASVSALRTYLLISESSAPDTIELDFPDIS FNHKWSINDFNAITEDQVNSQKLAKAQQATDGLSQELVSLLDPLLAQLSESFHY HAAFCFLYMFVCLCPHAKNIKFSLKSTLPIGAGLGSSASISVSLALAMAYLGGLIG SNDLEKLSENDKHIVNQWAFIGEKCIHGTPSGIDNAVATYGNALLFEKDSHNGTI NTNNFKFLDDFPAIPMILTYTRIPRSTKDLVARVRVLVTEKFPEVMKPILDAMGE CALQGLEIMTKLSKCKGTDDEAVETNNELYEQLLELIRINHGLLVSIGVSHPGLEL IKNLSDDLRIGSTKLTGAGGGGCSLTLLRRDITQEQIDSFKKKLQDDFSYETFETD LGGTGCCLLSAKNLNKDLKIKSLVFQLFENKTTTKQQIDDLLLPGNTNLPWTS
[0076] A nucleic acid sequence for MvK2, phosphomevalonate kinase [EC:2.7.4.2], from Staphylococcus aureus subsp. aureus is provided below (SEQ ID NO:28):1 atgattcagg tcaaagcacc cggaaaactt tatattgctg gagaatatgc tgtaacagaa 61 ccaggatata aatctgtact tattgcgtta gatcgttttg taactgctac tattgaagaa 121 gcagaccaat ataaaggtac cattcattca aaagcattac atcataaccc agttacattt 181 agtagagatg aagatagtat tgtcatttca gatccacatg cagcaaaaca attaaattat 241 gtggtcacag ctattgaaat atttgaacag tatgcgaaaa gctgcgatat agcgatgaag 301 cattttcatc tgactattga tagtaattta gatgattcaa atggtcataa atatggatta 361 ggttcaagtg cagcagtact tgtatcagtt ataaaagtat taaatgaatt ttatgatatg 421 aagttatcta atttatacat ttataaacta gcagtgattg caaatatgaa gttacaaagt 481 ttaagttcat gcggagatat tgctgtgagt gtatatagtg gatggctagc gtatagtact541 tttgatcatg aatgggttaa gcatcaaatt gaagatacta cggttgaaga agttttaatc 601 aaaaactggc ctggattgca catcgaacca ttacaagcac ctgaaaatat ggaagtactt 661 atcggttgga ctggctcacc ggcgtcatca ccacactttg ttagcgaagt gaaacgtttg 721 aaatcagatc cttcatttta cggtgacttc ttagaagatt cacatcgttg tgttgaaaaa 781 cttattcatg cttttaaaac aaataacatt aaaggtgtgc aaaagatggt gcgtcagaat 841 cgtacaatta ttcaacgtat ggataaagaa gctacagttg atatagaaac tgaaaagcta 901 aaatatttgt gtgatattgc tgaaaagtat cacggcgcat ctaaaacatc aggcgctggt 961 ggtggagact gtggtattac aattatcaat aaagatgtag ataaagaaaa aatttatgat 1021 gaatggacaa aacatggtat taaaccatta aaatttaata tttatcatgg gcaa
[0077] An amino acid sequence for MvK2 encoded by nucleic acid sequence of SEQ ID NO:28 is provided below (SEQ ID NO:29):AIIQVKAPGKLYIAGEYAVTEPGYKSVLLYLDRFVTATIEEADQYKGTIHSKALHH NPVTFSRDEDSIVISDPHAAKQLNYVVTAIEIFEQYAKSCDIAMKHFHLTIDSNLD DSNGHKYGLGSSAAVLVSVIKVLNEFYDMKLSNLYIYKLAVIANMKLQSLSSCG DIAVSVYSGWLAYSTFDHEWVKHQIEDTTVEEEVLIKNWPGLHIEPLQAPENMEV LIGWTGSPASSPHFVSEVKRLKSDPSFYGDFLEDSHRCVEKLIHAFKTNNIKGVQ KWRQYRTHQRMDKEATVDIETEKLKYLCDIAEKYHGASKTSGAGGGDCGITH NKDVDKEKIYDEWTKHGIKPLKFNIYHGQ
[0078] A nucleic acid sequence for MvK2, phosphomevalonate kinase [EC:2.7.4.2], from Saccharomyces cerevisiae is provided below (SEQ ID NO: 30):1 atgtcagagt tgagagcctt cagtgcccca gggaaagcgt tactagctgg tggatattta 61 gttttagata caaaatatga agcatttgta gtcggattat cggcaagaat gcatgctgta 121 gcccatcctt acggttcatt gcaagggtct gataagtttg aagtgcgtgt gaaaagtaaa 181 caatttaaag atggggagtg gctgtaccat ataagtccta aaagtggctt cattcctgtt 241 tcgataggcg gatctaagaa ccctttcatt gaaaaagtta tcgctaacgt atttagctac 301 tttaaaccta acatggacga ctactgcaat agaaacttgt tcgttattga tattttctct 361 gatgatgcct accattctca ggaggatagc gttaccgaac atcgtggcaa cagaagattg 421 agttttcatt cgcacagaat tgaagaagtt cccaaaacag ggctgggctc ctcggcaggt 481 ttagtcacag ttttaactac agctttggcc tccttttttg tatcggacct ggaaaataat 541 gtagacaaat atagagaagt tattcataat ttagcacaag ttgctcattg tcaagctcag 601 ggtaaaattg gaagcgggtt tgatgtagcg gcggcagcat atggatctat cagatataga 661 agattcccac ccgcattaat ctctaatttg ccagatattg gaagtgctac ttacggcagt 721 aaactggcgc atttggttga tgaagaagac tggaatatta cgattaaaag taaccattta 781 ccttcgggat taactttatg gatgggcgat attaagaatg gttcagaaac agtaaaactg 841 gtccagaagg taaaaaattg gtatgattcg catatgccag aaagcttgaa aatatataca 901 gaactcgatc atgcaaattc tagatttatg gatggactat ctaaactaga tcgcttacac 961 gagactcatg acgattacag cgatcagata tttgagtctc ttgagaggaa tgactgtacc 1021 tgtcaaaagt atcctgaaat cacagaagtt agagatgcag ttgccacaat tagacgttcc 1081 tttagaaaaa taactaaaga atctggtgcc gatatcgaac ctcccgtaca aactagctta 1141 ttggatgatt gccagacctt aaaaggagtt cttacttgct taatacctgg tgctggtggt 1201 tatgacgcca ttgcagtgat tactaagcaa gatgttgatc ttagggctca aaccgctaat 1261 gacaaaagat tttctaaggt tcaatggctg gatgtaactc aggctgactg gggtgttagg 1321 aaagaaaaag atccggaaac ttatcttgat aaa
[0079] A nucleic acid sequence of SEQ ID NO:30 that is codon optimized for C. necator Hl 6 is provided below (SEQ ID NO:31):1 atgtccgagt tgcgtgcatt cagcgcaccc gggaaggcat tgctcgcagg tgggtatctg 61 gtgctcgata ccaagtatga agccttcgtg gtcggtctgt cggcacgcat gcatgcagtc 121 gcccatccct acggttcgct gcagggctcg gataagtttg aagtgcgcgt gaagtcgaag 181 cagttcaagg atggcgagtg gctgtaccat atctcgccga agtccggctt catccccgtg 241 tcgatcggtg gctcgaagaa cccgttcatc gagaaggtga tcgcgaacgt gttcagctac 301 ttcaagccga acatggatga ctactgcaat cgcaacctgt tcgtgatcga tatcttctcg 361 gacgatgcat accattcgca ggaggatagc gtgaccgagc atcgtggcaa ccggcgcctg 421 tcgtttcatt cgcaccgcat cgaggaagtg cccaagaccg gtctgggctc ctcggcaggc 481 ctggtcaccg tgctgaccac cgcactggca tccttctttg tgagcgacct ggagaacaac 541 gtggacaagt atcgcgaagt gatccataac ctggcacagg tggcgcattg ccaggcacag 601 ggcaagatcg gtagcggctt cgatgtggcc gcagcggcct atggctcgat ccggtatcgc 661 cgcttcccgc ccgcactgat ctcgaatctg cccgatatcg gctcggcgac ctacggctcg 721 aagctggcac atctggtgga tgaagaggac tggaacatca cgatcaagtc gaaccatctg 781 ccgtcgggcc tgacgctgtg gatgggcgat atcaagaatg gcagcgagac cgtgaagctg 841 gtccagaagg tgaagaactg gtatgattcg catatgccgg agagcctgaa gatctacacc 901 gagctggatc atgccaactc gcgcttcatg gatggcctgt cgaagctgga tcgtctgcac 961 gagacccatg acgattacag cgatcagatc ttcgagtcgc tggagcgcaa cgactgcacc 1021 tgccagaagt atccggagat caccgaagtg cgcgatgccg tggccaccat ccggcgctcc 1081 ttccgcaaga tcaccaagga gtcgggtgcc gatatcgaac cgcccgtgca gaccagcctg 1141 ctggatgatt gccagaccct gaagggcgtg ctgacctgcc tgatccccgg tgcaggtggc 1201 tatgacgcca tcgccgtgat caccaagcag gatgtggatc tgcgggcaca gacggccaat 1261 gacaagcgct tctcgaaggt gcagtggctg gatgtgaccc aggcagactg gggtgtccgc 1321 aaggagaagg atccggagac ctatctggat aag
[0080] An amino acid sequence for MvK2 encoded by nucleic acid sequence of SEQ ID NO:31 is provided below (SEQ ID NO:32):MSELRAFSAPGKALLAGGYLVLDTKYEAFVVGLSARMHAVAHPYGSLQGSDKFE VRVKSKQFKDGEWLYHISPKSGFIPVSIGGSKNPFIEKVIANVFSYFKPNMDDYCN RNLFVIDIFSDDAYHSQEDSVTEHRGNRRLSFHSHRIEEVPKTGLGSSAGLVTVLTT ALASFFVSDLENNVDKYREVIHNLAQVAHCQAQGKIGSGFDVAAAAYGSIRYRR FPPALISNLPDIGSATYGSKLAHLVDEEDWNITIKSNHLPSGLTLWMGDIKNGSET VKLVQKVKNWYDSHMPESLKIYTELDHANSRFMDGLSKLDRLHETHDDYSDQIF ESLERNDCTCQKYPEITEVRDAVATIRRSFRKITKESGADIEPPVQTSLLDDCQTLK GVLTCLIPGAGGYDAIAVITKQDVDLRAQTANDKRFSKVQWLDVTQADWGVRK EKDPETYLDK
[0081] A nucleic acid sequence for MvaD, diphosphomevalonate decarboxylase [EC:4.1.1.33], from Staphylococcus aureus subsp. aureus is provided below (SEQ ID NO:33):1 atgattaaaa gtggcaaagc acgtgcacat acgaatattg cacttataaa atattggggt 61 aaaaaagatg aagcactaat cattccaatg aataatagca tatctgttac attagaaaaa 121 ttttacactg aaacgaaagt cacttttaac gaccagttaa cacaggatca attttggttg 181 aatggtgaaa aggttagtgg caaagaatta gagaaaattt caaaatatat ggatattgtc241 agaaatagag ctggcatcga ttggtatgca gaaattgaaa gcgacaattt tgtaccaaca 301 gcagcaggat tggcttcatc ggcaagcgca tatgcagctt tagcagcagc ttgtaatcaa 361 gcactagact tgcagctgtc agataaggat ttatcgagat tggcgcgaat tggttcgggt 421 tctgcatcgc gtagtattta tggtggattt gcagaatggg aaaaagggta taatgatgaa 481 acgtcatatg ccgttccact tgaatcgaat cattttgaag atgaccttgc catgatattt 541 attgtgatta atcaacattc taaaaaggta cctagtcgat atggtatgtc gttgacacga 601 aacacatcaa ggttttacca atactggtta gatcatattg atgaagattt agctgaagca 661 aaagcagcga ttcaagacaa agattttaaa cgccttggtg aagtaattga agaaaatggt 721 ttgcgtatgc atgccacgaa tctaggatca acaccgccgt tcacatatct tgtgcaagaa 781 agttatgatg taatggcgct cgttcacgaa tgccgagaag cgggatatcc gtgttatttt 841 acgatggatg cgggtcctaa tgtgaaaata cttgtagaaa agaaaaacaa gcaacagatt 901 atagataaat tattaacaca gtttgataat aaccaaatta ttgatagtga tattattgcc 961 acaggaattg aaataattga g
[0082] A nucleic acid sequence of SEQ ID NO:33 that is codon optimized for C. necalor H16 is provided below (SEQ ID NO:34):1 atgatcaagt cgggcaaggc ccgcgcccat acgaatatcg ccctgatcaa gtattggggc 61 aagaaggatg aagccctgat catcccgatg aataatagca tctcggtgac cctggaaaag 121 ttttacaccg aaacgaaggt cacctttaac gaccagctga cccaggatca gttttggctg 181 aatggcgaaa aggtgtcggg caaggaactg gagaagatct cgaagtatat ggatatcgtc 241 cgcaatcgcg ccggcatcga ttggtatgcc gaaatcgaaa gcgacaattt tgtgccgacc 301 gccgccggcc tggcctcgtc ggccagcgcc tatgccgccc tggccgccgc ctgcaatcag 361 gccctggacc tgcagctgtc ggataaggat ctgtcgcgcc tggcgcgcat cggctcgggc 421 tcggcctcgc gctcgatcta tggcggcttt gccgaatggg aaaagggcta taatgatgaa 481 acgtcgtatg ccgtgccgct ggaatcgaat cattttgaag atgacctggc catgatcttt 541 atcgtgatca atcagcattc gaagaaggtg ccgtcgcgct atggcatgtc gctgacccgc 601 aacacctcgc gcttttacca gtactggctg gatcatatcg atgaagatct ggccgaagcc 661 aaggccgcga tccaggacaa ggattttaag cgcctgggcg aagtgatcga agaaaatggc 721 ctgcgcatgc atgccacgaa tctgggctcg accccgccgt tcacctatct ggtgcaggaa 781 tcgtatgatg tgatggcgct ggtgcacgaa tgccgcgaag cgggctatcc gtgctatttt 841 acgatggatg cgggcccgaa tgtgaagatc ctggtggaaa agaagaacaa gcagcagatc 901 atcgataagc tgctgaccca gtttgataat aaccagatca tcgattcgga tatcatcgcc 961 accggcatcg aaatcatcga g
[0083] An amino acid sequence for MvaD encoded by nucleic acid sequence of SEQ ID NO:34 is provided below (SEQ ID NO:35):MIKSGKARAHTNIALIKYWGKKDEALIIPMNNSISVTLEKFYTETKVTFNDQLTQ DQFWLNGEKVSGKELEKISKYMDIVRNRAGIDWYAEIESDNFVPTAAGLASSAS AYAALAAACNQALDLQLSDKDLSRLARIGSGSASRSIYGGFAEWEKGYNDETSY AVPLESNHFEDDLAMIFIVINQHSKKVPSRYGMSLTRNTSRFYQYWLDHIDEDLA EAKAAIQDKDFKRLGEVIEENGLRMHATNLGSTPPFTYLVQESYDVMALVHECR EAGYPCYFTMDAGPNVKILVEKKNKQQIIDKLLTQFDNNQIIDSDIIATGIEIIE
[0084] A nucleic acid sequence for MvaD (also called MvD), diphosphomevalonate decarboxylase [EC:4.1.1.33], from Saccharomyces cerevisiae is provided below (SEQ ID NO:36):1 atgaccgttt acacagcatc cgttaccgca cccgtcaaca tcgcaaccct taagtattgg 61 gggaaaaggg acacgaagtt gaatctgccc accaattcgt ccatatcagt gactttatcg 121 caagatgacc tcagaacgtt gacctctgcg gctactgcac ctgagtttga acgcgacact 181 ttgtggttaa atggagaacc acacagcatc gacaatgaaa gaactcaaaa ttgtctgcgc 241 gacctacgcc aattaagaaa ggaaatggaa tcgaaggacg cctcattgcc cacattatct 301 caatggaaac tccacattgt ctccgaaaat aactttccta cagcagctgg tttagcttcc 361 tccgctgctg gctttgctgc attggtctct gcaattgcta agttatacca attaccacag 421 tcaacttcag aaatatctag aatagcaaga aaggggtctg gttcagcttg tagatcgttg 481 tttggcggat acgtggcctg ggaaatggga aaagctgaag atggtcatga ttccatggca 541 gtacaaatcg cagacagctc tgactggcct cagatgaaag cttgtgtcct agttgtcagc 601 gatattaaaa aggatgtgag ttccactcag ggtatgcaat tgaccgtggc aacctccgaa 661 ctatttaaag aaagaattga acatgtcgta ccaaagagat ttgaagtcat gcgtaaagcc 721 attgttgaaa aagatttcgc cacctttgca aaggaaacaa tgatggattc caactctttc 781 catgccacat gtttggactc tttccctcca atattctaca tgaatgacac ttccaagcgt 841 atcatcagtt ggtgccacac cattaatcag ttttacggag aaacaatcgt tgcatacacg 901 tttgatgcag gtccaaatgc tgtgttgtac tacttagctg aaaatgagtc gaaactcttt 961 gcatttatct ataaattgtt tggctctgtt cctggatggg acaagaaatt tactactgag 1021 cagcttgagg ctttcaacca tcaatttgaa tcatctaact ttactgcacg tgaattggat 1081 cttgagttgc aaaaggatgt tgccagagtg attttaactc aagtcggttc aggcccacaa 1141 gaaacaaacg aatctttgat tgacgcaaag actggtctac caaaggaa
[0085] An amino acid sequence for MvaD encoded by nucleic acid sequence of SEQ ID NO:36 is provided below (SEQ ID NO:37):MTVYTASVTAPVNIATLKYWGKRDTKLNLPTNSSISVTLSQDDLRTLTSAATAPE FERDTLWLNGEPHSIDNERTQNCLRDLRQLRKEMESKDASLPTLSQWKLHIVSEN NFPTAAGLASSAAGFAALVSAIAKLYQLPQSTSEISRIARKGSGSACRSLFGGYVA WEMGKAEDGHDSMAVQIADSSDWPQMKACVLVVSDIKKDVSSTQGMQLTVAT SELFKERIEHVVPKRFEVMRKAIVEKDFATFAKETMMDSNSFHATCLDSFPPIFY MNDTSKRIISWCHTINQFYGETIVAYTFDAGPNAVLYYLAENESKLFAFIYKLFGS VPGWDKKFTTEQLEAFNHQFESSNFTARELDLELQKDVARVILTQVGSGPQETN ESLID AKTGLPKE
[0086] A nucleic acid sequence for IDI, isopentenyl-diphosphate Delta-isom erase [EC5.3.3.2], from Saccharomyces cerevisiae is provided below (SEQ ID NO:38):1 atgactgccg acaacaatag tatgccccat ggtgcagtat ctagttacgc caaattagtg 61 caaaaccaaa cacctgaaga cattttggaa gagtttcctg aaattattcc attacaacaa 121 agacctaata cccgatctag tgagacctca aatgacgaaa gcggagaaac atgtttttct 181 ggtcatgatg aggagcaaat taagttaatg aatgaaaatt gtattgtttt ggattgggac 241 gataatgcta ttggtgccgg taccaagaaa gtttgtcatt taatggaaaa tattgaaaag 301 ggtttactac atcgtgcatt ctccgtcttt attttcaatg aacaaggtga attactttta 361 caacaaagag ccactgaaaa aataactttc cctgatcttt ggactaacac atgctgctct 421 catccactat gtattgatga cgaattaggt ttgaagggta agctagacga taagattaag 481 ggcgctatta ctgcggcggt gagaaaacta gatcatgaat taggtattcc agaagatgaa 541 actaagacaa ggggtaagtt tcacttttta aacagaatcc attacatggc accaagcaat 601 gaaccatggg gtgaacatga aattgattac atcctatttt ataagatcaa cgctaaagaa 661 aacttgactg tcaacccaaa cgtcaatgaa gttagagact tcaaatgggt ttcaccaaat 721 gatttgaaaa ctatgtttgc tgacccaagt tacaagttta cgccttggtt taagattatt781 tgcgagaatt acttattcaa ctggtgggag caattagatg acctttctga agtggaaaat 841 gacaggcaaa ttcatagaat gctagggaga ggc
[0087] A nucleic acid sequence of SEQ ID NO:38 that is codon optimized for C. necator H16 is provided below (SEQ ID NO:39):1 atgaccgccg acaacaattc gatgccccat ggcgccgtgt cgtcgtacgc caagctggtg 61 cagaaccaga ccccggaaga catcctggaa gagtttccgg aaatcatccc gctgcagcag 121 cgcccgaata cccgctcgtc ggagacctcg aatgacgaaa gcggcgaaac ctgcttttcg 181 ggccatgatg aggagcagat caagctgatg aatgaaaatt gcatcgtgct ggattgggac 241 gataatgcca tcggcgccgg caccaagaag gtgtgccatc tgatggaaaa tatcgaaaag 301 ggcctgctgc atcgcgcctt ctccgtcttt atcttcaatg aacagggcga actgctgctg 361 cagcagcgcg ccaccgaaaa gatcaccttc ccggatctgt ggaccaacac ctgctgctcg 421 catccgctgt gcatcgatga cgaactgggc ctgaagggca agctggacga taagatcaag 481 ggcgccatca ccgcggcggt gcgcaagctg gatcatgaac tgggcatccc ggaagatgaa 541 accaagaccc gcggcaagtt tcactttctg aaccgcatcc attacatggc cccgagcaat 601 gaaccgtggg gcgaacatga aatcgattac atcctgtttt ataagatcaa cgccaaggaa 661 aacctgaccg tcaacccgaa cgtcaatgaa gtgcgcgact tcaagtgggt gtcgccgaat 721 gatctgaaga ccatgtttgc cgacccgtcg tacaagttta cgccgtggtt taagatcatc 781 tgcgagaatt acctgttcaa ctggtgggag cagctggatg acctgtcgga agtggaaaat 841 gaccgccaga tccatcgcat gctgggccgc ggctag
[0088] An amino acid sequence for IDI encoded by nucleic acid sequence of SEQ ID NO:39 is provided below (SEQ ID NO:40):MTADNNSMPHGAVSSYAKLVQNQTPEDILEEFPEIIPLQQRPNTRSSETSNDESGETC F SGHDEEQIKLMNENCIVLDWDDNAIGAGTKKVCHLMENIEKGLLHRAF SVFIFNEQ GELLLQQRATEKITFPDLWTNTCCSHPLCIDDELGLKGKLDDKIKGAITAAVRKLDH ELGIPEDETKTRGKFHFLNRIHYMAPSNEPWGEHEIDYILFYKTNAKENLTVNPNVNE VRDFKWVSPNDLKTMFADPSYKFTPWFKIICENYLFNWWEQLDDLSEVENDRQIHR MLGRG
[0089] A nucleic acid sequence for IspA, geranyl diphosphate / farnesyl diphosphate synthase [EC:2.5.1.1 / 2.5.1.10], from A. coli K-12 MG1655: b0421 is provided below (SEQ ID NO:41):1 atggactttc cgcagcaact cgaagcctgc gttaagcagg ccaaccaggc gctgagccgt 61 tttatcgccc cactgccctt tcagaacact cccgtggtcg aaaccatgca gtatggcgca 121 ttattaggtg gtaagcgcct gcgacctttc ctggtttatg ccaccggtca tatgttcggc 181 gttagcacaa acacgctgga cgcacctgct gccgccgttg agtgtatcca cgcttactca 241 ttaattcatg atgatttacc ggcaatggat gatgacgatc tgcgtcgcgg tttgccaacc 301 tgccatgtga agtttggcga agcaaacgcg attctcgctg gcgacgcttt acaaacgctg 361 gcgttctcga ttttaagcga tgccgatatg ccggaagtgt cggaccgcga cagaatttcg 421 atgatttctg aactggcgag cgccagtggt attgccggaa tgtgcggtgg tcaggcatta 481 gatttagacg cggaaggcaa acacgtacct ctggacgcgc ttgagcgtat tcatcgtcat 541 aaaaccggcg cattgattcg cgccgccgtt cgccttggtg cattaagcgc cggagataaa 601 ggacgtcgtg ctctgccggt actcgacaag tatgcagaga gcatcggcct tgccttccag 661 gttcaggatg acatcctgga tgtggtggga gatactgcaa cgttgggaaa acgccagggt721 gccgaccagc aacttggtaa aagtacctac cctgcacttc tgggccttga gcaagcccgg 781 aagaaagccc gggatctgat cgacgatgcc cgtcagtcgc tgaaacaact ggctgaacag 841 tcactcgata cctcggcact ggaagcgcta gcggactaca tcatccagcg taataaa
[0090] A nucleic acid sequence of SEQ ID NO:41 that is codon optimized for C. necalor Hl 6 is provided below (SEQ ID NO: 42):1 atggattttc cgcagcagct ggaagcctgc gtgaagcagg ccaatcaggc gctgagccgc 61 tttatcgccc cgctgccctt tcagaatacc cccgtggtcg aaaccatgca gtatggcgcc 121 ctgctgggcg gcaagcgcct gcgcccgttt ctggtgtatg ccaccggcca tatgtttggc 181 gtgagcacca atacgctgga tgccccggcc gccgccgtgg aatgcatcca tgcctattcg 241 ctgatccatg atgatctgcc ggccatggat gatgatgatt tgcgccgcgg cctgccgacc 301 tgccatgtga agtttggcga agccaatgcg atcctggccg gcgatgccct gcagacgctg 361 gcgttttcga tcctgagcga tgccgatatg ccggaagtgt cggatcgcga tcgcatctcg 421 atgatctcgg aactggcgag cgcctcgggc atcgccggca tgtgcggcgg ccaggccctg 481 gatctggatg cggaaggcaa gcatgtgccg ctggatgcgc tggaacgcat ccatcgccat 541 aagaccggcg ccctgatccg cgccgccgtg cgcctgggcg ccctgagcgc cggcgataag 601 ggccgccgcg ccctgccggt gctggataag tatgccgaaa gcatcggcct ggcctttcag 661 gtgcaggatg atatcctgga tgtggtgggc gataccgcca cgctgggcaa gcgccagggc 721 gccgatcagc agctgggcaa gtcgacctat ccggccctgc tgggcctgga acaggcccgg 781 aagaaggccc gggatctgat cgatgatgcc cgccagtcgc tgaagcagct ggccgaacag 841 tcgctggata cctcggccct ggaagcgctg gcggattata tcatccagcg caataag
[0091] An amino acid sequence for IspA encoded by nucleic acid sequence of SEQ ID NO:42 is provided below (SEQ ID NO:43):MDFPQQLEACVKQANQALSRFIAPLPFQNTPVVETMQYGALLGGKRLRPFLVYA TGHMFGVSTNTLDAPAAAVECIHAYSLIHDDLPAMDDDDLRRGLPTCHVKFGEA NAILAGDALQTLAFSILSDADMPEVSDRDRISMISELASASGIAGMCGGQALDLD AEGKHVPLDALERIHRHKTGALIRAAVRLGALSAGDKGRRALPVLDKYAESIGL AFQVQDDILDVVGDTATLGKRQGADQQLGKSTYPALLGLEQARKKARDLIDDA RQ SLKQLAEQSLDT S ALE AL AD YIIQRNK
[0092] A nucleic acid sequence for crtE, Geranylgeranyl diphosphate synthase [EC2.5.1.29] from, Pseudescherichia vulneris is provided below (SEQ ID NO:44):1 atggtgagtg gcagtaaagc gggcgtttcg cctcatcgcg aaatagaagt aatgagacaa 61 tccattgacg atcacctggc tggcctgtta cctgaaaccg acagccagga tatcgtcagc 121 cttgcgatgc gtgaaggcgt catggcaccc ggtaaacgga tccgtccgct gctgatgctg 181 ctggccgccc gcgacctccg ctaccagggc agtatgccta cgctgctcga tctcgcctgc 241 gccgttgaac tgacccatac cgcgtcgctg atgctcgacg acatgccctg catggacaac 301 gccgagctgc gccgcggtca gcccactacc cacaaaaaat ttggtgagag cgtggcgatc 361 cttgcctccg ttgggctgct ctctaaagcc tttggtctga tcgccgccac cggcgatctg 421 ccgggggaga ggcgtgccca ggcggtcaac gagctctcta ccgccgtggg cgtgcagggc 481 ctggtactgg ggcagtttcg cgatcttaac gatgccgccc tcgaccgtac ccctgacgct 541 atcctcagca ccaaccacct caagaccggc attctgttca gcgcgatgct gcagatcgtc 601 gccattgctt ccgcctcgtc gccgagcacg cgagagacgc tgcacgcctt cgccctcgac 661 ttcggccagg cgtttcaact gctggacgat ctgcgtgacg atcacccgga aaccggtaaa 721 gatcgcaata aggacgcggg aaaatcgacg ctggtcaacc ggctgggcgc agacgcggcc781 cggcaaaagc tgcgcgagca tattgattcc gccgacaaac acctcacttt tgcctgtccg 841 cagggcggcg ccatccgaca gtttatgcat ctgtggtttg gccatcacct tgccgactgg 901 tcaccggtca tgaaaatcgc c
[0093] A nucleic acid sequence of SEQ ID NO:44 that is codon optimized for C. necalor H16 is provided below (SEQ ID NO:45):1 atggtgagcg gctcgaaggc gggcgtgtcg ccgcatcgcg aaatcgaagt gatgcgccag 61 tccatcgacg atcatctggc cggcctgctg ccggaaaccg atagccagga tatcgtcagc 121 ctggcgatgc gcgaaggcgt catggcgccc ggcaagcgga tccgcccgct gctgatgctg 181 ctggccgccc gcgatctgcg ctaccagggc tcgatgccga cgctgctgga tctggcctgc 241 gccgtggaac tgacccatac cgcgtcgctg atgctggatg atatgccctg catggacaac 301 gccgaactgc gccgcggcca gcccaccacc cataagaagt ttggcgaaag cgtggcgatc 361 ctggcctccg tgggcctgct gtcgaaggcc tttggcctga tcgccgccac cggcgatctg 421 ccgggcgagc gccgcgccca ggcggtcaac gaactgtcga ccgccgtggg cgtgcagggc 481 ctggtgctgg gccagtttcg cgatctgaac gatgccgccc tggatcgcac cccggatgcc 541 atcctgagca ccaaccatct gaagaccggc atcctgttca gcgcgatgct gcagatcgtc 601 gccatcgcct ccgcctcgtc gccgagcacg cgcgaaaccc tgcatgcctt cgccctggat 661 ttcggccagg cgtttcagct gctggacgat ctgcgcgacg atcatccgga aaccggcaag 721 gatcgcaata aggatgcggg caagtcgacg ctggtcaacc ggctgggcgc cgatgcggcc 781 cggcagaagc tgcgcgaaca tatcgattcc gccgataagc atctgacctt tgcctgcccg 841 cagggcggcg ccatccgcca gtttatgcat ctgtggtttg gccatcatct ggccgattgg 901 tcgccggtca tgaagatcgc c
[0094] An amino acid sequence for crtE encoded by nucleic acid sequence of SEQ ID NO:45 is provided below (SEQ ID NO:46):MVSGSKAGVSPHREIEVMRQSIDDHLAGLLPETDSQDIVSLAMREGVMAPGKRI RPLLMLLAARDLRYQGSMPTLLDLACAVELTHTASLMLDDMPCMDNAELRRG QPTTHKKFGESVAILASVGLLSKAFGLIAATGDLPGERRAQAVNELSTAVGVQGL VLGQFRDLNDAALDRTPDAILSTNHLKTGILFSAMLQIVAIASASSPSTRETLHAF ALDFGQAFQLLDDLRDDHPETGKDRNKDAGKSTLVNRLGADAARQKLREHIDS ADKHLTFACPQGGAIRQFMHLWFGHHLADWSPVMKIA
[0095] A nucleic acid sequence for crtB, phytoene synthase [EC:2.5.1.32], from Lamprocystis purpurea is provided below (SEQ ID NO:47):1 atgagccaac cgccgctgct tgaccacgcc acgcagacca tggccaacgg ctcgaaaagt 61 tttgccaccg ctgcgaagct gttcgacccg gccacccgcc gtagcgtgct gatgctctac 121 acctggtgcc gccactgcga tgacgtcatt gacgaccaga cccacggctt cgccagcgag 181 gccgcggcgg aggaggaggc cacccagcgc ctggcccggc tgcgcacgct gaccctggcg 241 gcgtttgaag gggccgagat gcaggatccg gccttcgctg cctttcagga ggtggcgctg 301 acccacggta ttacgccccg catggcgctc gatcacctcg acggctttgc gatggacgtg 361 gctcagaccc gctatgtcac ctttgaggat acgctgcgct actgctatca cgtggcgggc 421 gtggtgggtc tgatgatggc cagggtgatg ggcgtgcggg atgagcgggt gctggatcgc 481 gcctgcgatc tggggctggc cttccagctg acgaatatcg cccgggatat tattgacgat 541 gcggctattg accgctgcta tctgcccgcc gagtggctgc aggatgccgg gctgaccccg 601 gagaactatg ccgcgcggga gaatcgggcc gcgctggcgc gggtggcgga gcggcttatt 661 gatgccgcag agccgtacta catctcctcc caggccgggc tacacgatct gccgccgcgc721 tgcgcctggg cgatcgccac cgcccgcagc gtctaccggg agatcggtat taaggtaaaa 781 gcggcgggag gcagcgcctg ggatcgccgc cagcacacca gcaaaggtga aaaaattgcc 841 atgctgatgg cggcaccggg gcaggttatt cgggcgaaga cgacgagggt gacgccgcgt 901 ccggccggtc tttggcagcg tcccgtt
[0096] A nucleic acid sequence of SEQ ID NO:47 that is codon optimized for C. necator HI 6 is provided below (SEQ ID NO:48):1 atgagccagc cgccgctgct ggatcatgcc acgcagacca tggccaacgg ctcgaagtcg 61 tttgccaccg ccgcgaagct gttcgatccg gccacccgcc gcagcgtgct gatgctgtat 121 acctggtgcc gccactgcga tgacgtcatc gatgatcaga cccacggctt cgccagcgaa 181 gccgcggcgg aggaggaagc cacccaacgc ctggcccggc tgcgcacgct gaccctggcg 241 gcgtttgaag gcgccgaaat gcaagatccg gcctttgccg cctttcaaga agtggcgctg 301 acccatggca tcacgccccg catggcgctg gatcatctgg atggctttgc gatggatgtg 361 gcccaaaccc gctatgtcac ctttgaagat acgctgcgct attgctatca tgtggcgggc 421 gtggtgggcc tgatgatggc ccgcgtgatg ggcgtgcggg atgaacgggt gctggatcgc 481 gcctgcgatc tgggcctggc ctttcaactg acgaatatcg cccgggatat catcgatgat 541 gcggccatcg atcgctgcta tctgcccgcc gaatggctgc aagatgccgg cctgaccccg 601 gaaaattatg ccgcgcggga aaatcgggcc gcgctggcgc gggtggcgga acggctgatc 661 gatgccgccg aaccgtatta tatctcctcc caagccggcc tgcatgatct gccgccgcgc 721 tgcgcctggg cgatcgccac cgcccgcagc gtctatcggg aaatcggcat caaggtgaag 781 gcggcgggcg gcagcgcctg ggatcgccgc caacatacca gcaagggcga aaagatcgcc 841 atgctgatgg cggccccggg ccaagtgatc cgggcgaaga cgacgcgcgt gacgccgcgc 901 ccggccggcc tgtggcagcg tcccgtt
[0097] An amino acid sequence for crtB encoded by nucleic acid sequence of SEQ ID NO:48 is provided below (SEQ ID NO:49):MSQPPLLDHATQTMANGSKSFATAAKLFDPATRRSVLMLYTWCRHCDDVIDDQ THGFASEAAAEEEATQRLARLRTLTLAAFEGAEMQDPAFAAFQEVALTHGITPR MALDHLDGFAMDVAQTRYVTFEDTLRYCYHVAGVVGLMMARVMGVRDERVL DRACDLGLAFQLTNIARDIIDDAAIDRCYLPAEWLQDAGLTPENYAARENRAAL ARVAERLIDAAEPYYISSQAGLHDLPPRCAWAIATARSVYREIGIKVKAAGGSAW DRRQHTSKGEKIAMLMAAPGQVIRAKTTRVTPRPAGLWQRPV
[0098] A nucleic acid sequence for crtl, phytoene desaturase [EC: 1.3.99.3], from Lamprocystis purpurea is provided below (SEQ ID NO: 50):1 atgaaaaaaa ccgttgtgat tggcgcaggc tttggtggcc tggcgctggc gattcgcctg 61 caggcggcag ggatcccaac cgtactgctg gagcagcggg acaagcccgg cggtcgggcc 121 tacgtctggc atgaccaggg ctttaccttt gacgccgggc cgacggtgat caccgatcct 181 accgcgcttg aggcgctgtt caccctggcc ggcaggcgca tggaggatta cgtcaggctg 241 ctgccggtaa aacccttcta ccgactctgc tgggagtccg ggaagaccct cgactatgct 301 aacgacagcg ccgagcttga ggcgcagatt acccagttca acccccgcga cgtcgagggc 361 taccggcgct ttctggctta ctcccaggcg gtattccagg agggatattt gcgcctcggc 421 agcgtgccgt tcctctcttt tcgcgacatg ctgcgcgccg ggccgcagct gcttaagctc 481 caggcgtggc agagcgtcta ccagtcggtt tcgcgcttta ttgaggatga gcatctgcgg 541 caggccttct cgttccactc cctgctggta ggcggcaacc ccttcaccac ctcgtccatc 601 tacaccctga tccacgccct tgagcgggag tggggggtct ggttccctga gggcggcacc661 ggggcgctgg tgaacggcat ggtgaagctg tttaccgatc tgggcgggga gatcgaactc 721 aacgcccggg tcgaggagct ggtggtggcc gataaccgcg taagccaggt ccggctggcg 781 gatggtcgga tctttgacac cgacgccgta gcctcgaacg ctgacgtggt gaacacctat 841 aaaaagctgc tcggccacca tccggtgggg cagaagcggg cggcagcgct ggagcgcaag 901 agcatgagca actcgctgtt tgtgctctac ttcggtctga accagcctca ttcccagctg 961 gcgcaccata ccatctgttt tggtccccgc taccgggagc tgatcgacga gatctttacc 1021 ggcagcgcgc tggcggatga cttctcgctc tacctgcact cgccctgcgt gaccgatccc 1081 tcgctcgcgc ctcccggctg cgccagcttc tacgtgctgg ccccggtgcc gcatcttggc 1141 aacgcgccgc tggactgggc gcaggagggg ccgaagctgc gcgaccgcat ctttgactac 1201 cttgaggagc gctatatgcc cggcctgcgt agccagctgg tgacccagcg gatctttacc 1261 ccggcagact tccacgacac gctggatgcg catctgggat cggccttctc catcgagccg 1321 ctgctgaccc aaagcgcctg gttccgcccg cacaaccgcg acagcgacat tgccaacctc 1381 tacctggtgg gcgcaggtac tcaccctggg gcgggcattc ctggcgtagt ggcctcggcg 1441 aaagccaccg ccagcctgat gattgaggat ctgcaa
[0099] A nucleic acid sequence of SEQ ID NO: 50 that is codon optimized for C. necator H16 is provided below (SEQ ID NO:51):1 atgaagaaga ccgtggtgat cggcgccggc tttggcggcc tggcgctggc gatccgcctg 61 caggcggccg gcatcccgac cgtgctgctg gaacagcggg ataagcccgg cggccgggcc 121 tacgtctggc atgatcaggg ctttaccttt gatgccggcc cgacggtgat caccgatccg 181 accgcgctgg aagcgctgtt caccctggcc ggccgccgca tggaagatta tgtccgcctg 241 ctgccggtga agcccttcta tcgcctgtgc tgggaatccg gcaagaccct ggactatgcc 301 aacgatagcg ccgaactgga agcgcagatc acccagttca acccccgcga tgtcgaaggc 361 tatcggcgct ttctggccta ctcccaggcg gtgttccagg aaggctatct gcgcctgggc 421 agcgtgccgt tcctgtcgtt tcgcgatatg ctgcgcgccg gcccgcagct gctgaagctg 481 caggcgtggc agagcgtcta ccagtcggtg tcgcgcttta tcgaagatga acatctgcgg 541 caggccttct cgttccattc cctgctggtg ggcggcaacc ccttcaccac ctcgtccatc 601 tacaccctga tccatgccct ggaacgggaa tggggcgtct ggttcccgga aggcggcacc 661 ggcgcgctgg tgaacggcat ggtgaagctg tttaccgatc tgggcggcga aatcgaactg 721 aacgcccggg tcgaggaact ggtggtggcc gataaccgcg tgagccaggt ccggctggcg 781 gatggccgga tctttgacac cgatgccgtg gcctcgaacg ccgatgtggt gaacacctat 841 aagaagctgc tgggccacca tccggtgggc cagaagcggg cggccgcgct ggagcgcaag 901 agcatgagca actcgctgtt tgtgctgtac ttcggcctga accagccgca ttcccagctg 961 gcgcaccata ccatctgctt tggcccccgc tatcgggaac tgatcgatga aatctttacc 1021 ggcagcgcgc tggcggatga tttctcgctg tatctgcatt cgccctgcgt gaccgatccc 1081 tcgctggcgc cgcccggctg cgccagcttc tatgtgctgg ccccggtgcc gcatctgggc 1141 aacgcgccgc tggattgggc gcaggaaggc ccgaagctgc gcgatcgcat ctttgattat 1201 ctggaggaac gctatatgcc cggcctgcgc agccagctgg tgacccagcg gatctttacc 1261 ccggccgatt tccatgacac gctggatgcg catctgggct cggccttctc catcgaaccg 1321 ctgctgaccc agagcgcctg gttccgcccg cacaaccgcg acagcgatat cgccaacctg 1381 tatctggtgg gcgccggcac ccatccgggc gcgggcatcc cgggcgtggt ggcctcggcg 1441 aaggccaccg ccagcctgat gatcgaagat ctgcaa
[0100] An amino acid sequence for crtl encoded by nucleic acid sequence of SEQ ID NO:51 is provided below (SEQ ID NO:52):MKKTVVIGAGFGGLALAIRLQAAGIPTVLLEQRDKPGGRAYVWHDQGFTFDAG PTVITDPTALEALFTLAGRRMEDYVRLLPVKPFYRLCWESGKTLDYANDSAELE AQITQFNPRDVEGYRRFLAYSQAVFQEGYLRLGSVPFLSFRDMLRAGPQLLKLQ AWQ S VYQ S VSRFIEDEHLRQ AF SFHSLL VGGNPFTT S SIYTLIHALEREWGVWFPEGGTGALVNGMVKLFTDLGGEIELNARVEELVVADNRVSQVRLADGRIFDTDA VASNADVVNTYKKLLGHHPVGQKRAAALERKSMSNSLFVLYFGLNQPHSQLAH HTICFGPRYRELIDEIFTGSALADDFSLYLHSPCVTDPSLAPPGCASFYVLAPVPHL GNAPLDWAQEGPKLRDRIFDYLEERYMPGLRSQLVTQRIFTPADFHDTLDAHLG S AF SIEPLLTQ S AWFRPHNRD SDIANL YL VGAGTHPGAGIPGVVAS AKATASLMI EDLQ
[0101] A nucleic acid sequence for VidS(del.91), truncated viridiflorol synthase [EC: 4.2.3.88], from Serendipita indica is provided below (SEQ ID NO: 53):1 atgccatctg tatcaccagc cactattcgt ttgcccgata tcttaggcgc tatggatcgt 61 tttgaacttc gtactcatcc cgacgaacgt gaggtcactc gtgcgtccaa cgaatggttc 121 aattcgtaca atatgatgcc tccagcgttg tttgaaaaat ttgtgaagtg tgattttggc 181 ttgatgacgg gtatgtcgta tcccgacacg gacgctaccc gtttacgcat tacgtgtgac 241 tacatgagta tcttattcgc gtatgacgat ttaatggact taccgtccag cgatttgatg 301 cacgacaaaa tcgcaagtga caaggctgct aaaattatga tgggtgtcct gactcatccc 361 cataagtttc gcccctacgc tggattgcct gtggccactg cgtttcatga tttctggact 421 cgcttttgcg caactagcac acctaaaatg caaaaacgtt tcactgacac cacttacgag 481 tacgtgatgg cagtaaaaaa tcagtgcggc aaccgccagt catcccgttg tccaacaatc 541 gaggagtacg tggcattgcg ccgtgacact agcgccatca aagtgacata tgcatgtatt 601 gaatactgcc tgaacattga tgttcccgat gaggccttct atcacccgag tgtagctgcg 661 ttgcaagagg cagggaacaa cattttaagc tgggctaatg atgtctattc attcgataac 721 gagcaatcca gtggagactg tcacaacctt gttgcaatcg tcgctattaa caaaaacatc 781 accgtacagg ctgctatgga gtatgttatg ggaatgatcg acagtgcaat tgaacgcttt 841 ttcgaggagt gcgcaaatgt gcctagtttt gggcccgaag tggatccact tgtgcaagca 901 tatattaaag gtgtggaact ttaccttagc ggctcggttt tctggcatct tgagagcgaa 961 cgttatttcg gcgcgcgtgt ccaacacgta aaggatacgt tgatggtcga attgcgccct 1021 ttggatgaag gtgcgaagcc agcattcgac ttgatgtaca aattgcccag caacctgacg 1081 cccgaagtgc tttcggcggc tgcagtttcc gcggcccctg ctgcaccagc tccagtggca 1141 tcgcccgctc cgcagccaga aattcttagt cctaccccta tcagtcctat taacgtgaac 1201 ttccccctgg ggaacgttgc atgccctccg ccctcatatg agacccaacg tgtattggca 1261 aagatggtcg eg
[0102] A nucleic acid sequence of SEQ ID NO: 53 that is codon optimized for C. necaior HI6 is provided below (SEQ ID NO:54):1 atgccgtcgg tgtcgccggc caccatccgc ctgcccgata tcctgggcgc catggatcgc 61 ttcgaactgc gcacccattc cgacgaacgc gaggtcaccc gcgcgtccaa cgaatggttc 121 aattcgtaca atatgatgcc gccggcgctg ttcgagaagt tcgtgaagtg cgacttcggc 181 ctgatgacgg gcatgtcgta tcccgacacg gacgccaccc gcctgcgcat cacgtgcgac 241 tacatgtcga tcctgttcgc gtatgacgat ctgatggacc tgccgtccag cgatctgatg 301 cacgacaaga tcgcctcgga caaggccgcc aagatcatga tgggcgtcct gacccatccc 361 cataagttcc gcccctacgc cggcctgccg gtggccaccg cgttccatga cttctggacc 421 cgcttctgcg ccaccagcac cccgaagatg cagaagcgct tcaccgacac cacctacgag 481 tacgtgatgg ccgtgaagaa tcagtgcggc aaccgccagt cgtcccgctg cccgaccatc 541 gaggagtacg tggccctgcg ccgcgacacc agcgccatca aggtgaccta tgcctgcatc 601 gaatactgcc tgaacatcga tgtgcccgat gaggccttct atcacccgtc ggtggccgcg 661 ctgcaggagg ccggcaacaa catcctgagc tgggccaatg atgtctattc gttcgataac 721 gagcagtcct cgggcgactg ccacaacctg gtggccatcg tcgccatcaa caagaacatc 781 accgtgcagg ccgccatgga gtatgtgatg ggcatgatcg actcggccat cgaacgcttc841 ttcgaggagt gcgccaatgt gccgtcgttc ggccccgaag tggatccgct ggtgcaggcc 901 tatatcaagg gcgtggaact gtacctgagc ggctcggtgt tctggcatct ggagagcgaa 961 cgctacttcg gcgcgcgcgt ccagcacgtg aaggatacgc tgatggtcga actgcgcccg 1021 ctggatgaag gcgcgaagcc ggccttcgac ctgatgtaca agctgcccag caacctgacg 1081 cccgaagtgc tgtcggcggc cgccgtgtcc gcggccccgg ccgccccggc cccggtggcc 1141 tcgcccgccc cgcagccgga gatcctgtcg ccgaccccga tctcgccgat caacgtgaac 1201 ttccccctgg gcaacgtggc ctgcccgccg ccctcgtatg agacccagcg cgtgctggcc 1261 aagatggtcg cc
[0103] An amino acid sequence for VidS(del.91) encoded by nucleic acid sequence of SEQ ID NO:54 is provided below (SEQ ID NO:55):MPSVSPATIRLPDILGAMDRFELRTHSDEREVTRASNEWFNSYNMMPPALFEKFV KCDFGLMTGMSYPDTDATRLRITCDYMSILFAYDDLMDLPSSDLMHDKIASDKA AKIMMGVLTHPHI< FRPYAGLPVATAFHDFWTRFCATSTPI< MQI< RFTDTTYEYV MA VKNQCGNRQ S SRCPTIEEYVALRRDT S AIKVT YACIEYCLNID VPDEAF YHP S VAALQEAGNNILSWANDVYSFDNEQSSGDCHNLVA1VAINKN1TVQAAMEYVM GMIDSAIERFFEECANVPSFGPEVDPLVQAYIKGVELYLSGSVFWHLESERYFGA RVQHVKDTLMVELRPLDEGAKPAFDLMYKLPSNLTPEVLSAAAVSAAPAAPAP VASPAPQPEILSPTPISPINVNFPLGNVACPPPSYETQRVLAKMVA
[0104] A nucleic acid sequence for VidS(del.91)Eus. IspA, chimeric viridiflorol synthase and IspAfromE. coli [EC:4.2.3.88 / 2.5.1.1 / 2.5.1.10] is provided below (SEQ IDNO:56): 1 atgccatctg tatcaccagc cactattcgt ttgcccgata tcttaggcgc tatggatcgt 61 tttgaacttc gtactcatcc cgacgaacgt gaggtcactc gtgcgtccaa cgaatggttc 121 aattcgtaca atatgatgcc tccagcgttg tttgaaaaat ttgtgaagtg tgattttggc 181 ttgatgacgg gtatgtcgta tcccgacacg gacgctaccc gtttacgcat tacgtgtgac 241 tacatgagta tcttattcgc gtatgacgat ttaatggact taccgtccag cgatttgatg 301 cacgacaaaa tcgcaagtga caaggctgct aaaattatga tgggtgtcct gactcatccc 361 cataagtttc gcccctacgc tggattgcct gtggccactg cgtttcatga tttctggact 421 cgcttttgcg caactagcac acctaaaatg caaaaacgtt tcactgacac cacttacgag 481 tacgtgatgg cagtaaaaaa tcagtgcggc aaccgccagt catcccgttg tccaacaatc 541 gaggagtacg tggcattgcg ccgtgacact agcgccatca aagtgacata tgcatgtatt 601 gaatactgcc tgaacattga tgttcccgat gaggccttct atcacccgag tgtagctgcg 661 ttgcaagagg cagggaacaa cattttaagc tgggctaatg atgtctattc attcgataac 721 gagcaatcca gtggagactg tcacaacctt gttgcaatcg tcgctattaa caaaaacatc 781 accgtacagg ctgctatgga gtatgttatg ggaatgatcg acagtgcaat tgaacgcttt 841 ttcgaggagt gcgcaaatgt gcctagtttt gggcccgaag tggatccact tgtgcaagca 901 tatattaaag gtgtggaact ttaccttagc ggctcggttt tctggcatct tgagagcgaa 961 cgttatttcg gcgcgcgtgt ccaacacgta aaggatacgt tgatggtcga attgcgccct 1021 ttggatgaag gtgcgaagcc agcattcgac ttgatgtaca aattgcccag caacctgacg 1081 cccgaagtgc tttcggcggc tgcagtttcc gcggcccctg ctgcaccagc tccagtggca 1141 tcgcccgctc cgcagccaga aattcttagt cctaccccta tcagtcctat taacgtgaac 1201 ttccccctgg ggaacgttgc atgccctccg ccctcatatg agacccaacg tgtattggca 1261 aagatggtcg ccgtgaggct aatggacttt ccgcagcaac tcgaagcctg cgttaagcag 1321 gccaaccagg cgctgagccg ttttatcgcc ccactgccct ttcagaacac tcccgtggtc 1381 gaaaccatgc agtatggcgc attattaggt ggtaagcgcc tgcgaccttt cctggtttat 1441 gccaccggtc atatgttcgg cgttagcaca aacacgctgg acgcacctgc tgccgccgtt1501 gagtgtatcc acgcttactc attaattcat gatgatttac cggcaatgga tgatgacgat 1561 ctgcgtcgcg gtttgccaac ctgccatgtg aagtttggcg aagcaaacgc gattctcgct 1621 ggcgacgctt tacaaacgct ggcgttctcg attttaagcg atgccgatat gccggaagtg 1681 tcggaccgcg acagaatttc gatgatttct gaactggcga gcgccagtgg tattgccgga 1741 atgtgcggtg gtcaggcatt agatttagac gcggaaggca aacacgtacc tctggacgcg 1801 cttgagcgta ttcatcgtca taaaaccggc gcattgattc gcgccgccgt tcgccttggt 1861 gcattaagcg ccggagataa aggacgtcgt gctctgccgg tactcgacaa gtatgcagag 1921 agcatcggcc ttgccttcca ggttcaggat gacatcctgg atgtggtggg agatactgca 1981 acgttgggaa aacgccaggg tgccgaccag caacttggta aaagtaccta ccctgcactt 2041 ctgggccttg agcaagcccg gaagaaagcc cgggatctga tcgacgatgc ccgtcagtcg 2101 ctgaaacaac tggctgaaca gtcactcgat acctcggcac tggaagcgct agcggactac 2161 atcatccagc gtaataaa
[0105] An amino acid sequence for VidS(del.91). Fus. IspA encoded by nucleic acid sequence of SEQ ID NO:56 is provided below (SEQ ID NO:57):MPSVSPATIRLPDILGAMDRFELRTHPDEREVTRASNEWFNSYNMMPPALFEKFV KCDFGLMTGMSYPDTDATRLRITCDYMSILFAYDDLMDLPSSDLMHDKIASDKA AKIMMGVLTHPHKFRPYAGLPVATAFHDFWTRFCATSTPKMQKRFTDTTYEYV MAVKNQCGNRQ S SRCPTIEEYVALRRDT S AIKVT YACIEYCLNID VPDEAF YHP S V A ALQE AGNN IL SW AND VYSFDNEQ S SGDCHNL VAIVAINKNITVQ AAMEYVM GMIDSAIERFFEECANVPSFGPEVDPLVQAYIKGVELYLSGSVFWHLESERYFGA RVQHVKDTLMVELRPLDEGAKPAFDLMYKLPSNLTPEVLSAAAVSAAPAAPAP VASPAPQPEILSPTPISPINVNFPLGNVACPPPSYETQRVLAKMVAATVEEKQRLA YSQPAEQYYSPAPQYYPSQPVEKFQQTNVLETAFKGSNSELTNILVIASVLMAGS PMALVPFVPLLALLLLPNETPVAPVAVRLMDFPQQLEACVKQANQALSRFIAPLP FQNTPVVETMQYGALLGGKRLRPFLVYATGHMFGVSTNTLDAPAAAVECIHAY SLIHDDLPAMDDDDLRRGLPTCHVKFGEANAILAGDALQTLAFSILSDADMPEVS DRDRISMISELASASGIAGMCGGQALDLDAEGKHVPLDALERIHRHKTGALIRAA VRLGALSAGDKGRRALPVLDKYAESIGLAFQVQDDILDVVGDTATLGKRQGAD QQLGKSTYPALLGLEQARKKARDLIDDARQSLKQLAEQSLDTSALEALADYIIQR NK
[0106] A nucleic acid sequence for dapA, 4-hydroxy-tetrahydrodipicolinate synthase [EC4.3.3.7] from E. coli K-12 MG1655: b2478 is provided below (SEQ ID NO:58): 1 atgttcacgg gaagtattgt cgcgattgtt actccgatgg atgaaaaagg taatgtctgt 61 cgggctagct tgaaaaaact gattgattat catgtcgcca gcggtacttc ggcgatcgtt 121 tctgttggca ccactggcga gtccgctacc ttaaatcatg acgaacatgc tgatgtggtg 181 atgatgacgc tggatctggc tgatgggcgc attccggtaa ttgccgggac cggcgctaac 241 gctactgcgg aagccattag cctgacgcag cgcttcaatg acagtggtat cgtcggctgc301 ctgacggtaa ccccttacta caatcgtccg tcgcaagaag gtttgtatca gcatttcaaa 361 gccatcgctg agcatactga cctgccgcaa attctgtata atgtgccgtc ccgtactggc 421 tgcgatctgc tcccggaaac ggtgggccgt ctggcgaaag taaaaaatat tatcggaatc 481 aaagaggcaa cagggaactt aacgcgtgta aaccagatca aagagctggt ttcagatgat 541 tttgttctgc tgagcggcga tgatgcgagc gcgctggact tcatgcaatt gggcggtcat 601 ggggttattt ccgttacggc taacgtcgca gcgcgtgata tggcccagat gtgcaaactg 661 gcagcagaag ggcattttgc cgaggcacgc gttattaatc agcgtctgat gccattacac 721 aacaaactat ttgtcgaacc caatccaatc ccggtgaaat gggcatgtaa ggaactgggt 781 cttgtggcga ccgatacgct gcgcctgcca atgacaccaa tcaccgacag tggtcgtgag 841 accgtcagag cggcgcttaa gcatgccggt ttgctgtaa
[0107] An amino acid sequence for dapA encoded by nucleic acid sequence of SEQ ID NO:58 is provided below (SEQ ID NO:59):MFTGSIVAIVTPMDEKGNVCRASLKKLIDYHVASGTSAIVSVGTTGESATLNHDE HADVVMMTLDLADGRIPVIAGTGANATAEAISLTQRFNDSGIVGCLTVTPYYNR PSQEGLYQHFKAIAEHTDLPQILYNVPSRTGCDLLPETVGRLAKVKNIIGIKEATG NLTRVNQIKELVSDDFVLLSGDDASALDFMQLGGHGVISVTANVAARDMAQMC KLAAEGHFAEARVINQRLMPLHNKLFVEPNPIPVKWACKELGLVATDTLRLPMT PITDSGRETVRAALKHAGLL
[0108] Assembly of a lower MVA pathway and high-value terpene module:
[0109] There are two bypass steps for the lower mevalonate pathway (Fig. 5, LMVP), the native and the modified pathway. The first enzyme in both is the mevalonate kinase (MvaKl) responsible for phosphorylation of mevalonate. Previous studies reported that MvaKl from S. aureus and the archaeon Methanosarcina mazei show the highest catalytic efficiency as well as the lowest feedback inhibition ( Voynova et al. 2004, Primak et al. 2011)
[0110] The compositions and method described herein use the native pathway. For this, the genes coding for mevalonate kinase (MvaKl), phospho-mevalonate kinase (MvK2), diphosphomevalonate decarboxylase (MvD), and isopentenyl diphosphate isomerase (IDI) from Saccharomyces aureas (S. aureus) and Saccharomyces cerevisiae (5. cerevisiae) were used to identify suitable candidates to increase IPP and DMAPP formation. Five different constructs were designed by varying the source of key candidate genes (Fig. 6). In all of these, MvaKl is from 5. aureus and IDI is from the yeast S. cerevisiae. The first and last constructs carry MvaKl, MvK2, and MvD from X aureus but differ in the order of MvK2 and MvD. The other three constructs are combinations of candidates from S. aureus and S. cerevisiae.
[0111] Previous studies demonstrated that the promoter strength for UMVP and LMVP requires balancing to support the growth of the cells and to decrease the accumulation of toxicintermediate, or metabolic byproducts (Huibin et al. 2017). In our previous work, a TolCl-promoter demonstrated an expression level in E. coli at 75% compared to the 7c7 / J-pro otcr (Aboulnaga et al. 2018), so TolCl may provide a better balance and reduced burden on the cells. Here, three different promoters (LacP, TetP, TolCl) that provide different strengths were tested as well as three different origins of replication (ColEl, pl A, BBRl-OriV) to balance the overall expression of the pathway.
[0112] A nucleic acid sequence for a TolCl -promoter is provided below (SEQ ID NO:60): TCACATGACCCGACACCATCGAATGGGATGTCGTGAGCTTACACGTTGAATCGTC ATTGACACTCTATCATTGATAGAGTATTATGGTCACAAGTCACGAAAAAAGTTCG ACCATGTTATTCCGGCAGCCCATGATCTAGAAATAATTTTGTTTAACTTTAAGAA GGAGATATACTAATG
[0113] A nucleic acid sequence for a TetP -promoter is provided below (SEQ ID NO:61): AATGGCCAGATGATTAATTCCTAATTTTTGTTGACACTCTATCATTGATAGAGTT ATTTTACCACTCCCTATCAGTGATAGAGAAAAGTGAAATGAATAGTTCGACAAA AATCTAGAAATAATTTTGTTTAACTTTAAGAAGGAGATATACAAATG
[0114] A nucleic acid sequence for a LacUV5 -promoter is provided below (SEQ ID NO:62):
[0115] tcgtttaggcaccccaggctttacactttatgcttccggctcgtataatgtgtggaattgtgagcggataacaatttcaga attcaaaagatctaaagGAGGCC ATCCTGGCC AT G
[0116] A nucleic acid sequence for a T7-promoter is provided below (SEQ ID NO:63): TCTCCCCGCGCGTTGGCCGATTCATTAATGCAGGATCTCGATCCCGCGAAATTAA TACGACTCACTATAGGGAGGCCACAACGGTTTCCCTCTAGAAATAATTTTGTTTA ACTTTAAGAAGGAGATATACAAATG
[0117] A nucleic acid sequence for a dapA-promoter is provided below (SEQ ID NO:64): TAGAAACAGAAGCCACTGATCACCAGATAATGTTGCGATGACAGTGTCAAACTG GTTATTCCTTTAAGGGGTGAGTTGTTCTTAAGGAAAGCATAAAAAAAACATGCAT ACAACAATCAGAACGGTTCTGTCTGCTTGCTTTTAATGCCATACCAAACGTACCA TTGAGACACTTGTTTGCACAGAGGATGGCCCATG
[0118] To simplify the screening, lycopene was chosen, which is a symmetrical tetraterpene providing a red color readout, to visualize production.
[0119] The chemical structure of lycopene is provided below:Exact Mass: 536.44Molecular Weight: 536.89m / z: 536.44 (100.0%), 537.44 (43.3%), 538.44 (9.1%)Elemental Analysis: C, 89.49; H, 10.51
[0120] The full pathway from acetyl-CoA to lycopene consists of 12 steps catalyzed by an 11 -enzyme cascade. The pathway was split into three different operons to facilitate controlling of the individual parts (Fig. 7A): the upper-MVA (UMVP), lower-MVA (LMVP), and lycopene pathway (Lyco). Initial screening of lycopene investigated the operons in individual plasmids with different copy number. High expression of lycopene was found when the lycopene operon was combined with the ColEl -origin (pESG-vector) under the TetP-promoter, the LMVP-operon in a pl5A-Origin (pACTR-vector), and the UMVP-operon in BBRl-OriV (pEF-BR42-vector) which is the map and sequence presented in Fig. 7B. Three independent plasmids and the respective selection represent a significant burden on the cells. To decrease the number of plasmids and antibiotics we combined both LMVP and UMVP in a pl5A-Origin (pACTR-vector). To balance the flux to lycopene (Fig. 7B) the UMVP-operon was expressed driven by the TolCl -promoter while the LMVP-operon was expressed under control of the TetP- or LacP -promoter.
[0121] E. coli carrying only the lycopene operon showed production of 10 mg / 1 lycopene (Fig. 8A). Transforming this strain with the UMVP and LMVP modules, increased lycopeneaccumulation about 20-fold using either the pPCTR-vector (LMVP under A / cL-promoter (170 mg / 1)) or pACTR-vector (LMVP under 7c / / J-promoter (200 mg / 1)). In comparison, a previous published module (pBbA5C-MevT-MBIS) (Peralta-Yahya et al. 2011) was tested, which reached 20 mg / 1 (Fig. 8B). This established the combination of the three modules for efficient production of lycopene, yet significant differences were observed between the five different combinations of the LMVP and the two induction agents.
[0122] To determine whether the LMVP operon, consuming mevalonate and providing the C5 building blocks for the lycopene formation, may represent a bottle neck, the residual mevalonate in culture media was measured (Fig. 9A). We found LMVP modules with combi or comb3, with highest lycopene production, show near complete mevalonate consumption which indicates that the flux moves efficiently through this pathway. In contrast, the other modules show significant residual accumulation of mevalonate, supporting that these are limiting lycopene formation.
[0123] Suboptimal expression of components of multi-step terpenoid pathways can lead to accumulation of intermediate metabolites, or metabolic byproducts which may be detrimental for E. coli (Ajikumar et al. 2010, Huibin et al. 2017). Early activation of the MEV pathway in E. coli was shown to suppress growth and productivity. Here, the effect of different LMVP modules and the observed mevalonate accumulation is inversely correlated with the cell growth and may provide further evidence for such inhibition (Fig. 9B). Specifically, intermediates of the LMVP such as isoprenyl diphosphates were suggested as toxic (Martin et al. 2003).
[0124] A previous study that compared the LMVP between S. aureus and Streptococcus pneumoniae in combination with the LMVP from E. faecalis, reported the lowest production of protoilludene and the highest accumulation of mevalonate when MvD and MvK2 from S. aureus were expressed (Yang et al. 2016). In contrast, the highest production was achieved when MvaKl (5. aureus), MvD and MvK2 (S. pneumoniae), and IDI from E. coli were coupled. Our data suggests a comparable capacity of MvD from both S. aureus and S. cerevisiae to accept mevalonate when the coupled HmgR is from S. aureus. Additionally, we found that MvK2 from S. cerevisiae supported optimal pathway flux in contrast to the MvK2 from S. aureus. To further investigate the impact of the order, we changed the position of the genes Mvk2 and MvD from. aureus in the LMVP operon (“S” and “comb4” modules) anddetected a significant increase in lycopene production when MvK2 is in the second rather than in the third position.
[0125] A nucleic acid sequence for a ColEl-origin of replication is provided below (SEQ ID NO: 67):TTGAGATCCTTTTTTTCTGCGCGTAATCTGCTGCTTGCAAACAAAAAAACCAC CGCTACCAGCGGTGGTTTGTTTGCCGGATCAAGAGCTACCAACTCTTTTTCCG AAGGTAACTGGCTTCAGCAGAGCGCAGATACCAAATACTGTCCTTCTAGTGT AGCCGTAGTTAGGCCACCACTTCAAGAACTCTGTAGCACCGCCTACATACCT CGCTCTGCTAATCCTGTTACCAGTGGCTGCTGCCAGTGGCGATAAGTCGTGTC TTACCGGGTTGGACTCAAGACGATAGTTACCGGATAAGGCGCAGCGGTCGGG CTGAACGGGGGGTTCGTGCACACAGCCCAGCTTGGAGCGAACGACCTACACC GAACTGAGATACCTACAGCGTGAGCTATGAGAAAGCGCCACGCTTCCCGAAG GGAGAAAGGCGGACAGGTATCCGGTAAGCGGCAGGGTCGGAACAGGAGAGC GCACGAGGGAGCTTCCAGGGGGAAACGCCTGGTATCTTTATAGTCCTGTCGG GTTTCGCCACCTCTGACTTGAGCGTCGATTTTTGTGATGCTCGTCAGGGGGGC GGAGCCTATGGAAAA
[0126] A nucleic acid sequence for a p!5A -origin of replication is provided below (SEQ ID NO: 68):TCAGCGCTAGCGGAGTGTATACTGGCTTACTATGTTGGCACTGATGAGGGTGT CAGTGAAGTGCTTCATGTGGCAGGAGAAAAAAGGCTGCACCGGTGCGTCAGC AGAATATGTGATACAGGATATATTCCGCTTCCTCGCTCACTGACTCGCTACGC TCGGTCGTTCGACTGCGGCGAGCGGAAATGGCTTACGAACGGGGCGGAGATT TCCTGGAAGATGCCAGGAAGATACTTAACAGGGAAGTGAGAGGGCCGCGGC AAAGCCGTTTTTCCATAGGCTCCGCCCCCCTGACAAGCATCACGAAATCTGA CGCTCAAATCAGTGGTGGCGAAACCCGACAGGACTATAAAGATACCAGGCGT TTCCCCTGGCGGCTCCCTCGTGCGCTCTCCTGTTCCTGCCTTTCGGTTTACCGG TGTCATTCCGCTGTTATGGCCGCGTTTGTCTCATTCCACGCCTGACACTCAGTT CCGGGTAGGCAGTTCGCTCCAAGCTGGACTGTATGCACGAACCCCCCGTTCA GTCCGACCGCTGCGCCTTATCCGGTAACTATCGTCTTGAGTCCAACCCGGAAA GACATGCAAAAGCACCACTGGCAGCAGCCACTGGTAATTGATTTAGAGGAGT TAGTCTTGAAGTCATGCGCCGGTTAAGGCTAAACTGAAAGGACAAGTTTTGGTGACTGCGCTCCTCCAAGCCAGTTACCTCGGTTCAAAGAGTTGGTAGCTCAGA GAACCTTCGAAAAACCGCCCTGCAAGGCGGTTTTTTCGTTTTCAGAGCAAGA GATTACGCGCAGACCAAAACGATCTCAAG
[0127] A nucleic acid sequence for a Fl-origin of replication is provided below (SEQ ID NO: 69):ACGCGCCCTGTAGCGGCGCATTAAGCGCGGCGGGTGTGGTGGTTACGCGCAG CGTGACCGCTACACTTGCCAGCGCCCTAGCGCCCGCTCCTTTCGCTTTCTTCC CTTCCTTTCTCGCCACGTTCGCCGGCTTTCCCCGTCAAGCTCTAAATCGGGGG CTCCCTTTAGGGTTCCGATTTAGTGCTTTACGGCACCTCGACCCCAAAAAACT TGATTAGGGTGATGGTTCACGTAGTGGGCCATCGCCCTGATAGACGGTTTTTC GCCCTTTGACGTTGGAGTCCACGTTCTTTAATAGTGGACTCTTGTTCCAAACT GGAACAACACTCAACCCTATCTCGGTCTATTCTTTTGATTTATAAGGGATTTT GCCGATTTCGGCCTATTGGTTAAAAAATGAGCTGATTTAACAAAAATTTAAC GCGAATTTTAACAAAATATTAACGCTTACAATTTC
[0128] A nucleic acid sequence for a BBRl-OriV origin of replication and its replicase is provided below (SEQ ID NO:70):CTACCGGCGCGGCAGCGTGACCCGTGTCGGCGGCTCCAACGGCTCGCCATCG TCCAGAAAACACGGCTCATCGGGCATCGGCAGGCGCTGCTGCCCGCGCCGTT CCCATTCCTCCGTTTCGGTCAAGGCTGGCAGGTCTGGTTCCATGCCCGGAATG CCGGGCTGGCTGGGCGGCTCCTCGCCGGGGCCGGTCGGTAGTTGCTGCTCGC CCGGATACAGGGTCGGGATGCGGCGCAGGTCGCCATGCCCCAACAGCGATTC GTCCTGGTCGTCGTGATCAACCACCACGGCGGCACTGAACACCGACAGGCGC AACTGGTCGCGGGGCTGGCCCCACGCCACGCGGTCATTGACCACGTAGGCCG ACACGGTGCCGGGGCCGTTGAGCTTCACGACGGAGATCCAGCGCTCGGCCAC CAAGTCCTTGACTGCGTATTGGACCGTCCGCAAAGAACGTCCGATGAGCTTG GAAAGTGTCTTCTGGCTGACCACCACGGCGTTCTGGTGGCCCATCTGCGCCAC GAGGTGATGCAGCAGCATTGCCGCCGTGGGTTTCCTCGCAATAAGCCCGGCC CACGCCTCATGCGCTTTGCGTTCCGTTTGCACCCAGTGACCGGGCTTGTTCTT GGCTTGAATGCCGATTTCTCTGGACTGCGTGGCCATGCTTATCTCCATGCGGT AGGGTGCCGCACGGTTGCGGCACCATGCGCAATCAGCTGCAACTTTTCGGCA GCGCGACAACAATTATGCGTTGCGTAAAAGTGGCAGTCAATTACAGATTTTCTTTAACCTACGCAATGAGCTATTGCGGGGGGTGCCGCAATGAGCTGTTGCGT ACCCCCCTTTTTTAAGTTGTTGATTTTTAAGTCTTTCGCATTTCGCCCTATATC TAGTTCTTTGGTGCCCAAAGAAGGGCACCCCTGCGGGGTTCCCCCACGCCTTC GGCGCGGCTCCCCCTCCGGCAAAAAGTGGCCCCTCCGGGGCTTGTTGATCGA CTGCGCGGCCTTCGGCCTTGCCCAAGGTGGCGCTGCCCCCTTGGAACCCCCGC ACTCGCCGCCGTGAGGCTCGGGGGGCAGGCGGGCGGGCTTCGCCTTCGACTG CCCCCACTCGCATAGGCTTGGGTCGTTCCAGGCGCGTCAAGGCCAAGCCGCT GCGCGGTCGCTGCGCGAGCCTTGACCCGCCTTCCACTTGGTGTCCAACCGGC AAGCGAAGCGCGCAGGCCGCAGGCCGGAGGCTTTTCCCCAGAGAAAATTAA AAAAATTGATGGGGCAAGGCCGCAGGCCGCGCAGTTGGAGCCGGTGGGTAT GTGGTCGAAGGCTGGGTAGCCGGTGGGCAATCCCTGTGGTCAAGCTCGTGGG CAGGCGCAGCCTGTCCATCAGCTTGTCCAGCAGGGTTGTCCACGGGCCGAGC GAAGCGAGCCAGCCGGTGGCCGC
[0129] Variants in sequences can occur amongst members of a species. In many cases such sequence variants still retain good enzyme activity. Enzymes described herein can have one or more deletions, insertions, replacements, or substitutions in a part of the enzyme. The enzyme(s) described herein can have, for example, at least 60%, or at least 70%, or at least 80%, or at least 90%, or at least 93%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99% sequence identity to a sequence described herein.
[0130] In some cases, enzymes can have conservative changes such as one or more deletions, insertions, replacements, or substitutions that have no significant effect on the activities of the enzymes. Examples of conservative substitutions are provided below in Table 1A.
[0131] Nucleic acids encoding the enzymes can also have sequence variations. For example, nucleic acid sequences described herein can be modified to express enzymes that do not have modifications. Most amino acids can be encoded by more than one codon. When an amino acid is encoded by more than one codon, the codons are referred to as degenerate codons.
[0132] A listing of degenerate codons is provided in Table IB below.Amino Acid Three Nucleotide CodonAla / A GCT, GCC, GCA, GCGArg / R CGT, CGC, CGA, CGG, AGA, AGGAsn / N AAT, AACAsp / D GAT, GACCys / C TGT, TGCGln / Q CAA. CAGGlu / E GAA, GAGGly / G GGT, GGC, GGA, GGGHis / H CAT, CACIle / I ATT, ATC, ATALeu / L TTA, TTG. CTT, CTC, CTA, CTGLys / K AAA, AAGMet / M ATGPhe / F TTT, TTCPro / P CCT, CCC, CCA, CCGSer / S TCT. TCC. TCA. TCG, AGT, AGCThr / T ACT, ACC, ACA, ACGTrp / W TGGTyr / Y TAT, TACVal / V GTT, GTC, GTA, GTGSTART ATG STOP TAG, TGA, TAA
[0133] Different organisms may translate different codons more or less efficiently (e.g., because they have different ratios of tRNAs) than other organisms. Hence, when some amino acids can be encoded by several codons, a nucleic acid segment can be designed to optimize the efficiency of expression of an enzyme by using codons that are preferred by an organism of interest. For example, the nucleotide coding regions of the enzymes described herein can be codon optimized for expression in various plant species.
[0134] An optimized nucleic acid can have less than 98%, less than 97%, less than 96%, less than 95%, or less than 94%, or less than 93%, or less than 92%, or less than 91%, or less than 90%, or less than 89%, or less than 88%, or less than 85%, or less than 83%, or less than80%, or less than 75% nucleic acid sequence identity to a corresponding non-optimized (e.g., a non-optimized parental or wild type enzyme nucleic acid) sequence.
[0135] The enzymes described herein can be expressed from an expression cassette and / or an expression vector. Such an expression cassette can include a nucleic acid segment that encodes an enzyme operably linked to a promoter to drive expression of the enzyme. Convenient vectors, or expression systems can be used to express such enzymes. In some instances, the nucleic acid segment encoding an enzyme is operably linked to a promoter and / or a transcription termination sequence. The promoter and / or the termination sequence can be heterologous to the nucleic acid segment that encodes an enzyme. Expression cassettes can have a promoter operably linked to a heterologous open reading frame encoding an enzyme. The invention therefore provides expression cassettes or vectors useful for expressing one or more enzyme(s).
[0136] Constructs, e.g., expression cassettes, and vectors comprising the isolated nucleic acid molecule, e.g., with optimized nucleic acid sequence, as well as kits comprising the isolated nucleic acid molecule, construct or vector are also provided.
[0137] The nucleic acids described herein can also be modified to improve or alter the functional properties of the encoded enzymes. Deletions, insertions, or substitutions can be generated by a variety of methods such as, but not limited to, random mutagenesis and / or sitespecific recombination-mediated methods. The mutations can range in size from one or two nucleotides to hundreds of nucleotides (or any value there between). Deletions, insertions, and / or substitutions are created at a desired location in a nucleic acid encoding the enzyme(s).
[0138] Nucleic acids encoding one or more enzyme(s) can have one or more nucleotide deletions, insertions, replacements, or substitutions. For example, the nucleic acids encoding one or more enzyme(s) can, for example, have less than 95%, or less than 94.8%, or less than 94.5%, or less than 94%, or less than 93.8%, or less than 94.50% nucleic acid sequence identity to a corresponding parental or wild-type sequence. In some cases, the nucleic acids encoding one or more enzyme(s) can have, for example, at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at 90% sequence identity to a corresponding parental or wild-type sequence. Examples of parental or wild type nucleic acid sequences for unmodified enzyme(s) with amino acid sequences SEQ IDNOs:19, 22, 24, 27, 29, 32, 35, 37, 40, 43, 46, 49, 52, and 55 include nucleic acid sequencesSEQ ID N0s:18, 20, 23, 25, 28, 30, 33, 36, 38, 41, 44, 47, 50, and 53, respectively. Any of these nucleic acid or amino acid sequences can, for example, encode or have enzyme sequences with less than 100%, less than 99%, less than 98%, less than 97%, less than 96%, less than 95%, less than 94.8%, less than 94.5%, less than 94%, less than 93.8%, less than 93.5%, less than 93%, less than 92%, less than 91%, or less than 90% sequence identity to a corresponding parental or wild-type sequence.
[0139] Also provided are nucleic acid molecules (polynucleotide molecules) that can include a nucleic acid segment encoding an enzyme with a sequence that is optimized for expression in at least one selected host organism or host cell. Optimized sequences include sequences which are codon optimized, i.e., codons which are employed more frequently in one organism relative to another organism. In some cases, the balance of codon usage is such that the most frequently used codon is not used to exhaustion. Other modifications can include addition or modification of Kozak sequences and / or introns, and / or to remove undesirable sequences, for instance, potential transcription factor binding sites.
[0140] An enzyme useful for synthesis of terpenes, diterpenes, diterpenoid alkaloids, and terpenoids may be expressed on the surface of, or within, a prokaryotic or eukaryotic cell. In some cases, expressed enzyme(s) can be secreted by that cell.
[0141] Techniques of molecular biology, microbiology, and recombinant DNA technology which are within the skill of the art can be employed to make and use the enzymes, expression systems, and terpene products described herein. Such techniques available in the literature. See, e.g., Sambrook, Fritsch & Maniatis, Molecular Cloning: A Laboratory Manual, Second Edition (1989); DNA Cloning, Vols. I and II (D. N. Glover ed. 1985); Oligonucleotide Synthesis (M. J. Gait ed. 1984); Nucleic Acid Hybridization (B. D. Hames & S. J. Higgins eds.1984); Animal Cell Culture (R. K. Freshney ed. 1986); Immobilized Cells and Enzymes (IRL press, 1986); Perbal, B., A Practical Guide to Molecular Cloning (1984); the series Methods In Enzymology (S. Colowick and N. Kaplan eds., Academic Press, Inc.); Current Protocols In Molecular Biology (John Wiley & Sons, Inc), Current Protocols In Protein Science (John Wiley & Sons, Inc), Current Protocols In Microbiology (John Wiley & Sons, Inc), Current Protocols In Nucleic Acid Chemistry (John Wiley & Sons, Inc), and Handbook of Experimental Immunology, Vols. I-IV (D. M. Weir and C. C. Blackwell eds., 1986, Blackwell Scientific Publications).
[0142] Promoters: The nucleic acids encoding enzymes can be operably linked to a promoter, which provides for expression of mRNA from the nucleic acids encoding the enzymes. The promoter can be a promoter functional in plants and can be a promoter functional during plant growth and development. A nucleic acid segment encoding an enzyme is operably linked to the promoter when it is located downstream from the promoter. The combination of a coding region for an enzyme operably linked to a promoter forms an expression cassette, which can optionally include other elements as well.
[0143] Promoter regions are found in the flanking DNA upstream from the coding sequence in both the prokaryotic and eukaryotic cells. A promoter sequence provides for regulation of transcription of the downstream gene sequence and can include from about 50 to about 2,000 nucleotide base pairs. Promoter sequences also contain regulatory sequences such as enhancer sequences that can influence the level of gene expression. Some isolated promoter sequences can provide for gene expression of heterologous DNAs, that is a DNA different from the native or homologous DNA.
[0144] Promoter sequences are also known to be strong or weak, or inducible. A strong promoter provides for a high level of gene expression, whereas a weak promoter provides for a very low level of gene expression. An inducible promoter is a promoter that provides for the turning gene expression on and off in response to an exogenously added agent, or to an environmental or developmental stimulus. For example, a bacterial promoter such as the Ptac promoter can be induced to varying levels of gene expression depending on the level of isopropyl-beta-D-thiogalactoside added to the transformed cells. Promoters can also provide for tissue specific or developmental regulation. An isolated promoter sequence that is a strong promoter for heterologous DNAs is advantageous because it provides for a sufficient level of gene expression for easy detection and selection of transformed cells and provides for a high level of gene expression when desired.
[0145] A nucleic acid encoding an enzyme can be combined with the promoter by standard methods to yield an expression cassette, for example, as described in Sambrook et al. (MOLECULAR CLONING: A LABORATORY MANUAL. Second Edition (Cold Spring Harbor, NY: Cold Spring Harbor Press (1989); MOLECULAR CLONING: A LABORATORY MANUAL. Third Edition (Cold Spring Harbor, NY: Cold Spring Harbor Press (2000)). Briefly, a plasmid containing a promoter such as the 35S CaMV promoter orthe CYP71D16 trichome-specific promoter can be constructed as described in Jefferson (Plant Molecular Biology Reporter 5:387405 (1987)) or obtained from Clontech Lab in Palo Alto, California (e.g., pBI121 or pBI221). These plasmids can be constructed to have multiple cloning sites having specificity for different restriction enzymes downstream from the promoter.
[0146] The nucleic acid sequence encoding for the enzyme(s) can be subcloned downstream from the promoter using restriction enzymes and positioned to ensure that the DNA is inserted in proper orientation with respect to the promoter so that the DNA can be expressed as sense RNA. Once the nucleic acid segment encoding the enzyme is operably linked to a promoter, the expression cassette so formed can be subcloned into a plasmid or other vector (e.g., an expression vector).
[0147] The following non-limiting Examples describe some procedures that can be performed to facilitate making and using the invention.EXAMPLES
[0148] Example 1: Modifications to the lycopene operon:
[0149] The data in Fig. 4C shows that the expression system described herein has the capacity to produce up to 2.6 g / 1 mevalonate which reflect the functionality of the UMVP. Despite this high mevalonate yield, lycopene yields only reached up to 200 mg / 1 (Fig. 8B).
[0150] The lycopene operon (LYC-IDI) consists of five (5) genes: three LYC-genes: geranylgeranyl pyrophosphate synthase (crtE), phytoene synthase (crtl), and lycopene synthase (crtB),' and two isopentyl diphosphate isomerases (IDF). To investigate whether the lycopene operon can be modified to produce improved levels of terpenoids, several modifications were introduced.
[0151] In embodiments, a first module carried the E. coli IspA gene in addition to the 5 genes in the lycopene operon, as shown in plasmid pESG-LYC-IDI-IspA. As IDI is already encoded in the LMVP, the additional copies were removed from the full lycopene operon. A second module contained the three genes for lycopene (crtE, cril, crtB) fused with RBSws (pESG-crtEIB). Another module carries the lycopene genes crtE crtl crtB) with E. coli IspA IpESG-crtEIB-IspA ).
[0152] The lycopene accumulation was determined with and without addition of an inducer (Figs. 10A-B). Upon induction, lycopene accumulated to nearly a gram per liter after 24 hours in the best strain carrying pACTR-TetP.Comb3-T1.MVA and pESG-crtEIB-IspA (Fig. 10B). To further evaluate these lycopene-operon modifications, we found that the lycopene module (pESG-crtEIB-IspA') not only has the highest lycopene yield but also has no impact on biomass production while the other modules have lower lycopene production and display decreased growth (Fig. 10C). To test the scalability of our system (E. coli strain carrying pACTR-TetP. Comb3-Tl. MVA and pESG-crtEIB-IspA), induction was repeated in 250 ml shake flasks with 100 ml media (at 50-volume of the previous conditions). Robust expression was found and consistent accumulation of lycopene. This is the highest lycopene production in E. coli under shake flask conditions and the first improvement of UMVP and LMVP using combination candidates from S. aureus for terpenoid production.
[0153] Modification of the inducer types and concentration:
[0154] The effects of the inducer types, AHT and DOC, as well as their concentration on lycopene production were measured in both mineral media (M9) or LB-media supplemented with 1% glycerol as a carbon source (Fig. 11A). Using the balanced system (pACTR-TetP. Comb3-Tl. MVA and pESG-crtEIB-IspA'), the improved expression of lycopene regardless of the media was achieved with a concentration between 25- 50ng / ml of AHT, or DOC between 12.5-25 ng / ml. At, or above 50 ng / ml DOC an impact on both lycopene production and cell growth was observed (Fig. 11B). This demonstrates induction of a terpenoid pathway using sensitive DOC-dependent promoters (TetP and TolCT).
[0155] Example 2: Validation of the system in other E. coli strains.
[0156] To further validate the engineered system in different E. coli strains, we tested cloning (DH5-alpha MC1061, (DH)) and expression (BL21 C41, (C41)) strains with our dual vector system (pACTR-TetP. Comb3-Tl. MVA and pESG-crtEIB-IspA). While DH showed moderate expression, C41 did not grow after induction, plausibly indicating a high burden for the cells and incompatibility with our dual vector system. Thus, we further constructed the full pathway with the three operons in a single plasmid (Fig. 12A). Three different plasmids were obtained using vectors with different copy number. Both Lyco- and LMVP-operons were controlled via TetP-promoter, while UMVP was controlled under TolC1-promoter. When we used the ColEl -dependent vector (pESG-high copy number), little growth after induction wasobserved with both strains. In contrast, both pACTR- (medium copy number) and pBBRl -vectors (low copy number) demonstrated growth and lycopene production, with the former resulting in highest yields. No residual mevalonate could be detected, indicating efficient turnover. The E. coli DH-strain exceeded lycopene compared with the C41 -strain (Fig. 12B).With this one-plasmid strategy lycopene production reached up to 75% compared with the two plasmids and demonstrated compatibility with the most used E. coli C21 expression strain.
[0157] Auxotrophic growth of E. coli:
[0158] The requirement for antibiotics for selection and plasmid maintenance is economically unfavorable compared with auxotrophic selection, i.e., strains engineered for dependence on a specific media component, or presence of a vital gene on a plasmid. Here, we targeted a key gene of the lysine biosynthesis pathway. This route forms of 2,6-diaminopimelic acid (DAP) vital for peptidoglycan biosynthesis (cell wall formation). After removing the 4-hydroxy-tetrahydrodipicolinate synthase (EC: 4.3.3.7; gene name: dapA' gene locus: b2478) we obtained a DAP dependent strain (Fig. 13A). This auxotrophic E. coli strain (Hus 10) only grows when either the media is supplemented with DAP or it carries a plasmid expressing dapA (Fig. 13B). Using this strategy, our current strain yielded 509 mg / 1 lycopene, DAP dependent, which is 71% of the yield with antibiotic selection.
[0159] Example 3: Validation of the system for other terpene production.
[0160] To validate our system for production of other terpenoids, we selected the sesquiterpene viridiflorol (C15, U. S. Patent Application Serial No.: 62 / 899,391), synthesized by the viridiflorol synthase (VidS) from Serendipita indica Ntana et al. 2021).
[0161] The chemical structure of viridiflorol (is provided below:Chemical Formula: C15H26OExact Mass: 222.20Molecular Weight: 222.37m / z: 222.20 (100.0%), 223.20 (16.2%), 224.21 (1.2%) Elemental Analysis: C, 81.02; H, 11.79; O, 7.19
[0162] The enzyme catalyzes formation of viridiflorol and viridiflorin in equal amounts. We expressed VidS alone or co-expressed with IspA in the pESG-vector (high copy number plasmid) to create the VidS-operon (Fig. 14A). When VidS was introduced into our platform strain expressing UMVP and LMVP (pACTR-LMVP-UMVP) 5.5 mg / 1 and 4.0 mg / 1 of viridiflorol and viridiflorin were reached, respectively (Fig. 14B). When we co-expressed E. coli IspA, production was increased up to 3-fold. We hypothesized that further modification and improvement of VidS, the apparent bottleneck, may resolve this apparent bottleneck.
[0163] A homology-based search against the non-redundant protein sequences at NCBI indicated that VidS carries an unusually long C-terminal extension. A previous study investigating the unrelated VidS from the fungal species Agrocybe aegerita reported that deletion of the first 82 N-terminal amino acids improved the yield 7-fold (Shukal et al. 2019}.
[0164] To guide possible truncations, we investigated the number and the place of disulfide bonds in VidS using the predictive DiANNA 1.1 web server. We detect 10 cysteines in the open reading frame of VidS (Fig. 15A) with five predicted disulfide bonds (Fig. 15B). We designed three different truncation variants of VidS (Fig. 15C). The first version had deleted the first 40-amino acids (N-A40), the second version lacked the last 91 -amino acids (C-A91), while the final version had the C-terminal 140-amino acids removed (C-A140) including Cys408.
[0165] The two versions with the A40 and AMO amino acid deletions were inactive, while the A91 of VidS improved viridiflorol production up to 6.3-fold and viridiflorin up to 7.5-fold (Fig. 16). Further engineering of VidS by translational fusion of the VidS-A91 variant with the E. coli IspA further increased the yields of viridiflorol and viridiflorin up to 20- and 11-fold,respectively (Fig. 16). With these two strategies, truncation and fusion combined, production levels of viridiflorol of 110 mg / 1 and viridiflorin up to 43 mg / 1 were reached.
[0166] Example 4: Genetically engineered Cupriavidus necator H16 as a chassis host for high-capacity lycopene production
[0167] Carotenoids involve a diverse group of natural pigments and colorants with significant commercial relevance, and which have potential antioxidant and anti-inflammatory activities. These benefits increased the market demands for carotenoids especially from microbial sources. Microorganisms naturally producing carotenoids suffer from low yield as well as their non-pure carotenoids which need more modification. Genetically engineered tools open the door for improving the quantity, quality, and diversification of carotenoid products. A promising bacterial species is Cupriavidus necator H16 (also known as Ralstonia eutropha Hl 6) that is one of the best studied model organisms for autotrophic growth on H2 and CO2.H16 also has an excellent tolerance of numerous carbon sources during heterotrophic cultivation and can reach in culture cell densities far bypassing other bacterial hosts. The mevalonate pathway is one of two known routes to carotene (terpene) precursors. In this work, high production of the key pathway intermediate S-mevalonate was achieved in Cupriavidus necator Hl 6 with the mevalonate pathway encoding a gene from Staphylococcus aureus. After screening the conditions, the engineered strain yielded up to 2.5 g / 1 mevalonate in test tubes after 72 hours with the cost-efficient carbon source vegetable oil. As H16 cannot utilize glycerol as a native carbon source, we next set out to develop a strain capable of using the industrial byproduct glycerol for production of mevalonate. Under optimal conditions, titers identical with those using oil were reached. To demonstrate feasibility of Hl 6 for mevalonate derived high value bioproducts, we implemented previous technology, i.e., a lower mevalonate pathway providing terpene precursors combined with the route to lycopene, a natural pigment and antioxidant of industrial relevance. 3-12 mg / 1 of lycopene were reached. The technology to reach this level is described below.
[0168] Two routes are currently deployed for production of carotenoids, formal chemical synthesis using petrochemistry and extraction from natural sources such as plants, algae, fungi, yeast, or bacteria (Foong et al. 2021). Microbial carotenoids are considered the most promising source of natural carotenoids. Several microorganisms are industrially relevant for carotenoids:(i) bacteria such Dietzia natronolimnaea, Paracoccus carotinifaciens, or Pseudomonas putida, (ii) yeasts like Xanthophyllomyces dendrorhous, Yarrowia lipolylica. or Pichia pastoris, (iii) filamentous fungi such as Blakeslea trispora, and (iv) algae such as Haematococcus pluvialis, Chlor el la zofingiensis, Dimaliella salina and Coelastrella striolata (Martinez-Camara et al.2021). The titer ranges from 1 mg / 1 up to 350 mg / 1 with a yield up to 90 mg / g cell dry weight (9%) after several days (Martinez-Camara et al. 2021). The major challenges posed collectively by these systems are slow growth, little biomass production, and accumulation of products in complex mixtures of unwanted byproducts, requiring their purification. Genetically engineered microorganisms promise to overcome these problems. The major two chassis hosts previously engineered for carotenoid production are Escherichia coli and Saccharomyces cerevisiae. For example, genetically modified E. coli yielded 240 mg / 1 lycopene in shake flask, and 925 mg / 1 in batch fermentation after 40 h, while yields reached 2.3 to 3.5 g / 1 in fed-batch fermentation after 5 days (Xu et al. 2018). Genetically modified S. cerevisiae afforded 310 mg / 1 lycopene in shake-flask fermentation and up to 3.3 g / 1 after 7 days in fed-batch fermentation (Shi et al. 2019). In other studies, lycopene production reached from 1.6 to 6.0 g / 1 lycopene in fed-batch fermentation after 5 to 24 days (Shi et al. 2019, Li et al. 2020). A yield up to 320 mg / 1 was achieved for astaxanthin in engineered E. coli (Zhang et al. 2018) which is comparable with industrial Haematococcus pluvialis strains (Shah et al. 2016, Hu et al. 2021, Oslan et al. 2021).
[0169] From a biotechnology perspective Cupriavidus necator H16 (further termed H16), also known as Ralstonia eutropha is highly attractive as alternative host. Hl 6 is a strictly respiratory facultative lithoautotrophic beta-proteobacterium (Cramm 2009) and one of the best studied model organisms for growth on H2 and CO2. H16 can easily adapt between autotrophic and heterotrophic lifestyles. Organic carbon and energy sources for heterotrophic growth include a broad spectrum of TCA cycle intermediate metabolites, sugar acids, plant oils and fatty acids, amino acids, alcohols, and aromatic compounds, while utilization of sugars is restricted to fructose and N-acetylglucosamine. Hl 6 produces exceptional high biomass in flask and in batch fermentation up to 112 g / 1 cell dry weight (CDW) (Huschner et al. 2015). When nitrogen or phosphorus is limited, Hl 6 can direct all the carbon into formation and sequestration of the bioplastic 3 -polyhydroxy butyrate (PHB) which accumulates up to 80% PHB / CDW (Fukui et al. 2014, Przybylski et al. 2014, Sznajder et al. 2014, Huschner et al.2015, Obruca et al. 2015). This is consistent an excess of the reducing equivalent NADPH (needed for biosynthetic processes) and no byproduct from acetyl-CoA in Hl 6. The genome sequence of H16 has been sequenced to study this unique lifestyle. This has revealed genes for several isoenzymes, permitted assignment of well-known physiological functions to previously unidentified genes, and suggested the presence of additional components of energy metabolism. The respiratory chain is fueled by two NADH dehydrogenases, two hydrogenases for H2 uptake, and at least three formate dehydrogenases. The presence of five genes encoding for quinol oxidases and three encoding for cytochrome oxidases indicates that the aerobic respiration chain adapts to varying concentrations of 02 (Schiffels et al. 2013). Meanwhile, it contains four genes encoding for cytochromes P450 (H16 B1009, H16 B1279, H16 B1743, and H16 B2406) indicating a potential for further engineering toward functionalized isoprenoids / high value isoprenoid bioproducts. Several cloning and expression systems were recently established for H16 (Gruber et al. 2014, Gruber et al. 2016, Hanko et al. 2017, Sydow et al. 2017, Aboulnaga et al. 2018, Alagesan et al. 2018), which open the door for metabolic engineering.
[0170] Hl 6 has been investigated for terpenoid production (e.g., isoprene, alpha-humulene, famesene, lycopene) (Krieg et al. 2018, Lee et al. 2019, Milker and Holtmann 2021, Milker et al. 2021, Wu et al. 2022) with little productivity in comparison with E. coli. The autotrophic production of alpha-humulene from CO2 and electric energy with engineered Hl 6 reached 17 mg / g CDW (Krieg et al. 2018), while a gram scale of alpha-humulene (2 g / 1) was produced under heterotrophic growth with fructose as a carbon source but only after extensive cultivation in fed-batch fermentation (Milker et al. 2021) with a limited conversion of carbon to products (2% of fructose to alpha-humulene). The specific production was 6.15 mg / g CDW when the three genes encoding alpha-humulene synthase, FPP synthase, and IPP isomerase were expressed. Coupling to the mevalonate pathway from Myxococcus xanthus providing precursors increased the yield to 7.4 mg / g CDW. It was suggested that the 3-hydroxymethylglutaryl-CoA reductase and the pyrophosphomevalonate decarboxylase may represent the bottleneck and that replacement of the reductase might increase the yield of alpha-humulene (Milker et al. 2021). Another study obtained 24 mg / 1 isoprene in engineered H16 via the mevalonate pathway after 60 h batch cultivation (Lee et al. 2019). These examples demonstrate that there may be ample headway for efficient production of terpenes in Hl 6 atindustrially relevant levels and that multiple strategies can be combined to adapt Hl 6 for conversion of alternative, more cost-efficient carbon sources including industrial waste products.
[0171] Cupriavidus necator H16 was genetically modified for high yield production of carotenoids via an engineered mevalonate pathway. Fig. 30A shows the LMVP and lycopene operons under the 7b / C -promoter. Fig. 30B is a plasmid map of pEF-BR42-Tl. Comb3-Tl. MVA which was constructed with the native gene codon usage and C. necator Hl 6 codon usage for the LMVP. Fig.30C shows Images of lycopene production by C. necator Hl 6 when mevalonate was supplied. Fig. 31A shows images of plated (31A) and centrifuged cultures (31B) of C. necator Hl 6 expressing the full lycopene pathway split into 3 operons (mevalonate pathway “UMVA”, intermediate pathway “LMVA”, and lycopene pathway “Lyc”). Fig. 31C is an image of centrifuged culture of C. necator Hl 6 expressing the split lycopene operon when the culture is grown in three different media (Luria Broth “LB”, minimal media “M9”, and minimal media +yeast Extract “M9Y”). Cell growth for 24 hr was good in both M9 and M9Y.
[0172] With a yield of the intermediate mevalonate reaching a commercially relevant level of 2.5 g / 1 lycopene, the engineered strain was adapted to use the industrial waste product glycerol and finally demonstrated accumulation of lycopene.
[0173] References1. Aboulnaga, E. A., H. Zou, T. Selmer and M. Xian (2018). "Development of a plasmidbased, tunable, tolC-derived expression system for application in Cupriavidus necator H16." JBiotechnol 274: 15-27.2. Ajikumar, P. K., W.-H. Xiao, K. E. J. Tyo, Y. Wang, F. Simeon, E. Leonard, O.Mucha, T. H. Phon, B. Pfeifer and G. Stephanopoulos (2010). "Isoprenoid Pathway Optimization for Taxol Precursor Overproduction in Escherichia coli." Science 330(6000): 70-74.3. Campbell, P. H., TX, US), Bredow, Sebastian (Houston, TX, US), Zhou, Huaijin (Houston, TX, US), Doneske, Stephanie (Katy, TX, US), Monticello, Daniel J. (The Woodlands, TX, US) (2014). Microorganisms and processes for the production of isoprene. United States, Glycos Biotechnologies, Inc. (Houston, TX, US).4. Chatzivasileiou, A. O., V. Ward, S. M. Edgar and G. Stephanopoulos (2019). "Two- step pathway for isoprenoid synthesis." Proc Natl Acad Sci U S A 116(2): 506-511.5. Chotani, G. K. C., CA, US), Mcauliffe, Joseph C. (Sunnyvale, CA, US), Miller, Michael C. (San Francisco, CA, US), Muir, Rachel E. (Redwood City, CA, US), Vaviline, Dmitrii V. (Palo Alto, CA, US), Weyler, Walter (San Francisco, CA, US)(2013). Isoprene production using the DXP and MVA pathway. United States, Danisco US Inc. (Palo Alto, CA, US), The Goodyear Tire & Rubber Company (Akron, OH, US).6. Guo, J., Y. Cao, H. Liu, R. Zhang, M. Xian and H. Liu (2019). "Improving the production of isoprene and 1,3 -propanediol by metabolically engineered Escherichia coli through recycling redox cofactor between the dual pathways." (1432-0614 (Electronic)).7. Huibin, Z., H. Liu, E. Aboulnaga, H. Liu, T. Cheng and M. Xian (2017). "Microbial Production of Isoprene: Opportunities and Challenges." Industrial Biotechnology: 473-504.8. Kazieva, E., Y. Yamamoto, Y. Tajima, K. Yokoyama, J. Katashkina and Y. Nishio (2017). "Characterization of feedback-resistant mevalonate kinases from the methanogenic archaeons Methanosaeta concilii and Methanocella paludicola." Microbiology (Reading) 163(9): 1283-1291.9. Kim, J.-H., C. Wang, H.-J. Jang, M.-S. Cha, J.-E. Park, S.-Y. Jo, E.-S. Choi and S.-W. Kim (2016). "Isoprene production by Escherichia coli through the exogenous mevalonate pathway with reduced formation of fermentation byproducts." Microbial Cell Factories 15(1): 214.10. Li, M., H. Chen, C. Liu, J. Guo, X. Xu, H. Zhang, R. Nian and M. Xian (2019)."Improvement of isoprene production in Escherichia coli by rational optimization of RBSs and key enzymes screening." Microb Cell Fact 18(1): 4.11. Lim, J. H., S. W. Seo, S. Y. Kim and G. Y. Jung (2013). "Model-driven rebalancing of the intracellular redox state for optimization of a heterologous n-butanol pathway in Escherichia coli." Metab Eng 20: 56-62.12. Liu, C. L., H. R. Bi, Z. Bai, L. H. Fan and T. W. Tan (2019). "Engineering and manipulation of a mevalonate pathway in Escherichia coli for isoprene production." Appl Microbiol Biotechnol 103(1): 239-250.13. Liu, C. L., H. G. Dong, J. Zhan, X. Liu and Y. Yang (2019). "Multi-modular engineering for renewable production of isoprene via mevalonate pathway in Escherichia coli." J Appl Microbiol 126(4): 1128-1139.14. Liu, H., Y. Wang, Q. Tang, W. Kong, W. J. Chung and T. Lu (2014). "MEP pathway-mediated isopentenol production in metabolically engineered Escherichia coli." Microb Cell Fact 13: 135.15. Lv, X., W. Xie, W. Lu, F. Guo, J. Gu, H. Yu and L. Ye (2014). "Enhanced isoprene biosynthesis in Saccharomyces cerevisiae by engineering of the native acetyl-CoA and mevalonic acid pathways with a push-pull-restrain strategy." J Biotechnol 186: 128-136.16. Ma, S. M., D. E. Garcia, A. M. Redding- Johanson, G. D. Friedland, R. Chan, T. S. Batth, J. R. Haliburton, D. Chivian, J. D. Keasling, C. J. Petzold, T. S. Lee and S. R. Chhabra (2011). "Optimization of a heterologous mevalonate pathway through the use of variant HMG-CoA reductases." Metab Eng 13(5): 588-597.17. Martin, V. J. J., D. J. Pitera, S. T. Withers, J. D. Newman and J. D. Keasling (2003). "Engineering a mevalonate pathway in Escherichia coli for production of terpenoids." Nature Biotechnology 21(7): 796-802.18. Miziorko, H. M. (2011). "Enzymes of the mevalonate pathway of isoprenoid biosynthesis." Arch Biochem Biophys 505(2): 131-143.19. Ntana, F., W. W. Bhat, S. R. Johnson, H. J. L. Jorgensen, D. B. Collinge, B. Jensen and B. Hamberger (2021). "A Sesquiterpene Synthase from the Endophytic Fungus Serendipita indica Catalyzes Formation of Viridiflorol." Biomolecules 11(6).20. Peralta- Yahya, P. P., M. Ouellet, R. Chan, A. Mukhopadhyay, J. D. Keasling and T. S. Lee (2011). "Identification and microbial production of a terpene-based advanced biofuel." Nat Commun 2: 483.21. Primak, Y. A., M. Du, M. C. Miller, D. H. Wells, A. T. Nielsen, W. Weyler and Z. Q. Beck (2011). "Characterization of a feedback-resistant mevalonate kinase from the archaeon Methanosarcina mazei." Appl Environ Microbiol 77(21): 7772-7778.22. Ramos, K. R., K. N. Valdehuesa, H. Liu, G. M. Nisola, W. K. Lee and W. J. Chung (2014). "Combining De Ley-Doudoroff and methylerythritol phosphate pathways for enhanced isoprene biosynthesis from D-galactose." Bioprocess Biosyst Eng 37(12): 2505-2513.23. Shukal, S., X. Chen and C. Zhang (2019). "Systematic engineering for high-yield production of viridiflorol and amorphadiene in auxotrophic Escherichia coli." Metab Eng 55: 170-178.24. Tietze, L. and R. Laie (2021). "Importance of the 5' regulatory region to bacterial synthetic biology applications." Microbial Biotechnology 14(6): 2291-2315.25. Vickers, C. E. and S. Sabri (2015). "Isoprene." Adv Biochem Eng Biotechnol.Voynova, N. E., S. E. Rios and H. M. Miziorko (2004). "Staphylococcus aureus mevalonate kinase: isolation and characterization of an enzyme of the isoprenoid biosynthetic pathway." J Bacteriol 186(1): 61-67.26. Wang, C„ M. Liwei, J. B. Park, S. H. Jeong, G. Wei, Y. Wang and S. W. Kim (2018). "Microbial Platform for Terpenoid Production: Escherichia coli and Yeast." Front Microbiol 9: 2460.27. Wang, Q., S. Quan and H. Xiao (2019). "Towards efficient terpenoid biosynthesis: manipulating IPP and DMAPP supply." Bioresources and Bioprocessing 6(1): 6.28. Xu, J., X. Xu, Q. Xu, Z. Zhang, L. Jiang and H. Huang (2018). "Efficient production of lycopene by engineered E. coli strains harboring different types of plasmids." Bioprocess Biosyst Eng 41(4): 489-499.29. Xu, Y., H. Chu, C. Gao, F. Tao, Z. Zhou, K. Li, L. Li, C. Ma and P. Xu (2014)."Systematic metabolic engineering of Escherichia coli for high-yield production of fuel bio-chemical 2, 3 -butanediol." Metab Eng 23: 22-33.30. Yang, J., M Xian, S. Su, G. Zhao, Q. Nie, X. Jiang, Y. Zheng and W. Liu (2012). "Enhancing production of bio-isoprene using hybrid MVA pathway and isoprene synthase in E. coli." PLoS One 7(4): e33509.31. Yang, J., G. Zhao, Y. Sun, Y. Zheng, X. Jiang, W. Liu and M. Xian (2012). "Bio-isoprene production using exogenous MVA pathway and isoprene synthase in Escherichia coli." Bioresour Technol 104: 642-647.32. Yang, L., C. Wang, J. Zhou and S.-W. Kim (2016). "Combinatorial engineering of hybrid mevalonate pathways in Escherichiacoli for protoilludene production." Microbial Cell Factories 15: 14.33. Yao, L., F. Qi, X. Tan and X. Lu (2014). "Improved production of fatty alcohols in cyanobacteria by metabolic engineering." Biotechnol Biofuels 7: 94.34. Yao, Z., P. Zhou, B. Su, S. Su, L. Ye and H. Yu (2018). "Enhanced Isoprene Production by Reconstruction of Metabolic Balance between Strengthened Precursor Supply and Improved Isoprene Synthase in Saccharomyces cerevisiae." ACS Synth Biol 7(9): 2308-2316.35. Zhang, C. and K. Hong (2020). "Production of Terpenoids by Synthetic Biology Approaches." Front Bioeng Biotechnol 8: 347.36. Zhao, Y., J. Yang, B. Qin, Y. Li, Y. Sun, S. Su and M. Xian (2011). "Biosynthesis of isoprene in Escherichia coli via methylerythritol phosphate (MEP) pathway." Appl Microbiol Biotechnol 90(6): 1915-1922.37. Zheng, Y., Q. Liu, L. Li, W. Qin, J. Yang, H. Zhang, X. Jiang, T. Cheng, W. Liu, X. Xu and M. Xian (2013). "Metabolic engineering of Escherichia coli for high-specificity production of isoprenol and prenol as next generation of biofuels." Biotechnol Biofuels 6:57.38. Aboulnaga, E. A., H. Zou, T. Selmer and M. Xian (2018). "Development of a plasmidbased, tunable, tolC-derived expression system for application in Cupriavidus necator Hl 6." J Biotechnol 274: 15-27.39. Alagesan, S., E. K. R. Hanko, N. Malys, M. Ehsaan, K. Winzer and N. P. Minton (2018). "Functional Genetic Elements for Controlling Gene Expression in Cupriavidus necator Hl 6." Appl Environ Microbiol 84(19).40. Cramm, R. (2009). "Genomic view of energy metabolism in Ralstonia eutropha Hl 6." J Mol Microbiol Biotechnol 16(1-2): 38-52.41. Foong, L. C., C. W. L. Loh, H. S. Ng and J. C.-W. Lan (2021). "Recent development in the production strategies of microbial carotenoids." World Journal of Microbiology and Biotechnology 37(1): 12.42. Fukui, T., M. Mukoyama, I. Orita and S. Nakamura (2014). "Enhancement of glycerol utilization ability of Ralstonia eutropha Hl 6 for production of polyhydroxyalkanoates." Appl Microbiol Biotechnol 98(17): 7559-7568.43. Gruber, S., J. Hagen, H. Schwab and P. Koefinger (2014). "Versatile and stable vectors for efficient gene expression in Ralstonia eutropha Hl 6 " J Biotechnol 186: 74-82.44. Gruber, S., D. Schwendenwein, Z. Magomedova, E. Thaler, J. Hagen, H. Schwab and P. Hei dinger (2016). "Design of inducible expression vectors for improved protein production in Ralstonia eutropha H16 derived host strains." (1873-4863 (Electronic)). 45. Hanko, E. K. R., N. P. Minton and N. Malys (2017). "Characterisation of a 3-hydroxypropionic acid-inducible system from Pseudomonas putida for orthogonal gene expression control in Escherichia coli and Cupriavidus necator." Scientific Reports 7(1): 1724.46. Hu, Q., M. Song, D. Huang, Z. Hu, Y. Wu and C. Wang (2021). "Haematococcus pluvialis Accumulated Lipid and Astaxanthin in a Moderate and Sustainable Way by the Self-Protection Mechanism of Salicylic Acid Under Sodium Acetate Stress." Frontiers in Plant Science 12.47. Huschner, F., E. Grousseau, C. J. Brigham, J. Plassmeier, M. Popovic, C. Rha and A. J. Sinskey (2015). "Development of a feeding strategy for high cell and PHA density fed-batch fermentation of Ralstonia eutropha Hl 6 from organic acids and their salts." Process Biochemistry 50(2): 165-172.48. Kalia, V. C., J. Prakash and S. Koul (2016). "Biorefinery for Glycerol Rich Biodiesel Industry Waste." Indian J Microbiol 56(2): 113-125.49. Krieg, T., A. Sydow, S. Faust, I. Huth and D. Holtmann (2018). "CO2 to Terpenes: Autotrophic and Electroautotrophic alpha-Humulene Production with Cupriavidus necator." Angew Chem Int Ed Engl 57(7): 1879-1882.50. Lee, H. W., J. H. Park, H. S. Lee, W. Choi, S. H. Seo, I. D. Anggraini, E. S. Choi and H. W. Lee (2019). "Production of Bio-Based Isoprene by the Mevalonate Pathway Cassette in Ralstonia eutropha." J Microbiol Biotechnol 29(10): 1656-1664.51. Li, M., Q. Xia, H. Zhang, R. Zhang and J. Yang (2020). "Metabolic Engineering of Different Microbial Hosts for Lycopene Production." J Agric Food Chem.52. Ma, S. M., D. E. Garcia, A. M. Redding-Johanson, G. D. Friedland, R. Chan, T. S. Batth, J. R. Haliburton, D. Chivian, J. D. Keasling, C. J. Petzold, T. S. Lee and S. R. Chhabra (2011). "Optimization of a heterologous mevalonate pathway through the use of variant HMG-CoA reductases." Metab Eng 13(5): 588-597.53. Martinez-Camara, S., A. R. Ibanez, S.; Barreiro, C.; and J.-L. Barredo (2021). "Main Carotenoids Produced by Microorganisms." Encyclopedia 1: 1223-1245.54. Milker, S. and D. Holtmann (2021). "First time beta-farnesene production by the versatile bacterium Cupriavidus necator." Microb Cell Fact 20(1): 89.55. Milker, S., A. Sydow, I. Torres-Monroy, G. Jach, F. Faust, L. Kranz, L. Tkatschuk and D. Holtmann (2021). "Gram-scale production of the sesquiterpene alpha-humulene with Cupriavidus necator." Biotechnol Bioeng 118(7): 2694-2702.56. Miziorko, H. M. (2011). "Enzymes of the mevalonate pathway of isoprenoid biosynthesis." Arch Biochem Biophys 505(2): 131-143.57. Obruca, S., P. Benesova, D. Kucera, S. Petrik and I. Marova (2015). "Biotechnological conversion of spent coffee grounds into polyhydroxyalkanoates and carotenoids." N Biotechnol.58. Oslan, S. A.-O., J. S. Tan, S. A.-O. X. Oslan, P. Matanjun, R. A.-O. Mokhtar, R. Shapawi and N. A.-O. Huda (2021). "Haematococcus pluvialis as a Potential Source of Astaxanthin with Diverse Applications in Industrial Sectors: Current Research and Future Directions. LID - 10.3390 / molecules26216470 [doi] LID - 6470." (1420-3049 (Electronic)).59. Przybylski, D., T. Rohwerder, C. Dilssner, T. Maskow, H. Harms and R. H. Muller (2014). "Exploiting mixtures of H2, CO 2, and O 2 for improved production of methacrylate precursor 2-hydroxyisobutyric acid by engineered Cupriavidus necator strains." Appl Microbiol Biotechnol 99(5): 2131-2145.60. Schiffels, J., O. Pinkenburg, M. Schelden, H. A. Aboulnaga el, M. E. Baumann and T. Selmer (2013). "An innovative cloning platform enables large-scale production and maturation of an oxygen -tolerant [NiFe]-hydrogenase from Cupriavidus necator in Escherichia coli." PLoS One 8(7): e68812.61. Shah, M. M. R., Y. Liang, J. J. Cheng and M. Daroch (2016). "Astaxanthin-Producing Green Microalga Haematococcus pluvialis: From Single Cell to High Value Commercial Products." Frontiers in Plant Science 7.62. Shi, B., T. Ma, Z. Ye, X. Li, Y. Huang, Z. Zhou, Y. Ding, Z. Deng and T. Liu (2019). "Systematic Metabolic Engineering of Saccharomyces cerevisiae for Lycopene Overproduction." J Agric Food Chem 67(40): 11148-11157.63. Sydow, A., A. Pannek, T. Krieg, I. Huth, S. E. Guillouet and D. Holtmann (2017). "Expanding the genetic tool box for Cupriavidus necator by a stabilized L-rhamnose inducible plasmid system." (1873-4863 (Electronic)).64. Sznajder, A., D. Pfeiffer and D. Jendrossek (2014). "Comparative Proteome Analysis Reveals Four Novel Polyhydroxybutyrate (PHB) Granule- Associated Proteins in Ralstonia eutropha H16." Appl Environ Microbiol 81(5): 1847-1858.65. Wu, H., H. Pan, Z. Li, T. Liu, F. Liu, S. Xiu, J. Wang, H. Wang, Y. Hou, B. Yang, L. Lei and J. Lian (2022). "Efficient production of lycopene from CO2 via microbial electrosynthesis." Chemical Engineering Journal 430: 132943.66. Xu, J., X. Xu, Q. Xu, Z. Zhang, L. Jiang and H. Huang (2018). "Efficient production of lycopene by engineered E. coli strains harboring different types of plasmids." Bioprocess Biosyst Eng 41(4): 489-499.67. Yang, J., M. Xian, S. Su, G. Zhao, Q. Nie, X. Jiang, Y. Zheng and W. Liu (2012). "Enhancing production of bio-isoprene using hybrid MVA pathway and isoprene synthase inE. coli." PLoS One 7(4): e33509.68. Zhang, C., V. Y. Seow, X. Chen and H. P. Too (2018). "Multidimensional heuristic process for high-yield production of astaxanthin and fragrance molecules in Escherichia coli." Nat Commun 9(1): 1858.
[0174] The following embodiments are intended to describe and summarize various embodiments of the invention according to the foregoing description in the specification and figures.
[0175] Embodiments:1. An expression system comprising;a first operon comprising a first promoter operably linked to a nucleic acid segment encoding a modified ribosomal binding site, a Hmg-CoA reductase (HmgR) polypeptide, a 3 -hydroxy-3 -methylglutaryl CoA synthase (HmgS) polypeptide, and a - ketothiolase (PhaA) polypeptide; anda second operon comprising a second promoter operably linked to a nucleic acid segment encoding a modified ribosomal binding site, a mevalonate kinase (MvaKl) polypeptide, a phospho-mevalonate kinase (MvK2) polypeptide, and a diphosphomevalonate decarboxylase (MvD) polypeptide.2. The expression system of embodiment 1, wherein the ribosomal binding site has a nucleic acid sequence of SEQ ID NO: 1.3. The expression system of embodiment 1, further comprising a nucleic acid segment encoding a ColEl, p!5A, BBRl-OriV origin of replication, or combination thereof.4. The expression system of embodiment 1, wherein the first promoter is further operably linked to a nucleic acid segment encoding an isopentenyl diphosphate isomerase (IDI) polypeptide.5. The expression system of embodiment 4, wherein the second operon comprises at least one of the following arrangements of the nucleic acid segments:(a) 5’ Promoter - RBSHUS - MvaKl - MvK2 - MvD - IDI 3’; or(b) 5’ Promoter - RBSHUS - MvaKl - MvD - MvK2 - IDI 3’.6. The expression system of embodiment 5, wherein the MvK2 has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO: A 30, 31, or 28, the MvD has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO: 36, 33, or 34, the MvaKl has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO: 25 or 26, and the IDI has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO: 38 or 39.7. The expression system of embodiment 1, wherein the first operon and the second operon are in separate plasmids or in the same plasmid.8. The expression system of embodiment 1, further comprising a third operon comprising a third promoter operably linked to a nucleic acid segment encoding a modified ribosomal binding site, a geranylgeranyl diphosphate synthase (crtE) polypeptide, a phytoene desaturase (crtT) polypeptide, and a phytoene synthase (crtB) polypeptide.9. The expression system of embodiment 8, wherein crtE has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO:44 or 45, crtl has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO:50 or 51, and crtB has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO:47 or 48.10. The expression system of embodiment 1, further comprising a third operon comprising a third promoter operably linked to a nucleic acid segment encoding a modified ribosomal binding site and viridiflorol synthase (VidS) polypeptide.11. The expression system of embodiment 10, wherein the third promoter is further operably linked to a nucleic acid segment encoding a geranyl diphosphate / farnesyl diphosphate synthase (IspA) polypeptide.12. The expression system of embodiment 1, wherein the first promoter and the second promoter are inducible.13. The expression system of embodiment 1, wherein the first promoter is a TolCl-promoter.14. The expression system of embodiment 1, wherein the second promoter is a TetP-promoter or a / / c / ’-promoter.15. The expression system of embodiment 14, wherein the Zc / P-promoter is inducible by anhydrotetracycline or doxycycline.16. The expression system of embodiment 14, wherein the 7c / / -repressor comprises a modified ribosomal binding site with increased doxycycline sensitivity compared to a Tb P-repressor with an unmodified ribosomal binding site.17. The expression system of embodiment 16, wherein the modified ribosomal binding site has a nucleic acid sequence of SEQ ID NO: 1.18. The expression system of embodiment 1, wherein the nucleic acid segment encoding the HmgR polypeptide is from species Enterococcus faecalis, Saccharomyces cerevisiae, Pseudomonas mevalonii, Staphylococcus aureus, Bordetella petrii, or Delftia acidovorans.19. The expression system of embodiment 1, wherein the nucleic acid segment encoding the HmgS polypeptide is from species Enterococcus faecalis.20. The expression system of embodiment 1, wherein the nucleic acid segment encoding the PhaA polypeptide is from species Ralstonia eutropha.21. The expression system of embodiment 1, wherein the nucleic acid segment encoding the MvaKl polypeptide is from species Staphylococcus aureus or Methanosarcina mazei.22. The expression system of embodiment 5, wherein the nucleic acid segment encoding the IDI polypeptide is from species Saccharomyces cerevisiae.23. The expression system of embodiment 1, wherein the nucleic acid segment encoding the MvK2 and MvD polypeptides are from species Staphylococcus aureus.24. The expression system of embodiment 22, wherein the nucleic acid segment encoding the (IDI) polypeptide had at least 90% sequence identity to SEQ ID NO:38 or 39.25. The expression system of embodiment 5, wherein the ribosomal binding site has a nucleic acid sequence of SEQ ID NO: 1 [RBSHUS].26. The expression system of embodiment 21 or 25, wherein the third operon is in a separate plasmid from the first operon and the second operon or the third operon is in the same plasmid as the first operon and the second operon.27. The expression system of embodiment 10, wherein the nucleic acid segment encoding viridiflorol synthase (VidS) polypeptide has a C-terminal truncation of 91 amino acids.28. The expression system of embodiment 27, wherein the VidS polypeptide has an amino acid sequence of SEQ ID NO: 56.29. A host cell comprising the expression system of any one of embodiments 1-28, which is heterologous to the host cell.30. The host cell of embodiment 29, wherein the host cell is a plant cell, an algae cell, a fungal cell, a bacterial cell, or an insect cell.31. The host cell of embodiment 30, wherein the host cell is Saccharomyces cerevisiae, Escherichia coli, Cupriavidus necator, Corynebacterium glutamicum, Pichia pastoris, Pseudomonas putida, or Yarrowia lipolytica.32. A method for synthesizing a terpenoid comprising incubating a host cell comprising a heterologous expression system that includes at least one expression cassette having a heterologous promoter operably linked to a nucleic acid segment encoding an enzyme with at least 90% sequence identity to SEQ ID NO:18, 20, 21, 23, 25, 26, 28, 30, 31, 33, 34, 36, 38, 39, 41, 42, 44, 45, 47, 48, 50, 51, 53, 54, 56, 58, or a combination thereof.33. The method of embodiment 32, wherein the terpenoid is viridiflorol or lycopene.34. A method for synthesizing a terpenoid comprising incubating a host cell comprising the heterologous expression system of any of embodiments 1-28.
[0176] The specific methods, devices and compositions described herein are representative of preferred embodiments and are exemplary and not intended as limitations on the scope of the invention. Other objects, aspects, and embodiments will occur to those skilled in the art upon consideration of this specification, and are encompassed within the spirit of the invention as defined by the scope of the claims. It will be readily apparent to one skilled in the art that varying substitutions and modifications may be made to the invention disclosed herein without departing from the scope and spirit of the invention.
[0177] The invention illustratively described herein suitably may be practiced in the absence of any element or elements, or limitation or limitations, which is not specifically disclosed herein as essential. The methods and processes illustratively described herein suitably may be practiced in differing orders of steps, and the methods and processes are not necessarily restricted to the orders of steps indicated herein or in the claims.
[0178] Under no circumstances may the patent be interpreted to be limited to the specific examples or embodiments or methods specifically disclosed herein. Under no circumstances may the patent be interpreted to be limited by any statement made by any Examiner or any other official or employee of the Patent and Trademark Office unless such statement is specifically and without qualification or reservation expressly adopted in a responsive writing by Applicants.
[0179] The terms and expressions that have been employed are used as terms of description and not of limitation, and there is no intent in the use of such terms and expressions to exclude any equivalent of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, it willbe understood that although the present invention has been specifically disclosed by preferred embodiments and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims and statements of the invention.
[0180] The invention has been described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also form part of the invention. This includes the generic description of the invention with a proviso or negative limitation removing any subject matter from the genus, regardless of whether or not the excised material is specifically recited herein. In addition, where features or aspects of the invention are described in terms of Markush groups, those skilled in the art will recognize that the invention is also thereby described in terms of any individual member or subgroup of members of the Markush group.
Claims
CLAIMSWhat is Claimed:
1. An expression system comprising;a first operon comprising a first promoter operably linked to a nucleic acid segment encoding a modified ribosomal binding site, aHmg-CoA reductase (HmgR) polypeptide, a 3-hydroxy-3-methylglutaryl CoA synthase (HmgS) polypeptide, and a P-ketothiolase (PhaA) polypeptide; anda second operon comprising a second promoter operably linked to a nucleic acid segment encoding a modified ribosomal binding site, a mevalonate kinase (MvaKl) polypeptide, a phospho-mevalonate kinase (MvK2) polypeptide, and a diphosphomevalonate decarboxylase (MvD) polypeptide.
2. The expression system of claim 1, wherein the ribosomal binding site has a nucleic acid sequence of SEQ ID NO: 1.
3. The expression system of claim 1, further comprising a nucleic acid segment encoding a ColEl, pl5A, BBRl-OriV origin of replication, or combination thereof.
4. The expression system of claim 1, wherein the first promoter is further operably linked to a nucleic acid segment encoding an isopentenyl diphosphate isomerase (IDI) polypeptide.
5. The expression system of claim 4, wherein the second operon comprises at least one of the following arrangements of the nucleic acid segments:(a) 5’ Promoter - RBSHUS - MvaKl - MvK2 - MvD - IDI 3’; or(b) 5’ Promoter - RBSHUS - MvaKl - MvD - MvK2 - IDI 3’.
6. The expression system of claim 5, wherein the MvK2 has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO: A 30, 31, or 28, the MvD has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO: 36, 33, or 34, the MvaKl has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO: 25 or 26, and the IDI has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO: 38 or 39.
7. The expression system of claim 1, wherein the first operon and the second operon are in separate plasmids or in the same plasmid.
8. The expression system of claim 1, further comprising a third operon comprising a third promoter operably linked to a nucleic acid segment encoding a modified ribosomal binding site, a geranylgeranyl diphosphate synthase (crtE) polypeptide, a phytoene desaturase (crtl) polypeptide, and a phytoene synthase (crtB) polypeptide.
9. The expression system of claim 8, wherein crtE has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO:44 or 45, crtl has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO:50 or 51, and crtB has a nucleic acid sequence with at least 90% sequence identity to SEQ ID NO:47 or 48.
10. The expression system of claim 1, further comprising a third operon comprising a third promoter operably linked to a nucleic acid segment encoding a modified ribosomal binding site and viridiflorol synthase (VidS) polypeptide.
11. The expression system of claim 10, wherein the third promoter is further operably linked to a nucleic acid segment encoding a geranyl diphosphate / famesyl diphosphate synthase (IspA) polypeptide.
12. The expression system of claim 1, wherein the first promoter and the second promoter are inducible.
13. The expression system of claim 1, wherein the first promoter is a / b / C / -promoter.
14. The expression system of claim 1, wherein the second promoter is a ' / b / P-promoter or a ZczcP-promoter.
15. The expression system of claim 14, wherein the Ze / Z-promoter is inducible by anhydrotetracycline or doxycycline.
16. The expression system of claim 14, wherein the Ze / Z-repressor comprises a modified ribosomal binding site with increased doxycycline sensitivity compared to a 7bzZ-repressor with an unmodified ribosomal binding site.
17. The expression system of claim 16, wherein the modified ribosomal binding site has a nucleic acid sequence of SEQ ID NO: 1.
18. The expression system of claim 1, wherein the nucleic acid segment encoding the HmgR polypeptide is from species Enterococcus faecalis, Saccharomyces cerevisiae, Pseudomonas mevalonii, Staphylococcus aureus, Bordetella petrii, or Delftia acidovorans.
19. The expression system of claim 1, wherein the nucleic acid segment encoding the HmgS polypeptide is from species Enterococcus faecalis.
20. The expression system of claim 1, wherein the nucleic acid segment encoding the PhaA polypeptide is from species Ralstonia eutropha.
21. The expression system of claim 1, wherein the nucleic acid segment encoding the MvaKl polypeptide is from species Staphylococcus aureus or Methanosarcina mazei.
22. The expression system of claim 5, wherein the nucleic acid segment encoding the IDI polypeptide is from species Saccharomyces cerevisiae.
23. The expression system of claim 1, wherein the nucleic acid segment encoding the MvK2 and MvD polypeptides are from species Staphylococcus aureus.
24. The expression system of claim 22, wherein the nucleic acid segment encoding the (IDI) polypeptide had at least 90% sequence identity to SEQ ID NO:38 or 39.
25. The expression system of claim 5, wherein the ribosomal binding site has a nucleic acid sequence of SEQ ID NO: 1 [RBSHUS].
26. The expression system of claims 21 or 25, wherein the third operon is in a separate plasmid from the first operon and the second operon or the third operon is in the same plasmid as the first operon and the second operon.
27. The expression system of claim 10, wherein the nucleic acid segment encoding viridiflorol synthase (VidS) polypeptide has a C-terminal truncation of 91 amino acids.
28. The expression system of claim 27, wherein the VidS polypeptide has an amino acid sequence of SEQ ID NO: 56.
29. A host cell comprising the expression system of any one of claims 1-28, which is heterologous to the host cell.
30. The host cell of claim 29, wherein the host cell is a plant cell, an algae cell, a fungal cell, a bacterial cell, or an insect cell.
31. The host cell of claim 30, wherein the host cell is Saccharomyces cerevisiae, Escherichia coli, Cupriavidus necator, Corynebacterium glutamicum, Pichia pastoris, Pseudomonas pntida, or Yarrowia lipolytica.
32. A method for synthesizing a terpenoid comprising incubating a host cell comprising a heterologous expression system that includes at least one expression cassette having a heterologous promoter operably linked to a nucleic acid segment encoding an enzyme with at least 90% sequence identity to SEQ ID NO: 18, 20, 21, 23, 25, 26, 28, 30, 31, 33, 34, 36, 38, 39, 41, 42, 44, 45, 47, 48, 50, 51, 53, 54, 56, 58, or a combination thereof.
33. The method of claim 32, wherein the terpenoid is viridiflorol or lycopene.
34. A method for synthesizing a terpenoid comprising incubating a host cell comprising the heterologous expression system of any of claims 1-28.