Recombinant microorganisms for βeta-myrcene and nerol production
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BP CORP NORTH AMERICA INC
- Filing Date
- 2025-12-04
- Publication Date
- 2026-07-30
AI Technical Summary
Existing methods for producing β-myrcene and nerol using recombinant microorganisms face inefficiencies due to interference with endogenous biochemical pathways and the use of relatively inefficient terpene synthase enzymes, leading to low yields.
Engineering recombinant microorganisms to express neryl diphosphate synthase (NPPS) and myrcene synthase (MyrS) or nerol synthase (NerS) polypeptides, derived from plant species, to catalyze the production of β-myrcene and nerol from DMAPP and IPP via neryl diphosphate intermediates, and enhancing flux through the DXP or MVA pathways.
This approach enhances the production efficiency and yield of β-myrcene and nerol by bypassing the competition for substrates with endogenous pathways, utilizing alternative pathways that improve the recombinant microorganisms' ability to produce these compounds.
Smart Images

Figure US2025058190_30072026_PF_FP_ABST
Abstract
Description
Attorney Docket No. BPC-026WO RECOMBINANT MICROORGANISMS FOR BETA-MYRCENE AND NEROL PRODUCTION 1. CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the priority benefit of United States provisional application no.63 / 728,226, filed on December 5, 2024, the contents of which are incorporated herein in its entirety.2. SEQUENCE LISTING
[0002] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML Sequence Listing, created on November 14, 2025, is named BPC-026WO_SL.xml and is 160,282 bytes in size.3. BACKGROUND
[0003] p-myrcene is a commercial isoprenoid (also known as terpene, monoterpene, and monoterpenoid) compound with many uses. For example, p-myrcene can be used for the synthesis of flavors like menthol, geraniol, nerol and linalool. It can also be used to make products for a variety of applications, including polymers, pharmaceuticals, insect repellents, flavors and fragrances, vitamins, biodegradable surfactants, fuels, and lubricants.
[0004] p-myrcene can be extracted from natural sources or produced by pyrolysis of p-pinene obtained from turpentine. In the interest of sustainability and economic efficiency, it would be desirable to produce p-myrcene and other isoprenoids from renewable resources.
[0005] Nerol is another commercially important isoprenoid that has several uses, for example, as a fragrance ingredient and a food preservative. It is also useful as a chemical building block for other useful compounds and can be converted into other isoprenoids like carveol, carvone, limonene, para-cymene, and a-terpinene.
[0006] Attempts to produce isoprenoids using renewable resources have included engineering recombinant microorganisms to convert geranyl diphosphate (GPP) into isoprenoids. However, GPP is used as a substrate for endogenous biochemical pathways in the most common microorganisms used for fermentation processes, and engineered pathways for production of isoprenoids therefore interfere with endogenous pathways, limiting the efficiency and ultimate yield of isoprenoid production. In addition, in some cases biosynthetic pathways used for isoprenoid production used relatively inefficient terpene synthase enzymes, leading to low titers.
[0007] Accordingly, there is a need in the art for improved and sustainable methods for p-myrcene and nerol production.4. SUMMARY
[0008] The present disclosure addresses this need and provides novel recombinant microorganisms engineered to produce p-myrcene or nerol.
[0009] In certain aspects, the present disclosure provides recombinant microorganisms engineered to express a neryl diphosphate synthase (NPPS) polypeptide and a myrcene synthase (MyrS) polypeptide. In some embodiments, the NPPS and MyrS polypeptides are based on or derived from plant species. As shown herein, these polypeptides are capable of effectively catalyzing the production of β-myrcene from DMAPP and IPP via a neryl diphosphate (NPP) intermediate.
[0010] The present disclosure also provides recombinant microorganisms engineered to express an NPPS polypeptide and a nerol synthase (NerS) polypeptide. As shown herein, these polypeptides are capable of effectively catalyzing the production of nerol from DMAPP and IPP via an NPP intermediate.
[0011] Examples of recombinant microorganisms of the present disclosure are described in Section 6.2 and numbered embodiments 1 to 175 and 180 to 182.
[0012] Recombinant microorganisms can be further engineered enhance flux through the 1-deoxy-D-xylulose 5-phosphate (DXP) pathway (also known as the MEP pathway) or the mevalonate pathway (MVA) to HMBPP, e.g., as described in Sections 6.5, 6.6, and 6.7 and numbered embodiments 167 to 171.
[0013] Examples of NPPS polypeptides that can be expressed by the recombinant microorganisms of the present disclosure are described in Section 6.3.1 and numbered embodiments 2 to 11.
[0014] Examples of MyrS polypeptides that can be expressed by the recombinant microorganisms of the present disclosure are described in Section 6.3.2 and numbered embodiments 12 to 70.
[0015] Examples of NerS polypeptides that can be expressed by the recombinant microorganisms of the present disclosure are described in Section 6.4.1 and numbered embodiments 71 to 144.
[0016] Recombinant microorganisms of the present disclosure can be engineered from parental microorganisms using various methods. Examples of parental microorganisms and methods of engineering recombinant microorganisms of the present disclosure therefrom are described in Section 6.9 and numbered embodiments 161 to 166 and 172.
[0017] Recombinant microorganisms of the present disclosure can be used to produce |3-myrcene or nerol by culturing the recombinant microorganisms in appropriate culture media and under appropriate conditions. Examples of culture media and culture conditions for culturing recombinant microorganisms of the present disclosure are described in Section 6.10 and numbered embodiment 177. Examples of the production of p-myrcene and / or products thereof by recombinant microorganisms of the present disclosure are described in Section 6.11 and numbered embodiments 176 to 179. Examples of the production of nerol and / or products thereof by recombinant microorganisms of the present disclosure are described in Section 6.12 and numbered embodiments 183 to 186.5. BRIEF DESCRIPTION OF THE FIGURES
[0018] FIG. 1 depicts two alternative pathways for production of -myrcene from DMAPP and IPP by recombinant microorganisms.
[0019] FIG. 2 shows production of p-myrcene by recombinant E. coli strains engineered to include sequences encoding the indicated MyrS polypeptide and either GPPS or NPPS, as described in Example 1. AgMyrS: Abies grandis MyrS; AmMyrS: Antirrhinum majus MyrS; HIMyrS: Humulus lupulus MyrS; ObMyrS: Ocimum basilicum MyrS; PaMyrS: Picea abis MyrS; PfMyrS: Perilla frutescens MyrS; QiMyrS: Quercus ilex MyrS.
[0020] FIG. 3 shows production of p-myrcene, limonene, or nerol by recombinant E. coli strains engineered to include sequences encoding NPPS and the indicated MyrS polypeptide. QiMyrS: Quercus ilex MyrS; CsMyrS: Cannabis sativa MyrS.
[0021] FIG. 4 shows a diagram of the 2 gene operon designed to express terpene synthases and the neryl pyrophosphate (NPP) synthase.
[0022] FIG. 5 schematically depicts the DXP pathway, in context with pathways leading from glucose or xylose to DXP, and with pathways leading from DMAPP and IPP to isoprene and / or isoprenoids, with certain enzymes assigned numbers and certain reactants and products assigned letters for ease of reference and convenience. Abbreviations and / or assignments used: glucose (A); gluconate (B); KDG (C), 2-keto-3-deoxygluconate; KDGP (D), 2-keto-3-deoxy-6-phosphogluconate; GAP (E), glyceraldehyde-3-phosphate; pyruvate (F); ribulose-5-phosphate (G); DXP (H), 1-deoxyxylulose-5-phosphate; MEP (J), 2-C-methylerythritol 4-phosphate; CDP-ME (K), 4-diphosphocytidyl-2-C-methylerythritol; CDP-MEP (L), 4-diphosphocytidyl-2-C-methyl-D-erythritol 2-phosphate; MEcPP (M), 2-C-methyl-D-erythritol 2,4-cyclodiphosphate; HMBPP (N), (E)-4-Hydroxy-3-methyl-but-2-enyl pyrophosphate; DMAPP (P), dimethylallyl pyrophosphate; IPP (Q), isopentenyl pyrophosphate; xylose (R); xylulose (S); xylulose-5-phosphate (T); 1 -deoxyxylulose (DX) (U); CTP, cytidine triphosphate; CMP, cytidine monophosphate; Dxs (1), 1-deoxy-d-xylulose-5-phosphate synthase (EC 2.2.1.7); Dxr (2), 1-deoxy-D-xylulose 5-phosphate reductoisomerase (EC 1.1.1.267); IspD (3), 2-C-methyl-D-erythritol 4-phosphate cytidylyltransferase (EC 2.7.7.60); IspE (4), 4-(cytidine 5'-diphospho)-2-C-methyl-D-erythritol kinase (EC 2.7.1.148); IspF (5), 2-C-methyl-D-erythritol 2,4-cyclodiphosphate synthase (EC 4.6.1.12); IspG (6), 4-hydroxy-3-methylbut-2-en-1-yl diphosphate synthase (EC 1.17.7.1); IspH (7), 4-Hydroxy-3-methylbut-2-enyl diphosphate reductase (EC 1.17.1.2); Idi (8), isopentenyl-diphosphate Delta-isomerase (EC 5.3.3.2); PEC (9), a protein electron carrier (e.g., ferredoxin or flavodoxin), in reduced (red) and oxidized (ox) forms; Fpr (10), flavodoxin / ferredoxin--NADP reductase EC 1.19.1.1 or EC 1.18.1.2), in reduced (red) and oxidized (ox) forms; RibB, 3,4-dihydroxy-2-butanone 4-phosphate synthase (EC 4.1.99.12); YajO, 1-deoxyxylulose-5-phosphate synthase (EC 1.1.-.-); and XylB, xylulose kinase (EC 2.7.1.17). The use of multi-headed arrows (e.g.,indicates multiple enzymatic activities are involved in converting the substrate (or one or more of multiple substrates) to one or more of the reactants in the depicted step. Not all enzymatic activities are shown, and single-headed arrows can be indicative of multi-step processes.6. DETAILED DESCRIPTION6.1. Definitions
[0023] Unless otherwise defined herein, scientific and technical terms used in connection with the present disclosure shall have the meanings that are commonly understood by those of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Throughout this specification and embodiments, the words “have" and “comprise,” or variations such as “has,” “having,” “comprises,” or “comprising,” will be understood to imply the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers. All publications and other references mentioned herein are incorporated by reference in their entirety. Although a number of documents are cited herein, this citation does not constitute an admission that any of these documents forms part of the common general knowledge in the art.
[0024] Corresponding amino acid residues: An amino acid residue of a query amino acid sequence “corresponds” to an amino acid residue of a reference amino acid sequence when, upon alignment of the query and reference sequences for optimal comparison purposes (e.g., gaps can be introduced in one or both of the query and the reference sequence for optimal alignment), the amino acid residue in the query sequence is aligned with the amino acid residue in the reference sequence. An alignment of a query amino acid sequence and a reference amino acid sequence can be generated using the computerprogram ClustalW (version 1.83, default parameters), which allows alignments of polypeptide sequences to be carried out across their entire length (global alignment).ClustalW calculates the best match between a query and one or more reference sequences and aligns them so that identities, similarities and differences can be determined. Gaps of one or more residues can be inserted into a query sequence, a reference sequence, or both, to maximize sequence alignments. For fast pairwise alignment of amino acid sequences, the following parameters are used: word size: 1; window size: 5; scoring method: percentage; number of top diagonals: 5; gap penalty: 3.
[0025] Expression system: An “expression system” refers to a combination of the following operably linked components: (a) one or more regulatory elements ora regulatory system; and (b) and one or more coding nucleotide sequences.
[0026] From: A coding sequence may be referred to herein as being “from” an organism if the coding sequence and a polypeptide encoded thereby have a high degree of sequence identity and functional similarity to the coding sequence and encoded polypeptide as isolated from the organism. A coding sequence “from” an organism does not have to directly obtained from the organism. A coding sequence “from” an organism also does not have to have 100% sequence identity at the nucleotide level, nor does the encoded polypeptide have to have 100% sequence identity at the polypeptide level, to coding sequence and encoded polypeptide as isolated from the organism. A coding sequence “from” an organism encompasses variants, truncated versions, etc. Typically, a coding sequence “from” an organism encodes a polypeptide having at least 90% sequence identity to that organism's native polypeptide. In some embodiments, the sequence identity to the organism’s native polypeptide is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%.
[0027] Heterologous: As used herein, the term “heterologous,” when used to describe a first element in reference to a second element indicates that the first element and second element do not exist in nature disposed as described. For example, a heterologous nucleic acid molecule, construct, or sequence refers to (a) a polypeptide, nucleic acid molecule or portion of a polypeptide or nucleic acid molecule sequence that is not native to a cell in which it is expressed, (b) a polypeptide or nucleic acid molecule or portion of a polypeptide or nucleic acid molecule that has been altered or mutated relative to its native state, (c) a polypeptide or nucleic acid molecule with an altered expression as compared to the native expression levels under similar conditions, or (d) any combination of two or all of (a), (b) and (c). For example, a heterologous regulatory sequence (e.g., promoter, enhancer) can be used to regulate expression of a coding sequence in a way that is different than how the coding sequence is normally expressed in nature. In certain embodiments, a heterologousnucleic acid molecule may exist in a native host cell genome but may have an altered expression level or have a different sequence or both. In other embodiments, heterologous nucleic acid molecules may not be endogenous to a host cell or host genome but instead may have been introduced into a host cell by transformation, wherein the added molecule may integrate into the host genome or can exist as extra-chromosomal genetic material either transiently or semi-stably for more than one generation (e.g., episomal vector, plasmid or other self-replicating vector).
[0028] Myrcene Synthase (MyrS): As used herein, a “myrcene synthase” or “MyrS” refers to a polypeptide that catalyzes the conversion of neryl diphosphate (NPP) or geranyl diphosphate (GPP) to p-myrcene, including fragments and / or variants of parental or full-length MyrS polypeptides that are capable of catalyzing the conversion of NPP to or GPP to |3-myrcene.
[0029] Neryl Diphosphate Synthase (NPPS): As used herein, a “neryl diphosphate synthase,” or “NPPS,” refers to a polypeptide that catalyzes the condensation of dimethylallyl diphosphate (DMAPP) and isopentenyl diphosphate (IPP) to form neryl diphosphate (NPP) (EC 2.5.1.28), including fragments and / or variants of parental or full-length NPPS polypeptides that are capable of catalyzing the condensation of DMAPP and IPP to form NPP.
[0030] Nerol Synthase (NerS): As used herein, “nerol synthase,” or “NerS,” refers to a polypeptide that catalyzes the conversion of neryl diphosphate (NPP) to nerol, including fragments and / or variants of parental or full-length NerS polypeptides that are capable of catalyzing the conversion of NPP to nerol.
[0031] Nucleic Acid: The term “nucleic acid” is used herein interchangeably with the term “polynucleotide” and refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single- or double-stranded form.
[0032] Operably Linked: The term “operably linked,” when used to describe the relationship between a first nucleic acid or nucleotide sequence and a second nucleic acid or nucleotide sequence, indicates that the first nucleic acid or nucleotide sequence is placed in a functional relationship with the second nucleic acid or nucleotide sequence. For instance, a promoter or enhancer is operably linked to a coding sequence if the promoter or enhancer affects the transcription or expression of the coding sequence. Operably linked DNA sequences may be contiguous or non-contiguous. Where necessary to join two proteincoding sequences, operably linked sequences may be in the same reading frame.
[0033] Operon: An “operon” as used herein refers to a nucleic acid sequence encoding multiple coding sequences which are transcribed in a single transcript. The coding sequences of the operon thus share regulatory sequences that are 5’-ward of the most upstream coding sequence (which may be termed “operon upstream regulatory sequences”) and 3’-ward of the most downstream coding sequence (which may be termed “operon downstream regulatory sequences”).
[0001] Parental Cell, Parental Microorganism: The terms “parental cell” or “parental microorganism” are used interchangeably to refer to unicellular organisms from which a recombinant microorganism can be derived by one or more engineering steps, even if the recombinant microorganism is not directly obtained through such engineering steps. The adjective “parental” indicates that a recombinant cell or recombinant microorganism can be engineered by the introduction into a parental cell or parental microorganism of a heterologous nucleic acid or plurality of heterologous nucleic acids, such as nucleic acid(s) each comprising a coding region or plurality of coding regions each encoding a heterologous polypeptide, and / or by insertion, deletion, substitution, or other modification of coding regions or regulatory sequences in the genome of the parental microorganism.
[0034] A parental microorganism can be a microorganism found in nature or a microorganism that is non-naturally occurring. In other words, a parental microorganism can comprise one or more genetic modifications (e.g., insertion, deletion, or modification of one or more coding regions and / or regulatory sequences) relative to a strain thereof found in nature. In relationship to a recombinant microorganism of the disclosure generated through a series of engineering steps, the terms “parental cell” and “parental microorganism” can refer to an ancestral cell or organism incorporating any of the engineering steps, as well as a cell or microorganism without any of the engineering steps. Sometimes, for ease of reference and comparison, the terms “parental cell” and “parental microorganism” refer to a cell or microorganism which, if having genetic modifications, the genetic modification(s) do not relate to one or more aspects of the microorganism engineering described in Section 6.2. Further, the term “parental cell” and “parental microorganism” is intended for use as a reference cell or microorganism and does not imply that the cell or organism was used as a starting point for engineering a microorganism of the disclosure.
[0035] Polypeptide, Peptide, Protein: The terms “polypeptide,” “peptide” and “protein” are used interchangeably herein to refer to polymers of amino acids of any length. A polypeptide herein may be identified by a name or by a percentage of sequence identity to a reference amino acid sequence. When a polypeptide is identified by a name indicative of an activityperformed or enabled by the polypeptide, the name refers to any polypeptide capable of performing or enabling the activity.
[0036] Promoter: A “promoter” as used herein refers to a nucleic acid sequence which is capable of interacting with an RNA polymerase such that transcription of a sequence of interest begins. A typical prokaryotic promoter includes a -35 sequence (a region of about 6 nucleotides, the 5’ end of which is located from 30 to 40 nucleotides, such as 35 nucleotides, upstream ( / .e., 5’-ward) of the transcription start site) and a -10 sequence, also known as a Pribnow box (a region of about 6 nucleotides, the 5’ end of which is located from 5 to 15 nucleotides, such as 10 nucleotides, upstream of the transcription initiation site). A prokaryotic promoter typically has from 12 to 22 nucleotides, and in some embodiments 17 ± 3 (e.g., 14, 15, 16, 17, 18, 19, or 20 nucleotides), intervening between the -35 sequence and the -10 sequence. A promoter may include at least a portion of a repressor binding site and / or an activator binding site.
[0037] Recombinant Microorganism: The terms “recombinant cell’’ and “recombinant microorganism” are used interchangeably to refer to a cell that has been genetically engineered. It should be understood that this term refers not only to the particular subject cell but to the progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, a recombinant counterpart of a parental cell or parental microorganism includes progeny that are not identical to the initial recombinant cell or microorganism engineered from the parent cell or parental microorganism, but are still included within the scope of the terms “recombinant cell” or “recombinant microorganism” as used herein.
[0038] Redox polypeptide: A “redox polypeptide” is used herein to refer to any polypeptide capable of directly (via a single step) or indirectly (via more than one step) transferring electrons from NADPH to other polypeptides, e g., to the iron-sulfur clusters of iron-sulfur cluster polypeptides. In some embodiments, a redox polypeptide is capable of transferring electrons to ispG and / or ispH. Examples of redox polypeptides include ferredoxins, flavodoxins, and flavodoxin / ferredoxin-NADP reductases (fpr; EC:1.18.1.2).
[0039] Regulatory Element, Regulatory Sequence: A “regulatory element” or “regulatory sequence” as used herein refers to non-coding sequences that influence the expression (e.g., transcription or translation) of a transcribed sequence. Regulatory sequences include different types of regulatory elements such as promoters, operator regions, terminator sequences, intergenic sequences encoding small regulatory RNAs (sRNAs), Shine-Dalgarno (SD) sequences, etc.
[0040] Sequence Identity: “Sequence identity” in relation to nucleotide or amino acid sequence of a nucleic acid or polypeptide molecule, refers to the overall relatedness between two such sequences. Calculation of the percent sequence identity (nucleotide or amino acid sequence identity) of two sequences, for example, can be performed by aligning the two sequences for optimal comparison purposes (e.g., gaps can be introduced in one or both of a first and a second nucleic acid or amino acid sequence for optimal alignment). The nucleotides or amino acids at corresponding positions are then compared. When a position in the first sequence is occupied by the same nucleotide or amino acid as the corresponding position in the second sequence, then the molecules are identical at that position. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps, and the length of each gap, which needs to be introduced for optimal alignment of the two sequences. Percent sequence identity can be determined manually once an alignment of nucleotide or amino acid sequences is generated. An alignment of query nucleotide or amino acid sequence and a reference nucleotide or amino acid sequence can be generated using the computer program ClustalW (version 1.83, default parameters), which allows alignments of nucleic acid or protein sequences to be carried out across their entire length (global alignment). ClustalW calculates the best match between a query and one or more reference sequences and aligns them so that identities, similarities and differences can be determined. Gaps of one or more residues can be inserted into a query sequence, a reference sequence, or both, to maximize sequence alignments. For fast pair wise alignment of nucleotide sequences, the following default parameters are used: word size: 2; window size: 4; scoring method: percentage; number of top diagonals: 4; and gap penalty: 5. For fast pairwise alignment of amino acid sequences, the following parameters are used: word size: 1; window size: 5; scoring method: percentage; number of top diagonals: 5; gap penalty: 3. Unless indicated otherwise, the percent sequence identity between a reference nucleotide or amino acid sequence (e.g. a sequence with a defined SEQ ID NO as disclosed herein) and a query nucleotide or amino acid sequence is calculated across the entire length of the reference sequence.
[0041] Transformation: The term “transformation” refers to the introduction of nucleic acid molecules into cells. In the context of the present disclosure, the term “transformation” encompasses any method known to the skilled person for introducing nucleic acid molecules into cells, such as into bacterial or fungal cells. Such methods encompass, for example, electroporation, calcium phosphate precipitation, or nanoparticle-based transformation, among other techniques known to the person of ordinary skill in the art having the benefit of the present disclosure.
[0042] Wild-type: The term “wild-type” as used herein to describe a microorganism species or strain refers to a defined species or strain, e.g., as deposited with a depositary such as the American Type Culture Collection (Manassas, Virginia). When describing a nucleic acid or polypeptide, “wild-type” indicates the nucleic acid or polypeptide has a sequence identical to that of the corresponding nucleic acid or polypeptide in a wild-type species or strain. An exemplary microorganism strain that is sometimes referenced herein as a “wild-type” strain is E. coli K12 substrain MG 1655. Another “wild-type” E. coli strain is E. coli K12 substrain BW25113. Exemplary Saccharomyces cerevisiae “wild-type” strains include HAO and W303 strains.6.2. Recombinant Microorganisms Engineered to Express NPPS and MyrS or NerS
[0043] The present disclosure relates to recombinant microorganisms engineered to express NPPS and MyrS, which can be used, for example, for production of p-myrcene. In some embodiments, the recombinant microorganisms comprise one or more heterologous nucleic acids comprising nucleotide sequences encoding NPPS and MyrS polypeptides.
[0044] The present disclosure also relates to recombinant microorganisms engineered to express NPPS and NerS polypeptides, which can be used, for example, for production of nerol. In some embodiments, the recombinant microorganisms comprise one or more heterologous nucleic acids comprising nucleotide sequences encoding NPPS and NerS polypeptides.
[0045] The recombinant microorganisms can be any unicellular organisms, such as archaea, bacteria, or fungi, among others. In some embodiments, the bacteria are E. coli. In some embodiments, the E. coli are of strains MG1655, W3110, DH5alpha, W, or BL21. In some embodiments, the microorganism is a yeast. In some embodiments, the yeast is Saccharomyces cerevisiae.
[0046] The recombinant microorganisms can comprise one or more copies (e.g., one, two, three, four, or more copies, such as 1-6 copies, 1-10 copies, or 11-20 copies) of an NPPS and / or MyrS coding sequence or an NPPS and / or NerS coding sequence. Where one or more NPPS, MyrS, and / or NerS coding sequences are present in the recombinant microorganisms, the NPPS, MyrS, and / or NerS coding sequences can be the same or different. In some embodiments, a recombinant microorganism comprises more than one copy of the same NPPS coding sequence. In some embodiments, a recombinant microorganism comprises more than one copy of the same MyrS coding sequence. In some embodiments, a recombinant microorganism comprises more than one copy of the same NerS coding sequence.
[0047] In some embodiments, the recombinant microorganism comprises one or more NPPS coding sequences integrated into its genome. In some embodiments, genomic integrations of one or more NPPS coding sequences are at positions that leave unchanged the expression or activity of native genes and polypeptides encoded by coding sequences thereof. In some embodiments, genomic integrations of one or more NPPS coding sequences change the expression or activity of one or more native genes and / or polypeptides encoded by coding sequences thereof. In some embodiments, a change in expression or activity of a native gene comprises replacement of a native gene or a coding sequence thereof with an NPPS coding sequence.
[0048] In some embodiments, the recombinant microorganism comprises one or more MyrS coding sequences integrated into its genome. In some embodiments, genomic integrations of one or more MyrS coding sequences are at positions that leave unchanged the expression or activity of native genes and polypeptides encoded by coding sequences thereof. In some embodiments, genomic integrations of one or more MyrS coding sequences change the expression or activity of one or more native genes and / or polypeptides encoded by coding sequences thereof. In some embodiments, a change in expression or activity of a native gene comprises replacement of a native gene or a coding sequence thereof with a MyrS coding sequence.
[0049] In some embodiments, the recombinant microorganism comprises one or more NerS coding sequences integrated into its genome. In some embodiments, genomic integrations of one or more NerS coding sequences are at positions that leave unchanged the expression or activity of native genes and polypeptides encoded by coding sequences thereof. In some embodiments, genomic integrations of one or more NerS coding sequences change the expression or activity of one or more native genes and / or polypeptides encoded by coding sequences thereof. In some embodiments, a change in expression or activity of a native gene comprises replacement of a native gene or a coding sequence thereof with a NerS coding sequence.
[0050] When the NPPS, MyrS, and / or NerS coding sequences are integrated into the genome, in some embodiments, the recombinant microorganisms comprise one, two, three, four, five, or six copies of the NPPS, MyrS, and / or NerS coding sequences.
[0051] Additionally or alternatively, the recombinant microorganism may comprise one or more NPPS and / or MyrS coding sequences on one or more extrachromosomal nucleic acids. In some embodiments, the recombinant microorganism comprises one or more extrachromosomal nucleic acids comprising one or more NPPS coding sequences and one or more MyrS coding sequences. In some embodiments, the recombinant microorganismcomprises multiple extrachromosomal nucleic acids having different coding sequences. In some embodiments, the recombinant microorganism comprises a first extrachromosomal nucleic acid comprising one or more NPPS coding sequences and a second extrachromosomal nucleic acid comprising one or more MyrS coding sequences.
[0052] Additionally or alternatively, the recombinant microorganism may comprise one or more NPPS and / or NerS coding sequences on one or more extrachromosomal nucleic acids. In some embodiments, the recombinant microorganism comprises one or more extrachromosomal nucleic acids comprising one or more NPPS coding sequences and one or more NerS coding sequences. In some embodiments, the recombinant microorganism comprises multiple extrachromosomal nucleic acids having different coding sequences. In some embodiments, the recombinant microorganism comprises a first extrachromosomal nucleic acid comprising one or more NPPS coding sequences and a second extrachromosomal nucleic acid comprising one or more NerS coding sequences.
[0053] In some embodiments, extrachromosomal nucleic acids disclosed herein may be used to generate recombinant microorganisms having genomically integrated NPPS and / or MyrS coding sequences. In addition to the desired coding sequences, such extrachromosomal nucleic acids may include nucleotide sequences that provide for genomic integration of NPPS and / or MyrS, for example by homologous recombination, site-specific recombination, or transposon-mediated gene transposition.
[0054] In some embodiments, extrachromosomal nucleic acids disclosed herein may be used to generate recombinant microorganisms having genomically integrated NPPS and / or NerS coding sequences. In addition to the desired coding sequences, such extrachromosomal nucleic acids may include nucleotide sequences that provide for genomic integration of NPPS and / or NerS, for example by homologous recombination, site-specific recombination, or transposon-mediated gene transposition.
[0055] In some embodiments, the one or more extrachromosomal nucleic acids comprise a plasmid or a portion thereof.
[0056] In some embodiments, the one or more extrachromosomal nucleic acids comprise a bacterial artificial chromosome (BAC) or a portion thereof.
[0057] In some embodiments, the one or more extrachromosomal nucleic acids comprise a yeast artificial chromosome (YAC) or a portion thereof.
[0058] Plasmids, BACs, and YACs can be selected to incorporate NPPS coding sequence(s), MyrS coding sequence(s), and / or NerS coding sequences based at least in part on compatibility with the recombinant microorganism, the inclusion of selection markers,stability over multiple generations of the recombinant microorganism, and / or ease of insertion of the coding sequence therein, among other parameters that will be known to skilled persons.
[0059] When the NPPS, MyrS, and / or NerS coding sequences are on extrachromosomal nucleic acids, in some embodiments, the recombinant microorganisms comprise one, two, three, four, five, six, seven, eight, nine, ten, eleven, or twelve copies of the NPPS, MyrS, and / or NerS coding sequences. In some embodiments, when the NPPS, MyrS, and / or NerS coding sequences are on extrachromosomal nucleic acids, the recombinant microorganisms comprise eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, twenty, 21, or 22 copies of the NPPS, MyrS, and / or NerS coding sequences.
[0060] Suitable regulatory sequences, including promoters, to regulate expression of NPPS and / or MyrS coding sequences are described in Section 6.8.
[0061] In addition to being engineered to express NPPS and MyrS polypeptides or to express NPPS and NerS polypeptides, the recombinant microorganisms of the disclosure can be engineered to enhance flux through the DXP pathway or the MVA pathway.Engineering a microorganism to increase flux through the DXP and / or MVA pathways is described in Section 6.5. Further engineering to enhance performance of the DXP pathway is described in Section 6.6 and Section 6.7.
[0062] Specific examples of NPPS and MyrS polypeptides which can be expressed by recombinant microorganisms engineered for that purpose are described in Sections 6.3.1 and 6.3.2, respectively.
[0063] Specific examples of NPPS and NerS polypeptides which can be expressed by recombinant microorganisms engineered for that purpose are described in Sections 6.3.1 and 6.4.1, respectively.
[0064] Recombinant microorganisms can be engineered from parental microorganisms, such as those described in Section 6.9, using methods such as those described in Section 6.9.1.
[0065] Recombinant microorganisms, such as those described in this section, can be used to produce p-myrcene. Recombinant microorganisms can be cultured in media as described in Section 6.10.1 and under conditions as described in Section 6.10.2. p-myrcene and products thereof can be produced as described in Section 6.11.6.3. Engineering of -Myrcene Production Pathway
[0066] Previous approaches for engineering microorganisms to produce p-myrcene rely on biosynthetic pathways that use GPP as a substrate (see for example Kim E., et al., 2015, J. Agric. Food Chem. 63:4606-4612). Microorganisms useful for synthetic isoprenoid production, for example E. coli and S. cerevisiae, typically rely on GPP as a substrate in native biochemical pathways necessary to support growth. Engineered pathways for isoprenoid production may interfere with the native GPP-dependent biosynthetic pathways, and competition for use of GPP by endogenous pathways can reduce efficiency and yield of engineered p-myrcene production pathways.
[0067] FIG. 1 depicts two alternative pathways for producing p-myrcene in engineered microorganisms. In the pathway labeled “GPP Pathway,” a heterologous MyrS enzyme uses GPP as a substrate for production of p-myrcene. FIG. 1 also shows that GPP is a substrate for essential cellular components and can be a substrate for promiscuous phosphatases, leading to production of geraniol. Therefore, production of p-myrcene via this pathway competes with endogenous pathways for use of GPP. The alternative pathway labeled in FIG. 1 as the “NPP pathway,” does not rely on GPP for production of p-myrcene, but instead relies on NPP. Microorganisms can be engineered to produce p-myrcene via the NPP pathway by introducing heterologous npps and myrS genes encoding NPPS and MyrS enzymes, respectively.6.3.1. NPPS Polypeptides
[0068] Recombinant microorganisms of the disclosure are engineered to express one or more heterologous NPPS polypeptides. NPPS polypeptides catalyze the condensation of DMAPP and IPP to form NPP, which can then act as a substrate for production of p-myrcene catalyzed by MyrS.
[0069] In some embodiments, the NPPS polypeptide comprises a plant NPPS. In some embodiments, a plant NPPS polypeptide and / or coding sequence thereof is modified to facilitate expression in the recombinant microorganism. Expression of heterologous plant enzymes in microbes in some cases requires removal of amino acid sequences from the N-terminal region of the plant enzyme such as, for example, a transit peptide. The nucleotide sequence of a nucleic acid encoding a plant NPPS may also be codon optimized for expression in the particular microorganism based on well-known techniques.
[0070] Exemplary NPPS polypeptides suitable for use in the recombinant microorganisms and p-myrcene production methods of the disclosure include, for example, NPPS polypeptides native to Solanum lycopersicum. In some embodiments, the native NPPS amino acid sequence is altered to remove N-terminal presequences such as, for example,transit peptides and / or signal peptides. In some embodiments, the native NPPS coding sequence is modified by being codon-optimized for expression in the recombinant microorganism. In some embodiments, the NPPS coding sequence is operably linked to a heterologous promoter.
[0071] In some embodiments, the NPPS is an S. lycopersicum NPPS. SEQ ID NO:2 is the amino acid sequence of an S. lycopersicum NPPS with an N-terminal signal peptide removed. In some embodiments, the NPPS comprises or consists of the amino acid sequence of SEQ ID NO:2. In some embodiments, the NPPS comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:2. In some embodiments, the NPPS is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:12.
[0072] In some embodiments, recombinant microorganisms of the present disclosure comprise a means for condensing DMAPP and IPP to form NPP. Examples of polypeptides capable of condensing DMAPP and IPP to form NPP include the NPPS polypeptides disclosed in this Section 6.3.1. In some embodiments, recombinant microorganisms of the present disclosure are configured to condense DMAPP and IPP to form NPP.6.3.2. MyrS Polypeptides
[0073] In some embodiments, recombinant microorganisms of the disclosure are engineered to express heterologous MyrS polypeptides. In some embodiments, MyrS polypeptides catalyze the conversion of NPP to p-myrcene. The catalytic activity of a MyrS polypeptide for the conversion NPP to p-myrcene can be assayed according to the procedure of Example 1 below.
[0074] Exemplary MyrS polypeptides suitable for use in the recombinant microorganisms and p-myrcene production methods of the disclosure include, for example, MyrS polypeptides native to plant species. In some embodiments, the MyrS is native to Antirrhinium majus, Ocimum basilicum, Picea abis, Quercus ilex, or Cannabis sativa. In some embodiments, the native MyrS amino acid sequence is altered to remove N-terminal presequences such as, for example, transit peptides and / or signal peptides. In some embodiments, the native NPPS coding sequence is modified by being codon-optimized for expression in the recombinant microorganism. In some embodiments, the NPPS coding sequence is operably linked to a heterologous promoter. In some embodiments, recombinant microorganisms of the present disclosure comprise a means for converting NPP to p-myrcene. Examples of polypeptides capable of converting NPP to p-myrceneinclude the MyrS polypeptides disclosed in this Section 6.3.2. In some embodiments, recombinant microorganisms of the present disclosure are configured to convert NPP to β-myrcene.
[0075] The catalytic activity of a MyrS polypeptide can be assayed according to the procedure of Example 1 below. In addition to detecting p-myrcene production directly, efficiency of the p-myrcene production pathway can also be monitored by detecting and quantifying nerol. In the absence of heterologous nerol synthases, nerol is produced from NPP due to activity of endogenous E. coli enzymes. Accumulation of nerol in such strains indicates that NPP is not being efficiently converted to p-myrcene by the MyrS polypeptide. Thus, nerol concentrations are inversely proportional to p-myrcene production.
[0076] In some embodiments, the MyrS is an A. majus MyrS. SEQ ID NO:4 is the amino acid sequence of wild-type A. majus MyrS with an N-terminal signal peptide removed. In some embodiments, the MyrS comprises or consists of the amino acid sequence of SEQ ID NO:4. In some embodiments, the MyrS comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:4 or a functional fragment thereof. In some embodiments, the MyrS is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:14.
[0077] In some embodiments, the MyrS is an O. basilicum MyrS. SEQ ID NO:6 is the amino acid sequence of wild-type O. basilicum MyrS with an N-terminal signal peptide removed. In some embodiments, the MyrS comprises or consists of the amino acid sequence of SEQ ID NO:6. In some embodiments, the MyrS comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:6 or a functional fragment thereof. In some embodiments, the MyrS is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:16.
[0078] In some embodiments, the MyrS is an P. abis MyrS. SEQ ID NO:7 is the amino acid sequence of wild-type P. abis MyrS with an N-terminal signal peptide removed. In some embodiments, the MyrS comprises or consists of the amino acid sequence of SEQ ID NO:7. In some embodiments, the MyrS comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:7 or a functional fragment thereof. In some embodiments, the MyrS is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%,80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:17.
[0079] In some embodiments, the MyrS is a Q. ilex MyrS. SEQ ID NO:9 is the amino acid sequence of wild-type Q. ilex MyrS with an N-terminal signal peptide removed. In some embodiments, the MyrS comprises or consists of the amino acid sequence of SEQ ID NO:9. In some embodiments, the MyrS comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:9 or a functional fragment thereof. In some embodiments, the MyrS is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:19.
[0080] In some embodiments, the MyrS is a C. sativa MyrS. SEQ ID NO: 10 is the amino acid sequence of wild-type C. sativa MyrS with an N-terminal signal peptide removed. In some embodiments, the MyrS comprises or consists of the amino acid sequence of SEQ ID NO: 10. In some embodiments, the MyrS comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:10 or a functional fragment thereof. In some embodiments, the MyrS is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:20.
[0081] In some embodiments, the MyrS is an A. grandis MyrS. SEQ ID NO:3 is the amino acid sequence of wild-type A. grandis MyrS with an N-terminal signal peptide removed. In some embodiments, the MyrS comprises or consists of the amino acid sequence of SEQ ID NO:3. In some embodiments, the MyrS comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:3 or a functional fragment thereof. In some embodiments, the MyrS is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:13.
[0082] In some embodiments, the MyrS is an H. lupulus MyrS. SEQ ID NO:5 is the amino acid sequence of wild-type H. lupulus MyrS with an N-terminal signal peptide removed. In some embodiments, the MyrS comprises or consists of the amino acid sequence of SEQ ID NO:5. In some embodiments, the MyrS comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:5 or a functional fragment thereof. In some embodiments, the MyrS isencoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:15.
[0083] In some embodiments, the MyrS is a P. frutescens MyrS. SEQ ID NO:8 is the amino acid sequence of wild-type P. frutescens MyrS with an N-terminal signal peptide removed. In some embodiments, the MyrS comprises or consists of the amino acid sequence of SEQ ID NO:8. In some embodiments, the MyrS comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:8 or a functional fragment thereof. In some embodiments, the MyrS is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:18.
[0084] MyrS polypeptides used in recombinant microorganisms and methods of the present disclosure may be referred to by names other than myrcene synthase or MyrS. Certain enzymes referred to herein as MyrS polypeptides have been known by names other than myrcene synthase or MyrS. In some instances, this is due to terpene synthase enzymes conventionally being named based on the major product produced when using GPP as a substrate. However, the present disclosure shows that some enzymes surprisingly produce a different terpene (e.g., 0-myrcene) as the major product when using NPP as a substrate.
[0085] In some embodiments, the MyrS is an enzyme from Antirrhinum majus known as 0-ocimene synthase (AmOcis; GenBank: AAO42614.1). SEQ ID NO:33 is the amino acid sequence of wild-type AmOcis. In some embodiments, the MyrS comprises or consists of the amino acid sequence of SEQ ID NO:32. In some embodiments, the MyrS comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:32 or a functional fragment thereof. In some embodiments, the MyrS is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:47.
[0086] In some embodiments, the MyrS is an enzyme from Phaseolus lunatus known as p-ocimene synthase (PIQcis; GenBank: ABY65110.1). SEQ ID NO:35 is the amino acid sequence of wild-type PlOcis. In some embodiments, the MyrS comprises or consists of the amino acid sequence of SEQ ID NO:35. In some embodiments, the MyrS comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:35 or a functional fragment thereof. In some embodiments, the MyrS is encoded by a nucleotide sequence comprising or consistingof a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:51.
[0087] In some embodiments, recombinant microorganisms of the present disclosure comprise a means for converting NPP to p-myrcene. Examples of polypeptides capable of converting NPP to 0-myrcene include the MyrS polypeptides disclosed in this Section 6.3.2. In some embodiments, recombinant microorganisms of the present disclosure are configured to condense DMAPP and IPP to form NPP.6.4. Engineering of Nerol Production Pathway
[0088] Production of nerol can be advantageously performed using NPP as a substrate in recombinant microorganisms that comprise heterologous genes encoding an NPPS enzyme and a NerS enzyme.
[0089] The NPPS included in recombinant microorganisms of the disclosure for engineered for production of nerol can be any of the NPPS polypeptides disclosed in Section 6.3.1 and can be encoded by any of the NPPS-encoding genes disclosed in Section 6.3.1. Any of the NPPS polypeptides and / or NPPS-encoding genes disclosed in Section 6.3.1 can be combined with any of the NerS polypeptides or NerS-encoding genes disclosed in Section 6.4.1.6.4.1. NerS Polypeptides
[0090] In some embodiments, recombinant microorganisms of the disclosure are engineered to express heterologous NerS polypeptides. In some embodiments, NerS polypeptides catalyze the conversion of NPP to nerol. The catalytic activity of a NerS polypeptide for the conversion NPP to nerol can be assayed according to the procedure of Example 2 below. NerS polypeptides used in recombinant microorganisms and methods of the present disclosure may be referred to by names other than nerol synthase or NerS. Certain enzymes referred to herein as NerS polypeptides, i.e., polypeptides that catalyze conversion of NPP to nerol, have been known by names other than nerol synthase or NerS. In some instances, this is due to terpene synthase enzymes conventionally being named based on the major product produced when using GPP as a substrate. However, the present disclosure shows that some enzymes surprisingly produce a different terpene (e.g., 0-myrcene or nerol) as the major product when using NPP as a substrate.
[0091] In some embodiments, the NerS polypeptide is an enzyme that produces nerol as the most abundant product when NPP is used as the substrate. In some embodiments, the NerS polypeptide produces 50, 55, 60, 70, 75, 80, 85, 90, or 95 mol% nerol using NPP as a substrate relative to all terpenes produced. In some embodiments, the NerS polypeptideproduces 50 to 70 mol%, 50 to 90 mol%, 60 to 80 mol%, or 60 to 90 mol% nerol using NPP as a substrate relative to all terpenes produced. The catalytic activity of a NerS polypeptide can be assayed according to the procedure of Example 2 below.
[0092] Exemplary NerS polypeptides suitable for use in the recombinant microorganisms and nerol production methods of the disclosure include, for example, a Spatholobus suberectus nerolidol synthase (SsNerS), a Cajanus cajan nerolidol synthase (CcNerS), a Cinnamomum tenuipile geraniol synthase (CtGerS), a Perilla citriadora geraniol synthase (PcGerS), a Camptotheca acuminata geraniol synthase (CaGerS), an Arabidopsis thaliana b-Ocimene Synthase (AtOciS), a Citrus unshiu b-Ocimene Synthase (CuOciS), or a Heliconius melpomene b-Ocimene Synthase (HmOciS).
[0093] In some embodiments, the NerS polypeptide is an enzyme from Spatholobus suberectus known as (3S,6E)-nerolidol synthase 1 (SsNerS; GenBank: TKY71585.1). SEQ ID NO:23 is the amino acid sequence of wild-type SsNerS. In some embodiments, the NerS polypeptide comprises or consists of the amino acid sequence of SEQ ID NO:23. In some embodiments, the NerS polypeptide comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:23 or a functional fragment thereof. In some embodiments, the NerS polypeptide is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:38.
[0094] In some embodiments, the NerS polypeptide is an enzyme from Cajanus cajan known as (3S,6E)-nerolidol synthase 1 (CsNerS; GenBank: XP_020234024.1). SEQ ID NO:24 is the amino acid sequence of wild-type CsNerS. In some embodiments, the NerS polypeptide comprises or consists of the amino acid sequence of SEQ ID NO:24. In some embodiments, the NerS polypeptide comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:24 or a functional fragment thereof. In some embodiments, the NerS polypeptide is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:39.
[0095] In some embodiments, the NerS polypeptide is an enzyme from Cinnamomum tenuipile known as geraniol synthase (CtGerS; GenBank: CAD29734.2). SEQ ID NO:25 is the amino acid sequence of wild-type CtGerS. In some embodiments, the NerS polypeptide comprises or consists of the amino acid sequence of SEQ ID NO:25. In some embodiments, the NerS polypeptide comprises or consists of an amino acid sequence having at least 80%,85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:25 or a functional fragment thereof. In some embodiments, the NerS polypeptide is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:40.
[0096] In some embodiments, the NerS polypeptide is an enzyme from Perilla citriadora known as geraniol synthase (PcGerS; GenBank: Q4JHG3). SEQ ID NO:26 is the amino acid sequence of wild-type PcGerS. In some embodiments, the NerS polypeptide comprises or consists of the amino acid sequence of SEQ ID NO:26. In some embodiments, the NerS polypeptide comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:26 or a functional fragment thereof. In some embodiments, the NerS polypeptide is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:41.
[0097] In some embodiments, the NerS polypeptide is an enzyme from Camptotheca acuminata known as geraniol synthase (CaGerS; GenBank: ALL56347.1). SEQ ID NO:27 is the amino acid sequence of wild-type CaGerS. In some embodiments, the NerS polypeptide comprises or consists of the amino acid sequence of SEQ ID NO:27. In some embodiments, the NerS polypeptide comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:27 or a functional fragment thereof. In some embodiments, the NerS polypeptide is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:42.
[0098] In some embodiments, the NerS polypeptide is an enzyme from Arabidopsis thaliana known as p-ocimene synthase (AtOciS; UniProtKB: A4FVP2.1). SEQ ID NO:33 is the amino acid sequence of wild-type AtOciS. In some embodiments, the NerS polypeptide comprises or consists of the amino acid sequence of SEQ ID NO:33. In some embodiments, the NerS polypeptide comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:33 or a functional fragment thereof. In some embodiments, the NerS polypeptide is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:48.
[0099] In some embodiments, the NerS polypeptide is an enzyme from Citrus unshiu known as p-ocimene synthase (CuOciS; GenBank: BAD91046.1). SEQ ID NO:34 is the amino acid sequence of wild-type CuOciS. In some embodiments, the NerS polypeptide comprises or consists of the amino acid sequence of SEQ ID NO:34. In some embodiments, the NerS polypeptide comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:34 or a functional fragment thereof. In some embodiments, the NerS polypeptide is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:49.
[0100] In some embodiments, the NerS polypeptide is an enzyme from Heliconius melpomene known as p-ocimene synthase (HmOciS; https: / / doi.org / 10.1371 / journal.pbio.3001022). SEQ ID NO:35 is the amino acid sequence of wild-type HmOciS. In some embodiments, the NerS polypeptide comprises or consists of the amino acid sequence of SEQ ID NO:35. In some embodiments, the NerS polypeptide comprises or consists of an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to the amino acid sequence of SEQ ID NO:35 or a functional fragment thereof. In some embodiments, the NerS polypeptide is encoded by a nucleotide sequence comprising or consisting of a nucleotide sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% sequence identity to the nucleotide sequence of SEQ ID NO:50.
[0101] In some embodiments, recombinant microorganisms of the present disclosure comprise a means for converting NPP to nerol. Examples of polypeptides capable of converting NPP to nerol include the NerS polypeptides disclosed in this Section 6.4.1. In some embodiments, recombinant microorganisms of the present disclosure are configured to condense DMAPP and IPP to form NPP.6.5. DXP & MVA Pathway Engineering
[0102] As can be seen from FIG. 1, NPPS polypeptides act on DMAPP and IPP to form NPP. In some microorganisms (e.g., E. coli), DMAPP and IPP are produced by an endogenous 1-deoxyxylulose-5-phosphate (DXP) pathway. In some microorganisms (e.g., Saccharomyces cerevisiae) DMAPP and IPP are produced by an endogenous mevalonate (MVA) pathway. The production of isoprenoids (e.g., p-myrcene and nerol) from DMAPP and IPP catalyzed by NPPS and terpene synthase (e.g., MyrS) polypeptides of the present disclosure can be increased by increasing flux through the DXP pathway or MVA pathway relative to a native microorganism comprising only native DXP pathway or MVA pathway enzymes expressed under native regulation.6.5.1. Heterologous DXP Pathway
[0103] In some embodiments, recombinant microorganisms of the present disclosure comprise an engineered DXP pathway comprising one or more heterologous DXP pathway polypeptides. FIG. 5 shows the DXP pathway in the context of other enzymatic pathways. The reference numbers in the description of the DXP pathway and DXP pathway polypeptides below refer to the numbers in FIG. 5.6.5.1.1. Dxs Polypeptides
[0104] Recombinant microorganisms of the present disclosure may comprise a first DXP pathway nucleotide sequence encoding a dxs polypeptide (1).
[0105] In some embodiments, a dxs polypeptide (1) is heterologous with respect to the parental microorganism and comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:53. In some embodiments, a dxs polypeptide (1) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:53. In some embodiments, a dxs polypeptide (1) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:53. In some embodiments, a dxs polypeptide (1) comprises an amino acid sequence having 100% sequence identity to the amino acid sequence of SEQ ID NO:53.
[0106] Dxs polypeptides can be obtained or derived from any source organism. In some embodiments, the source organism (1) produces at least twice the amount of pigments as E. coli on an average pigmentbiomass weight ratio; (2) is photosynthetic; (3) is distantly related to E. coli', (4) is characterized by any combination of two or all three of (1), (2) and (3) (e.g., (1) + (2), (1) + (3), (2) + (3), or (1) + (2) + (3)). Though not to be bound by theory, a dxs polypeptide from such a source organism may be less susceptible than E. coli dxs to regulation by E. coli and / or has greater DXP pathway activity than that of E. coli.
[0107] In various embodiments, a distantly related source organism is not a member of at least one of (a) order Enterobacterales: (b) class Gammaproteobacteria-, (c) phylum Pseudomonadota: and (d) domain Bacteria.
[0108] In some embodiments, dxs polypeptides (1) comprise amino acid sequences having less than 65% sequence identity to the amino acid sequence of SEQ ID NO:53. Dxs polypeptides (1) of these embodiments can comprise amino acid sequences having at least 35% sequence identity to the amino acid sequence of SEQ ID NO:53. Dxs polypeptides (1) of these embodiments can comprise amino acid sequences having at least 40% sequence identity to the amino acid sequence of SEQ ID NO:53. Dxs polypeptides (1) of theseembodiments can comprise amino acid sequences having at least 45% sequence identity to the amino acid sequence of SEQ ID NO:53.
[0109] Exemplary dxs polypeptides (1) able to complement Adxs E. coli and lead to increased isoprenoid production include those from Rhodobacter capsulatus (dxs1, UniProt Accession No. D5AP89, SEQ ID NO:54; dxs2, UniProt Accession No. D5ASU5, SEQ ID NO: 55).
[0110] In some embodiments, a dxs polypeptide (1) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:54. In some embodiments, a dxs polypeptide (1) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:54. In some embodiments, a dxs polypeptide (1) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:54. In some embodiments, a dxs polypeptide (1) comprises the amino acid sequence of SEQ ID NO:54.
[0111] In some embodiments, a dxs polypeptide (1) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:55. In some embodiments, a dxs polypeptide (1) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:55. In some embodiments, a dxs polypeptide (1) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:55. In some embodiments, a dxs polypeptide (1) comprises the amino acid sequence of SEQ ID NO:55.
[0112] In some embodiments, recombinant microorganisms are E. coli comprising first DXP pathway nucleotide sequences encoding dxs2 polypeptides (1) comprising the amino acid sequence of SEQ ID NO:55.
[0113] Another exemplary dxs polypeptide (1) is that from Pseudomonas fluorescens (UniProt Accession No. A0A379IJU9, SEQ ID NO:56).
[0114] In some embodiments, a dxs polypeptide (1) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:56. In some embodiments, a dxs polypeptide (1) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:56. In some embodiments, a dxs polypeptide (1) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:56. In some embodiments, a dxs polypeptide (1) comprises the amino acid sequence of SEQ ID NO:56.6.5.1.2. Dxr Polypeptides
[0115] Recombinant microorganisms of the present disclosure may comprise a second DXP pathway nucleotide sequence encoding a dxr polypeptide (2).
[0116] In some embodiments, a dxr polypeptide (2) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:57. In some embodiments, a dxr polypeptide (2) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:57. In some embodiments, a dxr polypeptide (2) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:57. In some embodiments, a dxr polypeptide (2) comprises an amino acid sequence having 100% sequence identity to the amino acid sequence of SEQ ID NO:57.
[0117] In some embodiments, a dxr polypeptide (2) comprises an amino acid sequence having less than 60% sequence identity to the amino acid sequence of SEQ ID NO:57. Dxr polypeptides (2) of these embodiments can comprise amino acid sequences having at least 30% sequence identity to the amino acid sequence of SEQ ID NO:57. Dxr polypeptides (2) of these embodiments can comprise amino acid sequences having at least 35% sequence identity to the amino acid sequence of SEQ ID NO:57. Dxr polypeptides (2) of these embodiments can comprise amino acid sequences having at least 40% sequence identity to the amino acid sequence of SEQ ID NO:57.
[0118] Exemplary dxr polypeptides (2) able to complement dxr E. coli and lead to increased isoprene and / or isoprenoid production include those from Rhodobacter capsulatus (dxr, UniProt Accession No. D5ATT5, SEQ ID NO:58).
[0119] In some embodiments, a dxr polypeptide (2) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:58. In some embodiments, a dxr polypeptide (2) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:58. In some embodiments, a dxr polypeptide (2) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:58. In some embodiments, dxr polypeptides (2) comprise the amino acid sequence of SEQ ID NO:58.
[0120] In some embodiments, recombinant microorganisms are E. coli comprising a second DXP pathway nucleotide sequence encoding dxr polypeptides (2) comprising the amino acid sequence of SEQ ID NO:58.
[0121] Another exemplary dxr polypeptide (2) is that from Pseudomonas fluorescens (UniProt Accession No. A0A0P9ALW5, SEQ ID NO:59).
[0122] In some embodiments, a dxr polypeptide (2) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:59. In some embodiments, a dxr polypeptide (2) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:59. In some embodiments, a dxr polypeptide (2) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:59. In some embodiments, a dxr polypeptide (2) comprises the amino acid sequence of SEQ ID NO:59.6.5.1.3. IspD Polypeptides
[0123] Recombinant microorganisms of the present disclosure may comprise a third DXP pathway nucleotide sequence encoding an ispD polypeptide (3).
[0124] In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:60. In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:60. In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:60. In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having 100% sequence identity to the amino acid sequence of SEQ ID NO:60.
[0125] In some embodiments, ispD polypeptides (3) can comprise amino acid sequences having less than 50% sequence identity to the amino acid sequence of SEQ ID NO:60. IspD polypeptides (3) of these embodiments can comprise amino acid sequences having at least 20% sequence identity to the amino acid sequence of SEQ ID NO:60. IspD polypeptides (3) of these embodiments can comprise amino acid sequences having at least 25% sequence identity to the amino acid sequence of SEQ ID NO:60. IspD polypeptides (3) of these embodiments can comprise amino acid sequences having at least 30% sequence identity to the amino acid sequence of SEQ ID NO:60.
[0126] Exemplary ispD polypeptides (3) able to complement AispD in E. coli and lead to increased isoprene and / or isoprenoid production include those from Rhodobacter capsulatus (UniProt Accession No. Q08113, SEQ ID NO:61; amino acids 1-222 of UniProt Accession No. Q08113, SEQ ID NO:62).
[0127] In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:61. In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:61. In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having at least97% sequence identity to the amino acid sequence of SEQ ID NO:61. In some embodiments, an ispD polypeptide (3) comprises the amino acid sequence of SEQ ID NO:61.
[0128] In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:62. In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:62. In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:62. In some embodiments, an ispD polypeptide (3) comprises the amino acid sequence of SEQ ID NO:62.
[0129] In some embodiments, recombinant microorganisms are E. coli comprising third DXP pathway nucleotide sequences encoding ispD polypeptides (3) as domains of ispDF fusion polypeptides, the ispD polypeptide (3) domains comprising the amino acid sequence of SEQ ID NO:61.
[0130] Another exemplary ispD polypeptide (3) is that from Pseudomonas (GenBank Accession No. QQU70697.1, SEQ ID NO:63).
[0131] In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:63. In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:63. In some embodiments, an ispD polypeptide (3) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:63. In some embodiments, an ispD polypeptide (3) comprises the amino acid sequence of SEQ ID NO:63.6.5.1.4. IspE Polypeptides
[0132] Recombinant microorganisms of the present disclosure may comprise a fourth DXP pathway nucleotide sequence encoding an ispE polypeptide (4).
[0133] In some embodiments, an ispE polypeptide (4) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:64. In some embodiments, an ispE polypeptide (4) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:64. In some embodiments, an ispE polypeptide (4) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:64. In someembodiments, an ispE polypeptide (4) comprises an amino acid sequence having 100% sequence identity to the amino acid sequence of SEQ ID NO:64.
[0134] In some embodiments, an ispE polypeptide (4) comprises an amino acid sequence having less than 60% sequence identity to the amino acid sequence of SEQ ID NO:64. IspE polypeptides (4) of these embodiments can comprise amino acid sequences having at least 20% sequence identity to the amino acid sequence of SEQ ID NO:64. IspE polypeptides (4) of these embodiments can comprise amino acid sequences having at least 25% sequence identity to the amino acid sequence of SEQ ID NO:64. IspE polypeptides (4) of these embodiments can comprise amino acid sequences having at least 30% sequence identity to the amino acid sequence of SEQ ID NO:64.
[0135] Exemplary ispE polypeptides (4) able to complement AispE E. coli and lead to increased isoprene and / or isoprenoid production include those from Rhodobacter capsulatus (UniProt Accession No. D5AMY7, SEQ ID NO:65).
[0136] In some embodiments, an ispE polypeptide (4) comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:65. In some embodiments, an ispE polypeptide (4) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:65. In some embodiments, an ispE polypeptide (4) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:65. In some embodiments, an ispE polypeptide (4) comprises the amino acid sequence of SEQ ID NO:65.
[0137] In some embodiments, recombinant microorganisms are E. coli comprising fourth DXP pathway nucleotide sequences encoding ispE polypeptides (4) comprising the amino acid sequence of SEQ ID NO:65.
[0138] Another exemplary ispE polypeptide (4) is that from Pseudomonas fluorescens (GenBank Accession No. QQU66123.1, SEQ ID NO:66).
[0139] In some embodiments, an ispE polypeptide (4) comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:66. In some embodiments, an ispE polypeptide (4) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:66. In some embodiments, an ispE polypeptide (4) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:66. In some embodiments, an ispE polypeptide (4) comprises the amino acid sequence of SEQ ID NO:66.6.5.1.5. IspF Polypeptides
[0140] Recombinant microorganisms of the present disclosure may comprise a fifth DXP pathway nucleotide sequence encoding an ispF polypeptide (5).
[0141] In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:67. In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:67. In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:67. In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having 100% sequence identity to the amino acid sequence of SEQ ID NO:67.
[0142] In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having less than 75% sequence identity to the amino acid sequence of SEQ ID NO:67. IspF polypeptides (5) of these embodiments can comprise amino acid sequences having at least 35% sequence identity to the amino acid sequence of SEQ ID NO:67. IspF polypeptides (5) of these embodiments can comprise amino acid sequences having at least 40% sequence identity to the amino acid sequence of SEQ ID NO:67. IspF polypeptides (5) of these embodiments can comprise amino acid sequences having at least 45% sequence identity to the amino acid sequence of SEQ ID NO:67.
[0143] Exemplary ispF polypeptides (5) able to complement AispF E. coli and lead to increased isoprene and / or isoprenoid production include those from Rhodobacter capsulatus (UniProt Accession No. Q08113, SEQ ID NO:61; amino acids 223-379 of UniProt Accession No. Q08113, SEQ ID NO:68).
[0144] In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having at least 90% identity to the amino acid sequence of SEQ ID NO:61. In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having at least 95% identity to the amino acid sequence of SEQ ID NO:61. In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having at least 97% identity to the amino acid sequence of SEQ ID NO:61. In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having 100% identity to the amino acid sequence of SEQ ID NO:61.
[0145] In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having at least 90% identity to the amino acid sequence of SEQ ID NO:68. In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having at least 95% identity to the amino acid sequence of SEQ ID NO:68. In some embodiments, an ispFpolypeptide (5) comprises an amino acid sequence having at least 97% identity to the amino acid sequence of SEQ ID NO:68. In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having 100% identity to the amino acid sequence of SEQ ID NO:68.
[0146] In some embodiments, third DXP pathway nucleotide sequences and fifth DXP pathway nucleotide sequences are operably linked so as to encode fusion proteins comprising ispD polypeptides (3) and ispF polypeptides (5).
[0147] Another exemplary ispF polypeptide (5) is that from Pseudomonas fluorescens (GenBank Accession No. QQU70693.1, SEQ ID NO:69).
[0148] In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having at least 90% identity to the amino acid sequence of SEQ ID NO:69. In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having at least 95% identity to the amino acid sequence of SEQ ID NO:69. In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having at least 97% identity to the amino acid sequence of SEQ ID NO:69. In some embodiments, an ispF polypeptide (5) comprises an amino acid sequence having 100% identity to the amino acid sequence of SEQ ID NO:69.6.5.1.6. IspG Polypeptides
[0149] Recombinant microorganisms of the present disclosure may comprise a sixth DXP pathway nucleotide sequence encoding an ispG polypeptide (6).
[0150] In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:70. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:70. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:70. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having 100% sequence identity to the amino acid sequence of SEQ ID NO:70.
[0151] In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having less than 100% sequence identity to the amino acid sequence of SEQ ID NO:70. IspG polypeptides (6) of these embodiments can comprise amino acid sequences having at least 30% sequence identity to the amino acid sequence of SEQ ID NO:70. IspG polypeptides (6) of these embodiments can comprise amino acid sequences having at least 35% sequence identity to the amino acid sequence of SEQ ID NO:70. IspG polypeptides (6)of these embodiments can comprise amino acid sequences having at least 40% sequence identity to the amino acid sequence of SEQ ID NO:70.
[0152] Exemplary ispG polypeptides (6) able to complement AispG E. coli and lead to increased isoprene and / or isoprenoid production include those from Pantoea ananatis (UniProt Accession No. A0A0H3KYT9, SEQ ID NO:71); Proteus mirabilis (UniProt Accession No. B4EZT3, SEQ ID NO:72); Serratia marcescens (UniProt Accession No. A0A0P0QHY9 with K365N substitution, SEQ ID NO:73); and Shewanella oneidensis (UniProt Accession No. Q8EC32, SEQ ID NO:74).
[0153] In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:71. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:71. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:71. In some embodiments, an ispG polypeptide (6) comprises the amino acid sequence of SEQ ID NO:71.
[0154] In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:72. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:72. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:72. In some embodiments, an ispG polypeptide (6) comprises the amino acid sequence of SEQ ID NO:72.
[0155] In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:73. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:73. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:73. In some embodiments, an ispG polypeptide (6) comprises the amino acid sequence of SEQ ID NO:73.
[0156] In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:74. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:74. In someembodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:74. In some embodiments, an ispG polypeptide (6) comprises the amino acid sequence of SEQ ID NO:74.
[0157] Another exemplary ispG polypeptide (6) is that from Synechococcus elongatus (SEQ ID NO:75).
[0158] In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:75. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:75. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:75. In some embodiments, an ispG polypeptide (6) comprises the amino acid sequence of SEQ ID NO:75.
[0159] Another exemplary ispG polypeptide (6) able to lead to increased isoprene and / or isoprenoid production is that from R. capsulatus (UniProt Accession No. D5AT87, SEQ ID NO: 76).
[0160] In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:76. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:76. In some embodiments, an ispG polypeptide (6) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:76. In some embodiments, an ispG polypeptide (6) comprises the amino acid sequence of SEQ ID NO:76.
[0161] IspG polypeptides (6) that are heterologous to E. coli can comprise one, any combination of two or more, or all of: (i) an arginine at the position corresponding to R14 of SEQ ID NO:70; (ii) an aspartic acid at the position corresponding D24 of SEQ ID NO:70; (iii) a glycine at the position corresponding to G25 of SEQ ID NO:70; (iv) a cysteine at the position corresponding to C124 of SEQ ID NO:70; (v) an asparagine at the position corresponding to N129 of SEQ ID NO:70; (vi) a glutamine at the position corresponding to Q175 of SEQ ID NO:70; (vii) a serine at the position corresponding to S191 of SEQ ID NO:70; (viii) an alanine at the position corresponding to A213 of SEQ ID NO:70; (ix) an arginine at the position corresponding to R364 of SEQ ID NO:70; and (x) an isoleucine at the position corresponding to I365 of SEQ ID NO:70.
[0162] In some embodiments, IspG polypeptides (6) that are heterologous to E. coli comprise an aspartic acid at the position corresponding D24 of SEQ ID NO:70 and a glycine at the position corresponding to G25 of SEQ ID NO:70, and optionally one, any combination of one or more, or all of: an arginine at the position corresponding to R14 of SEQ ID NO:70; a cysteine at the position corresponding to C124 of SEQ ID NO:70; an asparagine at the position corresponding to N129 of SEQ ID NO:70; a glutamine at the position corresponding to Q175 of SEQ ID NO:70; a serine at the position corresponding to S191 of SEQ ID NO:70; an alanine at the position corresponding to A213 of SEQ ID NO:70; an arginine at the position corresponding to R364 of SEQ ID NO:70; and an isoleucine at the position corresponding to I365 of SEQ ID NO:70.
[0163] In other embodiments, IspG polypeptides (6) that are heterologous to E. coli comprise an arginine at the position corresponding to R364 of SEQ ID NO:70 and an isoleucine at the position corresponding to I365 of SEQ ID NO:70, an optionally one, any combination of two or more, or all of an arginine at the position corresponding to R14 of SEQ ID NO:70; an aspartic acid at the position corresponding D24 of SEQ ID NO:70; a glycine at the position corresponding to G25 of SEQ ID NO:70; a cysteine at the position corresponding to C124 of SEQ ID NO:70; an asparagine at the position corresponding to N129 of SEQ ID NO:70; a glutamine at the position corresponding to Q175 of SEQ ID NO:70; a serine at the position corresponding to S191 of SEQ ID NO:70; an alanine at the position corresponding to A213 of SEQ ID NO:70; an arginine at the position corresponding to R364 of SEQ ID NO:70; and an isoleucine at the position corresponding to I365 of SEQ ID NO:70.
[0164] In further embodiments, IspG polypeptides (6) that are heterologous to E. coli comprise an arginine at the position corresponding to R14 of SEQ ID NO:70; an aspartic acid at the position corresponding D24 of SEQ ID NO:70; a glycine at the position corresponding to G25 of SEQ ID NO:70; a cysteine at the position corresponding to C124 of SEQ ID NO:70; an asparagine at the position corresponding to N 129 of SEQ ID NO:70; a glutamine at the position corresponding to Q175 of SEQ ID NO:70; an arginine at the position corresponding to R364 of SEQ ID NO:70; and an isoleucine at the position corresponding to I365 of SEQ ID NO:70, and optionally one or both of a serine at the position corresponding to S191 of SEQ ID NO:70 and an alanine at the position corresponding to A213 of SEQ ID NO:70.6.5.1.7. IspH Polypeptides
[0165] Recombinant microorganisms of the present disclosure may comprise a seventh DXP pathway nucleotide sequence encoding an ispH polypeptide (7).
[0166] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:77. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:77. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:77. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having 100% sequence identity to the amino acid sequence of SEQ ID NO:77.
[0167] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having less than 100% sequence identity to the amino acid sequence of SEQ ID NO:77. IspH polypeptides (7) of these embodiments can comprise amino acid sequences having at least 10% sequence identity to the amino acid sequence of SEQ ID NO:77. IspH polypeptides (7) of these embodiments can comprise amino acid sequences having at least 15% sequence identity to the amino acid sequence of SEQ ID NO:77. IspH polypeptides (7) of these embodiments can comprise amino acid sequences having at least 20% sequence identity to the amino acid sequence of SEQ ID NO:77. IspH polypeptides (7) of these embodiments can comprise amino acid sequences having at least 30% sequence identity to the amino acid sequence of SEQ ID NO:77. IspH polypeptides (7) of these embodiments can comprise amino acid sequences having at least 40% sequence identity to the amino acid sequence of SEQ ID NO:77. IspH polypeptides (7) of these embodiments can comprise amino acid sequences having at least 50% sequence identity to the amino acid sequence of SEQ ID NO:77. IspH polypeptides (7) of these embodiments can comprise amino acid sequences having at least 60% sequence identity to the amino acid sequence of SEQ ID NO:77. IspH polypeptides (7) of these embodiments can comprise amino acid sequences having at least 70% sequence identity to the amino acid sequence of SEQ ID NO:77. IspH polypeptides (7) of these embodiments can comprise amino acid sequences having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:77. IspH polypeptides (7) of these embodiments can comprise amino acid sequences having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:77.
[0168] Exemplary ispH polypeptides (7) able to complement AispH E. coli and lead to increased isoprene and / or isoprenoid production include those from Acinetobacter baylyi (UniProt Accession No. Q9RBJ0, SEQ ID NO:78); Burkholderia glumae BGR1 ispH 1, UniProt Accession No. C5AC36, SEQ ID NO:79); Caulobacter crescentus (UniProt Accession No. A0A0H3CCC9, SEQ ID NO:80); Pantoea ananatis (UniProt Accession No. D4GJM9 with D148E substitution, SEQ ID NO:81); Pseudomonas fluorescens (UniProt Accession No. C3KDX7, SEQ ID NO:82); Proteus mirabilis (UniProt Accession No. B4F2T8,SEQ ID NO:83); P. sitchensis (UniProt Accession No. C0PR44 with AM1-A44 truncation, SEQ ID NO:84); P. trichocarpa (UniProt Accession No. B3GEM6, with AC1, A2M, and A67T mutations, SEQ ID NO:85); Serratia marcescens (UniProt Accession No. A0A0P0QA59 with D149E and N188S substitutions, SEQ ID NO:86); Shewanella oneidensis (UniProt Accession No. Q8EBI7, SEQ ID NO:87); and Synechococcus sp. (UniProt Accession No. B1XPG7, SEQ ID NO:88).
[0169] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:78. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:78. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:78. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:78.
[0170] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:79. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:79. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:79. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:79.
[0171] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:80. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:80. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:80. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:80.
[0172] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:81. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:81. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:81. In someembodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:81.
[0173] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:82. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:82. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:82. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:82.
[0174] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:83. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:83. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:83. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:83.
[0175] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:84. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:84. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:84. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:84.
[0176] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:85. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:85. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:85. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:85.
[0177] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:86. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:86. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:86. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:86.
[0178] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:87. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:87. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:87. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:87.
[0179] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:88. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:88. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:88. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:88.
[0180] Other exemplary ispH polypeptides (7) include those from Nostocaceae (GenBank Accession No. RUR80075, SEQ ID NO:89); Synechococcus elongatus (GenBank Accession No. Q5N249.1, SEQ ID NO:90); and Pseudomonas (GenBank Accession No. QQU66087.1, SEQ ID NO:91).
[0181] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:89. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:89. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:89. In someembodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:89.
[0182] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:90. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:90. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:90. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:90.
[0183] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:91. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:91. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:91. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:91.
[0184] Another exemplary ispH polypeptide (7) able to lead to increased isoprene and / or isoprenoid production is that from R. capsulatus (UniProt Accession No. D5AT87, SEQ ID NO: 92).
[0185] In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:92. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:92. In some embodiments, an ispH polypeptide (7) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:92. In some embodiments, an ispH polypeptide (7) comprises the amino acid sequence of SEQ ID NO:92.
[0186] In some embodiments, ispH polypeptides (7) are not capable of directly converting HMBPP (N) to isoprene to an appreciable extent, e.g., the ispH polypeptides (7) have an activity level of HMBPP to isoprene conversion that is no more than 30% (e.g., no more than 20%, no more than 10% or no more than 5%) of their activity level of HMBPP conversion to DMAPP (P) and IPP (Q).6.5.1.8. Idi Polypeptides
[0187] Recombinant microorganisms of the present disclosure may comprise an eighth DXP pathway nucleotide sequence encoding an idi polypeptide (8).
[0188] In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:93. In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:93. In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:93. In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having 100% sequence identity to the amino acid sequence of SEQ ID NO:93.
[0189] In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having less than 40% sequence identity to the amino acid sequence of SEQ ID NO:93. Idi polypeptides (8) of these embodiments can comprise amino acid sequences having at least 15% sequence identity to the amino acid sequence of SEQ ID NO:93. Idi polypeptides (8) of these embodiments can comprise amino acid sequences having at least 20% sequence identity to the amino acid sequence of SEQ ID NO:93. Idi polypeptides (8) of these embodiments can comprise amino acid sequences having at least 25% sequence identity to the amino acid sequence of SEQ ID NO:93. Idi polypeptides (8) of these embodiments can comprise amino acid sequences having at least 30% sequence identity to the amino acid sequence of SEQ ID NO:93.
[0190] Exemplary idi polypeptides (8) expected to lead to increased isoprene and / or isoprenoid production include those from R. capsulatus (idi1, UniProt Accession No.D5AKF1, SEQ ID NO:94; idi2, UniProt Accession No. P26173, SEQ ID NO:95).
[0191] In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:94. In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:94. In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:94. In some embodiments, an idi polypeptide (8) comprises the amino acid sequence of SEQ ID NO:94.
[0192] In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:95. In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:95. In someembodiments, an idi polypeptide (8) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:95. In some embodiments, an idi polypeptide (8) comprises the amino acid sequence of SEQ ID NO:95.
[0193] In some embodiments, recombinant microorganisms are E. coli cell and eighth DXP pathway nucleotide sequences encode idi1 polypeptides (8) comprising the amino acid sequence of SEQ ID NO:94.
[0194] Another exemplary idi polypeptide (8) is that from Pseudomonas (GenBank Accession No. VVO21040.1, SEQ ID NO:96).
[0195] In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:96. In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:96. In some embodiments, an idi polypeptide (8) comprises an amino acid sequence having at least 97% sequence identity to the amino acid sequence of SEQ ID NO:96. In some embodiments, an idi polypeptide (8) comprises the amino acid sequence of SEQ ID NO:96.6.5.1.9. Recombinant Microorganisms Comprising Combinations of Polypeptides
[0196] Recombinant microorganisms, such as recombinant E. coli, can comprise various combinations of DXP pathway nucleotide sequences encoding polypeptides (1)-(8).
[0197] In some embodiments, the first DXP pathway nucleotide sequence encodes a dxs2 polypeptide (1) comprising the amino acid sequence of SEQ ID NO:55; and the second DXP pathway nucleotide sequence encodes a dxr polypeptide (2) comprising the amino acid sequence of SEQ ID NO:58. The first and second DXP pathway nucleotide sequences can be included in an expression cassette comprising, in order from 5’ to 3’, 5’ untranscribed and / or untranslated regulatory elements (e.g., a gluconate-inducible promoter), the first DXP pathway nucleotide sequence, and the second DXP pathway nucleotide sequence. The promoter can be a synthetic T5 gluconate-inducible promoter, e.g., a promoter comprising nucleotides 1 to 58 of SEQ ID NO:97. The promoter can be a “leaky” gluconate-inducible promoter, e.g., a promoter comprising SEQ ID NO:98. The promoter, the first DXP pathway nucleotide sequence, and the second DXP pathway nucleotide sequence can be integrated in the cell’s chromosome downstream of the gluconate repressor gene gntR.
[0198] In some embodiments, the eighth DXP pathway nucleotide sequence encodes an idi1 polypeptide (8) comprising the amino acid sequence of SEQ ID NO:94; the third and fifth DXP pathway nucleotide sequences encode an ispDF fusion polypeptide (3 + 5) comprising the amino acid sequence of SEQ ID NO:61; and the fourth DXP pathway nucleotidesequence encodes an ispE polypeptide (4) comprising the amino acid sequence of SEQ ID NO:65. The third, fourth, fifth, and eighth DXP pathway nucleotide sequence can be in an expression cassette comprising in order a promoter, the eighth DXP pathway nucleotide sequence, the third and fifth DXP pathway nucleotide sequences, and the fourth DXP pathway nucleotide sequence. The promoter can be a T7 A3 constitutive promoter. The promoter, the eighth DXP pathway nucleotide sequence, the third and fifth DXP pathway nucleotide sequences, and the fourth DXP pathway nucleotide sequence can be integrated in the cell’s chromosome at the insH10 locus.
[0199] In some embodiments, the sixth DXP pathway nucleotide sequence encodes an ispG polypeptide (6) comprising the amino acid sequence of SEQ ID NO:74; and the seventh DXP pathway nucleotide sequence encodes an ispH polypeptide (7) comprising the amino acid sequence of SEQ ID NO:87. The sixth and seventh DXP pathway nucleotide sequences can be in an expression cassette comprising in order a promoter, the sixth DXP pathway nucleotide sequence, the seventh DXP pathway nucleotide sequence, the eighth DXP pathway nucleotide sequence, the third and fifth DXP pathway nucleotide sequences, and the fourth DXP pathway nucleotide sequence. The promoter can be a T7 A3 constitutive promoter. The promoter, the eighth DXP pathway nucleotide sequence, the third and fifth DXP pathway nucleotide sequences, and the fourth DXP pathway nucleotide sequence can be integrated in the cell’s chromosome at the insH10 locus.
[0200] The combinations of DXP pathway nucleotide sequences and expression cassettes described in this section can further comprise regulatory components, including, but not limited to, anti-termination sites, gt10 sequences, ribosome binding sites / Shine-Dalgarno sequences, or spacer sequences, among others.
[0201] The particular combinations of DXP pathway nucleotide sequences and expression cassettes described in this section are exemplary. Other combinations, expression cassettes, promoters, and sites of integration are contemplated.
[0202] Examples of additional microorganisms engineered to increase flux through the DXP pathway include those disclosed in US 10,480,015; US 10,774,346; US 11,352,648; WO 2007 / 140339; WO 2008 / 128159; WO 2010 / 148150; WO 2012 / 088450; WO 2012 / 088462; WO 2018 / 140778; WO 2012 / 135591; WO 2014 / 138419; and Wu et al., 2021, Microbial Cell Factories 20:101, each of which is hereby incorporated by reference in their entireties.6.5.2. Heterologous MVA Pathway
[0203] The production of β-myrcene from DMAPP and IPP catalyzed by NPPS and MyrS polypeptides of the present disclosure can be increased by increasing flux through the MVApathway relative to a native microorganism comprising only native MVA pathway enzymes expressed under native regulation.
[0204] Examples of microorganisms engineered to increase flux through the MVA pathway include those disclosed in US 2018 / 0030481; US 7,659,097; US2013 / 0309741; Lu et al., 2022, Microbial Biotechnology 15:2292-2306; Liao et al., 2016, Biotechnology Advances 34:697-713; Hu et al., 2020, J Industrial Microbiology & Biotechnology 47:1083-1097; Broker et al., 2018, 102:6923-6934, each of which is hereby incorporated by reference in their entireties.6.6. Increasing Flux into DXP Pathway
[0205] In addition to comprising first through eighth DXP pathway nucleotide sequences encoding polypeptides (1)-(8), recombinant microorganisms of the present disclosure can further be engineered to increase flux into the DXP pathway. As seen in FIG. 5, various enzymatic pathways lead from glucose (A) or xylose (R) to DXP (H), the first intermediate in the DXP pathway. Increasing the activity of one or more polypeptides catalyzing one or more reactions leading to DXP (H), such as RibB*, YajO, XylB, and / or polypeptides active in the conversion of glucose (A) to GAP (E) and pyruvate (F), can increase flux into the DXP pathway and production rate and / or yield of isoprene and / or isoprenoids.
[0206] Increasing the activity of one or more polypeptides to increase flux into the DXP pathway can include one or more of increasing the copy number of genes comprising coding sequences for the polypeptides; operably linking coding sequences to promoters with higher activity and / or that are inducible at desired stages of the cell cycle compared to the promoters of parental microorganisms; operably linking coding sequences to other regulatory elements such that expression is increased relative to endogenous expression in parental microorganisms; codon optimization; modifying the amino acid sequences of polypeptides to have higher activity than wild-type; replacing endogenous polypeptide coding sequences with heterologous polypeptide coding sequences having higher activity; and adding genes comprising coding sequences for polypeptides to a parental microorganism lacking those genes, among other approaches that persons of ordinary skill in the art will be able to implement with the benefit of the present disclosure.
[0207] In some embodiments, flux into the DXP pathway can be increased by engineering recombinant microorganisms to have improved assimilation of 2-keto-3-deoxy-gluconate (KDG) relative to a parental strain. In some embodiments, improved KDG assimilation is described accomplished as described in PCT application publication number WO 2025 / 155822, which is hereby incorporated by reference in its entirety.6.7. Increasing IspG and / or IspH Activity by Expression of Redox Polypeptides
[0208] Both ispG and ispH (6)-(7) contain [4Fe-4S] iron-sulfur clusters which must be in reduced form for the polypeptides to transfer electrons to the substrates MEcPP (M) and HMBPP (N), respectively. After transferring electrons, the polypeptides are subsequently rereduced by interactions with redox polypeptides, such as flavodoxins (fld) or ferredoxins (fdx), in reduced form. The electron transfer oxidizes the redox polypeptides, which have to be re-reduced in order to be capable of transferring electrons to other polypeptides. In the case of flavodoxins and ferredoxins, flavodoxin / ferredoxin-NADP reductases (fpr;EC: 1.18.1.2) re-reduce the redox polypeptides from oxidized form to reduced form.
[0209] Microorganisms, such as E. coli, natively express one or more fld and / or fdx polypeptides along with fpr polypeptides. In principle, recombinant microorganisms expressing only native redox polypeptides may provide electrons to heterologous ispG and / or ispH polypeptides. However, not all heterologous ispG and / or ispH polypeptides provide desirable levels of activity in all such recombinant microorganisms. Though not to be bound by theory, low activity of at least some heterologous ispG and / or ispH polypeptides in at least some recombinant microorganisms may result from interactions between the heterologous polypeptides and native redox polypeptides that are weaker than those of the native redox polypeptides with native ispG and / or ispH.
[0210] The activity of heterologous ispG and / or heterologous polypeptides can be increased by engineering recombinant microorganisms to express one or more compatible heterologous redox polypeptides (e.g., ferredoxins, flavodoxins, and / or flavodoxin / ferredoxin--NADP reductases (fprs)) to improve iron-sulfur recycling in the heterologous ispG and / or ispH polypeptides in the microorganisms.
[0211] In some embodiments, the redox polypeptides comprise flavodoxins. In some embodiments, the redox polypeptides comprise ferrodoxins. In some embodiments, the redox polypeptides comprise fpr.
[0212] In some embodiments, the redox polypeptides do not comprise pyruvate:flavodoxin oxidoreductases (PFORs). PFORs are polypeptides capable of transferring electrons from pyruvate to flavodoxin (EC: 1.2.7.-). In some embodiments, the redox polypeptides do not comprise a PFOR capable of transferring electrons from NADP to flavodoxin or ferredoxin at a level that is greater than 30% (e.g., greater than 20%, greater than 10% or greater than 5%) of their activity of transferring in transferring electrons from pyruvate to ferrodoxin. Enzymatic assays for ferredoxins, flavodoxins and flavodoxin / ferredoxin-NADP reductases based on cytochrome c reduction by NADPH can be performed as described in, e.g., McIveret al., 1998, Eur. J. Biochem. 257:577-585). Enzymatic assays for PFORs can be performed as described in, e.g., Nakayama et al., 2013, Genes Genet. Syst. 88:175-188.
[0213] A heterologous redox polypeptide can be considered compatible with a heterologous ispG polypeptide and / or a heterologous ispH polypeptide if increases ispG and / or ispH activity as compared to the corresponding redox polypeptide native to a recombinant microorganism. For example, in the case of a flavodoxin or ferredoxin, a heterologous flavodoxin or ferrodoxin can be considered compatible with the heterologous ispG polypeptide and / or ispH polypeptide if it reduces the oxidized form of the ispG polypeptide and / or the ispH polypeptide to a greater degree than a native flavodoxin or ferrodoxin (e.g., where the recombinant microorganism is E. coll, wild-type E. coll flavodoxin or ferrodoxin). Alternatively or additionally, a heterologous redox polypeptide can be considered compatible with a heterologous ispG polypeptide and / or a heterologous ispH polypeptide if flux through the DXP pathway in a recombinant microorganism expressing the heterologous ispG and / or ispH polypeptide and the heterologous redox polypeptide is at least 50% greater (e.g., at least 2-fold greater, at least 3-fold greater, or at least 4-fold greater) than the flux in a recombinant microorganism expressing the heterologous ispG and / or ispH polypeptide but only native redox polypeptides.
[0214] In some embodiments, (a) heterologous redox polypeptides and (b) heterologous ispG and / or ispH polypeptides are derived from the same genus. In some embodiments, (a) heterologous redox polypeptides and (b) heterologous ispG and / or ispH polypeptides are derived from the same species of microorganism, whether the same strain or a different strain of the microorganism.
[0215] In some aspects, the present disclosure relates to recombinant microorganisms comprising engineered DXP pathways as described in Section 6.5.1, wherein sixth nucleotide sequences and / or seventh nucleotide sequences encode heterologous ispG and / or ispH polypeptides, respectively, and further comprising a nucleotide sequence (e.g., a ninth nucleotide sequence encoding a compatible heterologous redox polypeptide). In some embodiments, the recombinant microorganisms further comprise more than one additional nucleotide sequence (e.g., a ninth nucleotide sequence, a tenth nucleotide sequence, and optionally an eleventh nucleotide sequence) encoding compatible heterologous redox polypoeptides. In some embodiments, the one or more additional nucleotide sequences encode a flavodoxin, a ferrodoxin, a flavodoxin / ferredoxin-NADP reductase, or a combination of two or more of the foregoing (e.g., (a) a flavodoxin and a ferrodoxin or (b) a flavodoxin / ferredoxin--NADP reductases and either a flavodoxin, a ferrodoxin, or both).
[0216] In some embodiments, a recombinant microorganism comprises ninth (and optionally tenth and further optionally eleventh) DXP pathway nucleotide sequences encoding redox polypeptides from an organism in the same genus as the organism from which the sixth nucleotide sequence and / or the seventh nucleotide sequence are derived.
[0217] In some embodiments, a recombinant microorganism comprises ninth (and optionally tenth and further optionally eleventh) DXP pathway nucleotide sequences encoding redox polypeptides from an organism in the same species as the organism from which the sixth DXP pathway nucleotide sequence and / or the seventh DXP pathway nucleotide sequence are derived.
[0218] In some embodiments, a recombinant microorganism comprises ninth (and optionally tenth and further optionally eleventh) DXP pathway nucleotide sequences encoding redox polypeptides from the same strain from which the sixth DXP pathway nucleotide sequence and / or the seventh DXP pathway nucleotide sequence are derived.
[0219] In some aspects, a recombinant microorganism comprises a ninth DXP pathway nucleotide sequence encoding a flavodoxin. In some embodiments, the recombinant microorganism further comprises a tenth DXP pathway nucleotide sequence encoding a ferrodoxin and optionally an eleventh DXP pathway nucleotide sequence encoding a flavodoxin / ferredoxin-NADP reductase. In some embodiments, the recombinant microorganism further comprises a tenth DXP pathway nucleotide sequence encoding a flavodoxin / ferredoxin--NADP reductase and optionally an eleventh DXP pathway nucleotide sequence encoding ferrodoxin.
[0220] In other aspects, a recombinant microorganism comprises a ninth DXP pathway nucleotide sequence encoding a ferrodoxin. In some embodiments, the recombinant microorganism further comprises a tenth DXP pathway nucleotide sequence encoding a flavodoxin and optionally an eleventh DXP pathway nucleotide sequence encoding a flavodoxin / ferredoxin-NADP reductase. In some embodiments, the recombinant microorganism further comprises a tenth nucleotide sequence encoding a flavodoxin / ferredoxin-NADP reductase and optionally an eleventh nucleotide sequence encoding flavodoxin.
[0221] In some embodiments, the heterologous flavodoxins and / or ferrodoxins cannot complement native fldA genes of the recombinant microorganisms. Though not to be bound by theory, redox polypeptides (e.g., flavodoxins and ferredoxins) encoded by the ninth DXP pathway nucleotide sequences of these embodiments may have minimal redox interactions with native polypeptides.
[0222] In other embodiments, the heterologous flavodoxins and / or ferrodoxins can complement native fldA genes of the recombinant microorganisms, e.g., where the heterologous ispG and / or ispH polypeptides are sufficiently dissimilar from native ispG and / or ispH polypeptides that native redox polypeptides have minimal interactions with the heterologous ispG and / or ispH polypeptides.
[0223] In either or both scenarios, negative impacts on cell growth or maintenance arising from increased or decreased supply of redox partners to native polypeptides may be reduced.
[0224] In some embodiments, the first DXP pathway nucleotide sequence, the second DXP pathway nucleotide sequence, the third DXP pathway nucleotide sequence, the fourth DXP pathway nucleotide sequence, the fifth DXP pathway nucleotide sequence, the sixth DXP pathway nucleotide sequence, the seventh DXP pathway nucleotide sequence, the eighth DXP pathway nucleotide sequence, the ninth DXP pathway nucleotide sequence and the optional tenth DXP pathway nucleotide sequence and / or eleventh DXP pathway nucleotide sequence are derived from genes from organisms in the same genus.
[0225] In some embodiments, the first DXP pathway nucleotide sequence, the second DXP pathway nucleotide sequence, the third DXP pathway nucleotide sequence, the fourth DXP pathway nucleotide sequence, the fifth DXP pathway nucleotide sequence, the sixth DXP pathway nucleotide sequence, the seventh DXP pathway nucleotide sequence, the eighth DXP pathway nucleotide sequence, the ninth DXP pathway nucleotide sequence and the optional tenth DXP pathway nucleotide sequence and / or eleventh DXP pathway nucleotide sequence are derived from genes from the same organism.
[0226] Though not to be bound by theory, in these embodiments, by being derived from the same genus and / or species, polypeptides (1)-(9) (and optionally (10) and / or (11) may be evolutionarily tuned for higher interaction than if polypeptides (1)-(9) (and optionally (10) and / or (11) were derived from different organisms.
[0227] In particular embodiments, a sixth DXP pathway nucleotide sequence encodes an ispG polypeptide derived from Rhodobacter capsulatus (e.g., a polypeptide comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO:76), a seventh DXP pathway nucleotide sequence encodes an ispH polypeptide also derived from R. capsulatus (e.g., a polypeptide comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO:92), and a ninth DXP pathway nucleotide sequence encodes a redox polypeptide also derived from R. capsulatus.
[0228] In particular embodiments, the sixth DXP pathway nucleotide sequence encodes an ispG polypeptide derived from Rhodobacter capsulatus (e.g., a polypeptide comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO:76), the seventh DXP pathway nucleotide sequence encodes an ispH polypeptide also derived from R. capsulatus (e.g., polypeptides comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO:92), the ninth DXP pathway nucleotide sequence encodes a ferredoxin or flavodoxin also derived from R. capsulatus (e.g., a polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs:75-81 of WO 2025 / 155822), and the tenth DXP pathway nucleotide sequence encodes a flavodoxin / ferredoxin--NADP reductase also derived from R. capsulatus. An exemplary flavodoxin / ferredoxin--NADP reductase derived from R. capsulatus is fpr, UniProt Accession No. D5ATP7 (SEQ ID NO:82 of WO 2025 / 155822). In some embodiments, the tenth DXP pathway nucleotide sequence encodes a flavodoxin / ferredoxin--NADP reductase comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO:82 of WO2025 / 155822.6.8. Regulatory Elements
[0229] The NPPS, MyrS, and NerS coding sequences in the recombinant microorganisms of the disclosure are typically operably linked to one or more regulatory elements to control the expression of the NPPS, MyrS, and NerS polypeptides.
[0230] In some embodiments, the regulatory elements include a promoter. The nucleotide sequences encoding NPPS, MyrS, and NerS polypeptides can be expressed under control of the same promoter or different promoters. Promoters can be identical to native promoters, identical to promoters native to viruses that infect the recombinant microorganisms, engineered variants of native or viral promoters, or fully synthetic promoters. In some embodiments, promoters are recognized by native transcriptional machinery of recombinant microorganisms.
[0231] A promoter can be chosen for operable linkage to any particular NPPS, MyrS, or NerS coding sequence in view of the species and / or parental strain of the recombinant microorganism, the desired expression level of the NPPS, MyrS, or NerS coding sequence, and / or other parameters known to skilled persons.
[0232] The promoters can be constitutive or inducible.
[0233] Non-limiting examples of constitutive promoters suitable for use in E. coli include T7, Pspc, PL, PRNAI, PRNAII, P1, P2, and Ptac promoters. Non-limiting examples of- M -constitutive promoters suitable for use in S. cerevisiae include the TEF1, GAP, ACT 1, PGK, and ENO promoters.
[0234] In some embodiments, the promoters are inducible promoters. Inducible promoters to allow production of the NPPS, NerS, and / or MyrS polypeptides by the recombinant microorganisms at desired times in the cell cycle and / or to desired level.
[0235] Inducers of the promoters can be compounds produced by recombinant microorganisms at one or more stages of recombinant microorganisms’ life cycle and / or can be compounds not produced by recombinant microorganisms and instead added to culture media comprising recombinant microorganisms.
[0236] In some embodiments, the inducible promoters are part of an expression regulatory system, e.g., an operon. Non-limiting examples of inducible promoter systems include lactose regulated systems (e.g., lactose operon systems), sugar regulated systems, metal regulated systems, steroid regulated systems, alcohol regulated systems, IPTG inducible systems, arabinose regulated systems (e.g., arabinose operon systems, e.g., an ARA operon promoter, pBAD, pARA, PARAE, ARAE, ARAR-ParaE, portions thereof, combinations thereof and the like), synthetic amino acid regulated systems (e.g., see Rovner et al., 2015, Nature 518(7537):89-93), fructose repressors, atac promoter / operator (pTac), tryptophan promoters, PhoA promoters, recA promoters, proll promoters, cst-1 promoters, tetA promoters, cadA promoters, nar promoters, PL promoters, cspA promoters, trc promoters the like or combinations thereof. In certain embodiments, an inducible promoter is not a lac operon promoter.
[0237] In some embodiments, promoters operably linked to NPPS, NerS, and / or MyrS coding sequences are part of an expression system that is inducible by 2-keto-3-deoxy-D-gluconate (KDG). In some embodiments, promoters operably linked to NPPS and / or MyrS coding sequences are part of an expression system that is inducible by gluconate. In some embodiments, the gluconate-inducible expression system comprises a binding site for the gluconate repressor GntR. In some embodiments, promoters operably linked to NPPS, NerS, and / or MyrS coding sequences are part of an expression system that is induced by anaerobic conditions.
[0238] In addition to promoters, the recombinant microorganisms can comprise other regulatory sequences that influence the expression of NPPS, NerS, and / or MyrS coding sequences. As with promoters, particular instances of other regulatory sequences can be chosen for any particular NPPS, NerS, and / or MyrS coding sequence in view of the species and / or parental strain of the recombinant microorganism, the desired expression level of the NPPS, NerS, and / or MyrS coding sequence, and / or other parameters.
[0239] In some embodiments, regulatory sequences operably linked to NPPS, NerS, and / or MyrS coding sequences include regulatory sequences native to the genome of the recombinant microorganism. For example, an NPPS, NerS, and / or MyrS coding sequence may be integrated into the genome of a parental microorganism such that the coding sequence is operably linked to a promoter, enhancer, and / or other regulatory element native to the genome of the parental microorganism. This could be the result, for example, of replacing a native ORF of the parental microorganism with a heterologous NPPS, NerS, and / or MyrS coding sequence. This may result in a heterologous NPPS, NerS, or MyrS coding sequence being expressed and regulated in the same way as the original native gene.6.9. Parental Microorganisms
[0240] Any microorganisms can be parental microorganisms engineered to yield recombinant microorganisms described herein. The parental microorganisms include, but are not limited to, prokaryotes, such as bacteria, e.g., E. coli. Suitable parental microorganisms may also include eukaryotes, such as, for example, S. cerevisiae.
[0241] In some embodiments, the parental microorganism is E. coli. In particular aspects, the E. coli is E. coli strain K12 or a strain derived therefrom, such as E. coli K12 substrain MG1655.
[0242] Other E. coli strains from which recombinant microorganisms of the present disclosure can be engineered include, but are not limited to, E. coli K12 substrain W3110, E. coli K12 substrain DH5alpha, and non-K12 strains, such as E. coli BL21 and E. coli W.
[0243] Recombinant microorganisms of the present disclosure can be engineered from bacteria other than E. coli. Examples of such bacteria include, but are not limited to, Rhodobacter capsulatus, Bacillus subtilis, Pantoea ananatis, Tatumella citrea, Pseudomonas fluorescens, and Pseudomonas putida.
[0244] In some embodiments, the parental microorganism is S. cerevisiae. In particular aspects, the S. cerevisiae is an HAO strain or a W303 strain.
[0245] In some embodiments, parental microorganisms can be engineered to express NPPS and MyrS polypeptides, e.g., as described in Section 6.2, through the introduction of nucleic acids comprising NPPS, NerS, and / or MyrS coding sequences.
[0246] In some embodiments, parental microorganisms are further engineered to increase flux through the DXP pathway or MVA pathway, e.g., as described in Sections 6.5, 6.6, and 6.7.6.9.1. Engineering Methods
[0247] Parental microorganisms can be engineered using techniques known in the art.
[0248] In some embodiments, nucleic acids are introduced into microorganisms by any appropriate transformation technique. Nucleic acids can be extrachromosomal, on a vector (such as a plasmid or a phage), such as a low copy number vector, an intermediate copy number vector, or a high copy number vector. Nucleic acids can be maintained episomally and thus comprise a sequence for autonomous replication, such as an autosomal replication sequence. Alternatively, nucleic acids can be integrated in one or more copies into the genome of the cell. Integration into the cell’s genome can occur at random by non-homologous recombination, or at selected locations by homologous recombination (e.g., to replace an endogenous coding sequence and / or regulatory sequence with a modified one, a replacement therefor, or a partial or complete deletion thereof), as is well known in the art.
[0249] When multiple coding sequences are to be introduced into a cell, such as for integration into the genome, one or more coding sequences can be included in each of one or more expression cassettes. For example, two expression cassettes can be used. The expression cassettes can be operons, e.g., one promoter and associated regulatory sequences can be operably linked to and enable expression of multiple coding sequences. Independently of one another, each expression cassette can be integrated into the cell’s genome. For example, if two expression cassettes are used, either or both can be integrated into the cell’s genome.
[0250] Various genome editing techniques, including but not limited to homologous recombination, CRISPR, zinc finger nucleases, and transcription activator-like effector nucleases (TALENs), can be used to delete or disrupt genes in a parental microorganism or to operably link a coding sequence to a regulatory sequence to which it is not operably linked in a parental microorganism (which may change promoter strength, change whether a promoter is constitutive or inducible, or change which inducer molecule induces transcription of a coding sequence from an inducible promoter), to reduce or increase enzymatic activity of polypeptides encoded by those genes.
[0251] RNAi techniques can be used to reduce activity of enzymes in prokaryotes by regulating gene expression (Waters et al., 2009, Cell 136(4):615-628) or interfering with translation of RNAs encoding the enzymes. Nucleic acids can be introduced into or engineered in recombinant microorganisms to produce regulatory RNAs, microRNAs (miRNAs), small interfering RNAs (siRNAs), antisense RNAs (asRNAs), and / or single guide RNAs (sgRNAs) for CRISPR interference.
[0252] Engineering methods can reduce activity of an endogenous polypeptide relative to a wild-type microorganism. For example, all or a portion of coding sequences can be mutated, optionally wherein the mutation can be a deletion; all or a portion of regulatory sequences can be mutated, optionally wherein the mutation can be a deletion; heterologous sequences can be introduced into endogenous loci; interfering RNA (RNAi) systems that reduce activity can be engineered into recombinant microorganisms; or any two, any three, or all four thereof, among other techniques.6.10. Methods of Use6.10.1. Culture Media
[0253] Generally, methods disclosed herein comprise growing cells of a recombinant microorganism in a growth medium suitable for growth to a desired cell density and culturing the cells in a production medium suitable for production of p-myrcene or nerol. Culturing can be in a batch mode or a continuous mode.
[0254] Examples of media that can be used in batch mode culturing include M9 medium and Hi-Def medium. In some embodiments, M9 medium comprises the following: sodium phosphate dibasic heptahydrate, 1.28 w / v%; potassium phosphate monobasic, 0.3 w / v%; sodium chloride, 0.05 w / v%; ammonium chloride, 0.1 w / v%; glucose, 0.4 w / v%; MgSO4, 0.024 w / v%; and CaCl2, 0.001 w / v%. In some embodiments, Hi-Def medium comprises ingredients known to the person of ordinary skill in the art, and it is commercially available (Teknova Inc. Hollister, CA). Other suitable growth media include, but are not limited to, MOPS, LB, and TB.
[0255] In some embodiments, a culture medium comprises at least 0.1 w / v% sucrose, at least 0.2 w / v% sucrose, at least 0.3 w / v% sucrose, at least 0.4 w / v% sucrose, at least 0.5 w / v% sucrose, at least 0.6 w / v% sucrose, at least 0.7 w / v% sucrose, at least 0.8 w / v% sucrose, at least 0.9 w / v% sucrose, or at least 1 w / v% sucrose. A culture medium typically comprises less than 5 w / v% sucrose, more typically less than 2 w / v% sucrose (e.g., in some embodiments, culture media comprise from 0.1 w / v% to 5 w / v% sucrose; from 0.1 w / v% to 2 w / v% sucrose; from 0.1 w / v% to 1 w / v% sucrose; or from 1 w / v% to 2 w / v% sucrose, among other possible ranges).
[0256] In some embodiments, culture media comprise at least 0.1 w / v% glucose, at least 0.2 w / v% glucose, at least 0.3 w / v% glucose, at least 0.4 w / v% glucose, at least 0.5 w / v% glucose, at least 0.6 w / v% glucose, at least 0.7 w / v% glucose, at least 0.8 w / v% glucose, at least 0.9 w / v% glucose, or at least 1 w / v% glucose. A culture medium typically comprises less than 5 w / v% glucose, more typically less than 2 w / v% glucose (e.g., in some embodiments, culture media comprise from 0.1 w / v% to 5 w / v% glucose; from 0.1 w / v% to 2w / v% glucose; from 0.1 w / v% to 1 w / v% glucose; or from 1 w / v% to 2 w / v% glucose, among other possible ranges).
[0257] In some embodiments, culture media comprise at least 0.1 w / v% gluconate, at least 0.2 w / v% gluconate, at least 0.3 w / v% gluconate, at least 0.4 w / v% gluconate, at least 0.5 w / v% gluconate, at least 0.6 w / v% gluconate, at least 0.7 w / v% gluconate, at least 0.8 w / v% gluconate, at least 0.9 w / v% gluconate, or at least 1 w / v% gluconate. A culture medium typically comprises less than 5 w / v% gluconate, more typically less than 2 w / v% gluconate (e.g., in some embodiments, culture media comprise from 0.1 w / v% to 5 w / v% gluconate; from 0.1 w / v% to 2 w / v% gluconate; from 0.1 w / v% to 1 w / v% gluconate; or from 1 w / v% to 2 w / v% gluconate, among other possible ranges).
[0258] In some embodiments, culture media comprise at least 0.1 w / v% cellulose-derived sugars, at least 0.2 w / v% cellulose-derived sugars, at least 0.3 w / v% cellulose-derived sugars, at least 0.4 w / v% cellulose-derived sugars, at least 0.5 w / v% cellulose-derived sugars, at least 0.6 w / v% cellulose-derived sugars, at least 0.7 w / v% cellulose-derived sugars, at least 0.8 w / v% cellulose-derived sugars, at least 0.9 w / v% cellulose-derived sugars, or at least 1 w / v% cellulose-derived sugars. A culture medium typically comprises less than 5 w / v% cellulose-derived sugars, more typically less than 2 w / v% cellulose-derived sugars (e.g., in some embodiments, culture media comprise from 0.1 w / v% to 5 w / v% cellulose-derived sugars; from 0.1 w / v% to 2 w / v% cellulose-derived sugars; from 0.1 w / v% to 1 w / v% cellulose-derived sugars; or from 1 w / v% to 2 w / v% cellulose-derived sugars, among other possible ranges). The concentrations of cellulose-derived sugars listed here are the sum of the concentrations of all cellulose-derived sugars (which may be one or more cellulose-derived sugars, e.g., glucose and / or xylose) in the media.
[0259] In some embodiments, a production medium comprises sucrose. In some embodiments, a production medium comprises glucose. In some embodiments, a production medium comprises gluconate. In some embodiments, a production medium comprises one or more cellulose-derived sugars. In some embodiments, a production medium comprises any two, and three, or all four of sucrose, glucose, gluconate, or cellulose-derived sugars. The inclusion of gluconate in production media can induce expression of sequences of interest in recombinant microorganisms comprising nucleic acids comprising coding sequences operably linked to gluconate-inducible promoters.
[0260] In some embodiments, a production medium comprises an inducer. In some embodiments, a production medium comprises KDG. The inclusion of KDG in production media can induce expression of sequences of interest in recombinant microorganismscomprising nucleic acids comprising coding sequences operably linked to KDG-inducible promoters.
[0261] In some embodiments in which culture media comprise two or more carbon sources, culture media comprise at least 0.5 w / v% total carbon sources, at least 0.6 w / v% total carbon sources, at least 0.7 w / v% total carbon sources, at least 0.8 w / v% total carbon sources, at least 0.9 w / v% total carbon sources, or at least 1 w / v% total carbon sources. A culture medium typically comprises less than 5 w / v% total carbon sources, more typically less than 2 w / v% total carbon sources (e.g., in some embodiments, culture media comprise from 0.1 w / v% to 5 w / v% total carbon sources; from 0.1 w / v% to 2 w / v% total carbon sources; from 0.1 w / v% to 1 w / v% total carbon sources; or from 1 w / v% to 2 w / v% total carbon sources, among other possible ranges).
[0262] In some embodiments of some methods described herein, it may be desirable to allow growth of a recombinant microorganism without expression of one or more genes until a desired biomass of the recombinant microorganism has been reached. For example, such growth can be encouraged or effected by use of a growth medium comprising glycerol, such as at least 0.1 w / v% glycerol, at least 0.2 w / v% glycerol, at least 0.3 w / v% glycerol, at least 0.4 w / v% glycerol, at least 0.5 w / v% glycerol, at least 0.6 w / v% glycerol, at least 0.7 w / v% glycerol, at least 0.8 w / v% glycerol, at least 0.9 w / v% glycerol, or at least 1 w / v% glycerol. A growth medium typically comprises less than 5 w / v% glycerol, more typically less than 2 w / v% glycerol (e.g., in some embodiments, growth media comprise from 0.1 w / v% to 5 w / v% glycerol; from 0.1 w / v% to 2 w / v% glycerol; from 0.1 w / v% to 1 w / v% glycerol; or from 1 w / v% to 2 w / v% glycerol, among other possible ranges).
[0263] Although glycerol can provide a carbon source for growth of a recombinant microorganism in a growth medium, glycerol can be included in a production medium. Typically, glycerol is included in a production medium at the same or lower concentration than in a growth medium.
[0264] In some embodiments, a growth medium lacks added glucose and / or sucrose, / .e., one or both of these sugars is not intentionally included in a growth medium. In particular embodiments, a growth medium comprises no more than 0.1 w / v% each of glucose and / or sucrose.
[0265] The selection of particular concentrations of sucrose, glucose and / or glycerol to include in a production medium and / or a growth medium can be made by the person of ordinary skill in the art having the benefit of the present disclosure as a routine matter.
[0266] For fed-batch and / or continuous mode culturing, the ranges of sucrose, glucose, gluconate, glycerol, or combinations thereof given above can be initially provided to the medium. The consumption of the carbon source(s) during culturing can be repeatedly or continuously monitored and additional carbon source(s) can be provided as needed to sustain a desired respiratory coefficient, growth rate, rate of expression of sequences of interest, and / or a rate of production of desired compound(s). The feed rate may be adjusted to avoid accumulation of carbon source(s), which may maximize output of desired compound(s) and minimize waste of carbon source(s). The person of ordinary skill in the art having the benefit of the present disclosure can select the medium composition and the amount of carbon source added thereto during the process to enable the expression of sequences of interest to a desired level and / or production of desired product(s) to a desired concentration, such as at least 20 g / L, at least 50 g / L, or at least 100 g / L.6.10.2. Culture Conditions
[0267] Recombinant cells comprising expression systems of the disclosure may be cultured under suitable conditions in a medium, such as a medium described in Section 6.10.1. In some embodiments, recombinant cells undergo fermentation. Fermentation conditions include batch, fed-batch and continuous fermentation. Classical batch fermentation is a closed system, wherein the composition of the medium is not subject to artificial alterations during fermentation. In fed-batch fermentation, the substrate is added in increments as fermentation progresses. In both classical batch fermentation and batch-fed fermentation, the product(s) remain in the bioreactor until the end of the process. Batch and fed-batch fermentation are common and well-known in the art. In continuous fermentation, a defined medium is added continuously to the bioreactor and an equal volume of product containing medium is removed simultaneously. Continuous fermentation aims to maintain steady state growth conditions. Methods for modulating nutrients and growth factors for continuous fermentation processes as well as techniques for maximizing the rate of product formation are well known in the art of industrial microbiology. The fermentation process is typically an aerobic fermentation process.
[0268] The fermentation process is typically run at a temperature that is optimal for growth of a recombinant microorganism. Fermentation for a mesophilic microorganism is typically carried out at a temperature within the range of from 20°C to 45°C, from 25°C to 40°C, from 35°C to 40°C, or from 30°C to 37°C. In some embodiments wherein a recombinant microorganism is derived from one of the exemplary microorganisms described herein, culturing comprises maintaining the recombinant microorganism at a mesophilic temperature. In some embodiments, the mesophilic temperature is selected from any of the foregoing ranges.
[0269] Fermentation is typically carried out at a pH in the range of 4 to 8, in the range of 5 to 7, or the range of 5.5 to 6.5. For example, fermentation can be carried out for a period of time within the range of from 8 to 240 hours, from 12 hours to 168 hours, from 16 hours to 144 hours, from 20 hours to 120 hours, from 24 hours to 72 hours, or from 36 to 48 hours.6.11. Methods for Producing p-myrcene
[0270] In some aspects, the present disclosure relates to methods for producing p-myrcene that makes use of recombinant microorganisms. In some embodiments, the methods comprising culturing recombinant microorganisms as described in Section 6.2 under conditions in which p-myrcene is produced. In some embodiments, the microorganisms are E. coli, including but not limited to E. coli described in Section 6.9. In some embodiments, the microorganisms are S. cerevisiae, including but not limited to S. cerevisiae described in Section 6.9.
[0271] The conditions can include culturing recombinant microorganisms in appropriate media, e.g., media comprising glucose.
[0272] Cells can be grown to desired cell concentrations in appropriate media. After growth to desired cell concentrations, cells in which one or more nucleotide sequences encoding NPPS and MyrS polypeptides are operably linked to inducible promoters can be cultured in media comprising the inducer. Induction of expression of NPPS and MyrS polypeptides leads to the enzymatic conversion of DMAPP and IPP to NPP and NPP to p-myrcene.
[0273] In some embodiments, methods further comprise purifying p-myrcene. p-myrcene can be purified using techniques known to the skilled person. Examples of purification methods include distillation and chromatography, p-myrcene purity can be assayed by any appropriate method, such as column chromatography, HPLC analysis, or GC-MS analysis.
[0274] In some embodiments, due to the very low solubility of myrcene in water, special methods may be used to ensure that the myrcene produced by the microorganisms is efficiently captured from the fermentation vessel. A common method is the use of two-phase extractive fermentation, which relies on the use of an organic solvent overlay, which is added to the fermentation vessel. This organic overlay does not mix with water, but is capable of dissolving and capturing myrcene, avoiding its evaporation. In other embodiments, myrcene present in the headspace of the fermenter is captured. This can be accomplished by passing the off gas of the fermenter through an organic solvent in which myrcene is soluble. Some examples of organic solvents suitable for myrcene capture are dodecane, tetradecane, hexadecane, isopropyl myristate, or isopropyl palmitate.6.12. Methods for Producing Nerol
[0275] In some aspects, the present disclosure relates to methods for producing nerol that makes use of recombinant microorganisms. In some embodiments, the methods comprising culturing recombinant microorganisms as described in Section 6.2 under conditions in which nerol is produced. In some embodiments, the microorganisms are E. coli, including but not limited to E. coli described in Section 6.9. In some embodiments, the microorganisms are S. cerevisiae, including but not limited to S. cerevisiae described in Section 6.9.
[0276] The conditions can include culturing recombinant microorganisms in appropriate media, e.g., media comprising glucose.
[0277] Cells can be grown to desired cell concentrations in appropriate media. After growth to desired cell concentrations, cells in which one or more nucleotide sequences encoding NPPS and NerS polypeptides are operably linked to inducible promoters can be cultured in media comprising the inducer. Induction of expression of NPPS and NerS polypeptides leads to the enzymatic conversion of DMAPP and IPP to NPP and NPP to nerol.
[0278] In some embodiments, methods further comprise purifying nerol. Nerol can be purified using techniques known to the skilled person. Examples of purification methods include distillation and chromatography. Nerol purity can be assayed by any appropriate method, such as column chromatography, HPLC analysis, or GC-MS analysis.
[0279] In some embodiments, due to the very low solubility of nerol in water, special methods may be used to ensure that the nerol produced by the microorganisms is efficiently captured from the fermentation vessel. A common method is the use of two-phase extractive fermentation, which relies on the use of an organic solvent overlay, which is added to the fermentation vessel. This organic overlay does not mix with water, but is capable of dissolving and capturing nerol, avoiding its evaporation. In other embodiments, nerol present in the headspace of the fermenter is captured. This can be accomplished by passing the off gas of the fermenter through an organic solvent in which nerol is soluble. Some examples of organic solvents suitable for nerol capture are dodecane, tetradecane, hexadecane, isopropyl myristate, or isopropyl palmitate.7. SPECIFIC EMBODIMENTS
[0280] The present disclosure is exemplified by the specific numbered embodiments below.1. A recombinant microorganism engineered to express:(a) a neryl diphosphate synthase (NPPS); and(b) a terpene synthase selected from:(i) a myrcene synthase (MyrS); or(ii) a nerol synthase (NerS).The recombinant microorganism of embodiment 1, wherein the NPPS is native to a plant species.The recombinant microorganism of embodiment 1 or 2, wherein the plant species is Solanum lycopersicum.The recombinant microorganism of any one of embodiments 1 to 3, wherein the NPPS comprises an amino acid sequence having at least 80% sequence identity to SEQ ID NO:2.The recombinant microorganism of any one of embodiments 1 to 3, wherein the NPPS comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:2.The recombinant microorganism of any one of embodiments 1 to 3, wherein the NPPS comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO:2.The recombinant microorganism of any one of embodiments 1 to 3, wherein the NPPS comprises the amino acid sequence of SEQ ID NO:2.The recombinant microorganism of any one of embodiments 1 to 7, wherein the recombinant microorganism comprises a first nucleotide sequence encoding the NPPS. The recombinant microorganism of embodiment 5, wherein the first nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO: 12.The recombinant microorganism of embodiment 5, wherein the first nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO: 12.The recombinant microorganism of embodiment 5, wherein the first nucleotide sequence comprises a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence of SEQ ID NO: 12.The recombinant microorganism of any one of embodiments 1 to 11, wherein the terpene synthase is a myrcene synthase (MyrS).The recombinant microorganism of any one of embodiments 1 to 12, wherein the recombinant microorganism comprises a second nucleotide sequence encoding the MyrS.The recombinant microorganism of any one of embodiments 1 to 12, wherein the MyrS is native to a plant species.The recombinant microorganism of any one of embodiments 1 to 14, wherein the MyrS is an Antirrhinium majus MyrS.The recombinant microorganism of embodiment 15, wherein the Antirrhinium majus MyrS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:4.The recombinant microorganism of embodiment 15, wherein the Antirrhinium majus MyrS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:4.The recombinant microorganism of embodiment 15, wherein the Antirrhinium majus MyrS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:4.The recombinant microorganism of embodiment 15, wherein the Antirrhinium majus MyrS comprises the amino acid sequence of SEQ ID NO:4.The recombinant microorganism of any one of embodiments 15 to 19, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO: 14.The recombinant microorganism of any one of embodiments 15 to 19, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO: 14.The recombinant microorganism of any one of embodiments 15 to 19, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:14. The recombinant microorganism of any one of embodiments 1 to 14, wherein the MyrS is an Ocimum basilicum MyrS.The recombinant microorganism of embodiment 23, wherein the Ocimum basilicum MyrS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:6.The recombinant microorganism of embodiment 23, wherein the Ocimum basilicum MyrS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:6.The recombinant microorganism of embodiment 23, wherein the Ocimum basilicum MyrS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:6.The recombinant microorganism of embodiment 23, wherein the Ocimum basilicum MyrS comprises the amino acid sequence of SEQ ID NO:6.The recombinant microorganism of any one of embodiments 23 to 27, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO: 16.The recombinant microorganism of any one of embodiments 23 to 27, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO: 16.The recombinant microorganism of any one of embodiments 23 to 27, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:16. The recombinant microorganism of any one of embodiments 1 to 14, wherein the MyrS is an Picea abis MyrS.The recombinant microorganism of embodiment 31, wherein the Picea abis MyrS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:7.The recombinant microorganism of embodiment 31, wherein the Picea abis MyrS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:7.The recombinant microorganism of embodiment 31, wherein the Picea abis MyrS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:7.The recombinant microorganism of embodiment 31, wherein the Picea abis MyrS comprises the amino acid of SEQ ID NO:7.The recombinant microorganism of any one of embodiments 31 to 35, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO: 17.The recombinant microorganism of any one of embodiments 31 to 35, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO: 17.The recombinant microorganism of any one of embodiments 31 to 35, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:17. The recombinant microorganism of any one of embodiments 1 to 14, wherein the MyrS is a Quercus ilex MyrS.The recombinant microorganism of embodiment 39, wherein the Quercus ilex MyrS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:9.The recombinant microorganism of embodiment 39, wherein the Quercus ilex MyrS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:9.The recombinant microorganism of embodiment 39, wherein the Quercus ilex MyrS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:9.The recombinant microorganism of embodiment 39, wherein the Quercus ilex MyrS comprises the amino acid sequence of SEQ ID NO:9.The recombinant microorganism of any one of embodiments 39 to 44, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO: 19.The recombinant microorganism of any one of embodiments 39 to 44, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO: 19.The recombinant microorganism of any one of embodiments 39 to 44, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:19. The recombinant microorganism of any one of embodiments 1 to 14, wherein the MyrS is a Cannabis sativa MyrS.The recombinant microorganism of embodiment 47, wherein the Cannabis sativa MyrS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO: 10.The recombinant microorganism of embodiment 47, wherein the Cannabis sativa MyrS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO: 10.The recombinant microorganism of embodiment 47, wherein the Cannabis sativa MyrS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO: 10.The recombinant microorganism of embodiment 47, wherein the Cannabis sativa MyrS comprises the amino acid sequence of SEQ ID NO:10.The recombinant microorganism of any one of embodiments 47 to 51, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO:20.The recombinant microorganism of any one of embodiments 47 to 51, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO:20.The recombinant microorganism of any one of embodiments 47 to 51, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:20. The recombinant microorganism of any one of embodiments 1 to 14, wherein the MyrS is an Antirrhinium majus OciS (AmOciS).The recombinant microorganism of embodiment 55, wherein the AmOciS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:32.The recombinant microorganism of embodiment 55, wherein the AmOciS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:32.The recombinant microorganism of embodiment 55, wherein the AmOciS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:32.The recombinant microorganism of embodiment 55, wherein the AmOciS comprises the amino acid sequence of SEQ ID NO:32.The recombinant microorganism of any one of embodiments 55 to 59, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO:47.The recombinant microorganism of any one of embodiments 55 to 59, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO:47.The recombinant microorganism of any one of embodiments 55 to 59, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:47. The recombinant microorganism of any one of embodiments 1 to 14, wherein the MyrS is a Phaseolus lunatus OciS (PlOciS).The recombinant microorganism of embodiment 63, wherein the PlOciS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:36.The recombinant microorganism of embodiment 63, wherein the PlOciS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:36.The recombinant microorganism of embodiment 63, wherein the PlOciS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:36.The recombinant microorganism of embodiment 63, wherein the PlOciS comprises the amino acid sequence of SEQ ID NO:36.The recombinant microorganism of any one of embodiments 63 to 67, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO:51.The recombinant microorganism of any one of embodiments 63 to 67, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO:51.The recombinant microorganism of any one of embodiments 63 to 67, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:51. The recombinant microorganism of any one of embodiments 1 to 11, wherein the terpene synthase is a nerol synthase (NerS).The recombinant microorganism of embodiment 71, wherein the recombinant microorganism comprises a second nucleotide sequence encoding the NerS.The recombinant microorganism of embodiment 71 or 72, wherein the NerS is from Spatholobus suberectus.The recombinant microorganism of embodiment 73, wherein the NerS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:23.The recombinant microorganism of embodiment 73, wherein the NerS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:23.The recombinant microorganism of embodiment 73, wherein the NerS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:23.The recombinant microorganism of embodiment 73, wherein the NerS comprises the amino acid sequence of SEQ ID NO:23.The recombinant microorganism of any one of embodiments 73 to 77, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO:38.The recombinant microorganism of any one of embodiments 73 to 77, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO:38.The recombinant microorganism of any one of embodiments 73 to 77, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:38. The recombinant microorganism of embodiment 71 or 72, wherein the NerS is from Cajanus cajan.The recombinant microorganism of embodiment 81, wherein the NerS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:24.The recombinant microorganism of embodiment 81, wherein the NerS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:24.The recombinant microorganism of embodiment 81, wherein the NerS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:24.The recombinant microorganism of embodiment 81, wherein the NerS comprises the amino acid sequence of SEQ ID NO:24.The recombinant microorganism of any one of embodiments 81 to 85, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO:39.The recombinant microorganism of any one of embodiments 81 to 85, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO:39.The recombinant microorganism of any one of embodiments 81 to 85, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:39. The recombinant microorganism of embodiment 71 or 72, wherein the NerS is from Cinnamomum tenuipile.The recombinant microorganism of embodiment 89, wherein the NerS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:25.The recombinant microorganism of embodiment 89, wherein the NerS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:25.The recombinant microorganism of embodiment 89, wherein the NerS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:25.The recombinant microorganism of embodiment 89, wherein the NerS comprises the amino acid sequence of SEQ ID NO:25.The recombinant microorganism of any one of embodiments 89 to 93, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO:40.The recombinant microorganism of any one of embodiments 89 to 93, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO:40.The recombinant microorganism of any one of embodiments 89 to 93, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:40. The recombinant microorganism of embodiment 71 or 72, wherein the NerS is from Peril I a citriadora.The recombinant microorganism of embodiment 97, wherein the NerS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:26.The recombinant microorganism of embodiment 97, wherein the NerS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:26.The recombinant microorganism of embodiment 97, wherein the NerS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:26.The recombinant microorganism of embodiment 97, wherein the NerS comprises the amino acid sequence of SEQ ID NO:26.The recombinant microorganism of any one of embodiments 97 to 101, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO:41.The recombinant microorganism of any one of embodiments 97 to 101, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO:41.The recombinant microorganism of any one of embodiments 97 to 101, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:41. The recombinant microorganism of embodiment 71 or 72, wherein the NerS is from Camptotheca acuminata.The recombinant microorganism of embodiment 105, wherein the NerS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:26.The recombinant microorganism of embodiment 105, wherein the NerS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:26.The recombinant microorganism of embodiment 105, wherein the NerS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:26.The recombinant microorganism of embodiment 105, wherein the NerS comprises the amino acid sequence of SEQ ID NO:26.The recombinant microorganism of any one of embodiments 105 to 109, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO:41.The recombinant microorganism of any one of embodiments 105 to 109, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO:41.The recombinant microorganism of any one of embodiments 105 to 109, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:41. The recombinant microorganism of embodiment 71 or 72, wherein the NerS is from Streptomyces clavuligerus.The recombinant microorganism of embodiment 113, wherein the NerS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:30.The recombinant microorganism of embodiment 113, wherein the NerS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:30.The recombinant microorganism of embodiment 113, wherein the NerS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:30.The recombinant microorganism of embodiment 113, wherein the NerS comprises the amino acid sequence of SEQ ID NO:30.The recombinant microorganism of any one of embodiments 113 to 117, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO:45.The recombinant microorganism of any one of embodiments 113 to 117, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO:45.The recombinant microorganism of any one of embodiments 113 to 117, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:45. The recombinant microorganism of embodiment 71 or 72, wherein the NerS is from Arabidopsis thaliana.The recombinant microorganism of embodiment 121, wherein the NerS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:33.The recombinant microorganism of embodiment 121, wherein the NerS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:33.The recombinant microorganism of embodiment 121, wherein the NerS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:33.The recombinant microorganism of embodiment 121, wherein the NerS comprises the amino acid sequence of SEQ ID NO:33.The recombinant microorganism of any one of embodiments 121 to 125, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO:48.The recombinant microorganism of any one of embodiments 121 to 125, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO:48.The recombinant microorganism of any one of embodiments 121 to 125, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:48. The recombinant microorganism of embodiment 71 or 72, wherein the NerS is from Citrus unshiu.The recombinant microorganism of embodiment 129, wherein the NerS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:34.The recombinant microorganism of embodiment 129, wherein the NerS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:34.The recombinant microorganism of embodiment 129, wherein the NerS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:34.The recombinant microorganism of embodiment 129, wherein the NerS comprises the amino acid sequence of SEQ ID NO:34.The recombinant microorganism of any one of embodiments 129 to 133, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO:49.The recombinant microorganism of any one of embodiments 129 to 133, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO:49.The recombinant microorganism of any one of embodiments 129 to 133, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:49. The recombinant microorganism of embodiment 71 or 72, wherein the NerS is from Heliconius melpomene.The recombinant microorganism of embodiment 137, wherein the NerS comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NO:35.The recombinant microorganism of embodiment 137, wherein the NerS comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:35.The recombinant microorganism of embodiment 137, wherein the NerS comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:35.The recombinant microorganism of embodiment 137, wherein the NerS comprises the amino acid sequence of SEQ ID NO:35.The recombinant microorganism of any one of embodiments 137 to 141, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80% sequence identity to the nucleotide sequence of SEQ ID NO:50.The recombinant microorganism of any one of embodiments 137 to 141, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 90% sequence identity to the nucleotide sequence of SEQ ID NO:50.The recombinant microorganism of any one of embodiments 137 to 141, wherein the second nucleotide sequence comprises the nucleotide sequence of SEQ ID NO:50. The recombinant microorganism of any one of embodiments 8 to 144, wherein the recombinant microorganism comprises a first nucleic acid comprising the first nucleotide sequence.The recombinant microorganism of embodiment 145, wherein the first nucleic acid is genomically integrated.The recombinant microorganism of embodiment 145, wherein the first nucleic acid is extrachromosomal.The recombinant microorganism of any one of embodiments 145 to 147, wherein the first nucleic acid comprises a first promoter that is operably linked to the first nucleotide sequence.The recombinant microorganism of embodiment 148, wherein the first promoter is an inducible promoter.The recombinant microorganism of embodiment 148, wherein the first promoter is a constitutive promoter.The recombinant microorganism of any one of embodiments 145 to 150, wherein the first nucleic acid further comprises the second nucleotide sequence.The recombinant microorganism of embodiment 151, wherein the first nucleic acid further comprises a second promoter that is operably linked to the second nucleotide sequence.The recombinant microorganism of embodiment 152, wherein the second promoter is an inducible promoter.The recombinant microorganism of embodiment 152, wherein the second promoter is a constitutive promoter.The recombinant microorganism of any one of embodiments 13 to 154, wherein the recombinant microorganism comprises a second nucleic acid comprising the second nucleotide sequence.The recombinant microorganism of embodiment 155, wherein the second nucleic acid is genomically integrated.The recombinant microorganism of embodiment 155, wherein the second nucleic acid is extrachromosomal.The recombinant microorganism of any one of embodiments 155 to 157, wherein the second nucleic acid further comprises a second promoter that is operably linked to the second nucleotide sequence.The recombinant microorganism of embodiment 158, wherein the second promoter is an inducible promoter.. The recombinant microorganism of embodiment 158, wherein the second promoter is a constitutive promoter.. The recombinant microorganism of any one of embodiments 1 to 160, wherein the recombinant microorganism is an E. coli.. The recombinant microorganism of embodiment 161, wherein the E. coli is an E. coli MG1655.. The recombinant microorganism of embodiment 161, wherein the E. coli is an E. coli K12 substrain W3110.. The recombinant microorganism of embodiment 161, wherein the E. coli is an E. coli K12 substrain DH5alpha.. The recombinant microorganism of embodiment 161, wherein the E. coli is an E. coli W.. The recombinant microorganism of embodiment 161, wherein the E. coli is an E. coli BL21.. The recombinant microorganism of any one of embodiments 1 to 166, wherein the recombinant microorganism comprises modifications that increase conversion of glucose or xylose to dimethylallyl pyrophosphate (DMAPP) and isopentenyl pyrophosphate (IPP) via the 1-deoxyxylulose-5-phosphate (DXP) pathway relative to a native microorganism or parental microorganism that does not comprise the modifications.. The recombinant microorganism of embodiment 167, wherein the recombinant microorganism is an E. coli cell comprising a third nucleotide sequence encoding a dxs2 polypeptide comprising the amino acid sequence of SEQ ID NO:55 and a fourth nucleotide sequence encoding a dxr polypeptide (2) comprising the amino acid sequence of SEQ ID NO:58.. The recombinant microorganism of embodiment 168, which further comprises:(a) a fifth nucleotide sequence encoding an ispD polypeptide that is not native to E. co / / ; (b) a sixth nucleotide sequence encoding an ispE polypeptide that is not native to E. coli; (c) a seventh nucleotide sequence encoding an ispF polypeptide that is not native to E.coli;(d) an eighth nucleotide sequence encoding an ispG polypeptide that is not native to E.coll;(e) a ninth nucleotide sequence encoding an ispH polypeptide that is not native to E. coli and(f) a tenth nucleotide sequence encoding an idi polypeptide that is not native to E. coli.. The recombinant microorganism of embodiment 169, wherein the ispD polypeptide comprises an amino acid sequence having at least 95% identity to SEQ ID NO:62, the ispE polypeptide comprises an amino acid sequence having at least 95% identity to SEQ ID NO:65, the ispF polypeptide comprises an amino acid sequence having at least 95% identity to SEQ ID NO:68, the ispG polypeptide comprises an amino acid sequence having at least 95% identity to SEQ ID NO:76, the ispD polypeptide comprises an amino acid sequence having at least 95% identity to SEQ ID NO:92.. The recombinant microorganism of any one of embodiments 1 to 170, wherein the recombinant microorganism comprises modifications that increase the conversion of glucose or xylose to dimethylallyl pyrophosphate (DMAPP) and isopentenyl pyrophosphate (IPP) via the mevalonate pathway relative to a native microorganism or parental microorganism that does not comprise the modifications.. The recombinant microorganism of any one of embodiments 1 to 166 or 171, wherein the recombinant microorganism is S. cerevisiae.. A recombinant microorganism configured to:(a) condense dimethylallyl diphosphate (DMAPP) and isopentenyl diphosphate (IPP) to form neryl diphosphate (NPP); and(b) convert NPP to β-myrcene.. A recombinant microorganism comprising:(a) a means for condensing DMAPP and IPP to form NPP; and(b) a means for converting NPP to β-myrcene.. A recombinant microorganism comprising one or more heterologous nucleic acids comprising:(a) a heterologous nucleotide sequence encoding a neryl diphosphate synthase (NPPS); and(b) a heterologous nucleotide sequence encoding a terpene synthase.. A method of producing β-myrcene, the method comprising culturing the recombinant microorganism of any one of embodiments 1 to 70 or 145 to 175 in a production medium.The method of embodiment 176, wherein the production medium comprises glucose and / or xylose.The method of embodiment 176 or 177, further comprising recovering the p-myrcene from the production medium.The method of anyone of embodiments 176 to 178, further comprising recovering the p-myrcene from the head space above the production medium.A recombinant microorganism configured to:(a) condense dimethylallyl diphosphate (DMAPP) and isopentenyl diphosphate (IPP) to form neryl diphosphate (NPP); and(b) convert NPP to nerol.A recombinant microorganism comprising:(a) a means for condensing DMAPP and IPP to form NPP; and(b) a means for converting NPP to nerol.A recombinant microorganism comprising one or more heterologous nucleic acids comprising:(a) a heterologous nucleotide sequence encoding a neryl diphosphate synthase (NPPS); and(b) a heterologous nucleotide sequence encoding a nerol synthase.A method of producing nerol, the method comprising culturing the recombinant microorganism of any one of embodiments 71 to 172 or 180 to 182 in a production medium.The method of embodiment 183, wherein the production medium comprises glucose and / or xylose.The method of embodiment 183 or 184, further comprising recovering the nerol from the production medium.The method of anyone of embodiments 183 to 185, further comprising recovering the nerol from the head space above the production medium.8. EXAMPLES8.1. Example 1 - p-Myrcene Production by Recombinant E. coli Engineered with MyrS polypeptides8.1.1. Cloning of MyrS variants
[0281] Codon-optimized genes encoding GPPS from Abies grandis (SEQ ID NO: 1), NPPS from Solanum lycopersicum (SEQ ID NO: 2), and MyrS variants from A. grandis (SEQ ID NO: 3), Antirrhinium majus (SEQ ID NO: 4), Humulus lupulus (SEQ ID NO: 5), Ocimum basilicum (SEQ ID NO: 6), Picea abis (SEQ ID NO: 7), Perilla frutescens (SEQ ID NO: 8), Quercus ilex (SEQ ID NO: 9), and Cannabis sativa (SEQ ID NO: 10) were synthesized by GenScript (Piscataway, NJ). These synthetic genes (SEQ ID NO: 13-20) were assembled into pTrcHis2B plasmid (SEQ ID NO: 21) at the Ncol and Hindlll sites such that myrS was upstream of gppS (SEQ ID NO: 11) or nppS (SEQ ID NO: 21), with both genes under control of the trc promoter and a lac operator. The resulting plasmid constructs were subsequently transformed into E. coli DH5alpha and purified using standard protocols. The sequences were confirmed by Sanger sequencing.8.1.2. Myrcene production of MyrS variants
[0282] Plasmids carrying myrS homologs were electroporated into an E. coli MG 1655 strain with a modified DXP pathway. The resulting strains were grown in Lysogeny Broth (LB) containing carbenicillin (100 mg / L) in a shaking incubator at 37°C overnight. The starting optical density at 600 nm (OD600) of each culture was normalized to 0.05 by diluting overnight cultures in a total volume of 1 ml Terrific Broth (TB) containing carbenicillin (100 mg / L). A 0.8-ml aliquot of each TB culture was transferred into a 20-ml headspace vial (Agilent, Santa Clara, CA), and 0.2 ml isopropyl palmitate was added as an overlay solvent to capture isoprenoid products. The vials were sealed and incubated in a shaking incubator at 37°C for 24 h. Isoprenoid production was measured using gas chromatography flame ionization detector (GC-FID) method and normalized to OD600of each culture.
[0283] FIG. 2 shows the normalized myrcene titers of E. coli strains carrying different MyrS variants and A. grandis GPPS or S. lycopersicum NPPS. FIG. 3 shows the normalized titers of myrcene, limonene, and nerol produced from NPP by Q. ilex or C. sativa MyrS variant.8.2. Example 2 - Products from NPP Produced by Monoterpene Synthases
[0284] To explore the production of monoterpenes of interest using NPP as a substrate, the amino acid sequences of 15 terpene synthase enzymes were retrieved from public databases and one scientific publication (see Table 2 and SEQ ID NOs.37-51). Limonene synthase from M. spicata (SEQ ID NO:22) was chosen as a positive control for monoterpene production because it has been shown to produce limonene efficiently using NPP as asubstrate. Of the remaining 14 enzymes, 12 have been studied in the literature using GPP as substrate, and their specificities for producing a particular terpenoid from GPP have been demonstrated, mainly using in vitro assays. The 2 remaining enzymes were annotated in GenBank as (3S,6E)-nerolidol synthases (SEQ ID NO: 38 and 39). Nerolidol is a C15terpene, while nerol is a C terpene. Table 2 shows the names of the terpene synthase enzymes chosen fortesting and the annotations indicating the primary products of the enzymes when using GPP as a substrate as reported in public databases and literature. The C. sativa myrcene synthase (SEQ ID NO:10) was also included in the testing, for a total of 16 enzymes tested.Attorney Docket No. BPC-026WO TABLE 2. Terpene synthases utilizedAttorney Docket No. BPC-026WO 8.2.1. Cloning of Terpene Synthases
[0285] The sequences of all 16 terpene synthase genes listed in Table 2 were E.coli codon-optimized (SEQ ID NOs: 20, 22-36). The terpene synthases were then each cloned into the expression plasmid pTAC2 PAR (SEQ ID NO:52) as Ncol-Hindlll DNA fragments comprising 2-gene operons, where the first gene was a terpene synthase and the second gene was NPP synthase (DNA sequence: SEQ ID NO:12), as shown in FIG. 4. In the expression plasmid, the terpene synthase and NPP synthase genes are under control of the TAC2 promoter, which is inducible by IPTG.
[0286] Expression plasmids were electroporated into an engineered E. coli strain containing a heterologous DXP pathway and reduced PDH activity. Seed cultures were grown overnight in 2 mL minimal media (3-(N-morpholino)propane-1-sulfonic acid (MOPS), 41.86 g / L; tricine, 0.717 g / L; sodium chloride, 50 mM; potassium phosphate, 1.32 mM; ammonium chloride, 9.5 mM; potassium sulfate, 0.276 mM; 0.5 pM; magnesium chloride, 0.525 mM; trace metals (Teknova, Inc.), 0.1X; vitamins (Teknova, Inc.)) supplemented with 5 g / L soytone (US Biological), 1.5 g / L glycerol, and 200 mg / L carbenicillin, then fermentation was initiated at OD600 0.1 in 5 mL of the same media and incubated at 37 °C with shaking at 200 rpm. After 5 hours, cultures (OD600 0.08-1.2) were induced with 1 mM IPTG, and 1 mL isopropyl myristate was added to extract terpenes. Cultures grew for another 24 hours at 30 °C with shaking at 200 rpm, after which terpenes were measured by GC-FID. Titers were corrected to adjust broth volume and normalized to cell density. Total terpenes made, specific titers, and percentages of each terpene made are reported in Table 3. As can be seen, when NPP was used as the substrate in vivo, only 5 terpene synthases (marked with *) produced as the main product the terpene that corresponds to their name as previously annotated. The 2 terpene synthases that have been annotated as (3S,6E)-nerolidol synthases (marked withA) produced nerol as the major product in vivo using NPP as a substrate. Nine terpene synthases (marked with t) produce a different main product from the one indicated by the annotated enzyme name in the public databases and literature. These results indicate that with NPP as a substrate in vivo, some enzymes have a different specificity and generate products different from the products indicated by their annotations, in some cases with a specificity higher than 80%. Finally, the results indicate that using NPP, the two putative nerolidol synthases produce nerol efficiently.Attorney Docket No. BPC-026WO TABLE 3. Terpene productionlojuaja©T.icoivun<£> CPauauotuneuaojXw t cooinineueuyooajoaujO’g'i,« S 'g S? C»,xs <s3 ■■§ 3 £ & '3 f < S>8 « "»> ® 3 $J s.^ © --g~VOo<t>jS & -TotalTerpenesMeasured <» i\ Terpene Total mg / mg 1 « WM< i:! * * »*••«\ Synthase Terpenes ceils JS ■ ®\ Annotation Measured (Specific 9 i Sto i 3Source \ mq / L titer) «• i O7.2 2.52 j S7.2f1.4 | 1.4 IacuOT8a! (Case®)i CineoleCitrus unshiu I Synthase 7.48 3.51 5.4 32.8 7.7 11.3 | 42.8* \_ i (CuSinS) _ - [. —i CineoteHypoxyfonsp. \ Synthase 4.39 4.29 79,0* | 1.7 2.7 | 14.7 2.0 |\ (HypSinS)Sraptoyras j5.65 4.94 15.0 | 80.3f1.4 3.4 { I davuHgerus ^cSj(5S)4.68 1.73 34.2 | 4.5 f rruutsiccoos 60.2* | 1.0saa i Synthase (SfSinS)1 Ocimene1 Sybase 19.26 7.37 \ 3.5 29,1 67,5t|majJS| (AmOcIS)_ _ IJ NleroMyrcene| LimoenenLillnaooGileranoTotal. J [Terpenes IoMeasured ® §Terpene Total mgfmg £ E O ■ Synthase Terpenes ceils <£ | 2 c Annotation Measured (Specific 2 fl?! Source j mg / L titerj | 1. I.. i pSTse 5.82 2.06 | 0.3 99.7*..(AtOciS)" Ocimene. I.Citrus unshiti Synthase 6.32 2.6 J | 5.0 79.2* 15.8I (CuOcIS)Helicanius Ocintene ISynthase 6.82 4.92 1 31.8 5 14.7 I ma / pomene |(HymOdS)3.5*n,. Ocimene10.18 3.47 36.3 63.7*teate... 1 (HOBS) _ I _ I _ _‘ Annotated enzyme name in the literature corresponds to the major product produced from NPPAEnzymes annotated as (SS. OEj-neroiidoi synthases., which produced mainly nerol from NPPt Enzymes that, when using NPP;produced a different primary product than indicated by the annotated enzyme name| Mycener[ Limonene| LllinaooAttorney Docket No. BPC-026WO 9. SEQUENCES
[0287] Exemplary sequences referred to herein are provided in Table 1 below (where “SEQ” refers to the SEQ ID NO).Table 1Description Sequence SEQ: A. grandis GPPS MFDFNKYMDSKAMTVNEALNKAIPLRYPQKIYESMRYSLLAGGKRVR 1 (AgGPPS) PVLCIAACELVGGTEELAIPTACAIEMIHTMSLMHDDLPCIDNDDLRRG KPTNHKIFGEDTAVTAGNALHSYAFEHIAVSTSKTVGADRILRMVSEL GRATGSEGVMGGQMVDIASEGDPSIDLQTLEWIHIHKTAMLLECSVV CGAIIGGASEIVIERARRYARCVGLLFQVVDDILDVTKSSDELGKTAGK DLISDKATYPKLMGLEKAKEFSDELLNRAKGELSCFDPVKAAPLLGLA DYVAFRQNS. lycopersicum NPPS MSARGLNKISCSLNLQTEKLCYEDNDNDLDEELMPKHIALIMDGNRR 2 (SINPPS) WAKDKGLEVYEGHKHIIPKLKEICDISSKLGIQIITAFAFSTENWKRSKE EVDFLLQMFEEIYDEFSRSGVRVSIIGCKSDLPMTLQKCIALTEETTKG NKGLHLVIALNYGGYYDILQATKSIVNKAMNGLLDVEDINKNLFDQELE SKCPNPDLLIRTGGEQRVSNFLLWQLAYTEFYFTNTLFPDFGEEDLKE AIMNFQQRHRRFGGHTYA. grandis MyrS MSISLATAAPDDGVQRRIGDYHSNIWDDDFIQSLSTPYGEPSYQERA 3 (AgMyrS) ERLIVEVKKIFNSMYLDDGRLMSSFNDLMQRLWIVDSVERLGIARHFK NEITSALDYVFRYWEENGIGCGRDSIVTDLNSTALGFRTLRLHGYTVS PEVLKAFQDQNGQFVCSPGQTEGEIRSVLNLYRASLIAFPGEKVMEE AEIFSTRYLKEALQKIPVSALSQEIKFVMEYGWHTNLPRLEARNYIDTL EKDTSAWLNKNAGKKLLELAKLEFNIFNSLQQKELQYLLRWWKESDL PKLTFARHRHVEFYTLASCIAIDPKHSAFRLGFAKMCHLVTVLDDIYDT FGTIDELELFTSAIKRWNSSEIEHLPEYMKCVYMWFETVNELTREAE KTQGRNTLNYVRKAWEAYFDSYMEEAKWISNGYLPMFEEYHENGKV SSAYRVATLQPILTLNAWLPDYILKGIDFPSRFNDLASSFLRLRGDTRC YKADRDRGEEASCISCYMKDNPGSTEEDALNHINAMVNDIIKELNWEL LRSNDNIPMLAKKHAFDITRALHHLYIYRDGFSVANKETKKLVMETLLE SMLFA. majus MyrS MAELPMDYEGKIKETRHLLHLKGENDPIESLIFVDATLRLGVNHHFQK 4 (AmMyrS) EIEEILRKSYATMKSPIICEYHTLHEVSLFFRLMRQHGRYVSADVFNNF KGESGRFKEELKRDTRGLVELYEAAQLSFEGERILDEAENFSRQILHG NLAGMEDNLRRSVGNKLRYPFHTSIARFTGRNYDDDLGGMYEWGKT LRELALMDLQVERSVYQEELLQVSKWWNELGLYKKLNLARNRPFEFY TWSMVILADYINLSEQRVELTKSVAFIYLIDDIFDVYGTLDELIIFTEAVN KWDYSATDTLPENMKMCCMTLLDTINGTSQKIYEKHGYNPIDSLKTT WKSLCSAFLVEAKWSASGSLPSANEYLENEKVSSGVYVVLVHLFCLM GLGGTSRGSIELNDTQELMSSIAIIFRLWNDLGSAKNEHQNGKDGSYL NCYKKEHINLTAAQAHEHALELVAIEWKRLNKESFNLNHDSVSSFKQA ALNLARMVPLMYSYDHNQRGPVLEEYVKFMLSDH. lupulus MyrS MTVVNNTDRRSANYEPSIWSFDYIQSLTSQYKGKSYSSRLNELKKEV 5 (HIMyrS) KMMEDGTKECLAQLDLIDTLQRLGISYHFEDEINTILKRKYINIQNNINH NYNLYSTALQFRLLRQHGYLVTQEVFNAFKDETGKFKTYLSDDIMGVL SLYEASFYAMKHENVLEEARVFSTECLKEYMMKMEQNKVLLDHDLD HNDNFNVNHHVLIINHALELPLHWRITRSEARWFIDVYEKKQDMDSTL LEFAKLDFNMVQSTHQEDLKHLSRWWRHSKLGEKLNFARDRLMEAFLWEVGLKFEPEFSYFKRISARLFVLITIIDDIYDVYGTLEELELFTKAVERWDVNAINELPEYMKMPFLVLHNTINEMAFDVLGDQNFLNIEYLKKSLV DLCKCYLQEAKWYYSGYQPTLQEYIEMAWLSIGGPVILVHAYFCFTN PITKESMKFFTEGYPNIIQQSCLIVRLADDFGTFSDELNRGDVPKSIQC YMYDTGASEDEAREHIKFLICETWKDMNKNDEDNSCFSETFVEVCKN LARTALFMYQYGDGHASQNCLSKERIFALIINPINFHERK0. basilicum MyrS MVEPRRSGNYQPSAWDFNYIQSLNNNHSKEERHLERKAKLIEEVKML 6 (ObMyrS) LEQEMAAVQQLELIEDLKNLGLSYLFQDEIKIILNSIYNHHKCFHNNHE QCIHVNSDLYFVALGFRLFRQHGFKVSQEVFDCFKNEEGSDFSANLA DDTKGLLQLYEASYLVTEDEDTLEMARQFSTKILQKKVEEKMIEKENL LSWTLHSLELPLHWRIQRLEAKWFLDAYASRPDMNPIIFELAKLEFNIA QALQQEELKDLSRWWNDTGIAEKLPFARDRIVESHYWAIGTLEPYQY RYQRSLIAKIIALTTVVDDVYDVYGTLDELQLFTDAIRRWDIESINQLPS YMQLCYLAIYNFVSELAYDIFRDKGFNSLPYLHKSWLDLVEAYFVEAK WFHDGYTPTLEEYLNNSKITIICPAIVSEIYFAFANSIDKTEVESIYKYHD ILYLSGMLARLPDDLGTSSFEMKRGDVAKAIQCYMKEHNASEEEARE HIRFLMREAWKHMNTAAAADDCPFESDLVVGAASLGRVANFVYVEG DGFGVQHSKIHQQMAELLFYPYQPicea abis MyrS MSTDELKPLPTTIPTRGMCGRRMSVTPSMSMSLNTVVSDNDAVQRRI 7 (PaMyrS) GDYHSNLWNDDFIQSLTTPYGAPSYIERADRLISEVKEMFNRMCMED GELMSPLNDLIQRLWTVDSVERLGIDRHFKNEIKASLDYVYSYWNEK GIGCGRQSVVTDLNSTALGLRILRQHGYTVSSEVLKVFEEENGQFAC SPSQTEGEIRSFLNLYRASLIAFPGEKVMEEAQIFSSRYLKEAVQKIPV SGLSREIGDVLEYGWHTNLPRWEARNYMDVFGQDTNTSFNKNKMQ YMNTEKILQLVKLEFNIFHSLQQRELQCLLRWWKESGLPQLTFARHR HVEFYTLASCIACEPKHSAFRLGFAKMCHLVTVLDDVYDTFGKMDEL ELFTAAVKRWDLSETERLPEYMKGLYVVVFETVNELAQEAEKTQGRN TLNYVRKAWEAYFDSYMKEAEWISTGYLPTFEEYCENGKVSSAYRVA ALQPILTLDVQLPDDILKGIDFPSRFNDLASSFLRLRGDTRCYEADRAR GEEASCISCYMKDNPGSTEEDALNHINAMINDIIRELNWEFLKPDSNIP MPARKHAFDITRALHHLYIYRDGFSVANKETKNLVEKTLLESMLFPerilla frutescens MyrS MQLSDQRRSGNYSPSFWNTDYILSLNCDYEDERRMRGAAGELVEQV 8 (PfMyrS) KMLMEKETDPIVQLELIDVLQKLALSHHFEKEFEGILFNISTIYDDKNRE RDLYSTTLAFRLLRQHGYQVPQELFECFKNDKGEFKESLSNDTKGLL QLYEASFLLTEGETTLELAREFATKFLQEKEKLNIDDDDDTNLISCVRH SLDMPIYWRIQRPNARWWIHAYNRRTHINPLVLELSKLDFNIIQAQYQ QELKQDLRWWRNTCIAEKLPFARDRLVESYFWSTGIIQPRQHENARI MMAKALALITTLDDVYDVYGTLEELELFIEAIRRWEISSIDQLPNYMQL CFLTINNFVDDTAYDVMKEKDINIIPYLRKSWVDLAEAYLVEAKWFYG GYKPNLEEYLNNGWISVSGPAILCHVFFGVTDSITMETVESLFKYHDLI RCSSTLVRLADDLATSLDEVSRGDVPKSIQCYMNDNNASEEEARLHV RWLIAETWKEMNVEMVSADSPFCKDFIACAADMGRMAQYMYHNGD GHGMQNSQIHQQMTDFLFQKLAVRDRASTARNQuercus ilex MyrS MVANKVSTSPDILRRSANYQPSIWNHDYIESLRIEYVGETCTRQINVLK 9 (QiMyrS) EQVRMMLHKWNPLEQLELIEILQRLGLSYHFEEEIKRILDGVYNNDH GGDTWKAENLYATALKFRLLRQHGYSVSQEVFNSFKDERGSFKACL CEDTKGMLSLYEASFFLIEGENILEEARDFSTKHLEEYVKQNKEKNLA TLVNHSLEFPLHWRMPRLEARWFINIYRHNQDVNPILLEFAELDFNIV QAAHQADLKQVSTWWKSTGLVENLSFARDRPVENFFWTVGLIFQPQ FGYCRRMFTKVFALITTIDDVYDVYGTLDELELFTDVVERWDINAMDQ LPDYMKICFLTLHNSVNEMALDTMKEQRFHIIKYLKKAWVDLCRYYLV EAKWYSNKYRPSLQEYIENAWISIGAPTILVHAYFFVTNPITKEALDCLEEYPNIIRWSSIIARLADDLGTSTDELKRGDVPKAIQCYMNETGASEEGAREYIKYLISATWKKMNKDRAASSPFSHIFIEIALNLARMAQCLYQHG DGHGLGNRETKDRILSLLIQPIPLNKDCannabis sativa MyrS MCYPIQCTVVNNSSPSSTIVRRSANYEPPIWSFDYIQSLSTQYKGESY 10 (CsMyrS) TGQLNKLKKEVKRMLLRMEINSLALLELIDTLQRLGISYHFKNEINTILK KKYNDNYINNNIITSPNYNNLYATALEFRLLRQHGYTVPQEIFNAFKDK RGKFKTSLSDDIMGVLCLYEASFYAMKHENILEEARIFSTKCLKKYME KIENEEEKKILLLNDNNINSNLLLINHAFELPLHWRITRSEARWFIDEIYE KKQDMNSTLFEFAKLDFNIVQSTHQEDLQHLSRWWRDCKLGGKLNF ARDRLMEAFLWDVGLKFEGEFSYFRRTNARLFVLITIIDDIYDVYGTLE ELELFTSAVERWDVKLINELPDYMKMPFFVLHNTINEMGFDVLVEQNF VNIEYLKKSWVDLCKCYLQEAKWYYSGYQPTLEEYTELGWLSIGASVI LMHAYFCFTNPITKQDLKSLQLQHHYPNIIKQACLITRLANDLGTSSDE LNRGDVPKSIQCYMYDNNATEDEAREHIKFLISETWKDMNKKDEDES CLSENFVEVCKNMARTALFIYENGDGHGSQNSLSKERISTLIITPINIPKA. grandis gppS ATGTTTGATTTTAACAAGTATATGGACTCGAAGGCGATGACCGTGA 11ACGAAGCTCTGAACAAGGCTATCCCACTGCGCTATCCCCAGAAAA TTTACGAGTCGATGCGCTACAGCTTGTTAGCTGGTGGTAAGCGTG TGCGCCCAGTCTTATGTATCGCTGCATGCGAATTAGTTGGGGGGA CTGAAGAGCTGGCAATTCCGACTGCCTGTGCAATTGAAATGATTC ATACGATGTCTCTGATGCACGATGACTTACCCTGCATTGATAATGA CGATCTGCGTCGTGGAAAGCCTACCAATCACAAGATTTTCGGGGA GGATACCGCAGTAACAGCCGGAAATGCGCTTCACTCCTACGCTTT CGAGCACATCGCAGTTTCTACTAGCAAGACGGTGGGGGCTGATC GCATCTTGCGTATGGTTAGCGAGCTGGGACGTGCAACGGGATCC GAGGGGGTCATGGGGGGGCAGATGGTGGACATTGCATCCGAGG GAGACCCTAGTATTGACCTGCAGACCCTGGAATGGATTCATATTCA TAAGACGGCTATGCTTTTAGAATGTAGCGTTGTCTGCGGTGCGATT ATCGGCGGAGCCTCCGAAATTGTGATCGAGCGCGCACGCCGTTA TGCACGTTGCGTGGGGCTTCTTTTCCAGGTTGTTGACGACATCTTA GATGTAACCAAGTCTTCTGACGAGCTGGGGAAGACCGCCGGAAA GGACCTGATTAGCGACAAGGCAACGTATCCAAAGTTAATGGGCTT GGAGAAAGCGAAAGAGTTCTCTGACGAGTTACTGAATCGCGCCAA AGGAGAACTGTCATGTTTTGATCCAGTGAAAGCCGCCCCCCTTTT GGGTCTTGCTGACTATGTGGCGTTCCGCCAAAATTAAS. lycopersicum nppS ATGTCCGCCCGTGGCCTTAATAAAATTTCCTGTTCACTGAACTTGC 12AGACGGAGAAATTGTGCTACGAGGACAATGACAACGACCTGGATG AAGAATTAATGCCCAAGCACATCGCCCTGATTATGGACGGCAACC GCCGTTGGGCAAAAGACAAAGGCCTGGAAGTGTACGAGGGTCAT AAACACATTATCCCAAAGCTGAAAGAAATCTGCGACATTAGCTCAA AGCTGGGGATTCAAATTATCACCGCTTTCGCTTTCTCGACGGAGA ACTGGAAGCGCTCTAAAGAAGAGGTCGACTTCTTGTTACAAATGTT CGAAGAAATCTATGATGAATTTAGTCGCAGTGGAGTTCGCGTCAG CATCATCGGATGTAAAAGCGATTTACCGATGACCCTGCAAAAGTGT ATTGCCCTTACAGAGGAGACTACGAAGGGGAACAAGGGCCTGCA CCTGGTAATTGCCCTTAATTACGGCGGATACTATGACATCCTTCAA GCTACCAAGAGTATTGTGAACAAAGCAATGAACGGACTTTTGGAT GTGGAGGATATTAACAAAAACTTATTCGACCAGGAATTAGAAAGTA AGTGCCCAAATCCCGACCTTTTAATCCGTACAGGGGGTGAACAAC GTGTATCAAATTTCCTGTTGTGGCAGTTGGCTTACACGGAGTTCTA CTTTACAAACACTTTATTTCCTGACTTTGGTGAGGAAGACTTAAAA GAGGCCATTATGAACTTCCAACAGCGCCATCGTCGCTTTGGCGGC CACACTTATTAAA. gran dis myrS ATGTCTATTTCTCTGGCTACGGCAGCCCCAGATGATGGAGTGCAG 13 CGTCGTATCGGGGACTATCATTCTAACATTTGGGATGATGATTTTA TCCAGAGCCTTTCCACGCCGTACGGCGAGCCCTCGTATCAAGAGC GTGCAGAACGTTTGATCGTAGAGGTCAAAAAAATTTTCAACAGTAT GTATTTAGATGACGGACGCCTTATGTCCAGTTTTAATGATCTGATG CAGCGCCTTTGGATTGTGGATTCAGTCGAACGTTTAGGTATCGCC CGTCATTTCAAAAACGAGATTACCTCTGCACTTGACTACGTGTTCC GCTACTGGGAAGAAAACGGAATCGGTTGTGGCCGCGACTCAATTG TAACGGACTTGAACTCCACAGCTTTGGGCTTCCGTACCTTGCGCT TACATGGTTATACAGTATCTCCGGAAGTGCTTAAAGCGTTTCAGGA CCAAAATGGACAATTTGTGTGCAGCCCAGGTCAAACAGAAGGTGA AATCCGCTCGGTTTTGAATTTATACCGCGCTTCATTGATTGCGTTC CCAGGAGAAAAAGTTATGGAAGAGGCTGAAATTTTTTCAACTCGCT ACCTTAAGGAAGCCCTGCAAAAGATCCCGGTGAGCGCCCTTAGCC AAGAAATCAAATTCGTGATGGAGTATGGGTGGCATACCAACTTACC TCGCCTTGAAGCGCGCAATTATATCGATACACTTGAGAAAGATACC TCCGCTTGGTTGAACAAAAACGCGGGCAAAAAATTACTGGAACTT GCTAAGTTAGAGTTTAACATTTTCAATAGTCTTCAACAAAAGGAACT TCAATATCTTTTACGTTGGTGGAAAGAATCCGACCTGCCAAAATTA ACCTTCGCCCGTCACCGCCACGTGGAGTTCTACACGCTGGCCTCA TGCATTGCAATCGACCCTAAACACTCGGCATTTCGCTTGGGCTTTG CGAAAATGTGCCACCTGGTTACAGTTTTGGATGACATCTACGACAC ATTTGGGACTATTGACGAACTGGAGCTGTTCACATCGGCCATCAA ACGTTGGAATTCATCAGAGATCGAGCATCTGCCTGAATATATGAAA TGTGTGTATATGGTGGTTTTCGAGACTGTAAACGAGCTGACCCGC GAAGCCGAGAAAACCCAAGGCCGCAACACATTAAACTATGTTCGC AAAGCGTGGGAGGCGTACTTCGACTCATACATGGAAGAAGCAAAA TGGATCAGTAATGGGTACTTACCCATGTTTGAAGAGTATCACGAGA ACGGAAAAGTGTCTAGCGCATACCGCGTCGCGACTTTGCAGCCGA TCCTGACTCTGAATGCTTGGTTGCCTGACTACATTTTGAAAGGAAT CGACTTTCCCTCTCGCTTTAACGACCTGGCAAGCTCATTTCTTCGC CTTCGTGGCGACACGCGCTGTTACAAGGCAGACCGTGATCGCGG GGAGGAAGCCTCTTGCATCTCTTGCTACATGAAGGACAACCCGGG GAGCACGGAGGAAGACGCATTGAATCATATCAATGCGATGGTGAA CGATATCATCAAAGAGCTTAACTGGGAATTGCTGCGCTCAAACGAT AACATCCCCATGTTGGCAAAAAAACATGCGTTTGACATTACCCGCG CCTTGCATCATCTGTATATCTACCGTGACGGCTTCTCGGTCGCCAA TAAAGAAACAAAAAAATTAGTGATGGAGACCTTATTAGAATCGATG CTTTTTTAAA. majus myrS ATGGCGGAGTTACCGATGGATTATGAGGGCAAGATCAAAGAAACA 14CGTCACTTATTGCACTTAAAGGGCGAGAATGATCCCATCGAGTCG TTAATTTTCGTAGATGCGACTTTGCGTCTTGGTGTGAACCACCATT TCCAGAAAGAGATTGAAGAGATTTTACGCAAAAGCTACGCTACCAT GAAATCGCCAATCATCTGCGAATATCACACCTTACATGAGGTTTCG CTGTTTTTCCGTTTGATGCGTCAACATGGACGCTATGTATCGGCG GACGTCTTCAATAATTTTAAGGGGGAATCAGGCCGCTTTAAGGAG GAATTAAAACGCGACACTCGCGGGTTAGTTGAGTTATATGAAGCC GCTCAGCTTTCTTTTGAGGGAGAGCGTATTTTAGATGAAGCAGAG AACTTCTCTCGCCAGATCTTACACGGTAACTTGGCTGGAATGGAA GACAACCTTCGCCGTTCCGTGGGTAATAAATTGCGTTACCCCTTC CATACATCAATCGCTCGTTTCACAGGGCGTAATTACGACGATGAC CTTGGGGGTATGTACGAGTGGGGCAAGACTTTACGTGAATTAGCA CTGATGGACCTTCAGGTCGAACGTTCTGTCTATCAAGAGGAACTGTTACAAGTAAGTAAATGGTGGAACGAATTGGGATTGTATAAAAAACTGAACTTAGCGCGTAATCGTCCCTTTGAATTCTATACGTGGAGTAT GGTAATTCTTGCCGACTACATCAACCTTTCTGAACAACGCGTCGAG CTTACAAAATCAGTCGCGTTCATCTATTTGATCGACGATATCTTCG ACGTCTACGGTACATTGGATGAGCTTATCATCTTTACCGAAGCCGT GAATAAGTGGGACTATAGCGCGACAGATACCTTGCCAGAGAATAT GAAAATGTGCTGTATGACACTTCTTGATACGATTAATGGAACCTCC CAGAAGATCTATGAAAAACACGGCTACAATCCTATCGATTCTCTGA AGACAACATGGAAAAGTCTGTGCAGTGCCTTCTTGGTGGAAGCTA AGTGGTCTGCCAGTGGCTCTTTGCCATCAGCCAATGAGTATTTAG AGAATGAAAAAGTCTCATCTGGGGTCTATGTTGTTTTAGTACACTT GTTCTGCTTGATGGGGCTGGGGGGTACCTCCCGCGGCTCAATTG AATTAAACGATACTCAGGAGCTGATGAGTTCTATCGCGATTATTTT CCGCTTGTGGAATGATTTAGGGTCTGCCAAGAATGAGCACCAAAA TGGAAAAGACGGCTCCTATCTGAACTGCTACAAGAAGGAACATATT AATCTTACCGCCGCACAAGCGCACGAGCACGCATTAGAACTTGTT GCTATTGAGTGGAAACGCTTGAATAAAGAATCATTTAATCTTAATCA CGACTCGGTATCCTCCTTCAAACAGGCCGCCTTGAATTTGGCCCG TATGGTTCCGCTGATGTATTCATATGACCATAACCAGCGCGGCCC GGTTTTGGAAGAGTACGTGAAATTCATGCTGTCTGATTAAH. lupulus myrS ATGACAGTAGTAAACAATACAGATCGTCGTAGTGCCAACTATGAGC 15CATCCATTTGGTCCTTCGACTATATCCAGTCGTTGACGTCGCAGTA TAAGGGCAAGAGTTACAGCTCACGCTTAAACGAATTAAAGAAAGAA GTAAAGATGATGGAGGATGGCACGAAAGAGTGTCTGGCCCAACTG GATCTTATTGACACCCTGCAACGTCTTGGTATCAGTTATCACTTTG AGGATGAGATTAACACTATTCTTAAGCGTAAATACATTAATATTCAA AATAATATTAATCACAACTATAACCTTTATTCGACCGCGCTGCAATT CCGTCTGCTGCGCCAACATGGTTATCTTGTAACTCAAGAGGTCTTC AATGCATTCAAGGACGAAACAGGGAAATTTAAGACGTATTTATCAG ACGATATCATGGGGGTGCTTAGCCTGTACGAAGCATCCTTCTACG CCATGAAACATGAGAACGTTTTAGAGGAAGCTCGCGTCTTTAGCA CAGAGTGTTTGAAAGAATACATGATGAAAATGGAGCAGAATAAAGT TCTTCTGGACCACGATCTTGATCACAACGATAATTTCAACGTAAAC CATCATGTGCTTATCATCAACCATGCCTTAGAGTTGCCCTTGCACT GGCGTATCACCCGCTCCGAGGCGCGTTGGTTCATTGATGTCTATG AGAAAAAGCAGGACATGGATTCCACATTATTGGAGTTTGCCAAGTT AGATTTTAATATGGTTCAGTCGACACATCAGGAGGATCTGAAGCAC CTGTCTCGCTGGTGGCGTCACTCTAAGTTGGGAGAGAAGTTAAAC TTTGCTCGCGATCGTCTTATGGAAGCATTTCTGTGGGAAGTTGGG CTGAAGTTTGAGCCAGAATTTTCTTATTTCAAACGCATCTCTGCTC GTTTGTTTGTACTGATCACAATCATCGACGATATCTACGATGTGTA CGGTACGCTGGAGGAGTTAGAACTTTTTACCAAAGCTGTAGAACG CTGGGATGTAAATGCCATCAACGAACTTCCGGAATATATGAAAATG CCATTCTTAGTACTGCATAACACCATTAACGAAATGGCTTTCGACG TATTGGGTGACCAAAACTTCTTAAATATCGAGTATCTGAAAAAATC CTTGGTTGACTTATGCAAGTGCTATTTACAAGAAGCAAAATGGTAT TACTCAGGCTATCAACCCACATTGCAGGAATATATCGAGATGGCTT GGTTGTCTATTGGGGGGCCGGTTATTTTAGTCCATGCGTACTTCT GTTTTACTAATCCTATCACAAAAGAGAGCATGAAGTTTTTTACCGAA GGCTATCCTAATATCATTCAACAGTCATGCCTTATCGTACGTCTTG CTGACGACTTTGGCACATTTTCTGATGAATTAAATCGCGGTGATGT ACCGAAATCGATTCAATGTTATATGTACGACACTGGTGCTTCTGAA GATGAGGCGCGTGAACACATTAAATTTCTGATTTGCGAAACCTGG AAAGACATGAATAAGAACGACGAAGACAATTCATGCTTTTCAGAAACATTCGTTGAAGTGTGTAAAAATTTGGCGCGTACGGCGTTATTCATGTACCAGTACGGTGATGGCCATGCTAGCCAGAATTGCCTTTCCAA AGAACGTATCTTTGCCTTGATTATCAACCCCATTAATTTTCATGAGC GTAAATAA0 basilicum myrS ATGGTTGAGCCCCGTCGCTCGGGCAATTATCAACCCAGCGCGTG 16GGATTTTAATTACATCCAGTCATTAAATAATAATCACTCGAAAGAGG AACGTCACCTGGAGCGTAAGGCGAAGTTGATCGAAGAAGTCAAAA TGTTATTGGAACAAGAAATGGCTGCCGTTCAGCAGCTGGAACTTAT TGAAGACCTGAAGAACCTGGGCTTATCTTACCTTTTTCAGGACGAG ATTAAGATTATCTTAAACAGCATCTATAACCACCATAAGTGCTTCCA CAACAATCATGAGCAGTGCATCCATGTCAACTCGGATCTGTACTTC GTAGCCTTAGGTTTCCGCCTGTTCCGCCAACATGGCTTCAAGGTA AGTCAGGAAGTATTCGACTGTTTTAAAAATGAGGAGGGTTCCGACT TCTCTGCAAACTTGGCTGATGACACCAAGGGGCTTCTTCAATTATA TGAAGCATCTTACCTGGTGACCGAGGATGAAGATACATTAGAGAT GGCTCGTCAGTTTTCAACGAAAATCCTTCAGAAAAAGGTTGAAGAG AAGATGATTGAGAAAGAAAATCTGCTTAGTTGGACACTGCACAGTC TGGAACTGCCACTTCATTGGCGTATCCAACGTCTTGAGGCCAAGT GGTTCTTAGATGCGTACGCTTCCCGTCCAGATATGAATCCAATCAT TTTTGAATTAGCGAAGTTGGAGTTTAATATCGCTCAGGCTTTACAA CAGGAGGAGTTAAAGGACCTGTCACGTTGGTGGAATGACACAGGA ATCGCGGAAAAATTGCCATTTGCCCGCGACCGCATTGTCGAGTCG CACTATTGGGCGATCGGGACACTGGAACCCTATCAATACCGTTAT CAGCGTAGCCTTATCGCAAAAATTATCGCGCTTACGACTGTAGTC GATGATGTATATGACGTTTATGGAACCTTGGACGAGCTTCAATTGT TCACGGATGCGATTCGCCGCTGGGACATCGAATCTATTAACCAGT TGCCTTCATACATGCAACTTTGTTATTTGGCTATTTATAACTTCGTA TCGGAGCTGGCGTATGACATTTTTCGCGATAAGGGCTTCAACTCT CTTCCATACCTGCACAAAAGCTGGCTTGACCTTGTTGAAGCATATT TTGTCGAGGCAAAATGGTTTCACGACGGATACACACCTACTCTTGA GGAGTACCTTAATAATTCGAAGATTACGATTATTTGCCCGGCAATT GTATCTGAAATCTATTTTGCGTTTGCTAATAGTATTGATAAGACTGA GGTTGAATCGATTTACAAGTACCATGACATCTTGTATCTTTCTGGG ATGCTGGCCCGTCTGCCGGACGACTTAGGAACTAGCAGCTTCGAA ATGAAGCGTGGAGACGTTGCTAAAGCTATTCAGTGTTATATGAAG GAACATAACGCTTCCGAGGAAGAGGCCCGTGAGCATATCCGTTTC TTAATGCGCGAGGCTTGGAAACATATGAATACTGCTGCCGCAGCG GATGATTGTCCATTTGAGAGCGATCTTGTGGTTGGGGCCGCGAGC CTTGGCCGTGTGGCTAACTTTGTTTACGTAGAGGGAGACGGCTTT GGCGTACAGCATTCGAAAATTCATCAGCAGATGGCCGAATTGTTG TTTTACCCTTATCAATAAPicea abis myrS ATGAGCACGGATGAACTGAAGCCCTTGCCCACTACGATTCCTACT 17CGTGGTATGTGCGGACGCCGCATGAGTGTCACCCCCTCGATGTC CATGTCGCTTAATACAGTGGTATCTGATAATGACGCAGTACAGCGT CGTATCGGCGACTACCATTCTAATTTGTGGAATGATGACTTCATCC AGTCTTTGACTACGCCTTATGGTGCACCTTCGTACATCGAGCGTG CCGATCGTCTTATTTCAGAGGTCAAAGAGATGTTCAACCGCATGTG TATGGAAGATGGTGAGCTGATGTCTCCTCTGAACGACCTGATCCA ACGTTTGTGGACGGTCGATTCCGTTGAGCGTCTGGGCATTGACCG TCACTTTAAGAATGAAATCAAAGCCAGCCTGGATTATGTCTACTCG TATTGGAATGAAAAGGGCATCGGGTGCGGACGTCAATCCGTAGTA ACAGATTTAAATTCGACAGCTTTAGGTCTTCGTATTTTGCGCCAGC ATGGTTACACGGTTTCATCAGAGGTCTTAAAAGTCTTTGAGGAAGA GAACGGACAGTTTGCGTGTTCTCCATCACAGACCGAAGGCGAAATCCGCAGTTTTTTAAACCTTTACCGCGCCTCACTGATCGCCTTCCCTGGAGAAAAGGTGATGGAGGAAGCTCAGATTTTCAGCTCCCGCTAT CTTAAGGAAGCAGTTCAAAAAATCCCCGTATCTGGTCTTTCTCGCG AAATTGGAGATGTTTTGGAGTATGGGTGGCACACTAACCTGCCCC GTTGGGAGGCGCGCAACTATATGGACGTTTTCGGGCAAGATACTA ACACGTCGTTCAACAAAAACAAGATGCAATATATGAATACCGAAAA AATTTTGCAGCTGGTGAAACTTGAATTCAACATTTTTCACTCCCTTC AACAACGCGAATTACAGTGCCTTTTACGTTGGTGGAAGGAATCGG GATTGCCGCAGTTGACGTTCGCGCGTCACCGTCACGTGGAATTCT ATACGCTGGCATCTTGTATTGCGTGCGAGCCCAAGCACAGTGCGT TTCGCCTGGGATTCGCAAAGATGTGCCACCTTGTAACTGTTCTTGA CGATGTATATGATACGTTTGGGAAAATGGATGAGCTGGAATTGTTC ACCGCAGCCGTTAAACGTTGGGATTTAAGTGAGACTGAACGCCTT CCAGAGTATATGAAGGGGCTGTATGTAGTTGTGTTCGAAACGGTG AATGAGCTTGCGCAAGAAGCCGAGAAAACACAAGGTCGCAATACG TTAAATTATGTTCGTAAGGCATGGGAGGCATACTTTGACAGCTATA TGAAAGAAGCGGAGTGGATCAGCACGGGATATCTGCCGACGTTTG AGGAATACTGCGAGAACGGGAAGGTAAGTAGCGCATATCGCGTC GCCGCACTTCAGCCGATCCTGACACTTGATGTCCAACTTCCCGAT GACATCCTTAAGGGCATCGATTTTCCGTCCCGTTTTAACGACTTGG CCTCTTCTTTCTTACGTCTGCGTGGCGATACGCGCTGTTACGAAG CCGACCGTGCACGTGGTGAGGAAGCGTCCTGTATTAGTTGCTATA TGAAGGATAACCCGGGGAGCACCGAGGAGGATGCTCTTAACCAC ATCAATGCCATGATTAACGATATCATCCGCGAGTTAAATTGGGAAT TTTTGAAGCCCGACTCAAATATTCCAATGCCTGCTCGCAAGCATGC ATTCGATATTACCCGCGCTTTGCACCATTTATATATCTATCGTGAC GGTTTTAGCGTTGCAAATAAAGAAACGAAGAACTTGGTTGAAAAAA CACTTTTAGAGAGTATGTTATTCTAAPerilla frutescens myrS ATGCAACTTTCGGATCAACGCCGTAGCGGCAATTATAGCCCGTCT 18TTTTGGAATACCGACTACATTTTATCTTTGAATTGTGACTACGAGGA CGAGCGCCGTATGCGCGGTGCAGCGGGTGAACTTGTCGAACAAG TCAAAATGTTGATGGAGAAGGAAACCGATCCTATTGTGCAATTGGA ACTGATTGATGTCCTTCAAAAACTGGCCCTTTCCCACCACTTCGAG AAGGAGTTCGAGGGTATCTTATTTAACATCTCTACCATTTACGACG ATAAGAACCGCGAACGTGATCTTTACTCTACAACATTAGCTTTCCG TTTGTTGCGCCAACACGGCTACCAGGTTCCCCAGGAGTTGTTTGA GTGTTTTAAGAATGACAAAGGTGAGTTTAAAGAAAGCCTGTCAAAT GACACGAAGGGGTTATTGCAGCTTTATGAAGCGAGCTTCTTATTGA CTGAAGGAGAAACAACGTTGGAACTTGCACGCGAGTTTGCAACTA AATTCCTTCAGGAGAAGGAAAAGCTTAATATCGATGATGACGACGA TACGAACTTAATCTCCTGCGTACGTCACTCGTTAGACATGCCGATC TACTGGCGTATTCAACGCCCAAATGCTCGCTGGTGGATTCATGCG TATAATCGCCGTACACACATCAACCCATTGGTCTTAGAGTTGTCAA AGTTAGACTTTAATATCATCCAGGCGCAGTACCAACAGGAATTAAA GCAAGACTTACGTTGGTGGCGTAACACATGTATTGCGGAGAAGCT TCCATTCGCACGTGACCGTCTGGTGGAATCCTATTTCTGGTCGAC CGGTATCATTCAACCGCGTCAGCACGAGAATGCACGTATTATGAT GGCCAAGGCCCTGGCGCTTATCACAACGCTTGACGACGTTTACGA CGTATATGGGACATTGGAGGAGTTGGAGCTGTTTATTGAGGCTAT TCGCCGTTGGGAAATTAGCAGTATTGATCAGTTACCCAATTATATG CAGCTTTGTTTTCTTACGATCAATAACTTCGTTGATGACACAGCTTA CGATGTAATGAAGGAGAAGGACATTAATATTATCCCCTATCTGCGT AAAAGTTGGGTGGATCTGGCTGAAGCTTATCTGGTCGAGGCTAAG TGGTTTTATGGAGGTTATAAGCCGAATTTGGAGGAGTACTTGAATAACGGCTGGATTTCGGTCTCGGGGCCAGCAATTCTTTGCCACGTCTTCTTTGGTGTCACGGACAGTATCACTATGGAGACGGTAGAATCGT TGTTTAAGTATCATGATTTAATTCGTTGCTCCTCAACTCTGGTGCG CCTGGCCGACGACTTGGCGACATCGCTGGACGAAGTCAGCCGCG GCGATGTCCCGAAAAGCATTCAATGTTACATGAATGACAACAACG CGTCCGAGGAGGAGGCACGCCTGCACGTACGCTGGTTAATTGCC GAGACTTGGAAAGAAATGAATGTTGAAATGGTGTCAGCAGATTCG CCATTTTGCAAGGACTTCATCGCGTGTGCTGCGGATATGGGTCGT ATGGCTCAATATATGTACCACAACGGTGATGGCCATGGAATGCAA AACTCACAAATTCATCAGCAGATGACGGACTTTTTGTTCCAAAAGC TGGCGGTGCGCGATCGCGCTTCAACCGCCCGTAACTAAQuercus ilex myrS ATGGTTGCGAATAAAGTGTCTACGTCCCCAGACATCTTGCGTCGTT 19CTGCTAACTACCAGCCTTCGATTTGGAATCACGATTACATTGAGTC TTTACGTATCGAGTATGTTGGCGAGACATGTACTCGCCAAATCAAT GTTTTGAAAGAACAAGTACGTATGATGCTTCATAAAGTCGTCAATC CGCTGGAACAATTAGAGCTGATCGAGATTTTGCAACGCCTTGGAC TTAGTTATCATTTTGAAGAAGAGATCAAGCGTATCTTGGACGGCGT GTATAACAATGACCACGGCGGGGATACTTGGAAAGCAGAGAACTT ATACGCGACGGCGTTGAAATTTCGCTTATTGCGCCAACACGGATA CAGTGTCTCACAGGAAGTTTTTAACAGCTTCAAGGACGAGCGTGG TTCCTTCAAAGCCTGCCTTTGCGAGGACACCAAAGGGATGCTGTC ATTATACGAGGCATCCTTTTTTCTTATCGAGGGTGAGAACATTCTT GAAGAGGCACGCGACTTTAGCACTAAGCATTTAGAGGAGTATGTG AAGCAAAACAAGGAGAAGAATCTTGCTACACTTGTGAATCACAGTC TTGAGTTCCCATTGCATTGGCGCATGCCGCGCTTAGAAGCACGCT GGTTTATCAATATTTACCGTCACAATCAGGACGTAAATCCGATCCT GTTGGAGTTCGCCGAGCTTGATTTCAATATTGTCCAAGCCGCCCA TCAAGCCGACTTAAAGCAAGTCTCGACCTGGTGGAAAAGTACCGG TTTGGTGGAGAACTTGTCATTTGCGCGTGACCGCCCGGTAGAGAA CTTTTTTTGGACTGTCGGTCTTATCTTTCAACCCCAATTCGGTTACT GTCGTCGTATGTTCACGAAAGTTTTCGCGCTTATTACAACGATTGA TGACGTTTATGACGTCTATGGGACGCTGGACGAACTGGAACTTTT CACCGACGTAGTTGAGCGTTGGGATATCAATGCTATGGACCAACT GCCGGATTACATGAAAATTTGTTTCTTAACTTTGCATAACTCCGTG AACGAGATGGCCTTGGATACCATGAAGGAGCAGCGCTTTCACATC ATTAAATACCTTAAGAAGGCCTGGGTCGATTTATGCCGTTACTACT TAGTGGAAGCAAAGTGGTACTCAAACAAGTATCGCCCCTCGTTGC AAGAGTACATCGAAAACGCATGGATTAGCATCGGAGCGCCCACGA TTTTAGTTCATGCATATTTTTTCGTAACGAATCCCATTACGAAGGAG GCCCTGGACTGCTTGGAGGAGTACCCTAATATCATCCGTTGGAGT TCAATTATCGCCCGTCTTGCTGACGATCTTGGGACTTCGACGGAC GAACTGAAGCGTGGAGATGTGCCTAAGGCTATCCAATGTTACATG AACGAAACAGGGGCCTCAGAGGAAGGCGCACGCGAATACATCAA GTACTTAATTTCTGCAACATGGAAAAAGATGAACAAGGACCGTGCT GCGTCGTCCCCGTTTTCCCACATTTTCATTGAAATTGCACTTAACC TTGCACGTATGGCTCAGTGCCTGTATCAACACGGCGATGGGCATG GACTTGGTAACCGTGAAACTAAAGATCGTATTCTGTCATTGTTAAT CCAACCGATTCCTTTGAACAAGGACTAACannabis sativa myrS ATGTGCTATCCCATACAATGTACAGTTGTAAACAATTCTTCCCCGA 20GCTCGACCATCGTGCGTCGTTCCGCGAATTACGAGCCGCCTATTT GGAGCTTTGATTATATCCAGAGCCTGTCCACCCAGTATAAAGGTG AAAGCTACACCGGTCAGCTGAACAAATTGAAGAAAGAGGTTAAAC GCATGCTGCTCCGTATGGAAATCAACTCCCTGGCGCTGTTGGAGC TGATCGACACCCTGCAACGTTTGGGCATTAGCTACCATTTTAAGAACGAAATTAACACCATTCTGAAGAAGAAGTATAACGACAACTACATCAACAACAACATAATTACGTCCCCGAATTACAATAACCTGTATGCTA CCGCTCTGGAGTTCCGCCTGTTACGTCAGCACGGCTATACCGTTC CGCAGGAGATCTTCAACGCGTTTAAGGACAAAAGAGGCAAATTTA AAACCAGCCTAAGTGACGACATCATGGGTGTACTGTGCCTGTACG AAGCGTCCTTTTATGCCATGAAACACGAAAACATCCTGGAGGAGG CGCGCATTTTTAGCACGAAGTGCCTGAAGAAGTACATGGAGAAAA TCGAGAATGAGGAGGAGAAGAAGATCTTGCTGCTGAATGATAACA ACATTAATAGCAATCTGCTCTTGATTAACCACGCCTTTGAGCTGCC GCTGCATTGGCGTATTACCCGTTCGGAGGCCCGCTGGTTTATCGA CGAGATCTACGAGAAGAAGCAAGATATGAACTCCACGTTGTTCGA GTTCGCGAAATTAGATTTCAACATCGTGCAGAGCACCCATCAGGA GGACTTACAACACCTGAGCCGTTGGTGGCGCGACTGTAAACTGG GCGGTAAATTGAATTTTGCTCGTGATCGTCTGATGGAAGCGTTCCT GTGGGATGTTGGTCTGAAGTTCGAGGGTGAATTCAGCTACTTCCG TCGTACGAATGCTCGTCTGTTCGTCCTCATTACCATCATCGACGAT ATTTACGACGTCTATGGTACTTTGGAGGAGTTAGAGCTTTTCACCA GCGCAGTTGAAAGATGGGATGTTAAACTGATTAACGAATTACCAGA TTACATGAAAATGCCGTTTTTTGTGCTTCACAATACCATCAACGAAA TGGGTTTTGACGTGTTGGTTGAACAGAACTTCGTGAATATTGAATA CCTGAAGAAGAGCTGGGTTGATTTATGCAAGTGCTACCTGCAAGA GGCGAAATGGTATTACTCTGGTTATCAGCCGACCCTGGAAGAGTA TACGGAATTGGGATGGCTGTCGATCGGCGCAAGCGTTATTTTAAT GCATGCGTATTTCTGCTTTACTAATCCGATTACCAAACAAGATCTG AAGTCTTTGCAACTGCAACACCATTATCCGAATATCATTAAGCAGG CATGTCTGATCACTCGCCTTGCGAATGACCTGGGCACCAGCTCTG ATGAATTGAACCGTGGTGATGTGCCGAAATCGATTCAATGTTACAT GTATGACAACAACGCGACCGAAGACGAGGCTCGCGAACATATTAA GTTCCTGATTAGCGAAACCTGGAAGGATATGAACAAAAAGGACGA GGACGAGAGCTGCCTGTCCGAAAATTTCGTGGAAGTGTGCAAAAA CATGGCACGTACCGCCCTGTTCATCTACGAAAACGGCGACGGCCA CGGCAGCCAGAACAGTCTGAGTAAAGAACGTATTTCTACGCTGAT CATCACTCCGATCAACATCCCGAAATAApTrcHis2B GTTTGACAGCTTATCATCGACTGCACGGTGCACCAATGCTTCTGG 21CGTCAGGCAGCCATCGGAAGCTGTGGTATGGCTGTGCAGGTCGT AAATCACTGCATAATTCGTGTCGCTCAAGGCGCACTCCCGTTCTG GATAATGTTTTTTGCGCCGACATCATAACGGTTCTGGCAAATATTC TGAAATGAGCTGTTGACAATTAATCATCCGGCTCGTATAATGTGTG GAATTGTGAGCGGATAACAATTTCACACAGGAAACAGCGCCGCTG AGAAAAAGCGAAGCGGCACTGCTCTTTAACAATTTATCAGACAATC TGTGTGGGCACTCGACCGGAATTATCGATTAACTTTATTATTAAAA ATTAAAGAGGTATATATTAATGTATCGATTAAATAAGGAGGAATAAA CCATGGATCCGAGCTCGAGATCTGCAGCTGGTACCATATGGGAAT TCGAAGCTTTCTAGAACAAAAACTCATCTCAGAAGAGGATCTGAAT AGCGCCGTCGACCATCATCATCATCATCATTGAGTTTAAACGGTCT CCAGCTTGGCTGTTTTGGCGGATGAGAGAAGATTTTCAGCCTGAT ACAGATTAAATCAGAACGCAGAAGCGGTCTGATAAAACAGAATTTG CCTGGCGGCAGTAGCGCGGTGGTCCCACCTGACCCCATGCCGAA CTCAGAAGTGAAACGCCGTAGCGCCGATGGTAGTGTGGGGTCTC CCCATGCGAGAGTAGGGAACTGCCAGGCATCAAATAAAACGAAAG GCTCAGTCGAAAGACTGGGCCTTTCGTTTTATCTGTTGTTTGTCGG TGAACGCTCTCCTGAGTAGGACAAATCCGCCGGGAGCGGATTTGA ACGTTGCGAAGCAACGGCCCGGAGGGTGGCGGGCAGGACGCCC GCCATAAACTGCCAGGCATCAAATTAAGCAGAAGGCCATCCTGACGGATGGCCTTTTTGCGTTTCTACAAACTCTTTTTGTTTATTTTTCTAAATACATTCAAATATGTATCCGCTCATGAGACAATAACCCTGATAA ATGCTTCAATAATATTGAAAAAGGAAGAGTATGAGTATTCAACATTT CCGTGTCGCCCTTATTCCCTTTTTTGCGGCATTTTGCCTTCCTGTT TTTGCTCACCCAGAAACGCTGGTGAAAGTAAAAGATGCTGAAGAT CAGTTGGGTGCACGAGTGGGTTACATCGAACTGGATCTCAACAGC GGTAAGATCCTTGAGAGTTTTCGCCCCGAAGAACGTTTTCCAATGA TGAGCACTTTTAAAGTTCTGCTATGTGGCGCGGTATTATCCCGTGT TGACGCCGGGCAAGAGCAACTCGGTCGCCGCATACACTATTCTCA GAATGACTTGGTTGAGTACTCACCAGTCACAGAAAAGCATCTTACG GATGGCATGACAGTAAGAGAATTATGCAGTGCTGCCATAACCATG AGTGATAACACTGCGGCCAACTTACTTCTGACAACGATCGGAGGA CCGAAGGAGCTAACCGCTTTTTTGCACAACATGGGGGATCATGTA ACTCGCCTTGATCGTTGGGAACCGGAGCTGAATGAAGCCATACCA AACGACGAGCGTGACACCACGATGCCTGTAGCAATGGCAACAAC GTTGCGCAAACTATTAACTGGCGAACTACTTACTCTAGCTTCCCGG CAACAATTAATAGACTGGATGGAGGCGGATAAAGTTGCAGGACCA CTTCTGCGCTCGGCCCTTCCGGCTGGCTGGTTTATTGCTGATAAA TCTGGAGCCGGTGAGCGTGGGTCTCGCGGTATCATTGCAGCACT GGGGCCAGATGGTAAGCCCTCCCGTATCGTAGTTATCTACACGAC GGGGAGTCAGGCAACTATGGATGAACGAAATAGACAGATCGCTGA GATAGGTGCCTCACTGATTAAGCATTGGTAACTGTCAGACCAAGTT TACTCAT ATATACTTTAGATTGATTTAAAACTTCATTTTTAATTTAAA AGGATCTAGGTGAAGATCCTTTTTGATAATCTCATGACCAAAATCC CTTAACGTGAGTTTTCGTTCCACTGAGCGTCAGACCCCGTAGAAA AGATCAAAGGATCTTCTTGAGATCCTTTTTTTCTGCGCGTAATCTG CTGCTTGCAAACAAAAAAACCACCGCTACCAGCGGTGGTTTGTTT GCCGGATCAAGAGCTACCAACTCTTTTTCCGAAGGTAACTGGCTT CAGCAGAGCGCAGATACCAAATACTGTCCTTCTAGTGTAGCCGTA GTTAGGCCACCACTTCAAGAACTCTGTAGCACCGCCTACATACCT CGCTCTGCTAATCCTGTTACCAGTGGCTGCTGCCAGTGGCGATAA GTCGTGTCTTACCGGGTTGGACTCAAGACGATAGTTACCGGATAA GGCGCAGCGGTCGGGCTGAACGGGGGGTTCGTGCACACAGCCC AGCTTGGAGCGAACGACCTACACCGAACTGAGATACCTACAGCGT GAGCTATGAGAAAGCGCCACGCTTCCCGAAGGGAGAAAGGCGGA CAGGTATCCGGTAAGCGGCAGGGTCGGAACAGGAGAGCGCACGA GGGAGCTTCCAGGGGGAAACGCCTGGTATCTTTATAGTCCTGTCG GGTTTCGCCACCTCTGACTTGAGCGTCGATTTTTGTGATGCTCGTC AGGGGGGCGGAGCCTATGGAAAAACGCCAGCAACGCGGCCTTTT TACGGTTCCTGGCCTTTTGCTGGCCTTTTGCTCACATGTTCTTTCC TGCGTTATCCCCTGATTCTGTGGATAACCGTATTACCGCCTTTGAG TGAGCTGATACCGCTCGCCGCAGCCGAACGACCGAGCGCAGCGA GTCAGTGAGCGAGGAAGCGGAAGAGCGCCTGATGCGGTATTTTC TCCTTACGCATCTGTGCGGTATTTCACACCGCATATGGTGCACTCT CAGTACAATCTGCTCTGATGCCGCATAGTTAAGCCAGTATACACTC CGCTATCGCTACGTGACTGGGTCATGGCTGCGCCCCGACACCCG CCAACACCCGCTGACGCGCCCTGACGGGCTTGTCTGCTCCCGGC ATCCGCTTACAGACAAGCTGTGACCGTCTCCGGGAGCTGCATGTG TCAGAGGTTTTCACCGTCATCACCGAAACGCGCGAGGCAGCAGAT CAATTCGCGCGCGAAGGCGAAGCGGCATGCATTTACGTTGACACC ATCGAATGGTGCAAAACCTTTCGCGGTATGGCATGATAGCGCCCG GAAGAGAGTCAATTCAGGGTGGTGAATGTGAAACCAGTAACGTTA TACGATGTCGCAGAGTATGCCGGTGTCTCTTATCAGACCGTTTCC CGCGTGGTGAACCAGGCCAGCCACGTTTCTGCGAAAACGCGGGA AAAAGTGGAAGCGGCGATGGCGGAGCTGAATTACATTCCCAACCGCGTGGCACAACAACTGGCGGGCAAACAGTCGTTGCTGATTGGCGTTGCCACCTCCAGTCTGGCCCTGCACGCGCCGTCGCAAATTGTCG CGGCGATTAAATCTCGCGCCGATCAACTGGGTGCCAGCGTGGTG GTGTCGATGGTAGAACGAAGCGGCGTCGAAGCCTGTAAAGCGGC GGTGCACAATCTTCTCGCGCAACGCGTCAGTGGGCTGATCATTAA CTATCCGCTGGATGACCAGGATGCCATTGCTGTGGAAGCTGCCTG CACTAATGTTCCGGCGTTATTTCTTGATGTCTCTGACCAGACACCC ATCAACAGTATTATTTTCTCCCATGAAGACGGTACGCGACTGGGC GTGGAGCATCTGGTCGCATTGGGTCACCAGCAAATCGCGCTGTTA GCGGGCCCATTAAGTTCTGTCTCGGCGCGTCTGCGTCTGGCTGG CTGGCATAAATATCTCACTCGCAATCAAATTCAGCCGATAGCGGAA CGGGAAGGCGACTGGAGTGCCATGTCCGGTTTTCAACAAACCATG CAAATGCTGAATGAGGGCATCGTTCCCACTGCGATGCTGGTTGCC AACGATCAGATGGCGCTGGGCGCAATGCGCGCCATTACCGAGTC CGGGCTGCGCGTTGGTGCGGATATCTCGGTAGTGGGATACGACG ATACCGAAGACAGCTCATGTTATATCCCGCCGTCAACCACCATCAA ACAGGATTTTCGCCTGCTGGGGCAAACCAGCGTGGACCGCTTGCT GCAACTCTCTCAGGGCCAGGCGGTGAAGGGCAATCAGCTGTTGC CCGTCTCACTGGTGAAAAGAAAAACCACCCTGGCGCCCAATACGC AAACCGCCTCTCCCCGCGCGTTGGCCGATTCATTAATGCAGCTGG CACGACAGGTTTCCCGACTGGAAAGCGGGCAGTGAGCGCAACGC AATTAATGTGAGTTAGCGCGAATTGATCTGMentha spicata 4S- ATGGAGCGCCGTTCAGGCAATTACAACCCCTCCCGCTGGGACGT 22 Limonene synthase CAATTTCATTCAGAGCCTTTTGAGTGACTACAAGGAGGATAAGCAC (LimS) GTCATCCGTGCTTCGGAATTAGTAACCCTGGTGAAGATGGAGTTA GAAAAGGAAACGGATCAGATTCGCCAGTTGGAATTAATCGATGAC CTGCAACGCATGGGTCTGAGTGACCATTTTCAAAACGAATTTAAGG AAATCCTTTCATCGATTTACTTGGATCATCATTATTATAAAAACCCG TTCCCAAAAGAAGAGCGCGATTTATATTCTACAAGCCTGGCATTTC GCCTGCTGCGCGAGCATGGTTTCCAGGTCGCTCAAGAAGTGTTCG ATTCATTTAAAAATGAAGAGGGTGAATTTAAAGAGTCGTTATCAGA CGACACACGCGGGCTGTTACAGTTATACGAAGCATCCTTTTTGTTA ACAGAAGGCGAGACGACACTTGAGAGTGCCCGTGAGTTTGCGAC AAAATTCCTTGAAGAAAAGGTAAACGAGGGGGGTGTAGACGGAGA CCTTCTTACACGTATTGCCTATTCTCTGGACATCCCGCTTCACTGG CGTATCAAACGTCCGAACGCGCCTGTGTGGATTGAATGGTACCGT AAGCGCCCCGACATGAACCCCGTTGTCTTAGAACTTGCTATTCTTG ACCTGAATATCGTTCAGGCCCAATTCCAAGAGGAATTGAAAGAATC GTTCCGTTGGTGGCGCAACACAGGCTTCGTCGAAAAGTTACCCTT TGCCCGCGATCGTCTGGTCGAGTGCTATTTCTGGAACACCGGAAT CATTGAACCTCGCCAGCACGCCTCAGCACGCATTATGATGGGAAA GGTAAACGCTTTGATCACGGTTATTGATGACATTTACGACGTTTAT GGCACTTTAGAAGAGCTGGAACAGTTTACCGACCTTATCCGCCGT TGGGACATCAATAGTATTGATCAACTGCCTGACTATATGCAACTGT GCTTCCTGGCGTTAAACAACTTTGTGGACGACACCTCATATGACGT TATGAAGGAAAAGGGAGTCAACGTCATTCCATATTTACGTCAAAGC TGGGTTGATTTGGCGGACAAGTATATGGTCGAAGCACGTTGGTTT TACGGCGGTCACAAACCGAGTCTTGAGGAGTATTTAGAAAATTCAT GGCAGAGTATCTCAGGACCGTGTATGTTGACACACATTTTCTTTCG CGTCACTGACTCTTTTACCAAAGAAACAGTCGATAGTCTGTACAAA TACCATGACCTTGTACGCTGGTCTAGTTTCGTTTTGCGCTTGGCCG ATGACTTAGGAACTTCTGTTGAGGAAGTATCACGCGGAGACGTGC CAAAGTCTTTACAGTGCTATATGAGTGACTACAATGCGTCCGAGG CCGAGGCCCGTAAACACGTGAAATGGTTGATTGCAGAGGTATGGAAGAAAATGAACGCGGAGCGTGTATCAAAGGATTCTCCCTTCGGCAAGGATTTCATCGGTTGTGCGGTAGATTTGGGTCGCATGGCTCAGT TGATGTATCATAACGGCGACGGACACGGGACTCAGCACCCCATTA TCCATCAACAAATGACACGTACATTGTTCGAGCCTTTCGCATAASpatholobus ATGGATGACATTTACATTCAACAGGCTCTTGTGTTGAAGGAAGTAA 23 suberectus (3S,6E)- AACACGTATTAAAAAAACTGATTTCAGAGGATCCAATGGAGTCACT nerolidol synthase 1 GTATATGGTAGATACGATTCAGCGCCTTGGAATCGAACATCATTTT (SsNerS) GAAGAGGAAATCGAGGCCGCTCTGCAAAAGCAACATTTAATTTTCA GCAGTCATTTGAGCGATTTCGCTAATAACCACAAATTATATGAAGT GGCTTTATTGTTTCGCCTGCTTCGTCAACGCGGTCATTACGTTCAT GCTGATGTGTTTGATTCTTTGAAAAGCAACAAGCGTGAATTTCGTG AAAAGCATGGGGAGGACGTCAAAGGACTGATCGCACTTTACGAGG CTACGCAGGTGTCTATTGAAGCTGATGATTCGCTGGACGATGCAG GTTATTTATCCTATCAATTGTTGCACGCCTGGCTTGCACGCAATAA AGAGCATCATGAAGCCATCTATGTCGCCAACACTTTACAAAATCCA CTGCATTATGGGCTGTCGCGTTTCATGGATCGTTCAACAATCTTAA CTAGTGATTTCAAGACAAAGAAGGAATGGAAATGTTTAGAGGAGTT AGCGGAGATCAACTCTTGTATCGTCAAGTTCATGAATCAGAATGAA ATCATCCAAGTTTATAAGTGGTGGAAGGACTTAGGGATGGCGAAC GAAGTCAAATTCGCACGTTATCAACCACTGAAGTGGTATATGTGGC CCATGGCATGCTTCACAGACCCACGTTTTTCAGATCAGCGCATTCA ATTAACGAAGCCTATTTCGTTAATTTACATTATTGATGATATTTTTGA CGTGCACGGGACATTAGACCAATTAACACTGTTCACGGATGCCGT GAACCGCTGGGAGTTAACGGGTACAGAGCAGTTACCAGATTTCAT GAAAATGTGCCTTTCCGTTTTGTACGATATCACAAACGATTTTGCA GAGATGATTAACAAAAAGCACGGACTTAACCCAATCGATACTCTTA AACGTTCTTGGGTCCGTCTGTTAAATGCTTTCCTTGAGGAAGCCCA TTGGCTTAATTCAGGACATTTGCCTCGCAGTGAGGAATACCTGAAC AACGGGATCGTGTCAACGGGGGTTCACGTCGTTTTGGTCCACGC GTTCTTCTTACTTGACCAGTCAATTAACAAAGAGATTGTTGCGATTA TCGACAACTTTCCTGAGATTATCCATTCCGTTGCCAAAATCCTGCG TCTTTCCGATGACTTAGAGGGAGCAAAATCTGAGGATCAAAACGG CTTAGATGGATCCTACTTAGATTGCTATATGAAGGAACATCAGGAC GTTAGTGCGGAGGATGCCCAGCGCCACGTAGCGCACCTGATCAG CTCCGAGTGGAAGCGTCTTAATCGTGAAATTTTGACACCCAACCC ATTGCCATCAAGTTTTACCAACTTTTGCCTTAACGCGGCTCGTATG GTCCCTTTAATGTACCACTACCGCTCAAATTCTTCTTTGTCGAACTT ACGCGAGCAGATTCAGATGTTACTGAATGTCGATGCCGGGCACAT CTAACajanus cajan (3S,6E)- ATGGATGACATCTACATGAAACAGGGGTTAGTCCTTAAAGAAGTAA 24 nerolidol synthase 1 AACATGTTTTTCAAAAACTTATTTCAGAGAACCCTATGGAATCGTTG (CcNerS) TGCATGGTCGATATTATCCAACGCTTGGGAATCGAGCACTTATTTG AGGAGGAAACAGAGGCGGTCTTGCAGAATCACCATTTTATTTTCTC CTCCCACTTGAACGATTTTTCCGACAATCATAAATTATACAAAGTAG CCCTGGCTTTCCGTCTGTTGCGTCAGCGTGGTCATTACGTTCATG CAGACCTTTTCGATTCGTTAAAGTCCAACAAGCGTGAATTTCGCGA AAAGTATGGCGAAGATGTTAAGTCGCTTATTGCGCTGTACGAGGC TACGCAGTTATCCATTGAGGGCGAAGACAGCGTGGACGAAGCTG GTTATTTATCGAACCAATTATTACATGCTTGGCTGACTACACACAAA GAACACCATGAAGCCATCTACGTTGCTAATACGCTTCAGTCGCCG TTACATTATGGATTATCCCGTTTCCGCGACAAGAATATGCATTTAA GCGATTTCAAGACCAAAAAGGAATGGACATGCCTTGAGGAGCTGG CCGAAATCAATTCCTGTATTGTTCGCTCGATGAACCAGAATGAAAT CAAACAAGTTTATAAATGGTGGAAAGACTTGGGCATGGCTAAAGAGGTTAAGTTCGCGCGTTATCAACCATTAAAGTGGTATATGTGGCCCATGGCATGCTTTACCGATCCGTCATTCTCCGACCAACGTATTGAGT TAACTAAACCGATCAGTCTGATTTACATTATTGATGATATCTTTGAT GTGTATGGGACCATTGACCAACTTACATTATTTACCGATGCGATCA ATCGCTGGGAACTGGCTGGTACCGAACAACTGCCAGACTTCATGA AGATGTGCTTGAGTGTTCTGTATGAAATCACCAACGATTTTGCGGA GAAAATTTACAAGAAACACGGCTTAAACCCAATTGACACGTTAAAA CGTTCGTGGGTTAAGTTGTTGAATGCGTTCCTGGAAGAGGCGCAC TGGTTGAACTCTGGGCATCTTCCGCGTTCAGAGGAATATCTTAACA ACGGGATTGTCACCACCGGCGTGCACGTTGTCTTGGTCCACGCTT TTTTTCTGCTTGACCAAACGATCAACAAAGAAATTGTCGCTATCGT CGATAATTTTCCTGAAATTATTCACAGTGTAGCAAAGATTTTACGCC TGTCAGACGATTTAGAGGGAGCGCAATCCGAAAATGAGAACGGGT TAGATGGATCTTATCTTGACTGCTACATGAACGAACATCAAGATAT CTCGACAGAGGACGCTCAGTCTCACGTCTGTCACTTGATTTCGCG CGAGTGGAAGCGCTTAAACGGACAAGTCTTGACCAGCACAGGCTT GCCTTCGAGTTTTACAAAGTTTTGCCTGAACGCAGCTCGCATGGT GCCCCTGATGTATCACTATCAGTCAAATCCCAGTCTTTCCCGCCTT CGCGAGCAAATCCAGACTCTGTTAGACGTTGACGCAGGACATATC TAACinnamomum tenuipile ATGCCTCGCCGCTCGGGCAATTACAAGCCGTCTATCTGGGATTAT 25 Geraniol synthase GACTTTGTTCAATCGTTAGGTTCAGGGTATAAAGTGGAGGCTCAC (CtGerS) GGGACACGCGTTAAGAAGCTGAAGGAAGTCGTTAAACACCTGTTA AAAGAGACGGATAGCTCACTTGCTCAGATCGAGCTGATCGATAAA CTGCGTCGTCTGGGCCTTCGCTGGCTGTTCAAGAATGAAATTAAG CAGGTCCTGTATACAATTTCTTCGGACAACACATCCATTGAAATGC GTAAAGATCTTCATGCAGTTAGTACCCGCTTTCGTCTGCTTCGCCA ACACGGATACAAGGTTTCTACCGATGTATTCAATGACTTCAAAGAC GAAAAAGGTTGTTTTAAACCTTCTCTGTCTATGGATATTAAGGGAA TGTTATCATTGTATGAAGCAAGTCACCTTGCGTTTCAGGGGGAGA CGGTATTGGACGAAGCCCGTGCGTTCGTTTCAACACATCTGATGG ACATTAAGGAAAATATTGACCCGATTTTGCACAAGAAGGTGGAACA TGCCCTTGATATGCCGTTACACTGGCGCCTTGAAAAGTTAGAGGC CCGCTGGTATATGGATATTTACATGCGTGAGGAGGGAATGAATAG TAGCCTTTTGGAACTGGCGATGTTGCATTTTAATATTGTGCAGACA ACATTCCAGACGAACCTGAAGTCCCTGTCCCGTTGGTGGAAAGAT CTTGGACTTGGAGAGCAGCTTTCGTTTACCCGTGACCGTCTGGTC GAATGTTTTTTCTGGGCTGCGGCAATGACGCCTGAACCTCAGTTC GGACGTTGCCAGGAAGTAGTTGCCAAAGTCGCTCAACTTATTATC ATTATCGACGACATCTACGACGTATATGGTACTGTAGACGAATTAG AATTGTTTACAAATGCGATCGACCGTTGGGACCTGGAGGCGATGG AACAGTTGCCAGAGTACATGAAGACTTGTTTCCTTGCTTTGTACAA TAGTATTAACGAAATTGGATATGATATTCTTAAAGAGGAGGGGCGT AATGTCATTCCGTATTTACGCAACACATGGACAGAACTTTGTAAAG CATTCCTGGTCGAAGCGAAGTGGTACAGTAGCGGCTATACCCCGA CGCTTGAGGAATACCTGCAGACGTCGTGGATTAGTATTGGTTCAC TGCCGATGCAAACCTATGTATTTGCCTTGTTAGGGAAGAATTTAGC GCCCGAAAGTAGTGACTTCGCTGAGAAAATTTCAGACATTCTTCGT TTAGGAGGCATGATGATCCGCCTGCCCGACGACCTGGGTACTTCG ACAGACGAGTTAAAGCGCGGAGACGTTCCTAAAAGTATCCAATGT TACATGCACGAGGCTGGTGTTACGGAGGACGTCGCCCGTGATCAT ATCATGGGTCTTTTTCAAGAAACTTGGAAAAAGCTGAATGAGTACC TTGTTGAATCTAGTTTACCACACGCATTCATTGATCATGCTATGAAT CTTGGACGCGTATCTTATTGTACGTATAAGCACGGTGACGGCTTCAGTGATGGCTTCGGCGATCCCGGTAGCCAAGAGAAGAAGATGTTTATGTCTCTGTTCGCTGAACCACTGCAAGTAGACGAGGCAAAGGGG ATTAGCTTCTATGTAGATGGCGGCAGCGCATAAPerilla citriadora ATGCGTCGCTCTGGAAATTACCAGCCATCGATCTGGGATTTCAATT 26 Geraniol synthase ACGTTCAAAGTTTAAACACGCCTTATAAAGAAGAGCGTTACCTTAC (PcGerS) TCGTCACGCGGAACTGATCGTACAGGTCAAGCCGTTACTTGAGAA GAAGATGGAGCCCGCGCAGCAATTGGAACTTATTGACGATTTGAA CAATTTAGGATTGAGTTACTTTTTTCAGGACCGTATCAAACAGATTC TGTCATTCATCTATGATGAAAATCAATGCTTTCATTCTAACATCAAC GATCAAGCAGAAAAGCGTGATCTTTACTTTACCGCGTTAGGATTCC GTCTGCTGCGTCAACATGGATTTGACGTATCACAGGAAGTCTTCG ATTGCTTCAAGAACGATAATGGTTCAGATTTTAAGGCTTCCTTATC CGACAATACAAAAGGGTTGTTGCAACTTTACGAGGCTTCGTTTTTG GTGCGCGAAGGCGAAGACACATTAGAACAGGCTCGTCAATTCGCA ACAAAATTTCTTCGTCGCAAGTTGGACGAGATTGACGACAACCACT TGCTGAGCTGTATTCATCACTCCCTGGAAATCCCTTTACACTGGCG TATTCAGCGCCTTGAAGCCCGCTGGTTTTTAGACGCGTATGCAAC ACGCCATGACATGAATCCAGTAATCCTGGAACTGGCGAAGTTAGA CTTTAATATTATCCAGGCAACGCACCAGGAAGAGTTGAAAGACGT CAGTCGCTGGTGGCAAAACACTCGTCTGGCCGAAAAACTGCCATT TGTGCGTGATCGTTTAGTAGAATCTTATTTTTGGGCCATCGCCCTG TTCGAACCTCACCAATACGGATATCAACGCCGCGTGGCGGCCAAG ATCATTACCCTTGCGACAAGTATCGATGACGTTTACGATATTTACG GGACCTTGGACGAACTTCAATTGTTCACGGACAATTTTCGCCGCT GGGATACCGAAAGTTTAGGTCGTCTGCCTTATTCTATGCAACTGTT TTATATGGTCATTCATAACTTTGTGAGCGAGCTTGCCTATGAGATC CTGAAAGAAAAGGGTTTTATCGTCATCCCATATTTACAGCGTTCGT GGGTCGACTTAGCGGAATCTTTCTTGAAAGAGGCGAACTGGTACT ATTCAGGTTATACGCCTAGCCTTGAGGAGTATATCGATAATGGGTC CATCTCGATCGGCGCGGTCGCTGTTCTTTCGCAAGTTTACTTCAC GTTGGCTAACTCAATCGAGAAGCCAAAGATCGAGTCTATGTACAAA TACCACCACATTCTGCGTTTATCTGGTTTATTAGTGCGTCTGCACG ACGATCTGGGCACTAGCCTTTTCGAGAAAAAGCGTGGCGACGTGC CCAAAGCAGTAGAAATCTGTATGAAAGAACGCAACGTCACAGAAG AGGAAGCCGAAGAGCACGTCAAGTATCTTATCCGCGAGGCCTGG AAAGAAATGAACACGGCCACTACAGCAGCGGGCTGCCCATTCATG GATGAATTGAACGTGGCAGCGGCTAATTTGGGTCGCGCCGCTCAA TTTGTCTACTTAGATGGCGACGGTCACGGCGTTCAGCACTCCAAG ATTCACCAACAAATGGGGGGGTTAATGTTCGAACCTTACGTCTAACamptotheca ATGGCAACGTCTACTGCCACAATTGGCGACACCGATTCTCTGCTG 27 acuminata Geraniol AAATCTCAACGTCAATTCACAGTCTATCTTCCCGCACACGAAGCGG synthase (CaGerS) ATAAGGATCGCAAGATTGAAGAGATCATGGAGAAAACCCAGGGTG AACTTGAAAAAACGAGCGACCCTACTTCGGTAATGAAATTCATCGA CACACTGGAGCGCTTAGGTATCGCGTATCATTTCGAGGAGGAAAT CAATTCTTTGCTTCAAGGTTTCCTTGCGAACGGTTATTCCCACTAT CCCCAGGACTTATTCACAACGGCACTTCGTTTTCGCCTTTTACGCC ACAATGGTTATCATATTAGTGCTGATGTTTTTCAGAAGTTTGTGGA CAAGAACGGGAAATTCAAAGAAAGTCTTCGCGAAGACACACAAGG CATGTTGTCCTTGTACGAAGCATCATACTTAGGCGCTAACGGTGA GGATATCTTGAGCCAGGCTATGGAGTTTACTGAAACACACTTTAAA CAAAGCATTCCGTTGATGGCTGCCGTTCCCCAGCTTGAGCAGGCA TTAGAACTGCCGCGTCATCTGCGTATGGCCCGTTTGGAGGCTCGC CGCTTCATTGAGGAGTATATTCGTGAGAGCGACCATTCATCGGCG CTTCTGGAGCTTGCGAAACTTGATTATAATAAGGTGCAGTTGCTGCACCAGTCAGAATTGAACGAGATTTCACGCTGGTGGAAACAGTTAGGTCTTGTCGAAAATTTGGGGTTCGGGCGCGATCGTCCCCTGGAAT GCTTCTTATGGACCGTGGGGATTTTGCCCGAACCCAAATATAGCG GATGTCGCATTGAACTTACTAAAACCATTGCGGTATTGTTGGTTTT GGACGATATTTTTGATTCCTTTGGCACCTTGGATGAGCTGGTTCGT TTTACGCACGCTATCCGTCGTTGGGATTTATCTGCGATGGAACAG CTGCCAGAATATATGAAGGTATGCTACATGGCGCTGTACAATACG ACTAACGAGATCGGATATAAGATCTTGAAAGAACACGGCTGGAAC GTCGTGCCCTACTTAAAGCGTACCTGGATTGACATGATCGAAGGA TTCCAAGCAGAGGCTAATTGGTGTAGTTCTGGCTACGTACCGTCA CTGGAGGAATACATTGAGAACGGCGTTACTACCGCTGGGTCTTAT ATGGCCTTGGTACACTTATTCTTTCTGATGGGGCAAGGTGTGACG GACGAGACCATTGGGATGCTTGAACCCTACCCTAAGTTTTTTTCTT CTTCCGGGCGCATCCTTCGCCTTTGGGACGATTTGGGCACAGCCT CAGAGGAGCAGGAGCGCGGGGATATCGCCTCATCGATCGAACTG TTCATGCGTGAAAAGGACTTGTCCAGTCAGGGAGAGGCACGTAAG TACGTAAAACAGGTGATCTACAGCTTGTGGAAGGAACTGAACGGC GAGTTAATGGCCAGCAAGGCGATGCCGTTGCCCTTGATTAAGGCG GCCTTTAATATGGCCCGTACCTCACAAGTGATTTACCAGCACGGG GATGATAATTCATTCCCTAGCGTGGATCAGTGTGTACAGTCCTTAT TTTTTACGCCTATTTTGTAACitrus unshiu 1,8 ATGGCGTCAACCACGACAATTAAGCCTGTGGACCAAACGATTATT 28 Cineole Synthase CGTCGCAGTGCAGACTACGGCCCCACAATTTGGAGTGTTGACTAC (CuSinS) ATTCAATCATTGGATTCGGAGTATAAAGAAAAATCATACGCACGTC AGTTACAAAAACTTAAAGAGCAGGTTTCCGCGATGTTGCAACAAGA TAATAAAGTCGTAGACTTGGACCCGCTTCATCAGTTAGAGTTAATT GATAACCTTCATCGCTTGGGCGTATCTTACCATTTCGAGGATGAAA TCAAACGCACATTAGATCGCATCCACAATAAAAATACGAACAAGTC ATTGTATGCACGTGCTCTTAAATTCCGTATCTTGCGTCAGTACGGA TATAATACTCCTGTGAAAGAGACATTTAGCCGTTTCATGGACGAGA AGGGTTCGTTTAAACTTTCCTCGCATAGCGATGATTGCAAGGGAAT GTTGGCCCTGTATGAGGCAGCTTACCTGCTGGTTGAGGAGGAATC CTCGATTTTCCGTGATGCAATCCGTTTCACGACTGCCTATCTTAAA GAATGGGTTGTAAAGCATGATATTGACAAAAACGACGATGAGTATC TGTGTACCCTGGTAAAGCACGCTTTAGAGTTGCCCCTGCACTGGC GCATGCGTCGTCTTGAAGCCCGTTGGTTTATCGACGTATACGAGT CCGGGCCCGACATGAACCCTATTCTGCTGGACCTGGCTAAATTAG ACTTTAATATCGTACAAGCCGTGCATCAGGAGGACATCAAGTACG CTTCTCGCTGGTGGAAGAAGATCGGACTGGGGGAGCGCCTTAAC TTTGCCCGTGACCGTATCATGGAAAATTTTTTCTGGACAGTTGGTG TTATTTTTGAACCGAATTTCGGCTATTGTCGTCGTATGTCAACGAT GGTAAACGCTCTGATTACAACAATCGATGACGTCTACGACGTTTAT GGCACCCTGGACGAACTGGAATTGTTTACGGATGCAGTGGAACGC TGGGACGCAACAACTATTGAGCAGCTGCCAGATTATATGAAGCTG TGCTTCCATGCATTGCATAATTCTATCAACGAGATGGCCTTCGATG CCTTGCGTGACCAGGGCGTGGGAATGGTAATTTCCTATCTGAAGA AGGCATGGGCCGATATCTGCAAGACTTACTTGGTAGAAGCTAAGT GGTATAATAATGGCTACATCCCAACGCTGCAGGAGTATATGGAAA ACGCATGGATCTCAATTTCAGCGCCTGTCATCCTGGTGCACGCCT ATACATACACAGCGAATCCGATTACAAAAGAGGGCTTGGAATTTGT CAAAGATTATCCAAACATTATCCGCTGGTCGTCGATTATTTTACGTT TAGCTGACGATCTGGGCACTAGTTCAGATGAATTGAAACGCGGCG ACGTACATAAGTCTATTCAATGCTATATGCATGAAGCGGGGGTCA GCGAGCGTGAGGCACGCGAGCATATCCACGATTTAATCGCCCAAACTTGGATGAAGATGAACCGTGACCGTTTCGGAAATCCACATTTCGTCAGTGACGTATTTGTCGGCATCGCTATGAATCTTGCACGCATGAG TCAATGTATGTATCAGTTCGGTGATGGTCATGGCCACGGGGTGCA AGAAATCACCAAAGCTCGTGTGCTGTCCTTGATTGTGGACCCAAT CGCGTAAHypoxylon sp. E7406B ATGCGTCCGATCACCTGCTCCTTCGATCCAGTCGGCATTTCTTTCC 29 1,8 Cineole Synthase AAACGGAATCCAAGCAAGAGAACTTTGAGTTTCTGCGCGAAGCAA (HypSinS) TTTCTCGTTCAGTGCCAGGGCTTGAGAACTGCAATGTCTTTGACCC TCGTTCCTTGGGAGTGCCTTGGCCGACTAGTTTCCCGGCGGCTGC TCAGTCTAAGTATTGGAAAGATGCAGAGGAAGCCGCAGCTGAATT AATGGACCAGATTGTCGCCGCAGCCCCAGGAGAACAAGGGTCAC TGCCAGCAGAGCTGGCAGTATCTGATAAGAAAGCCGCCAAGCGTC GTGAGCTTTTGGATACGTCCGTCTCCGCCCCAATGAACATGTTCC CCGCCGCGAACGCTCCTCGTGCTCGTATCATGGCAAAGGCAAAC CTGCTTATTTTCATGCACGATGATGTCTGCGAATATCAATCAGTCC AATCAACGATTATCGACTCGGCGTTGGCCGACACATCGACACCAA ATGGTAAGGGGGCGGACATCTTGTGGCAGAATCGCATTTTCAAGG AATTTAGTGAGGAAACCAACCGCGAGGATCCCGTCGTTGGACCGC AATTCCTGCAAGGAATCCTGAACTGGGTAGAACACACACGCAAAG CACTTCCTGCTTCCATGACGTTTCGTTCTTTTAACGAATACATTGAC TATCGTATTGGTGACTTTGCAGTTGATTTTTGCGATGCGGCTATTT TGCTGACGTGTGAGATTTTTTTGACGCCTGCAGATATGGAACCACT TCGTAAGCTACACCGTTTATACATGACGCATTTCAGTTTGACAAAT GATCTTTACTCTTTCAATAAGGAAGTCGTAGCCGAACAAGAGACAG GTAGCGCAGTTATTAATGCGGTTCGCGTGCTGGAACAGTTGGTCG ACACCTCAACGCGCTCAGCAAAAGTCCTGCTTCGCGCCTTCTTGT GGGACCTTGAACTGCAAATCCATGATGAACTTACGCGCTTGAAGG GTACGGACCTTACCCCGAGTCAATGGCGTTTCGCACGCGGTATGG TTGAGGTATGCGCTGGCAACATTTTCTACTCCGCTACCTGCTTACG TTATGCTAAGCCAGGTTTGCGCGGAATTTAAStreptomyces ATGCCAGCTGGGCATGAAGAATTTGACATCCCGTTCCCAAGCCGC 30 clavuligerus 1,8 GTAAATCCGTTTCACGCGCGTGCGGAGGATCGTCATGTGGCTTGG Cineole Synthase ATGCGCGCTATGGGCTTAATTACCGGTGATGCTGCAGAGGCTACG (ScSinS) TATCGCCGTTGGTCGCCAGCTAAGGTAGGCGCACGCTGGTTTTAC TTGGCGCAAGGGGAAGATCTTGACCTGGGGTGTGACATCTTTGGG TGGTTCTTTGCCTATGATGACCATTTCGATGGACCGACAGGCACG GACCCGCGTCAAACGGCAGCGTTTGTTAACCGCACAGTAGCCATG CTGGATCCCCGCGCTGACCCGACAGGCGAACATCCGCTGAATAT CGCTTTTCACGACCTGTGGCAACGTGAGAGCGCGCCCATGTCCC CCTTATGGCAACGTCGCGCCGTTGACCATTGGACTCAGTACTTGA CGGCCCATATTACTGAAGCGACGAACCGCACCCGCCATACCAGC CCAACAATTGCTGATTACCTGGAGCTGCGTCACCGTACGGGTTTT ATGCCGCCTCTGTTAGACCTGATTGAACGTGTCTGGCGTGCAGAG ATCCCCGCCCCAGTCTACACGACTCCTGAAGTGCAAACCCTTCTT CACACAACCAATCAAAATATCAACATCGTTAACGATGTGCTGAGTT TGGAAAAGGAAGAGGCACATGGGGACCCACATAATCTTGTGTTAG TAATTCAACATGAGCGTCAGTCCACACGCCAACAAGCTCTGGCGA CCGCCCGCCGTATGATCGACGAGTGGACGGACACATTTATTCGTA CCGAGCCGCGCTTACCTGCTCTTTGCGGACGTTTGGGCATTCCAC TTGCCGACCGCACGTCCTTATATACTGCAGTAGAAGGAATGCGCG CGGCAATTCGCGGGAATTACGACTGGTGTGCCGAAACAAACCGTT ATGCTGTACATCGTCCGACCGGCACAGGACGTGCTACGACGCCTT GGTAASalvia fruticose 1,8 ATGTCGCTGCAAACCGGAAACGAAATTCAAACCGAACGCCGTACT 31 Cineole Synthase GGAGGATATCAGCCAACTCTGTGGGATTTCAGCACAATCCAGAGC (SfSinS) TTTGATTCGGAATATAAGGAAGAGAAACATTTAATGCGTGCGGCC GGGATGATCGATCAGGTAAAAATGATGCTGCAAGAAGAGGTAGAC TCTATCCGTCGCTTAGAGCTGATCGACGACTTACGTCGTCTTGGC ATTAGTTGTCATTTTGAGCGCGAGATCGTTGAGATTCTGAACTCGA AGTACTATACTAACAATGAAATTGACGAGCGTGATCTTTATTCTAC CGCTTTACGTTTCCGCCTGTTGCGTCAATACGATTTCAGCGTCAGT CAAGAGGTCTTTGACTGTTTTAAAAATGCGAAAGGCACGGATTTCA AACCTAGCTTAGTAGATGATACCCGTGGACTTTTACAACTTTATGA AGCCAGCTTTTTGTCAGCGCAGGGGGAAGAGACATTACGTCTTGC TCGCGATTTTGCCACAAAGTTCTTGCAAAAGCGTGTCTTGGTGGAT AAAGATATTAATCTGCTGTCATCTATTGAACGCGCGCTTGAACTGC CCACGCATTGGCGTGTTCAGATGCCTAATGCACGCAGTTTCATCG ACGCCTACAAACGCCGCCCCGATATGAATCCCACGGTATTAGAGT TAGCAAAATTAGATTTTAATATGGTCCAAGCGCAATTTCAACAGGA GCTTAAGGAAGCCTCCCGCTGGTGGAATTCCACCGGCTTGGTCCA TGAATTACCATTTGTACGTGATCGTATTGTAGAATGCTATTATTGGA CTACGGGTGTGGTGGAACGTCGTCAACATGGGTATGAGCGCATTA TGTTAACAAAGATCAATGCATTGGTCACTACAATTGACGATGTCTT CGACATTTATGGCACTTTGGAGGAGTTGCAATTGTTCACCACCGC AATTCAACGCTGGGACATCGAATCAATGAAACAACTTCCTCCGTAT ATGCAGATTTGCTATCTGGCTCTTTTCAATTTCGTTAACGAGATGG CATATGACACCTTGCGCGACAAAGGATTTGATAGTACCCCCTACCT TCGCAAAGTCTGGGTGGGGCTGATCGAATCATACTTAATCGAAGC CAAATGGTACTACAAAGGCCATAAGCCCTCTTTAGAAGAGTACATG AAGAATTCTTGGATCTCCATCGGAGGGATCCCTATCTTGTCGCACT TATTTTTCCGCTTAACAGATTCAATTGAGGAGGAGGCAGCCGAGT CAATGCATAAGTACCATGACATCGTCCGCGCCTCGTGTACTATTTT GCGTTTAGCGGATGACATGGGAACCTCCTTAGATGAAGTTGAGCG CGGGGATGTACCAAAATCGGTACAGTGTTACATGAATGAGAAGAA TGCCTCTGAAGAAGAAGCCCGTGAACACGTCCGTAGCCTTATTGA CCAGACGTGGAAAATGATGAATAAAGAAATGATGACATCTTCATTT TCCAAGTATTTTGTAGAGGTGAGCGCAAATTTAGCTCGTATGGCAC AGTGGATTTACCAACATGAATCAGATGGATTTGGAATGCAACATTC ACTTGTCAATAAGATGCTTCGTGATCTTCTGTTTCATCGTTATGAAT AAAntirrhinum majus p- ATGGGAAGTACACCGCCACCGTCGAAGTTGCACCAGGCGTTGTGT 32 Ocimene Synthase CTGAATGCGCATTCGACTAGTTGCATGGCTGAGTTACCTATGGATT (AmOciS) ATGAAGGCAAGATTCAGGGAACTCGTCATTTGTTGCATCTGAAGG ATGAAAACGACCCAATTGAAAGCCTTATCTTTGTCGACGCTACGCA GCGTTTAGGTGTTAATCACCATTTTCAGAAAGAAATTGAAGAGATT CTGCGTAAGAGTTATGCCACTATGAAATCGCCGTCGATCTGTAAGT ACCACACGCTTCACGATGTGTCCCTGTTTTTCTGTTTGATGCGTCA ACACGGTCGTTATGTATCGGCGGATGTATTTAATAATTTTAAGGGA GAATCCGGACGTTTTAAAGAAGAGCTTAAACGTGACACCCGTGGG CTTGTGGAGTTATATGAGGCGGCCCAATTGTCATTTGAAGGTGAA CGCATCTTGGACGAGGCAGAAAATTTCAGTCGTCAGATCCTGCAT GGAAACTTGGCGTCAATGGAAGACAACCTTCGCCGTTCCGTGGG GAATAAGTTGCGTTACCCCTTCCATAAGTCGATCGCACGCTTCAC GGGAATTAACTACGACGATGACCTGGGCGGTATGTATGAATGGGG CAAAACCTTGCGTGAGCTTGCTTTAATGGATCTTCAAGTCGAGCGT TCGGTGTACCAGGAAGAATTGTTACAGGTTAGCAAGTGGTGGAATGAGTTGGGACTTTATAAGAAGTTAACCTTAGCTCGTAATCGCCCCTTCGAATTTTACATGTGGTCAATGGTCATTTTAACTGATTATATCAAC TTATCGGAGCAGCGCGTGGAGTTGACAAAATCCGTAGCATTTATC TACCTGATTGACGACATCTTCGACGTATATGGAACACTTGACGAAT TGATTATTTTTACTGAAGCTGTAAATAAGTGGGATTACTCCGCTACT GATACGTTACCAGACAATATGAAGATGTGCTATATGACCTTACTGG ATACAATCAATGGGACGTCACAAAAAATTTATGAAAAATATGGTCA TAACCCCATCGATAGTTTAAAAACGACATGGAAGAGCCTGTGCTCT GCATTTTTAGTTGAAGCGAAATGGTCGGCTAGTGGTTCATTGCCTT CAGCGAATGAATATCTGGAGAACGAGAAGGTATCATCTGGCGTAT ATGTCGTGTTAATTCATTTGTTTTTCCTGATGGGATTAGGCGGGAC AAATCGTGGTAGCATTGAACTTAATGACACCCGCGAGTTGATGTCT TCAATTGCCATCATCGTTCGCATTTGGAATGATCTGGGTTGTGCAA AAAACGAGCACCAGAACGGTAAGGATGGTTCTTATCTTGATTGTTA CAAAAAAGAACATATCAATTTGACAGCCGCACAAGTACATGAACAC GCACTGGAGTTGGTGGCGATCGAATGGAAACGCCTGAACAAAGAA TCATTCAATCTGAATCACGACTCAGTGTCGTCATTCAAACAAGCAG CACTTAATTTTGCCCGCATGGTCCCCTTGATGTACAGCTATGACAA TAACCGTCGCGGCCCTGTTCTGGAGGAGTACGTCAAGTTTATGTT ATCAGACTAAArabidopsis thaliana |3- ATGCCTAAGCGCCAGGCACAGCGTCGCTTTACACGCAAAACTGAT 33 Ocimene Synthase TCTAAAACGCCCTCTCAACCTCTTGTATCTCGCCGTTCGGCTAACT (AtOciS) ATCAGCCCTCCCTGTGGCAGCACGAGTATTTGCTTAGTCTTGGTA ATACATATGTTAAGGAGGATAACGTAGAACGCGTCACATTGTTGAA GCAAGAAGTAAGCAAAATGTTAAATGAAACGGAAGGATTGCTGGA GCAACTTGAACTGATCGATACACTGCAGCGTTTAGGGGTCAGCTA TCATTTTGAACAGGAGATCAAGAAGACCTTAACTAATGTTCATGTC AAGAACGTCCGCGCGCACAAAAACCGTATCGACCGTAATCGCTGG GGTGACTTATACGCGACAGCCTTGGAATTTCGCCTGTTACGTCAA CACGGTTTCTCTATTGCTCAGGATGTATTCGACGGGAACATCGGG GTCGATTTGGATGATAAGGATATCAAAGGGATTCTTTCGTTATACG AGGCGTCTTATTTGAGTACCCGTATCGATACAAAATTGAAAGAAAG CATTTACTACACGACCAAGCGCTTGCGTAAGTTCGTAGAGGTAAA CAAAAATGAGACCAAGTCCTATACATTACGTCGCATGGTAATTCAC GCCTTGGAGATGCCTTACCACCGCCGTGTGGGACGCCTGGAGGC TCGCTGGTATATTGAGGTCTACGGCGAACGCCACGACATGAATCC GATTTTATTGGAATTAGCGAAGCTGGATTTTAACTTCGTGCAAGCG ATCCATCAAGACGAGCTGAAATCTCTGTCCTCTTGGTGGAGTAAG ACCGGTTTAACTAAACATTTGGACTTCGTCCGCGACCGTATCACTG AGGGATATTTTTCGAGTGTAGGGGTGATGTACGAGCCTGAGTTTG CATACCACCGTCAGATGCTGACTAAGGTATTTATGTTAATCACGAC GATCGACGACATCTATGACATCTATGGCACCTTGGAAGAGTTACAA CTGTTTACCACCATCGTGGAGAAGTGGGACGTAAACCGTCTGGAA GAGCTTCCCAATTATATGAAGTTGTGCTTCTTGTGCCTTGTGAACG AGATCAATCAGATCGGATATTTCGTCCTGCGTGATAAAGGCTTTAA TGTTATTCCCTATTTGAAGGAGTCGTGGGCCGATATGTGCACCAC ATTTCTTAAAGAGGCTAAATGGTATAAGAGTGGCTATAAGCCTAAT TTCGAAGAATATATGCAGAATGGATGGATTAGCTCATCAGTTCCGA CCATCTTGCTGCACTTATTCTGCTTACTGTCCGACCAGACATTGGA CATTCTTGGCTCTTACAATCACAGTGTAGTTCGCTCATCTGCTACC ATCTTGCGCTTAGCAAACGATCTTGCCACTAGCTCTGAAGAACTG GCCCGTGGCGACACAATGAAATCTGTTCAGTGTCACATGCATGAA ACTGGAGCCTCGGAGGCGGAATCACGTGCGTACATTCAAGGCAT CATTGGAGTGGCTTGGGATGATCTGAATATGGAGAAGAAGAGCTGTCGCCTTCACCAGGGCTTCCTGGAGGCCGCTGCCAACTTAGGACGTGTTGCCCAATGTGTGTATCAATATGGTGATGGTCATGGATGTCC TGATAAAGCCAAAACTGTAAATCACGTACGCTCCTTGCTTGTGCAT CCCCTTCCACTGAACTAACitrus unshiu p- ATGTCTGCGACCAAAGTGCGTGACAAAGCTATCAACGATAACCGC 34 Ocimene Synthase CGTTCAGCCAATTATCAACCTAGTATGTGGAGCTATGATTACCTGC (CuOciS) AGTCCTTGTCAAACGGGTATGTGGGAGAGAGTTGTGCGCAGCGC ATTGAAAAATTGAAGGGAGAGGTCCGTTTAATGTTAGATAACTATA AGGAGGTAGACGACTACGTTGACGCTTTGCATTGCCTTGAAATCG TGGATAATCTGCAACGTTTGGGTGTATCATATCATTTTGAAGGCGA GATTAAACGTTTTCTTAATTCGATTTACAATAAGCGTAACTCACGTC GCTCATCTACGTATCACGCTAAGGAAAACCAGGAAAGTTTACTGTA TGCAGCTAGCCTTGAATTTCGTCTGTTACGCCAGCACGGGTACGA TATTCACGCTCATGGGACTCTTTCCTCCTTTATGGATGAGAAAGGA AAGTTTAAATCATGTTTGGGCGATGATATCAAGGGAATCCTGGCCT TATACGAAGCGGCGTATCTTCTGGGCGAGGAGGAGAGCACGATC TTTCATGAGGCAATTAACTTTACTACCACGCATCTGGAAGAATATG TAAAGAAGCATAACGACGACGACGGTTATTTCAGCGCCTTGGTAA ATCACGCTCTTGAGTTACCGCTGCACTGGCGCATGGTTCGTCTTG AAGCCCGTTGGTTCATTGATGTATATGAGCGCGGTACCGATATGA ACCCCGTACTGGTCGAACTGGCCAAACTGGACTTCAACTCCGTCC AAGCGGCCCACCAAGATGAATTGAAGTACGTCTCGTGGTGGTGGC GCAAAACTGGCTTAGGAGAATTACACTTCGCGCGCGACCGTATTT TAGAGAATTTCTTCTGGGCTTTAGGTGAAATCTGGGAACCGCAGTT TGGATATTGTCGTCGTATGTCGACTAAGGTAAATGCCCTTATCACG ACTATTGACGACGTCTATGACGTTTATGGCACTCTGGACGAATTGG AGCAGTTCACTAACGCTGTCGAGCGCTGGGACGTGAATGCTATGG ACCAGTTGCCCTATTACATGAAATTATGTTTCCACGTACTTCATAG CTCAACCAACGAAATGGCTTTTGATACTCTGAAGGATCAAGGGGT GCACGTCGTCCCATATTTGAAGAAAGCATGGGCTGATATGTGTAA GTCTTTCCTGCTGGAAGCTAAGTGGTATTCTAGCGGCTATATTCCT ACGCTTGACGAGTATATTGAAAACGCGTGGGTTAGCGTCAGCGGT CCAGTTATCTTGTTACACGCCTACACTCTGATTGCTAACCCCGCTA AAGAGGAGGCCCTTCAGTTCTTACAAGAATACCCTCACATCATCC GTTGGCCATCCATGATCTTTCGTCTTGCCAATGATCTGGCCACTAG TTCCGATGAAGTGAAGCGCGGTGATGTGCCCAAGGCCATTCAATG TTATATGCACGAAACGGGTGCCAGTGAAAGCGACGCTCGCCTTCA CATCCGCGACCTTATCACAGCAGCCTGGATGAAGATGAATAACAA GCGCGAAGGCGACGAAAACCCTGACCATCTGTTGCTTCCGAATAA TTTTGTGCAGTTTGCCATGAACTTAGCGCGTATGGCACAATGTACA TATCAAAACGGCGATCGTCACACGGTTCAAGATAACTCTAAGAACC GCGTTTTACCGTTATTAATCCATCCCATTAAGTCCTAAHeliconius melpomene ATGTCTGAGACAGAGGTCCACGTAATCAACGTCAACGGAAAAGAG 35 P-Ocimene Synthase AACGACGACCTTTTTCTTGAAAAAGAGATTTTAGCCCCTTTTTCTCA (HmOciS) CATCTGTCAAGTGAAAGGGAAACAATTACGCATCAAGATCATGCGT GCATTTAATCACTGGCTTCAAGCAAGCGAGGCCGATGTCATGAAA GCCCTTGGGATTGTAAACTCTCTGCACGTGGCAAGCCTGCTGATT GACGACATCCAGGACGATAGTACGTTGCGCCGTGGGATGCCAGC AGCTCATTGTGTGTACGGAGTACCTCTTACGGTGAATACAAGTTTG CACGCCACTTTCCTGGTATTAGAAGAGGCTTTCGCAACCGATCCG CGTACAGCAAAATTACTGGTTGAGGATTTCCTTGAAATGTGCCGTG GTCAAGGAATTGATATTTATTGGCGCGACCATTTAATTTGCCCGAC CGAGGGCCAATACTATAAAATGCTGGAGCAGAAGACGGGTCATTT CTTCCTGATGGGAGTGCGTATGATGCAACTGTTTAGCTGCAATAAGACGGACTACTCAGAATTAGTATTATTGATGGGGCGCTACTTCCAAATCCGTGATGATTACTGCAACTTAAGCCAGCAAGAAGCGTTGGAA GAATGGCCTGGGGCGGAAGACCTTCAGGTGTGTAAGAACGACTC ATTCTGCGAAGACATTACGGAGGGGAAAATCAGCTTACCCATTATT CATGCACTTCAAACTAAAAAAGCTGGAATCGTTATGAACATTCTGC GTCAAAAGACTCGTGATATGAACTTAAAGAAATACTGCGTGAGCAC ATTGAAGGAGATCGGGTCTTTGCAGTATACGCAAAACGTGCTTGA GAAGTTGGACCTGGAAATCCGCGCAGAGGTTGCACGTCTTGGCG GAAATCCATTGATCGACGAAGTGCTGCATTCTTTATTAAGTTGGAA GGACAACTAAPhaseolus lunatus 0- ATGAGCGAACGTAAATCTGCAAACTACCAACCTAATTTGTGGACAT 36 Ocimene Synthase ACGACTTTTTGCAAAGTCTTAAACACGCATACGCGGATACGCGCTA (PlOciS) TGAGGATCGTGCGAAACAGCTGCAGGAGGAAGTTCGCAAGATGAT CAAGGACGAAAACTCTGACATGTGGCTGAAACTTGAGCTTATCAAT GACGTAAAGCGTTTGGGATTGAGTTATCACTATGACAAGGAGATT GGGGAAGCCCTTTTGCGCTTCCATTCGAGTGCAACCTTCTCCGGT ACTATTGTCCATCGTTCACTGCACGAGACCGCGCTGTGCTTTCGC CTTTTACGCGAGTATGGTTATGATGTTACGGCTGACATGTTTGAGC GTTTTAAGGAGCGCAACGGCCATTTTAAGGCAAGTTTGATGTCGG ATGTCAAAGGGATGTTATCTCTTTACCAGGCCAGCTTTCTTGGATA CGAAGGGGAGCAGATTTTGGATGATGCGAAGGCTTTTTCTAGTTT CCATTTAAAAAGCGTTTTATCCGAAGGACGTAACAATATGGTCCTG GAAGAGGTTAACCATGCACTGGAACTGCCCTTACACCACCGTATT CAACGCCTGGAAGCGCGTTGGTACATTGAATACTATGCGAAGCAA CGTGATAGTAACCGTGTTTTGCTTGAAGCCGCCAAACTTGATTTCA ATATTTTACAGAGCACCTTGCAAAACGACCTTCAAGAAGTCAGTCG CTGGTGGAAAGGCATGGGCTTAGCCAGCAAATTATCTTTCTCTCG CGATCGTTTAATGGAATGTTTCTTTTGGGCTGCAGGCATGGTATTC GAACCGCAATTTTCAGATCTTCGCAAGGGTCTTACCAAAGTGGCTA GCCTGATCACCACTATTGACGACGTGTATGATGTTTATGGCACTCT GGAGGAATTAGAGCTGTTCACTGCCGCGGTTGAGAGTTGGGACG TCAAAGCGATTCAAGTTTTACCAGATTACATGAAAATCTGTTTTCTT GCACTTTATAACACAGTAAATGAGTTCGCGTATGACGCTTTAAAGG AACAGGGACAAGACATTCTGCCTTATCTTACGAAAGCTTGGTCTGA TCTTCTGAAAGCATTCTTACAGGAGGCTAAGTGGTCACGTGACCG TCACATGCCCCGTTTTAATGATTATTTGAACAATGCGTGGGTCTCA GTGTCCGGGGTAGTCTTATTGACTCACGCATATTTCTTGCTGAATC ACAGTATTACCGAGGAGGCTCTGGAGTCGTTGGACTCATATCATT CGTTGCTGCAAAATACTTCATTGGTGTTTCGTCTTTGTAATGACCTT GGAACGAGCAAGGCTGAGTTAGAGCGTGGCGAGGCCGCAAGCTC TATTTTATGTTATCGCCGTGAAAGTGGGGCCTCCGAGGAGGGAGC TTACAAACACATTTATTCATTATTGAACGAAACGTGGAAGAAGATG AACGAAGATCGCGTATCACAGAGTCCTTTTCCTAAGGCCTTTGTAG AGACTGCCATGAACTTGGCCCGCATTTCACACTGTACTTATCAGTA CGGCGATGGGCATGGGGCTCCTGACTCAACAGCAAAGAACCGTA TCCGCAGTTTAATTATCGAGCCTATTGCCTTGTACGAAACAGAGAT CAGTACGTCGTATTAAM. spicata LimS MALKVLSVATQMAIPSNLTTCLQPSHFKSSPKLLSSTNSSSRSRLRVY 37CSSSQLTTERRSGNYNPSRWDVNFIQSLLSDYKEDKHVIRASELVTLV KMELEKETDQIRQLELIDDLQRMGLSDHFQNEFKEILSSIYLDHHYYKN PFPKEERDLYSTSLAFRLLREHGFQVAQEVFDSFKNEEGEFKESLSD DTRGLLQLYEASFLLTEGETTLESAREFATKFLEEKVNEGGVDGDLLT RIAYSLDIPLHWRIKRPNAPVWIEWYRKRPDMNPVVLELAILDLNIVQA QFQEELKESFRWWRNTGFVEKLPFARDRLVECYFWNTGIIEPRQHASARIMMGKVNALITVIDDIYDVYGTLEELEQFTDLIRRWDINSIDQLPDYMQLCFLALNNFVDDTSYDVMKEKGVNVIPYLRQSWVDLADKYMVEA RWFYGGHKPSLEEYLENSWQSISGPCMLTHIFFRVTDSFTKETVDSL YKYHDLVRWSSFVLRLADDLGTSVEEVSRGDVPKSLQCYMSDYNAS EAEARKHVKWLIAEVWKKMNAERVSKDSPFGKDFIGCAVDLGRMAQ LMYHNGDGHGTQHPIIHQQMTRTLFEPFAS. suberectus SsNerS MDDIYIQQALVLKEVKHVLKKLISEDPMESLYMVDTIQRLGIEHHFEEEI 38EAALQKQHLIFSSHLSDFANNHKLYEVALLFRLLRQRGHYVHADVFDS LKSNKREFREKHGEDVKGLIALYEATQVSIEADDSLDDAGYLSYQLLH AWLARNKEHHEAIYVANTLQNPLHYGLSRFMDRSTILTSDFKTKKEW KCLEELAEINSCIVKFMNQNEIIQVYKWWKDLGMANEVKFARYQPLK WYMWPMACFTDPRFSDQRIQLTKPISLIYIIDDIFDVHGTLDQLTLFTD AVNRWELTGTEQLPDFMKMCLSVLYDITNDFAEMINKKHGLNPIDTLK RSWVRLLNAFLEEAHWLNSGHLPRSEEYLNNGIVSTGVHVVLVHAFF LLDQSINKEIVAIIDNFPEIIHSVAKILRLSDDLEGAKSEDQNGLDGSYLD CYMKEHQDVSAEDAQRHVAHLISSEWKRLNREILTPNPLPSSFTNFC LNAARMVPLMYHYRSNSSLSNLREQIQMLLNVDAGHIC. cajan CcNerS MDDIYMKQGLVLKEVKHVFQKLISENPMESLCMVDIIQRLGIEHLFEEE 39TEAVLQNHHFIFSSHLNDFSDNHKLYKVALAFRLLRQRGHYVHADLFD SLKSNKREFREKYGEDVKSLIALYEATQLSIEGEDSVDEAGYLSNQLL HAWLTTHKEHHEAIYVANTLQSPLHYGLSRFRDKNMHLSDFKTKKEW TCLEELAEINSCIVRSMNQNEIKQVYKWWKDLGMAKEVKFARYQPLK WYMWPMACFTDPSFSDQRIELTKPISLIYIIDDIFDVYGTIDQLTLFTDAI NRWELAGTEQLPDFMKMCLSVLYEITNDFAEKIYKKHGLNPIDTLKRS WVKLLNAFLEEAHWLNSGHLPRSEEYLNNGIVTTGVHVVLVHAFFLL DQTINKEIVAIVDNFPEIIHSVAKILRLSDDLEGAQSENENGLDGSYLDC YMNEHQDISTEDAQSHVCHLISREWKRLNGQVLTSTGLPSSFTKFCL NAARMVPLMYHYQSNPSLSRLREQIQTLLDVDAGHIC.tenuipile CtGerS MPRRSGNYKPSIWDYDFVQSLGSGYKVEAHGTRVKKLKEVVKHLLK 40ETDSSLAQIELIDKLRRLGLRWLFKNEIKQVLYTISSDNTSIEMRKDLHA VSTRFRLLRQHGYKVSTDVFNDFKDEKGCFKPSLSMDIKGMLSLYEA SHLAFQGETVLDEARAFVSTHLMDIKENIDPILHKKVEHALDMPLHWR LEKLEARWYMDIYMREEGMNSSLLELAMLHFNIVQTTFQTNLKSLSR WWKDLGLGEQLSFTRDRLVECFFWAAAMTPEPQFGRCQEVVAKVA QLIIIIDDIYDVYGTVDELELFTNAIDRWDLEAMEQLPEYMKTCFLALYN SINEIGYDILKEEGRNVIPYLRNTWTELCKAFLVEAKWYSSGYTPTLEE YLQTSWISIGSLPMQTYVFALLGKNLAPESSDFAEKISDILRLGGMMIR LPDDLGTSTDELKRGDVPKSIQCYMHEAGVTEDVARDHIMGLFQETW KKLNEYLVESSLPHAFIDHAMNLGRVSYCTYKHGDGFSDGFGDPGS QEKKMFMSLFAEPLQVDEAKGISFYVDGGSAP. citriadora PcGerS MRRSGNYQPSIWDFNYVQSLNTPYKEERYLTRHAELIVQVKPLLEKK 41MEPAQQLELIDDLNNLGLSYFFQDRIKQILSFIYDENQCFHSNINDQAE KRDLYFTALGFRLLRQHGFDVSQEVFDCFKNDNGSDFKASLSDNTK GLLQLYEASFLVREGEDTLEQARQFATKFLRRKLDEIDDNHLLSCIHH SLEIPLHWRIQRLEARWFLDAYATRHDMNPVILELAKLDFNIIQATHQE ELKDVSRWWQNTRLAEKLPFVRDRLVESYFWAIALFEPHQYGYQRR VAAKIITLATSIDDVYDIYGTLDELQLFTDNFRRWDTESLGRLPYSMQL FYMVIHNFVSELAYEILKEKGFIVIPYLQRSWVDLAESFLKEANWYYSG YTPSLEEYIDNGSISIGAVAVLSQVYFTLANSIEKPKIESMYKYHHILRL SGLLVRLHDDLGTSLFEKKRGDVPKAVEICMKERNVTEEEAEEHVKY LIREAWKEMNTATTAAGCPFMDELNVAAANLGRAAQFVYLDGDGHG VQHSKIHQQMGGLMFEPYVC. acuminata CaGerS MATSTATIGDTDSLLKSQRQFTVYLPAHEADKDRKIEEIMEKTQGELE 42 KTSDPTSVMKFIDTLERLGIAYHFEEEINSLLQGFLANGYSHYPQDLFT TALRFRLLRHNGYHISADVFQKFVDKNGKFKESLREDTQGMLSLYEA SYLGANGEDILSQAMEFTETHFKQSIPLMAAVPQLEQALELPRHLRMA RLEARRFIEEYIRESDHSSALLELAKLDYNKVQLLHQSELNEISRWWK QLGLVENLGFGRDRPLECFLWTVGILPEPKYSGCRIELTKTIAVLLVLD DIFDSFGTLDELVRFTHAIRRWDLSAMEQLPEYMKVCYMALYNTTNEI GYKILKEHGWN VVPYLKRTWI D Ml EGFQAEANWCSSGYVPSLEEYI E NGVTTAGSYMALVHLFFLMGQGVTDETIGMLEPYPKFFSSSGRILRL WDDLGTASEEQERGDIASSIELFMREKDLSSQGEARKYVKQVIYSLW KELNGELMASKAMPLPLIKAAFNMARTSQVIYQHGDDNSFPSVDQCV QSLFFTPILC unshiu CuCinS MASTTTIKPVDQTIIRRSADYGPTIWSVDYIQSLDSEYKEKSYARQLQK 43LKEQVSAMLQQDNKVVDLDPLHQLELIDNLHRLGVSYHFEDEIKRTLD RIHNKNTNKSLYARALKFRILRQYGYNTPVKETFSRFMDEKGSFKLSS HSDDCKGMLALYEAAYLLVEEESSIFRDAIRFTTAYLKEWVVKHDIDK NDDEYLCTLVKHALELPLHWRMRRLEARWFIDVYESGPDMNPILLDL AKLDFNIVQAVHQEDIKYASRWWKKIGLGERLNFARDRIMENFFWTV GVIFEPNFGYCRRMSTMVNALITTIDDVYDVYGTLDELELFTDAVERW DATTIEQLPDYMKLCFHALHNSINEMAFDALRDQGVGMVISYLKKAW ADICKTYLVEAKWYNNGYIPTLQEYMENAWISISAPVILVHAYTYTANP ITKEGLEFVKDYPNIIRWSSIILRLADDLGTSSDELKRGDVHKSIQCYM HEAGVSEREAREHIHDLIAQTWMKMNRDRFGNPHFVSDVFVGIAMNL ARMSQCMYQFGDGHGHGVQEITKARVLSLIVDPIAHypoxylon sp. E7406B MRPITCSFDPVGISFQTESKQENFEFLREAISRSVPGLENCNVFDPRS 44 HypCinS LGVPWPTSFPAAAQSKYWKDAEEAAAELMDQIVAAAPGEQGSLPAE LAVSDKKAAKRRELLDTSVSAPMNMFPAANAPRARIMAKANLLIFMH DDVCEYQSVQSTIIDSALADTSTPNGKGADILWQNRIFKEFSEETNRE DPVVGPQFLQGILNWVEHTRKALPASMTFRSFNEYIDYRIGDFAVDFC DAAILLTCEIFLTPADMEPLRKLHRLYMTHFSLTNDLYSFNKEVVAEQE TGSAVINAVRVLEQLVDTSTRSAKVLLRAFLWDLELQIHDELTRLKGT DLTPSQWRFARGMVEVCAGNIFYSATCLRYAKPGLRGIS. clavuligerus ScCinS MPAGHEEFDIPFPSRVNPFHARAEDRHVAWMRAMGLITGDAAEATY 45RRWSPAKVGARWFYLAQGEDLDLGCDIFGWFFAYDDHFDGPTGTD PRQTAAFVNRTVAMLDPRADPTGEHPLNIAFHDLWQRESAPMSPLW QRRAVDHWTQYLTAHITEATNRTRHTSPTIADYLELRHRTGFMPPLLD LIERVWRAEIPAPVYTTPEVQTLLHTTNQNINIVNDVLSLEKEEAHGDP HNLVLVIQHERQSTRQQALATARRMIDEWTDTFIRTEPRLPALCGRLG IPLADRTSLYTAVEGMRAAIRGNYDWCAETNRYAVHRPTGTGRATTP WS. fruticose SfCinS MSLQTGNEIQTERRTGGYQPTLWDFSTIQSFDSEYKEEKHLMRAAG 46MIDQVKMMLQEEVDSIRRLELIDDLRRLGISCHFEREIVEILNSKYYTN NEIDERDLYSTALRFRLLRQYDFSVSQEVFDCFKNAKGTDFKPSLVD DTRGLLQLYEASFLSAQGEETLRLARDFATKFLQKRVLVDKDINLLSSI ERALELPTHWRVQMPNARSFIDAYKRRPDMNPTVLELAKLDFNMVQ AQFQQELKEASRWWNSTGLVHELPFVRDRIVECYYWTTGVVERRQH GYERIMLTKINALVTTIDDVFDIYGTLEELQLFTTAIQRWDIESMKQLPP YMQICYLALFNFVNEMAYDTLRDKGFDSTPYLRKVWVGLIESYLIEAK WYYKGHKPSLEEYMKNSWISIGGIPILSHLFFRLTDSIEEEAAESMHKY HDIVRASCTILRLADDMGTSLDEVERGDVPKSVQCYMNEKNASEEEA REHVRSLIDQTWKMMNKEMMTSSFSKYFVEVSANLARMAQWIYQHE SDGFGMQHSLVNKMLRDLLFHRYEA. majus AmOciS MGSTPPPSKLHQALCLNAHSTSCMAELPMDYEGKIQGTRHLLHLKDE 47 NDPIESLIFVDATQRLGVNHHFQKEIEEILRKSYATMKSPSICKYHTLH DVSLFFCLMRQHGRYVSADVFNNFKGESGRFKEELKRDTRGLVELY EAAQLSFEGERILDEAENFSRQILHGNLASMEDNLRRSVGNKLRYPF HKSIARFTGINYDDDLGGMYEWGKTLRELALMDLQVERSVYQEELLQ VSKWWNELGLYKKLTLARNRPFEFYMWSMVILTDYINLSEQRVELTK SVAFIYLIDDIFDVYGTLDELIIFTEAVNKWDYSATDTLPDNMKMCYMT LLDTINGTSQKIYEKYGHNPIDSLKTTWKSLCSAFLVEAKWSASGSLP SANEYLENEKVSSGVYVVLIHLFFLMGLGGTNRGSIELNDTRELMSSI AIIVRIWNDLGCAKNEHQNGKDGSYLDCYKKEHINLTAAQVHEHALEL VAIEWKRLNKESFNLNHDSVSSFKQAALNFARMVPLMYSYDNNRRG PVLEEYVKFMLSDA thaliana AtOciS MPKRQAQRRFTRKTDSKTPSQPLVSRRSANYQPSLWQHEYLLSLGN 48TYVKEDNVERVTLLKQEVSKMLNETEGLLEQLELIDTLQRLGVSYHFE QEIKKTLTNVHVKNVRAHKNRIDRNRWGDLYATALEFRLLRQHGFSIA QDVFDGNIGVDLDDKDIKGILSLYEASYLSTRIDTKLKESIYYTTKRLRK FVEVNKNETKSYTLRRMVIHALEMPYHRRVGRLEARWYIEVYGERHD MNPILLELAKLDFNFVQAIHQDELKSLSSWWSKTGLTKHLDFVRDRIT EGYFSSVGVMYEPEFAYHRQMLTKVFMLITTIDDIYDIYGTLEELQLFT TIVEKWDVNRLEELPNYMKLCFLCLVNEINQIGYFVLRDKGFNVIPYLK ESWADMCTTFLKEAKWYKSGYKPNFEEYMQNGWISSSVPTILLHLFC LLSDQTLDILGSYNHSVVRSSATILRLANDLATSSEELARGDTMKSVQ CHMHETGASEAESRAYIQGIIGVAWDDLNMEKKSCRLHQGFLEAAAN LGRVAQCVYQYGDGHGCPDKAKTVNHVRSLLVHPLPLNC. unshiu CuOciS MSATKVRDKAINDNRRSANYQPSMWSYDYLQSLSNGYVGESCAQRI 49EKLKGEVRLMLDNYKEVDDYVDALHCLEIVDNLQRLGVSYHFEGEIKR FLNSIYNKRNSRRSSTYHAKENQESLLYAASLEFRLLRQHGYDIHAHG TLSSFMDEKGKFKSCLGDDIKGILALYEAAYLLGEEESTIFHEAINFTTT HLEEYVKKHNDDDGYFSALVNHALELPLHWRMVRLEARWFIDVYER GTDMNPVLVELAKLDFNSVQAAHQDELKYVSWWWRKTGLGELHFA RDRILENFFWALGEIWEPQFGYCRRMSTKVNALITTIDDVYDVYGTLD ELEQFTNAVERWDVNAMDQLPYYMKLCFHVLHSSTNEMAFDTLKDQ GVHVVPYLKKAWADMCKSFLLEAKWYSSGYIPTLDEYIENAWVSVSG PVILLHAYTLIANPAKEEALQFLQEYPHIIRWPSMIFRLANDLATSSDEV KRGDVPKAIQCYMHETGASESDARLHIRDLITAAWMKMNNKREGDE NPDHLLLPNNFVQFAMNLARMAQCTYQNGDRHTVQDNSKNRVLPLLI HPIKSH. melpomene HmOciS MSETEVHVINVNGKENDDLFLEKEILAPFSHICQVKGKQLRIKIMRAFN 50HWLQASEADVMKALGIVNSLHVASLLIDDIQDDSTLRRGMPAAHCVY GVPLTVNTSLHATFLVLEEAFATDPRTAKLLVEDFLEMCRGQGIDIYW RDHLICPTEGQYYKMLEQKTGHFFLMGVRMMQLFSCNKTDYSELVLL MGRYFQIRDDYCNLSQQEALEEWPGAEDLQVCKNDSFCEDITEGKIS LPIIHALQTKKAGIVMNILRQKTRDMNLKKYCVSTLKEIGSLQYTQNVL EKLDLEIRAEVARLGGNPLIDEVLHSLLSWKDNP. lunatus PlOciS MSERKSANYQPNLWTYDFLQSLKHAYADTRYEDRAKQLQEEVRKMI 51KDENSDMWLKLELINDVKRLGLSYHYDKEIGEALLRFHSSATFSGTIV HRSLHETALCFRLLREYGYDVTADMFERFKERNGHFKASLMSDVKG MLSLYQASFLGYEGEQILDDAKAFSSFHLKSVLSEGRNNMVLEEVNH ALELPLHHRIQRLEARWYIEYYAKQRDSNRVLLEAAKLDFNILQSTLQ NDLQEVSRWWKGMGLASKLSFSRDRLMECFFWAAGMVFEPQFSDL RKGLTKVASLITTIDDVYDVYGTLEELELFTAAVESWDVKAIQVLPDYMKICFLALYNTVNEFAYDALKEQGQDILPYLTKAWSDLLKAFLQEAKWSRDRHMPRFNDYLNNAWVSVSGVVLLTHAYFLLNHSITEEALESLDSY HSLLQNTSLVFRLCNDLGTSKAELERGEAASSILCYRRESGASEEGA YKHIYSLLNETWKKMNEDRVSQSPFPKAFVETAMNLARISHCTYQYG DGHGAPDSTAKNRIRSLIIEPIALYETEISTSYpTAC2_PAR GGAAGCTGTGGTATGGCTGTGCAGGTCGTAAATCACTGCATAATT 52CGTGTCGCTCAAGGCGCACTCCCGTTCTGGATAATGTTTTTTGCG CCGACATAATTGTGAGCGCTCACAATTTCTGAAATGAGCTGTTGAC AATTAATCATCCGGCTCGTATAATGTGTGGAATTGTGAGCGGATAA CAATTTCACACAGGAAACAGCGCCGCTGAGAAAAAGCGAAGCGG CACTGCTCTTTAACAATTTATCAGACAATCTGTGTGGGCACTCGAC CGGAATTATCGATTAACTTTATTATTAAAAATTAAATAGGAGAACAA ACCATGGATCCGAGCTCGAGATCTGCAGCTGGTACCATATGGGAA TTCGAAGCTTTCTAGAGTTTAAACCAAATAAAACGAAAGGCTCAGT CGAAAGACTGGGCCTTTCGTTTTATCTGTTGTTTGTCGGTGAACGC TCTCCTGAGTAGGACAAATCCGCCGGGAGCGGATTTGAACGTTGC GAAGCAACGGCCCGGAGGGTGGCGGGCAGGACGCCCGCCATAA ACTGCCAGGCATCAAATTAAGCAGAAGGCCATCCTGACGGATGGC CTTTTTGCGTTTCTACAAACTCTTTTTGTTTATTTTTCTAAATACATT CAAATATGTATCCGCTCATGAGACAATAACCCTGATAAATGCTTCA ATAATATTGAAAAAGGAAGAGTATGAGTATTCAACATTTCCGTGTC GCCCTTATTCCCTTTTTTGCGGCATTTTGCCTTCCTGTTTTTGCTCA CCCAGAAACGCTGGTGAAAGTAAAAGATGCTGAAGATCAGTTGGG TGCACGAGTGGGTTACATCGAACTGGATCTCAACAGCGGTAAGAT CCTTGAGAGTTTTCGCCCCGAAGAACGTTTTCCAATGATGAGCACT TTTAAAGTTCTGCTATGTGGCGCGGTATTATCCCGTGTTGACGCC GGGCAAGAGCAACTCGGTCGCCGCATACACTATTCTCAGAATGAC TTGGTTGAGTACTCACCAGTCACAGAAAAGCATCTTACGGATGGC ATGACAGTAAGAGAATTATGCAGTGCTGCCATAACCATGAGTGATA ACACTGCGGCCAACTTACTTCTGACAACGATCGGAGGACCGAAGG AGCTAACCGCTTTTTTGCACAACATGGGGGATCATGTAACTCGCCT TGATCGTTGGGAACCGGAGCTGAATGAAGCCATACCAAACGACGA GCGTGACACCACGATGCCTGTAGCAATGGCAACAACGTTGCGCAA ACTATTAACTGGCGAACTACTTACTCTAGCTTCCCGGCAACAATTA ATAGACTGGATGGAGGCGGATAAAGTTGCAGGACCACTTCTGCGC TCGGCCCTTCCGGCTGGCTGGTTTATTGCTGATAAATCTGGAGCC GGTGAGCGTGGGTCTCGCGGTATCATTGCAGCACTGGGGCCAGA TGGTAAGCCCTCCCGTATCGTAGTTATCTACACGACGGGGAGTCA GGCAACTATGGATGAACGAAATAGACAGATCGCTGAGATAGGTGC CTCACTGATTAAGCATTGGTAACTGTCAGACCAAGTTTACTCATAT ATACTTTAGATTGATTTAAAACTTCATTTTTAATTTAAAAGGATCTAG GTGAAGATCCTTTTTGATAATCTCATGACCAAAATCCCTTAACGTG AGTTTTCGTTCCACTGAGCGTCAGACCCCGTAGAAAAGATCAAAG GATCTTCTTGAGATCCTTTTTTTCTGCGCGTAATCTGCTGCTTGCA AACAAAAAAACCACCGCTACCAGCGGTGGTTTGTTTGCCGGATCA AGAGCTACCAACTCTTTTTCCGAAGGTAACTGGCTTCAGCAGAGC GCAGATACCAAATACTGTCCTTCTAGTGTAGCCGTAGTTAGGCCA CCACTTCAAGAACTCTGTAGCACCGCCTACATACCTCGCTCTGCTA ATCCTGTTACCAGTGGCTGCTGCCAGTGGCGATAAGTCGTGTCTT ACCGGGTTGGACTCAAGACGATAGTTACCGGATAAGGCGCAGCG GTCGGGCTGAACGGGGGGTTCGTGCACACAGCCCAGCTTGGAGC GAACGACCTACACCGAACTGAGATACCTACAGCGTGAGCTATGAG AAAGCGCCACGCTTCCCGAAGGGAGAAAGGCGGACAGGTATCCG GTAAGCGGCAGGGTCGGAACAGGAGAGCGCACGAGGGAGCTTCCAGGGGGAAACGCCTGGTATCTTTATAGTCCTGTCGGGTTTCGCCACCTCTGACTTGAGCGTCGA I I I I I G I GATGCTCGTCAGGGGGGC GGAGCCTATGGAAAAACGCCAGCAACGCGGCCTTTTTACGGTTCC TGGCCTTTTGCTGGCCTTTTGCTCACATGTTCTTTCCTGCGTTATC CCCTGATTCTGTGGATAACCGTATTACCGCCTTTGAGTGAGCTGAT ACCGCTCGCCGCAGCCGAACGACCGAGCGCAGCGAGTCAGTGAG CGAGGAAGCGGAAGAGCGCCTGATGCGGTATTTTCTCCTTACGCA TCTGTGCGGTATTTCACACCGCATATGGTGCACTCTCAGTACAATC TGCTCTGATGCCGCATAGTTAAGCCAGTATACCGGGCAAATCGCT GAATATTCCTTTTGTCTCCGACCATCAGGCACCTGAGTCGCTGTCT TTTTCGTGACATTCAGTTCGCTGCGCTCACGGCTCTGGCAGTGAA TGGGGGTAAATGGCACTACAGGCGCCTTTTATGGATTCATGCAAG GAAACTACCCATAATACAAGAAAAGCCCGTCACGGGCTTCTCAGG GCGTTTTATGGCGGGTCTGCTATGTGGTGCTATCTGACTTTTTGCT GTTCAGCAGTTCCTGCCCTCTGATTTTCCAGTCTGACCACTTCGGA TTATCCCGTGACAGGTCATTCAGACTGGCTAATGCACCCAGTAAG GCAGCGGTATCATCAACAGGCTTACCCGTCTTACTGTCGGCATGC ATTTACGTTGACACCATCGAATGGTGCAAAACCTTTCGCGGTATGG CATGATAGCGCCCGGAAGAGAGTCAATTCAGGGTGGTGAATGTGA AACCAGTAACGTTATACGATGTCGCAGAGTATGCCGGTGTCTCTTA TCAGACCGTTTCCCGCGTGGTGAACCAGGCCAGCCACGTTTCTGC GAAAACGCGGGAAAAAGTGGAAGCGGCGATGGCGGAGCTGAATT ACATTCCCAACCGCGTGGCACAACAACTGGCGGGCAAACAGTCGT TGCTGATTGGCGTTGCCACCTCCAGTCTGGCCCTGCACGCGCCG TCGCAAATTGTCGCGGCGATTAAATCTCGCGCCGATCAACTGGGT GCCAGCGTGGTGGTGTCGATGGTAGAACGAAGCGGCGTCGAAGC CTGTAAAGCGGCGGTGCACAATCTTCTCGCGCAACGCGTCAGTG GGCTGATCATTAACTATCCGCTGGATGACCAGGATGCCATTGCTG TGGAAGCTGCCTGCACTAATGTTCCGGCGTTATTTCTTGATGTCTC TGACCAGACACCCATCAACAGTATTATTTTCTCCCATGAAGACGGT ACGCGACTGGGCGTGGAGCATCTGGTCGCATTGGGTCACCAGCA AATCGCGCTGTTAGCGGGCCCATTAAGTTCTGTCTCGGCGCGTCT GCGTCTGGCTGGCTGGCATAAATATCTCACTCGCAATCAAATTCA GCCGATAGCGGAACGGGAAGGCGACTGGAGTGCCATGTCCGGTT TTCAACAAACCATGCAAATGCTGAATGAGGGCATCGTTCCCACTG CGATGCTGGTTGCCAACGATCAGATGGCGCTGGGCGCAATGCGC GCCATTACCGAGTCCGGGCTGCGCGTTGGTGCGGATATCTCGGT AGTGGGATACGACGATACCGAAGACAGCTCATGTTATATCCCGCC GTTAACCACCATCAAACAGGATTTTCGCCTGCTGGGGCAAACCAG CGTGGACCGCTTGCTGCAACTCTCTCAGGGCCAGGCGGTGAAGG GCAATCAGCTGTTGCCCGTCTCACTGGTGAAAAGAAAAACCACCC TGGCGCCCAATACGCAAACCGCCTCTCCCCGCGCGTTGGCCGAT TCATTAATGCAGCTGGCACGACAGGTTTCCCGACTGGAAAGCGGG CAGTAATAATTTAAATTE. coli dxs, UniProt MSFDIAKYPTLALVDSTQELRLLPKESLPKLCDELRRYLLDSVS 53 Accession No. RSSGHFASGLGTVELTVALHYVYNTPFDQLIWDVGHQAYPHKI P77488 LTGRRDKIGTIRQKGGLHPFPWRGESEYDVLSVGHSSTSISAGI GIAVAAEKEGKNRRTVCVIGDGAITAGMAFEAMNHAGDIRPDM LVILNDNEMSISENVGALNNHLAQLLSGKLYSSLREGGKKVFS GVPPIKELLKRTEEHIKGMWPGTLFEELGFNYIGPVDGHDVLG LITTLKNMRDLKGPQFLHIMTKKGRGYEPAEKDPITFHAVPKFD PSSGCLPKSSGGLPSYSKIFGDWLCETAAKDNKLMAITPAMRE GSGMVEFSRKFPDRYFDVAIAEQHAVTFAAGLAIGGYKPIVAIYSTFLQRAYDQVLHDVAIQKLPVLFAIDRAGIVGADGQTHQGAFDLSYLRCIPEMVIMTPSDENECRQMLYTGYHYNDGPSAVRYP RGNAVGVELTPLEKLPIGKGIVKRRGEKLAILNFGTLMPEAAKV AESLNATLVDMRFVKPLDEALILEMAASHEALVTVEENAIMGGA GSGVNEVLMAHRKPVPVLNIGLPDFFIPQGTQEEMRAELGLDA AGMEAKIKAWLARhodobacter capsulatus MSATPSRTPHLDRVTGPADLKAMSIADLTALASEVRREIVEWS 54 dxsi, UniProt QTGGHLGSSLGWELTVALHAVFNSPGDKLIWDVGHQCYPHKI Accession No. LTGRRSRMLTLRQAGGISGFPKRSESPHDAFGAGHSSTSISAA D5AP89 LGFAVGRELGQPVGDTIAIIGDGSITAGMAYEALNHAGHLKSR MFVILNDNDMSIAPPVGALQHYLNTIARQAPFAALKAAAEGIEM HLPGPVRDGARRARQMVTAMPGGATLFEELGFDYIGPVDGHD MAELVETLRVTRARASGPVLIHVCTTKGKGYAPAEGAEDKLHG VSKFDIETGKQKKSIPNAPNYTAVFGERLTEEAARDQAIVAVTA AMPTGTGLDIMQKRFPRRVFDVGIAEQHAVTFAAGMAAAGLK PFLALYSSFVQRGYDQLVHDVALQNLPVRLMIDRAGLVGQDG ATHAGAFDVSMLANLPNFTVMAAADEAELCHMWTAAAHDSG PIALRYPRGEGRGVEMPERGEVLEIGKGRVMTEGTEVAILSFG AHLAQALKAAEMLEAEGVSTTVADARFCRPLDTDLIDRLIEGHA ALITLEQGAMGGFGAMVLHYLARTGQLEKGRAIRTMTLPDCYI DHGSPEEMYAWAGLTANDIRDTALAAARPSKSVRIVHSARhodobacter capsulatus MSQPPVTPILDRVRLPSDMKGLSDRDLHRLADELRAETISAVS 55 dxs2, UniProt VTGGHLGAGLGWELTVALHAVFDCPRDKIIWDVGHQCYPHKI Accession No. LTGRRDRIRTLRTRGGLSGFTKRSESPYDAFGAGHSSTSISAA D5ASU5 LGFTMARELGSECGDAIAVIGDGAMTGGMAFEALNHGGHLGK RMFVILNDNEMSISPPVGALSSYLTRLYAEAPMQDLKAMAKGA VSLLPEPFQEGARRAKEMLKGMTIGGTLFEELGFDYVGPVDG HDLETLLHLFRTLKTRATGPVLIHALTKKGKGYAPAENRPDRG HATARFDVLTGAQVKAASNAPSYTKVFAESLIDQASRDEKIVAV TAAMPEGTGLNLFAERFPRRCFDVGIAEQHGVTFSAGLAAGG MKPFCAIYSTFLQRGYDQIVHDVAIQRLPVRFAIDRAGLVGADG ATHAGAFDIGFMANLPGMWMAAADEADLVHMVATAAAHDEG PIAFRYPRGEGMGVDMPTQGVPLEIGKGRIISEGARVAILSFGT RLAEVLKAREALAARGLAPTVADARFAKPLDKDLILRLVAEHEA LICIEEGAVGGFGSHVAQFLSEQGVFDRGYKFRSMVLPDTFID HANPEDMYAVARMNAADIEAKVLDLLGVAVAKRAPseudomonas MPTTFQEIPRKRPTTPLLDRANTPAGLRRLGEAELETLADELRL 56 fluorescens dxs, ELLYTVGQTGGHFGAGLGVIELTIALHYVFDTPDDRLVWDVGH UniProt Accession QAYPHKILTGRRERMATLRQKDGIAAFPRRSESEYDTFGVGHS No. A0A379IJU9 STSISAALGMAIAARLQNSDRKAIAVIGDGALTAGMAFEALNHA PEVDANMLVILNDNDMSISRNVGGLSNYLAKILSSRTYASMRE GSKKVLSRLPGAWEIARRTEEYAKGMLVPGTLFEELGWNYIGP IDGHDLPTLIATLRNMRDLKGPQFLHWTKKGKGFAPAEVDPIG YHAITKLEPLDAPVAAPKKTGGPKYSGVFGEWLCDMAAADPR LVGITPAMKEGSDLVAFSERFPLRYFDVAIAEQHAVTFAAGMA CEGAKPWAIYSTFLQRGYDQLVHDVAVQNLDVLFAIDRAGLV GEDGPTHAGSYDLSYLRCIPGMLVMTPSDENELRKMLSTGHL YNGPAAVRYPRGNGPNAVIEKDLEPIEIGKGIVRRQGSKTAFLVFGVQLAEALKVAEKIDATVIDMRFVKPLDEALVREIAGSHELLVTVEENAIMGGAGAAVSEFLARENMLKSVLHLGLPDIYVEHAKPA QM LAECGLDEAG I EAAVRQRMALLG LE. coli dxr, UniProt MKQLTILGSTGSIGCSTLDWRHNPEHFRWALVAGKNVTRMV 57 Accession No. EQCLEFSPRYAVMDDEASAKLLKTMLQQQGSRTEVLSGQQAA P45568 CDMAALEDVDQVMAAIVGAAGLLPTLAAIRAGKTILLANKESLV TCGRLFMDAVKQSKAQLLPVDSEHNAIFQSLPQPIQHNLGYAD LEQNGWSILLTGSGGPFRETPLRDLATMTPDQACRHPNWSM GRKISVDSATMMNKGLEYIEARWLFNASASQMEVLIHPQSVIH SMVRYQDGSVLAQLGEPDMRTPIAHTMAWPNRVNSGVKPLD FCKLSALTFAAPDYDRYPCLKLAMEAFEQGQAATTALNAANEI TVAAFLAQQIRFTDIAALNLSVLEKMDMREPQCVDDVLSVDAN AREVARKEVMRLASR. capsulatus dxr, MRRITVFGSTGSIGVNTLDLIARRPGAFQWALSGGRNIKLLAE 58 UniProt Accession QARAFRAEVAVTAHEDCLPALRDALAGTGIEATAGKAALIEAAL No. D5ATT5 RPADWIMSAIVGAAGLAPGFAALSQGTTLALANKESLVTAGPLL LAEAAKHNARLLPVDSEHSAVFQALVGEEMAAVERIIITASGGA FRDWPIEKLAHATVSQASTHPNWAMGQRITIDSASFFNKALELI ETKEYFGVRPEQIEVLVHPESLVHALVGFKDGALMAHLGAPDM RHAIGYALNWPERAELPVTRLDLAKIGALTFRAPDEDRFPALRL AREVMAMGGTAGAAFNAAKEAALDAFIGEEIGFLAMSRLVEAV LTEMTARHGLGKPLDSLDMVMDTDATARALARSLVPTVRTPseudomonas MSRPQQITVLGATGSIGLSTLDVIARHPDRYQVFALSGFTRLSE 59 fluorescens dxr, LLALCVRHVPRFAWPEVAAARTLQDDLRAAGLATRVLVGEEG UniProt Accession LCQVASDPEVDTLMAAIVGAAGLRPTLAAVEAGKKILLANKEAL No. A0A0P9ALW5 VMSGALFMQAVRKSGSVLLPIDSEHNAIFQCMPLDYARGLEAV GVRRILLTASGGPFRQTPMAELAHVSPDQACAHPNWSMGRKI SVDSASMMNKGLELIEACWLFDARPSQVEWIHPQSVIHSLVD YVDGSVLAQLGNPDMRTPIANALAWPERIDSGVAPLDLFAIARL DFEAPDEERFPCLRLARQAAEAGNSAPAMLNAANEVAVAAFL DERVRYPEIASIIEEVLNLEPWAVDDLDAVFTADAKARLLAGQ WLSRHGRE. coli ispD UniProt MATTHLDVCAWPAAGFGRRMQTECPKQYLSIGNQTILEHSVH 60 Accession No. ALLAHPRVKRWIAISPGDSRFAQLPLANHPQITWDGGDERAD Q46893 SVLAGLKAAGDAQWVLVHDAARPCLHQDDLARLLALSETSRT GG I LAAPVRDTM KRAEPG KNAI AHTVDRNGLWHALTPQFFPRE LLHDCLTRALNEGATITDEASALEYCGFHPQLVEGRADNIKVTR PEDLALAEFYLTRTIHQENTR. capsulatus MTVAVIIVAAGRGTRAGEGLPKQWRDLAGRPVLAQTVAAFAGL 61 ispDF, UniProt GRILWLHPDDMGLGMDLLGGSWLVAGGSTRSESVKNALEAL Accession No. EGSDVTRVLIHDGARPLVPASVTAAVLAALETTPGAAPALAVTD Q08113 ALWRGEAGLVAGTQDREGLYRAQTPQGFRFPEILAAHRAHPG GAADDVEVARHAGLSVAIVPGHEDNLKITYAPDFARAEAILRER KGLTMDVRLGNGYDVHAFCEGDHWLCGVKVPHVKALLGHSD ADVGMHALTDAIYGALAEGDIGRHFPPSDPQWKGAASWIFLD HAAKLAKSRGFRIGNADVTLICERPKVGPHAVAMAAELARIMEI EPSRVSVKATTSERLGFTGREEGIASIATVTLIGAR. capsulatus ispD, MTVAVIIVAAGRGTRAGEGLPKQWRDLAGRPVLAQTVAAFAGL 62 amino acids 1-222 GRILWLHPDDMGLGMDLLGGSWLVAGGSTRSESVKNALEAL of UniProt Accession EGSDVTRVLIHDGARPLVPASVTAAVLAALETTPGAAPALAVTD No. Q08113 ALWRGEAGLVAGTQDREGLYRAQTPQGFRFPEILAAHRAHPG GAADDVEVARHAGLSVAIVPGHEDNLKITYAPDFARAEAILRER KGLTMDPseudomonas ispD, MSHDLPAFWAVIPAAGVGARMAADRPKQYLQLGGRTILEHSL 63 GenBank Accession GCFLDHPRLKGLWSLAIDDPYWPTLACATDPRIQRADGGSER No. QQU70697.1 SGSVLNALLQLNALGASDDDWVLVHDAARPNLSRDDLDKLLA ELADDPVGGLLAVPARDTLKRVDKHGRWDTVDRSLIWQAYT PQMFRLGALHRALADSLVADAVITDEASAMEWSGQAPRLIEGR SDNIKVTRPEDLEWLKLRWANRRE. coli ispE, UniProt MRTQWPSPAKLNLFLYITGQRADGYHTLQTLFQFLDYGDTISIE 64 Accession No. LRDDGDIRLLTPVEGVEHEDNLIVRAARLLMKTAADSGRLPTG P62615 SGANISIDKRLPMGGGLGGGSSNAATVLVALNHLWQCGLSMD ELAEMGLTLGADVPVFVRGHAAFAEGVGEILTPVDPPEKWYLV AHPGVSIPTPVIFKDPELPRNTPKRSIETLLKCEFSNDCEVIARK RFREVDAVLSWLLEYAPSRLTGTGACVFAEFDTESEARQVLE QAPEWLNG FVAKG ANLSPLH RAM LR. capsulatus ispE, MRRIDVFAPAKVNLALHVTGQRADGYHLLDSLVAFSPVGDGLT 65 UniProt Accession LAPAEGLGLKVSGPEGAAVPEGPENLVLKAAALFGVGAAWLE No. D5AMY7 KCLPAASGIGGGSSDAAAALRGMAALTGQPLPEAGAVLNLGA DVPMCLDPRPARTRGIGDDLTPVTLPTLPAVLVNPRVEVPTPS VFKALRRKDNPPLPEI PAF ADAG AC I AWLAEQRN D LQAPAVAQ APVIAEVLAVLAALPGCQLARMSGSGATCFGLFGTEAEAAAAQ AALAASHPGWWSAHGALGDQAQKAAPQLRPseudomonas MSAQKLTLPSPAKLNLMLHILGRREDGYHELQTLFQFLDYGDE 66 fluorescens ispE, LTFAVREDGVIQLHTEFDGVPHDSNLIVRAAKKLQAQSGCPLGI GenBank Accession DIWIDKILPMGGGIGGGSSNAATTLLGLNHLWQLAWDHDRLAA No. QQU66123.1 LGLTLGADVPVFVRGHAAFAEGVGEKLTPVEPEEPWYWLVP QVSVSTAEIFSDPLLTRNSSPIKVRPVPKGNSRNDCLPWARR YPEVRNALNLLGKFTEAKLTGTGSCVFGGFPSKAEADKVSALL TETLTGFVAKGSNVSMLHRKLQSLLE. coli ispF, UniProt MRIGHGFDVHAFGGEGPIIIGGVRIPYEKGLLAHSDGDVALHAL 67 Accession No. TDALLGAAALGDIGKLFPDTDPAFKGADSRELLREAWRRIQAK P62617 GYTLGNVDVTIIAQAPKMLPHIPQMRVFIAEDLGCHMDDVNVK ATTTEKLGFTGRGEGIACEAVALLIKATKR. capsulatus ispF, VRLGNGYDVHAFCEGDHWLCGVKVPHVKALLGHSDADVGM 68 amino acids 223- HALTDAIYGALAEGDIGRHFPPSDPQWKGAASWIFLDHAAKLA 379 of UniProt KSRGFRIGNADVTLICERPKVGPHAVAMAAELARIMEIEPSRVS Accession No. VKATTSERLGFTGREEGIASIATVTLIGAQ08113Pseudomonas MRIGHGYDVHRFAEGDFITLGGVRIAHHHGLLAHSDGDWLHA 69fluorescens ispF, LSDALLGAAALGDIGKHFPDTDPTFKGADSRVLLRHWGLIHAKGenBank Accession G WKVG NVDNTI VAQAPKM APH I EAM RTAI AADLQ I ELDQVNVK No. QQU70693.1 ATTTEKLGFTGREEGIAVHSVALLLRAE. coli ispG UniProt MHNQAPIQRRKSTRIYVGNVPIGDGAPIAVQSMTNTRTTDVEA 70 Accession No. TVN Q I KALERVG AD I VRVSVPTM D AAE AF KL I KQ Q VN VP LVAD I P62620 HFDYRIALKVAEYGVDCLRINPGNIGNEERIRMWDCARDKNIPI RIGVNAGSLEKDLQEKYGEPTPQALLESAMRHVDHLDRLNFD QFKVSVKASDVFLAVESYRLLAKQIDQPLHLGITEAGGARSGA VKSAIGLGLLLSEGIGDTLRVSLAADPVEEIKVGFDILKSLRIRSR GINFIACPTCSRQEFDVIGTVNALEQRLEDIITPMDVSIIGCWN GPGEALVSTLGVTGGNKKSGLYEDGVRKDRLDNNDMIDQLEA RIRAKASQLDEARRIDVQQVEKPantoea ananatis MHNQAPIIRRKSKRIYVGQVPIGDGAPIAVQSMTNTRTTDVAAT 71 ispG, UniProt VNQIKALERVGVDIVRVSVPTMDAAEAFKLIKQQVNVPLVADIH Accession No. FDYRIALKVAEYGVDCLRINPGNIGNNERIRAWDCARDNNIPIR A0A0H3KYT9 IGVNAGSLEKDLQEKYGEPTPQALLESAMRHVDHLDRLNFDQF KVSVKASDVFLAVESYRLLAKQIEQPLHLGITEAGGARAGAVKS AIGLGLLLAEGIGDTLRISLAADPVEEVKVGFDILKSLRIRSRGIN FIACPTCSRQEFDVIGTVNALEQRLEDIITPMDVSIIGCWNGPG EATVSTLGVTGSNKKSGFYEDGVRQRERLDNDDMIDQLEARIR AKAAMLDETRRIDVQQLEKProteus mirabilis MHKESPIIRRKSTRIYVGNVPIGDGAPIAVQSMTNTRTTDVEAT 72 ispG, UniProt VNQIKSLERVGVDIVRVSVPTMDAAEAFKLIKQQVKVPLVADIH Accession No. FDYRIALKVAEYGVDCLRINPGNIGNEERIRQWDCARDRNIPIR B4EZT3 IGVNGGSLEKDIQEKYGEPTPEALVESAMRHVDILDKLNFDQFK VSVKASDVFLAVDSYRLLAKKIDQPLHLGITEAGGARAGSVKSA IGLGILLSEGIGDTLRISLAADPVEEVKVGFDILKSLRIRSRGINFI ACPTCSRQEFDVIGTVNELEQRLEDIITPMDVSIIGCWNGPGE AEVSTLGVTGAKTRSGFYEDGVRQKERLDNSNMIDLLEAKIRT KAAMLDGNLRININQLDKSerratia marcescens MHNQAPINRRKSTRIYVGKVPIGDGAPIAVQSMTNTRTTDVEA 73 ispG, UniProt TVNQIKALERVGVDIVRVSVPTMDAAEAFKLIKQQVNVPLVADI Accession No. HFDYRIALQVAEYGVDCLRINPGNIGNESRIRSWDCARDKNIPI A0A0P0QHY9 with RIGVNGGSLEKDLQEKYGEPTPEALLESAMRHVDILDRLNFDQ K365N substitution FKVSVKASDVFLAVQSYRLLASRIDQPLHLGITEAGGARSGSVK SAIGLGMLLSEGIGDTLRISLAADPVEEVKVGFDILKSLRIRARGI NFIACPTCSRQEFDVIGTVNALEQRLEDIITPMDVSIIGCWNGP GEALVSTMGVTGGHKKSGFYEDGVRQKERFDNEQMIDQLEA KIRAKAAMMDESNRITVNLLEKShewanella MYNETPIKRRPSTRIYVGNVPIGDGAPIAVQSMTNTKTTDVEAT 74 oneidensis ispG, VAQIRALEKVGADIVRVSVPTMDAAEAFKLIKQSVSVPLVADIHF UniProt Accession DYRIALKVAEYGVDCLRINPGNIGNEERIRSWECARDKNIPIRI No. Q8EC32 GVNGGSLEKDLMDKYKEPTPEALLESAMRHVDILDRLNFDQFK VSVKASDVF LAVESYRLLAKQ I RQ PLH LG ITEAG G ARAG AVKSA VGLGMLLAEGIGDTLRISLAADPVEEIKVGFDILKSLRIRSRGINFIACPSCSRQEFDVISTVNELERRLEDVTTAMDVSIIGCWNGPGEALVSHIGLTGGHRKSGYYDEGERQKERFDNDNLVDSLEAKIR AKASQMANRIQVKDTTESynechococcus MQTLSTPSTTATEFDTVIHRRPTRSVRVGDIWIGSRHPWVQS 75 elongatus ispG MINEDTLDIDGSVAAIRRLHEIGCEIVRVTVPSLGHAKAVGDIKK KLQDTYRDVPLVADVHHNGMKIALEVAKHVDKVRINPGLYVFE KPDPNRQGYTPEEFERIGKQIRDTLEPLVTSLREQDKAMRIGV NHGSLAERMLFTYGDTPEGMVESALEFLRLCEEMDFRNLVISM KASRAPVMM AAYRLM AKRM DDLG M DYPLHLGVTEAGDG DYG RIKSTVGIGTLLAEGIGDTIRVSLTEAPENEIPVCYSILQALGLRK TMVEYVACPSCGRTLFNLEEVLHKVRAATNHLVGLDIAVMGCI VNGPGEMADADYGYVGKTPGTIALYRGRDEIKRVPEEQGVEE LINLIKADGRWVEPEPIARhodobacter MSLNPHRPWRNITRRPSRQIWVGKVPVGGDAPISVQTMTNTIT 76 capsulatus ispG, SDVTATLGQVLRAADVGADIVRVSVPDEDATRALKEICRESPV UniProt Accession PIVADIHFHYKRAIEAAEAGAACLRINPGNIGDAARVREVIKAAK No. D5AT87 DHGCSMRIGVNAGSLEKHLLDKYGEPCPEAMLESGLDHIKILQ DNDFHEFKISVKASDVFLAAAAYTALAEATDAPIHLGITEAGGL MAGTVKSAVGLGNLLWAGIGDTIRVSLSADPVEEVKVGYEILKS LGLRTRGVQIISCPSCARQGFDVIKTVDALEKRLEHIKTPMSLSII GCWNGPGEALMTDIGFTGGGAGSGMVYLAGKQSHRLSNEQ M I DEI VRM VEDRAAQI AAEEQAAEE. coli ispH, UniProt MQILLANPRGFCAGVDRAISIVENALAIYGAPIYVRHEWHNRY 77 Accession No. WDSLRERGAIFIEQISEVPDGAILIFSAHGVSQAVRNEAKSRDL P62623 TVFDATCPLVTKVHMEVARASRRGEESILIGHAGHPEVEGTMG QYSNPEGGMYLVESPDDVWKLTVKNEEKLSFMTQTTLSVDDT SDVIDALRKRFPKIVGPRKDDICYATTNRQEAVRALAEQAEWL WGSKNSSNSNRLAELAQRMGKRAFLIDDAKDIQEEWVKEVK CVGVTAGASAPDILVQNWARLQQLGGGEAIPLEGREENIVFE VPKELRVDIREVDAcinetobacter baylyi MEIVLANPRGFCAGVDRAIAIVNRALECFNPPIYVRHEWHNKF 78 ispH, UniProt WDDLRQRGAIFVDELDQVPDDSIVIFSAHGVSKAVQQEAEHR Accession No. GLKVFDATCPLVTKVHIEVTKYAREGTEAILIGHEGHPEVEGTM Q9RBJ0 GQYDKSKGGHIYLVEDEADVEALAVNHPEKLAFVTQTTLSIDDT AKVIDALRTKFPQIQGPRKDDICYATQNRQDAVRDLASRCDW LWGSPNSSNSNRLRELAERMGKAAYLVDNADQLEQQWFDG VAKIGVTAGASAPEILIKQVIQRLQDWGAEAPKELDGREENITF SLPKELRIQVTQABurkholderia glumae M SSTDTLSG PTVAADAEI LLAQ PRG FCAG VDRAI El VERAI AM H 79 BGR1 ispH1, UniProt GSPIYVRHEIVHNKYWEDLKTKGAIFVEELEEVPSGNTVIFSAH Accession No. GVSKAVRDEAAVRGLRIYDATCPLVTKVHVEVAKMRQDGVDIV C5AC36 MIGHKGHPEVEGTMGQVERGMHLVESVEDVLALELPDPERVA LVTQTTLSVDDAAEIIAALKRKFPAIREPKKQDICYATQNRQDAV KFMAPQCDWIWGSPNSSNSSRLREVAEKRGVDAYMVDSPD QIDPAWVAGKRRIGVTAGASAPEVLAQAVIARLRELGVRNVRA LEGIEENVSFPLPRGLNLPPAGCaulobacter MNAQTPIRRPISLVLASPRGFCAGVDRAIQIVERAVEKFGAPVY 80 crescentus ispH, VRHEIVHNRHWDRLKALGAVFIEELEEAPDDRPWFSAHGVP UniProt Accession KSVPAEAKARQMIYLDATCPLVSKVHVEAQKHYDAGREIVLIGH No. A0A0H3CCC9 AGHPEVIGTMGQLPEGAVTLIEDLKDAAAWEPKDAANVAFLTQ TTLSVDDTADMVALLRERFPGIAAPHKEDICYATTNRQDAVKHL AEQSELILWGSKNSSNSVRLKEVGLKAGARDAHLIDDASGID WTWFDGISRVGLTAGASAPEDLVQGVIDAISARFDTTVEELVE ARETITFKLPRLLTAPantoea ananatis MQILLANPRGFCAGVDRAISIVENALTLFGAPVYVRHEWHNRY 81 ispH, UniProt WDSLRQRGAIFIEQIAEVPDGAILIFSAHGVSQAVRNEAKGRD Accession No. LTVFDATCPLVTKVHMEVARASRRGEESVLIGHAGHPEVEGT D4GJM9 with D148E MGQYNNPKGGMYLVESPEDVWKLDVKDDNRLSFMTQTTLSV substitution DDTSEVIDALRARFPKIVGPRKDDICYATTNRQEAVRAMAEQA DWLWGSKNSSNSNRLAELAQRMGKAAYLIDDAADIQPEWW SVDCVGVTAGASAPDILVQNVISRLQALGGGEARVLEGREENI VFEVPKELRVDVRNVQPseudomonas MQIKLANPRGFCAGVDRAIEIVNRALEVFGPPIYVRHEWHNKF 82 fluorescens ispH, WEDLRARGAIFVEELDQVPDDVIVIFSAHGVSQAVRTEAAGR UniProt Accession GLKVFDATCPLVTKVHIEVARYSRDGRECILIGHAGHPEVEGT No. C3KDX7 MGQYDASNGGAIYLVEDEKDVANLQVQNPERLAFVTQTTLSM DDTSRVIDALRSRFPAIGGPRKDDICYATQNRQDAVKQLADEC DWLWGSPNSSNSNRLRELAERMATPAYLIDGAEDMQRSWF DGVERIGITAGASAPEVLVRGVIQQLHAWGATGADELAGREEN ITFSMPKELRVRSLLProteus mirabilis MQILLANPRGFCAGVDRAISIVERALEIYGAPIYVRHEWHNRY 83 ispH, UniProt WNDLRERGAIFIEEISEVPDNAILIFSAHGVSQAIRQEARSRNL Accession No. TMLFDATCPLVTKVHMEVARASRKGKEAILIGHAGHPEVEGTM B4F2T8 GQYNNPEGGMYLVESPDDVWKLKVKDEDNLCFMTQTTLSVD DTSEVIDALNKRFPKIIGPRKDDICYATTNRQEAARELAERADV VFWGSKNSSNSNRLAELAQRAGKPSYLIDSAEDIDEIWVSHA NIVGVTAGASAPDILVQQVLARLKAFGAEEVIELSGREENIVFEV PKELRLDYKWEP. sitchensis ispH, MEADVCRNLQLRRTQVRCDAAPSAVDSATGEFDTKAFRRTLT 84 UniProt Accession RKENYNRKGFGHKEETLEAMDKEYTSDIIKTLKENNNEYTWGN No. C0PR44 with VTVKLAESFGFCWGVERAVQIAYEARKQFPDQKLWITNEIIHNP AM1-A44 truncation TVNQRLKEMQIEDIPVMEEGKKFDWNSDDWILPAFGAALSE MQILDEKSVKIVDTTCPWVSKVWNTVEKHKKESFTSVIHGKKG HEETVATSSFAGKYIIVKDIREAIYVCDYILAGKLDGSSGTKDEF LKKFEKAISRGFDPDCDLVKVGIANQTTMLKGETEEIGKLLEKT MMQKYGVEIINDHFMSFNTICDATQERQDAMYNLVKEKLDLILV VGGWNSSNTSHLQEIAEQNGIPTYWIDSEKRIGPGNRIAYKLS HGELVEKENWLPTGPLKIGITSGASTPDKILEDVLKWFKMKDE EALQTVP. trichocarpa ispH MGGDDSTSSVSLESEFDAKVFRHNLTRSKNYNRRGFGHKEET 85 UniProt Accession LELMNREYTSDIIKKLKENGYEYTWGNVTVKLAEAYGFCWGVE No. B3GEM6, with RAVQIAYEARKQFPDDKIWITNEIIHNPTVNKRLEEMEVENVPVEEG KKQFEWNGG DWI LPAFG AAVDEM LTLSSKNVQIVDTTCΔC1, A2M, and A67T PWVSKVWTTVEKHKKGDYTSIIHGKYAHEETVATASFAGKYIIV mutations KDMKEAMYVCDYILGGELNGSSSTREEFLEKFKNAVSKGFDP DSDLVKLGIANQTTMLKGETEDIGKLVERIMMRKYGVENVNDH FISFNTICDATQERQDAMYKLVEEKLDLMLWGGWNSSNTSHL QEIAEHHGIPSYWIDSEQRIGPGNKIAYKLNHGELVEKENWLPQ GPITIGVTSGASTPDKWEDALIKVFDIKRDEALQVASerratia marcescens MQILLANPRGFCAGVDRAISIVERALELYGAPIYVRHEWHNRY 86 ispH, UniProt WDSLRERGAVFIEEIAEVPDGSILIFSAHGVSQAVRAEAKARD Accession No. LTMLFDATCPLVTKVHMEVARASRRGTEAILIGHAGHPEVEGT A0A0P0QA59 with MGQYSNPQGGMYLVESPEDVWKLQVKDENNLCFMTQTTLSV D149E and N188S DDTSDVIDALRQRFPSIIGPRKDDICYATTNRQEAVRNLAGDAD substitutions WLWGSKNSSNSNRLAELAQRVGKPAYLIDSAADIQESWLSE ARNIGVTAGASAPDVLVQEVISRLKALGGLDVHEISGREENIVF EVPKELRVDVRQIDShewanella MKFDPTSPSLNIMLANPRGFCAGVDRAISIVERALELFSPPIYVR 87 oneidensis ispH, HEWHNRYWQNLKDRGAVFVEELDQVPDNNIVIFSAHGVSQA UniProt Accession VRAEAKARGLRVFDATCPLVTKVHLQVTRASRKGIECILIGHAG No. Q8EBI7 HPEVEGTMGQYDNPNGGVYLIESPADVETLEVRDPNNLCFVT QTTLSVDDTLDIISALLKRFPSIEGPRKDDICYATQNRQDAVRNL SADVDLLIWGSKNSSNSNRLRELALKTGTQSYLVDTADDIDSS WFENITKVAVTAGASAPEVLVQQWQAIAKLAPSWTEVEGRK EDTVFAVPAELRSynechococcus sp. MDTKAFKRSLNHSENYYRQGFGHKDEVAGMMTSEYQSSLIQE 88 ispH, UniProt IRDNNYELTRGDVTIYLAEAFGFCWGVERAVAMAYETRKHFPT Accession No. EKIWVTNEIIHNPSVNNRLKEMNVHFIEWDGDKDFSGVANGD B1XPG7 WILPAFGATVQEMQLLNDRGCTIVDTTCPWVSKVWNSVEKH KKKNYTSIIHGKYKHEETVATSSFAGTYLVLLNLEEAEYVRNYIL HGGDRQTFLDKFKNAYSEGFDPDKDLDRVGVANQTTMLKSET EQMGKLFEQTMLEKFGPTEINDHFMSFNTICDATQERQDAMF DLVEKDLDLMWIGGFNSSNTTHLQEIAVERNIPSYHIDSGDRL LGNNVIAHKPLDGEITTQTHWLPAGKLKIGVTSGASTPDKWED VIAKIFAEKAQPAALVNostocaceae ispH, MDTKTFKRTLQHSENYNRKGFGHQAEVATQLQSEYQSSLIQEI 89 GenBank Accession RDRNYTLQRGDVTIRLAQAFGFCWGVERAVAMAYETRKHFPT No. RUR80075 ERIWITNEIIHNPSVNQRMQEMQVGFIPVEAGNKDFSWGNND WILPAFGASVQEMQLLSEKGCKIVDTTCPWVSKVWNTVEKHK KGDHTSIIHGKYKHEETIATSSFAGKYLIVLNLKEAQYVADYILH GGNREEFLQKFAKACSAGFDPDRDLERVGIANQTTMLKGETE QIGKLFEHTMLQKYGPVELNQHFQSFNTICDATQERQDAMLEL VQENLDLMMGGFNSSNTTQLQQISQERGLPSYHIDWERIKSI NSIEHRQLNGELVTTENWLPAGKIWGVTSGASTPDKWEDVIE KIFALKATAAVFSynechococcus MDTKAFKRALHQSDRYNRKGFGKTTDVSGALESAYQSDLIQS 90 elongatus ispH, LRQNGYRLQRGEITIRLAEAFGFCWGVERAVAIAYETRQHFPQ GenBank Accession ERIWITNEIIHNPSVNQHLREMSVEFIPCERGEKDFSWDRGDV No. Q5N249.1 VILPAFGASVQEMQLLDEKGCHIVDTTCPWVSKVWNTVEKHKRGAHTSIIHGKYNHEETVATSSFAETYLWLNLEQAQYVCDYILNGGDRDEFMTRFGKACSAGFDPDRDLERIGIANQTTMLKSET EAIGKLFERTLLKKYGPQALNDHFLAFNTICDATQERQDAMFQL VEEPLDLIWIGGFNSSNTTHLQEIAIERQIPSFHIDAAERIGPGN RIEHKPLHTDLTTTEPWLPAGPLTIGITSGASTPDKWEDVIERL FDLQRSPseudomonas ispH, MQIKLANPRGFCAGVDRAIEIVNRALEVFGPPIYVRHEWHNKF 91 GenBank Accession WEDLRARGAIFVEELDQVPDDVIVIFSAHGVSQAVRTEAAGR No. QQU66087.1 GLKVFDATCPLVTKVHIEVARYSRDGRECILIGHAGHPEVEGT MGQYDASNGGAIYLVEDEKDVANLQVHNPDRLAFVTQTTLSM DDTSRVIDALRTRFPAIGGPRKDDICYATQNRQDAVKQLADEC DWLWGSPNSSNSNRLRELAERMATPAYLIDGAEDMQRSWF DGVERIGITAGASAPEVLVRGVIQQLQAWGATGADELAGREEN ITFSMPKELRVRSLLRhodobacter MKPALTLFLAAPRGFCAGVDRAVKIVEMSLEKWGAPVYVRHEI 92 capsulatus ispH VHNKFWDTLAAKGAVFVEELDECPVDRPVIFSAHGVPKAVPA UniProt Accession EAERRNMIYVDATCPLVSKVHLEAERHAEEGLQIVMIGHKGHP No. D5ASQ6 EVIGTMGQLPEGEVLLIETAADVATLEVRDPEKLAWITQTTLSV DDTAEWAALQARFPSLVGPGKDDICYATTNRQAAVKALAEKI EALLVIGAPNSSNSKRLVEVGRNAGCAVSHLVERASEIQWEAL EGLTKVGLTAGASAPQVLVDEVIEAFRDRYDLTVELVETAKERI EFKTPRIFRHDNE. coli idi, UniProt MQTEHVILLNAQGVPTGTLEKYAAHTADTRLHLAFSSWLFNAK 93 Accession No. GQLLVTRRALSKKAWPGVWTNSVCGHPQLGESNEDAVIRRC Q46822 RYELGVEITPPESIYPDFRYRATDPSGIVENEVCPVFAARTTSA LQ I N DD EVM DYQ WCD LADVLHG I DATP WAFSPWM VM QATN R EARKRLSAFTQLKR. capsulatus idi1, MSELIPAWVGDRLAPVDKLEVHLKGLRHKAVSVFVMDGENVLI 94 UniProt Accession QRRSEEKYHSPGLWANTCCTHPGWTERPEECAVRRLREELGI No. D5AKF1 TGLYPAHADRLEYRADVGGGMIEHEWDIYLAYAKPHMRITPD PREVAEVRWIGLYDLAAEAGRHPERFSKWLNIYLSSHLDRIFG SILRGR. capsulatus idi2, MAEEMIPAWVEGVLQPVEKLEAHRKGLRHLAISVFVTRGNKVL 95 UniProt Accession LQQRALSKYHTPGLWANTCCTHPYWGEDAPTCAARRLGQEL No. P26173 GIVGLKLRHMGQLEYRADVNNGMIEHEWEVFTAEAPEGIEPQ PDPEEVADTEWVRIDALRSEIHANPERFTPWLKIYIEQHRDMIF PPVTAPseudomonas idi, MNDKVILVDAHDIQIGVCDKRDAHLGTGRLHRAFSVHLIDSAGR 96 GenBank Accession HLIQRRAAGKMLWPGFWSNACCSHPAPGETVPDAASRRLRE No. VVO21040.1 ELGISAPCRSLYSFEYHAHFGSIGAEHELCHVLIAQSDAWSPT TEEVSEVAWLTRAQVSAQLSDPAVPFTPWFRMQWCRLLAQH PTFDTPACQEHybrid PH207-WT TTGATAAACGCTTCAAATTCTCGTATAATGTTACCGATAACA 97 promoter / optimized GTTACCCGTAACATTTTTAATTCTTGTATTGTGGAGGAATAA RBS and spacer / ACATGinitiation codonpT5-GntR-leaky TTTTTTAAAAAATTCATTTGATAAACGCTTCAAATTCTCGTATAATGT 98 promoter TACCGATAACAGTTIntergenic region for T AAT AAT AGGAGAACAAAC T T AT G 99 two-gene operon(FIG. 4)10. CITATION OF REFERENCES
[0288] All publications, patents, patent applications and other documents cited in this application are hereby incorporated by reference in their entireties for all purposes to the same extent as if each individual publication, patent, patent application or other document were individually indicated to be incorporated by reference for all purposes. In the event that there is an inconsistency between the teachings of one or more of the references incorporated herein and the present disclosure, the teachings of the present specification are intended.
Claims
WHAT IS CLAIMED IS:
1. A recombinant microorganism engineered to express:(a) a neryl diphosphate synthase (NPPS); and(b) a terpene synthase selected from:(i) a myrcene synthase (MyrS); or(ii) a nerol synthase (NerS).
2. The recombinant microorganism of claim 1, wherein the NPPS is native to a plant species.
3. The recombinant microorganism of claim 1 or 2, wherein the plant species is Solanum lycopersicum.
4. The recombinant microorganism of any one of claims 1 to 3, wherein the NPPS comprises an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to SEQ ID NO:2.
5. The recombinant microorganism of any one of claims 1 to 4, wherein the recombinant microorganism comprises a first nucleotide sequence encoding the NPPS.
6. The recombinant microorganism of claim 5, wherein the first nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 95% sequence identity to the nucleotide sequence of SEQ ID NO: 12.
7. The recombinant microorganism of any one of claims 1 to 6, wherein the terpene synthase is a myrcene synthase (MyrS).
8. The recombinant microorganism of any one of claims 1 to 7, wherein the recombinant microorganism comprises a second nucleotide sequence encoding the MyrS.
9. The recombinant microorganism of any one of claims 1 to 8, wherein the MyrS is an Antirrhinium majus MyrS, Ocimum basilicum MyrS, Picea abis MyrS, Quercus ilex MyrS, Cannabis sativa MyrS, Antirrhinium majus OciS (AmOciS), or Phaseolus lunatus OciS (PlOciS).
10. The recombinant microorganism of claim 9, wherein the MyrS is an Antirrhinium majus MyrS comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:4.
11. The recombinant microorganism of claim 9 or claim 10, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100%sequence identity to the nucleotide sequence of SEQ ID NO: 14.
12. The recombinant microorganism of claim 9, wherein the MyrS is an Ocimum basilicum MyrS comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:6.
13. The recombinant microorganism of claim 9 or 12, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO: 16.
14. The recombinant microorganism of claim 9 or 14, wherein the MyrS is an Picea abis MyrS comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:7.
15. The recombinant microorganism of claim 9 or 14, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO: 17.
16. The recombinant microorganism of claim 9, wherein the MyrS is a Quercus ilex MyrS comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:9.
17. The recombinant microorganism of claim 9 or 16, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO: 19.
18. The recombinant microorganism of claim 9, wherein the MyrS is a Cannabis sativa MyrS comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:10.
19. The recombinant microorganism of claim 9 or 18, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO:20.
20. The recombinant microorganism of claim 9 or 55, wherein the MyrS is an Antirrhinium majus OciS (AmOciS) comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:
32.
21. The recombinant microorganism of claim 9 or 20, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO:47.
22. The recombinant microorganism of claim 9, wherein the MyrS is a Phaseolus lunatus OciS (PlOciS) comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:36.
23. The recombinant microorganism of claim 9 or 22, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO:51.
24. The recombinant microorganism of any one of claims 1 to 6, wherein the terpene synthase is a nerol synthase (NerS).
25. The recombinant microorganism of claim 24, wherein the recombinant microorganism comprises a second nucleotide sequence encoding the NerS.
26. The recombinant microorganism of claims 24 or 25, wherein the NerS is from Spatholobus suberectus, Cajanus cajan, Cinnamomum tenuipile, Perilla citriadora, Camptotheca acuminata, Streptomyces clavuligerus, Arabidopsis thaliana, Citrus unshiu, or Heliconius melpomene.
27. The recombinant microorganism of claim 25 or 26, wherein the NerS is a Spatholobus suberectus comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:23.
28. The recombinant microorganism of claim 27, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO:38.
29. The recombinant microorganism of claim 25 or 26, wherein the NerS is a Cajanus cajan NerS comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:24.
30. The recombinant microorganism of claim 29, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO:39.
31. The recombinant microorganism of claim 25 or 26, wherein the NerS is a Cinnamomum tenuipile comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:25.
32. The recombinant microorganism of claim 31, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO:40.
33. The recombinant microorganism of claim 25 or 26, wherein the NerS is a Perilla citriadora comprises an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:26.
34. The recombinant microorganism of claim 33, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO:41.
35. The recombinant microorganism of claim 9, wherein the NerS is a Camptotheca acuminata NerS comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:26.
36. The recombinant microorganism of claim 35, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO:41.
37. The recombinant microorganism of claim 25 or 26, wherein the NerS comprises an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:30.
38. The recombinant microorganism of claim 37, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO:45.
39. The recombinant microorganism of claim 25 or 26, wherein the NerS is an Arabidopsis thaliana comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:33.
40. The recombinant microorganism of claim 39, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO:48.
41. The recombinant microorganism of claim 9, wherein the NerS is a Citrus unshiu comprising an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:34.
42. The recombinant microorganism of claim 41, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO:49.
43. The recombinant microorganism of claim 25 or 26, wherein the NerS comprises an amino acid sequence having at least 80%, 90%, 95%, or 100% sequence identity to the amino acid sequence of SEQ ID NO:35.
44. The recombinant microorganism of claim 43, wherein the second nucleotide sequence comprises a nucleotide sequence having at least 80%, 90%, or 100% sequence identity to the nucleotide sequence of SEQ ID NO:50.
45. The recombinant microorganism of any one of claims 5 to 44, wherein the recombinant microorganism comprises a first nucleic acid comprising the first nucleotide sequence.
46. The recombinant microorganism of claim 45, wherein the first nucleic acid is genomically integrated or extrachromosomal.
47. The recombinant microorganism of claim 45 or 46, wherein the first nucleic acid comprises a first promoter that is operably linked to the first nucleotide sequence.
48. The recombinant microorganism of claim 47, wherein the first promoter is an inducible promoter or a constitutive promoter.
49. The recombinant microorganism of any one of claim 47 or 48, wherein the first nucleic acid further comprises the second nucleotide sequence.
50. The recombinant microorganism of claim 49, wherein the first nucleic acid further comprises a second promoter that is operably linked to the second nucleotide sequence.
51. The recombinant microorganism of claim 50, wherein the second promoter is an inducible promoter or a constitutive promoter.
52. The recombinant microorganism of any one of claims 8 to 51, wherein the recombinant microorganism comprises a second nucleic acid comprising the second nucleotide sequence.
53. The recombinant microorganism of claim 52, wherein the second nucleic acid is genomically integrated or extrachromosomal.
54. The recombinant microorganism of claim 52 or 53, wherein the second nucleic acid further comprises a second promoter that is operably linked to the second nucleotide sequence.
55. The recombinant microorganism of claim 54, wherein the second promoter is an inducible promoter or a constitutive promoter.
56. The recombinant microorganism of any one of claims 1 to 55, wherein the recombinant microorganism is an E. coli.
57. The recombinant microorganism of claim 56, wherein the E. coli is an E. coli MG 1655, E. coli K12 substrain W3110, E. coli K12 substrain DH5alpha, E. coli W, E. coli BL21.
58. The recombinant microorganism of any one of claims 1 to 57, wherein the recombinant microorganism comprises modifications that increase conversion of glucose or xylose to dimethylallyl pyrophosphate (DMAPP) and isopentenyl pyrophosphate (IPP) via the1-deoxyxylulose-5-phosphate (DXP) pathway relative to a native microorganism or parental microorganism that does not comprise the modifications.
59. The recombinant microorganism of claim 58, wherein the recombinant microorganism is an E. coli cell comprising a third nucleotide sequence encoding a dxs2 polypeptide comprising the amino acid sequence of SEQ ID NO:55 and a fourth nucleotide sequence encoding a dxr polypeptide (2) comprising the amino acid sequence of SEQ ID NO:58.
60. The recombinant microorganism of claim 59, which further comprises:(g) a fifth nucleotide sequence encoding an ispD polypeptide that is not native to E. coli', (h) a sixth nucleotide sequence encoding an ispE polypeptide that is not native to E. coli; (i) a seventh nucleotide sequence encoding an ispF polypeptide that is not native to E.coli;(j) an eighth nucleotide sequence encoding an ispG polypeptide that is not native to E.coli;(k) a ninth nucleotide sequence encoding an ispH polypeptide that is not native to E. coli; and(l) a tenth nucleotide sequence encoding an idi polypeptide that is not native to E. coli.
61. The recombinant microorganism of claim 60, wherein the ispD polypeptide comprises an amino acid sequence having at least 95% identity to SEQ ID NO:62, the ispE polypeptide comprises an amino acid sequence having at least 95% identity to SEQ ID NO:65, the ispF polypeptide comprises an amino acid sequence having at least 95% identity to SEQ ID NO:68, the ispG polypeptide comprises an amino acid sequence having at least 95% identity to SEQ ID NO:76, the ispD polypeptide comprises an amino acid sequence having at least 95% identity to SEQ ID NO:92.
62. The recombinant microorganism of any one of claims 1 to 61, wherein the recombinant microorganism comprises modifications that increase the conversion of glucose or xylose to dimethylallyl pyrophosphate (DMAPP) and isopentenyl pyrophosphate (IPP) via the mevalonate pathway relative to a native microorganism or parental microorganism that does not comprise the modifications.
63. The recombinant microorganism of any one of claims 1 to 55 or 62, wherein the recombinant microorganism is S. cerevisiae.
64. A recombinant microorganism configured to:(a) condense dimethylallyl diphosphate (DMAPP) and isopentenyl diphosphate (IPP) to form neryl diphosphate (NPP); and(b) convert NPP to β-myrcene.
65. A recombinant microorganism comprising:(a) a means for condensing DMAPP and IPP to form NPP; and(b) a means for converting NPP to -myrcene.
66. A recombinant microorganism comprising one or more heterologous nucleic acids comprising:(a) a heterologous nucleotide sequence encoding a neryl diphosphate synthase (NPPS); and(b) a heterologous nucleotide sequence encoding a terpene synthase.
67. A method of producing p-myrcene, the method comprising culturing the recombinant microorganism of any one of claims 1 to 23 or 45 to 66 in a production medium, wherein the production medium comprises glucose and / or xylose.
68. The method of claim 67, further comprising recovering the p-myrcene from the production medium, optionally wherein the p-myrcene is recovered from the head space above the production medium.
69. A recombinant microorganism configured to:(a) condense dimethylallyl diphosphate (DMAPP) and isopentenyl diphosphate (IPP) to form neryl diphosphate (NPP); and(b) convert NPP to nerol.
70. A recombinant microorganism comprising one or more heterologous nucleic acids comprising:(a) a heterologous nucleotide sequence encoding a neryl diphosphate synthase (NPPS); and(b) a heterologous nucleotide sequence encoding a nerol synthase.
71. A method of producing nerol, the method comprising culturing the recombinant microorganism of any one of claims 24 to 63 or 69 to 70 in a production medium.
72. The method of claim 71, wherein the production medium comprises glucose and / or xylose.
73. The method of claim 71 or 72, further comprising recovering the nerol from the production medium.
74. The method of any one of claims 71 to 73, further comprising recovering the nerol from the head space above the production medium.