Bidirectional multienzyme scaffolds for biosynthesis of cannabinoids
By using a bidirectional multi-enzyme scaffold and a peptide scaffold to co-localize catalytic enzymes in recombinant host cells, we optimized cannabinoid biosynthesis, overcoming the scale and environmental limitations of traditional cannabinoid production and achieving efficient industrial-scale production.
Patent Information
- Application Number
- CN202510124649.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-19
- Filing Date
- 2019-11-25
- Publication Date
- 2025-09-09
AI Technical Summary
Traditional cannabinoid production relies on large-scale cultivation of cannabis, which is limited by environmental factors and scale, making it difficult to achieve efficient production on an industrial scale.
A bidirectional multi-enzyme scaffold was used for cannabinoid production in recombinant host cells to optimize the cannabinoid biosynthesis process by controlling the localization and stoichiometry of enzymes, including the use of multiple engineered enzymes and peptide scaffolds to co-localize catalytic enzymes to improve production.
The yield of cannabinoids and their precursors was significantly improved, high-throughput production in recombinant host cells was achieved, and the scale and environmental limitations of traditional production methods were resolved.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Application Serial No. 62 / 836,265, filed April 19, 2019, and U.S. Application Serial No. 62 / 771,839, filed November 27, 2018. The disclosure of the prior application is considered part of the disclosure of the present application and is incorporated herein in its entirety. Technical Field
[0003] This document relates to methods and materials for the biosynthesis of cannabinoids, and more particularly, to the biosynthesis of cannabinoids using a bidirectional multi-enzyme scaffold. Background Art
[0004] The emerging therapeutic potential of cannabinoids necessitates industrial-scale production to meet future comprehensive demand. Traditional cannabinoid production efforts rely on large-scale cultivation of cannabis (Cannabis sativa L). However, agricultural cannabinoid production is problematic due to uncontrollable environmental factors and scale limitations. Summary of the Invention
[0005] This disclosure is based, at least in part, on the discovery that a bidirectional multi-enzyme scaffold can be engineered to allow high-throughput cannabinoid production in recombinant host cells. By controlling the positioning, spatial orientation, and stoichiometry of enzymes that catalyze the biosynthesis of cannabinoids and cannabinoid precursors, the multi-enzyme scaffold described herein allows for throughput-optimized cannabinoid biosynthesis in genetically engineered host cells.
[0006] In one aspect, the present invention features a host cell capable of producing one or more cannabinoids selected from the group consisting of cannabigerolic acid, cannabidiolic acid, and cannabichromenic acid. The host cell comprises at least three different exogenous nucleic acids, wherein the first and second exogenous nucleic acids each encode a plurality of engineered enzymes selected from the group consisting of acetyl-CoA acetyltransferase, 3-hydroxybutyryl-CoA dehydrogenase, enoyl-CoA hydratase, β-ketothiolase, trans-enoyl-CoA reductase, HMG-CoA synthetase, HMG-CoA reductase, mevalonate kinase, phosphomevalonate kinase, diphosphomevalonate decarboxylase, isopentenyl diphosphate delta isomerase, geranyl diphosphate synthase, olivetol synthase, olivetate cyclase, and CBGA synthase; wherein each engineered enzyme comprises a heterologous interaction domain, wherein the heterologous interaction domain comprises a first and a second peptide motif, and wherein each heterologous interaction domain is different from each other; and wherein the third exogenous nucleic acid encodes a polypeptide scaffold comprising a plurality of peptide ligands, wherein each peptide ligand comprises an amino acid sequence capable of binding to the first or second peptide motif of one of the heterologous interaction domains. The plurality of engineered enzymes may further comprise ATP citrate lyase and acetyl-CoA carboxylase. The host cell may also include exogenous nucleic acids encoding cannabidiolic acid synthase (CBDAS) and cannabichromenic acid synthase (CBCAS). The host cell may include exogenous CBDAS. The host cell may include exogenous CBCAS. The host cell may include exogenous CBDAS and exogenous CBCAS. The host cell may include exogenous hexanoyl-CoA synthetase. The host cell may include at least four different exogenous nucleic acids, wherein the first, second, and fourth nucleic acids each encode a plurality of engineered enzymes. The host cell may include at least five different exogenous nucleic acids, wherein the first, second, fourth, and fifth nucleic acids each encode a plurality of engineered enzymes. The host cell may include at least six different exogenous nucleic acids, wherein the first, second, fourth, fifth, and sixth nucleic acids each encode a plurality of engineered enzymes. Each exogenous nucleic acid may include a constitutive promoter operably linked to a sequence encoding an engineered enzyme or a polypeptide scaffold, or an inducible promoter operably linked to a sequence encoding an engineered enzyme or a polypeptide scaffold. In some embodiments, the promoter is the GAL1-10 promoter. In some embodiments, the constitutive promoter used to express the polypeptide scaffold has a weaker constitutive activity level than the constitutive promoter used to express the engineered enzyme. In some embodiments, a constitutive promoter is used to express the engineered enzyme, while an inducible promoter is used to express the polypeptide scaffold. In some embodiments, an inducible promoter is used to express the engineered enzyme, while a constitutive promoter is used to express the polypeptide scaffold.
[0007] Any host cell can be a bacterium, yeast, algae or plant cell. The bacterial cell can be selected from the group consisting of Escherichia coli, Bacillus, Brevibacterium, Streptomyces and Pseudomonas cells. The yeast cell can be selected from the group consisting of Pichia pastoris, Saccharomyces cerevisiae, Yarrowia lipolytica, Kluyveromyces marxianus, and Komagataella phaffii cells. The algae cell can be Dunaliella sp., Chlorella variabilis, Euglena mutabilis, or Chlamydomonas reinhardtii cells. The plant cell may be a Cannabis or tobacco cell.
[0008] In some embodiments, each engineered enzyme has the formula: enzyme-linker 1- Spacer-Linker 2- motif 1- connector 3- Motif 2, wherein Linkers 1, 2, and 3 may be the same or different, Motif 1 and Motif 2 may be the same or different, and Motif 1 and Motif 2 form a heterologous interaction domain. The scaffold polypeptide may have the formula: N-terminus – [Ligand 1 – Linker – Ligand 2 – Spacer]n – (optionally labeled) C-terminus, where n is the number of heterologous interaction domains, and wherein Ligand 1 and Ligand 2 bind to Motif 1 and Motif 2, respectively, of the heterologous interaction domain. The scaffold polypeptide may be tagged with a MYC tag, a FLAG tag, or an HA tag. The host cell may further include a nucleic acid encoding a second polypeptide scaffold comprising a plurality of peptide ligands, wherein each peptide ligand comprises an amino acid sequence capable of binding to a different motif of the heterologous interaction domain. The linker may comprise a flexible GS-rich sequence flanked by rigid α-helical portions. The spacer may be a cTPR6 spacer.
[0009] The present invention also features a method for producing one or more cannabinoids selected from the group consisting of cannabigerolic acid, cannabidiolic acid, and cannabichromenic acid. The method can include culturing any host cell described herein under conditions such that the host cell produces the one or more cannabinoids. The host cell can be cultured in a medium supplemented with citrate, glucose, hexanoic acid, and / or other carbon sources and / or in a medium supplemented with malonyl-CoA. The method can also include extracting the one or more cannabinoids from the host cell.
[0010] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention pertains. Although methods and materials similar to or equivalent to those described herein may be employed in the practice or testing of the present invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references cited herein are incorporated herein by reference in their entirety. In the event of a conflict, the present specification, including definitions, will prevail. In addition, the materials, methods, and examples are illustrative only and are not intended to be limiting.
[0011] Other features and advantages of the present invention will become readily apparent from the following detailed description and appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1A Schematic diagram of a representative embodiment of a multi-enzyme cannabinoidergic scaffold in a cell. The multi-enzyme scaffold includes enzymes from the hexanoyl-CoA pathway, enzymes from the upper cannabinoid pathway, and enzymes from the mevalonate pathway. The schematic also depicts a second scaffold according to an embodiment comprising enzymes from the malonyl-CoA pathway, and depicts a non-scaffolded cannabidiolic acid synthase (CBDAS) and a non-scaffolded cannabichromenic acid synthase (CBCAS). ID refers to the enzyme-linked interaction domain; cTPR6 refers to the spacer sequence; and scaffolded ligand refers to a tandem peptide ligand that forms a scaffold binding site specific for each enzyme-linked ID. The target products, cannabigerolic acid (CBGA), cannabigerol (CBG), cannabidiolic acid (CBDA), cannabichromene (CBD), cannabichromenic acid (CBCA), and cannabichromene (CBC), are highlighted by box marks. CBG can be produced by decarboxylation of CBGA, CBD can be produced by decarboxylation of CBDA, and CBC can be produced by decarboxylation of CBCA. For each decarboxylation, the "Δ" symbol represents heat and the "hν" symbol represents light.
[0013] Figure 1BSchematic diagram of a representative embodiment of a bidirectional multi-enzyme scaffold in a cell (e.g., a yeast cell). The multi-enzyme scaffold (referred to as the intranuclear SCF gene cassette) includes enzymes of the hexanoyl-CoA pathway (referred to as the intranuclear HCA cassette), enzymes of the upper cannabinoid pathway (referred to as the intranuclear CAN cassette), and enzymes of the mevalonate pathway (referred to as the intranuclear GPP cassette). The schematic also depicts a second scaffold according to an embodiment comprising enzymes of the malonyl-CoA pathway, and depicts unscaffolded CBDAS and unscaffolded CBCAS. ID refers to the enzyme-linked interaction domain; cTPR6 refers to the spacer sequence; and scaffolded ligand refers to a tandem peptide ligand that forms a scaffold binding site specific for each enzyme-linked ID. The target products CBGA, CBG, CBDA, CBD, CBCA, and CBC are highlighted with box marks. CBG can be produced by decarboxylation of CBGA, CBD can be produced by decarboxylation of CBDA, and CBC can be produced by decarboxylation of CBCA. For each decarboxylation, the "Δ" symbol represents heat and the "hν" symbol represents light.
[0014] Figure 2A is a schematic diagram of a gene cassette according to one embodiment for engineering cannabinoidergic cells.
[0015] Figure 2B Schematic representation of the gene cassettes used in Examples 2-4 for biosynthesis of cannabinoids in yeast.
[0016] Figure 3 is an example of an enzyme-scaffold complex.
[0017] Figure 4 Schematic diagram of a representative embodiment of a multi-enzyme cannabinoid scaffold in a cell. The multi-enzyme scaffold includes enzymes from the hexanoyl-CoA pathway, the upper cannabinoid pathway, and the mevalonate pathway. The schematic also depicts a second scaffold according to an embodiment comprising enzymes from the malonyl-CoA pathway, and depicts unscaffolded CBDAS and unscaffolded CBCAS. Pyruvate dehydrogenase (E1) and dihydrolipoyl transacetylase (E2) replace ATP citrate lyase in the two depicted scaffolds. ID refers to the enzyme-linked interaction domain; cTPR6 refers to the spacer sequence; and scaffolded ligand refers to the tandem peptide ligands that form the scaffold binding site specific for each enzyme-linked ID. The target products CBGA, CBG, CBDA, CBD, CBCA, and CBC are highlighted with box marks. CBG can be produced by decarboxylation of CBGA, CBD can be produced by decarboxylation of CBDA, and CBC can be produced by decarboxylation of CBCA. For each decarboxylation, the "Δ" symbol represents heat and the "hν" symbol represents light.
[0018] Figure 5is a schematic diagram of a representative embodiment of a multi-enzyme cannabinoid scaffold in a cell. The multi-enzyme scaffold includes enzymes from the hexanoyl-CoA pathway, enzymes from the upper cannabinoid pathway, and enzymes from the MEP (2-C-methylerythritol 4-phosphate) pathway. The schematic also depicts a second scaffold according to an embodiment comprising enzymes from the malonyl-CoA pathway, and depicts unscaffolded CBDAS and unscaffolded CBCAS. ID refers to the enzyme-linked interaction domain; cTPR6 refers to the spacer sequence; and scaffolded ligand refers to a tandem peptide ligand that forms a scaffold binding site specific for each enzyme-linked ID. The target products CBGA, CBG, CBDA, CBD, CBCA, and CBC are highlighted with box marks. CBG can be produced by decarboxylation of CBGA, CBD can be produced by decarboxylation of CBDA, and CBC can be produced by decarboxylation of CBCA. For each decarboxylation, the "Δ" symbol represents heat and the "hν" symbol represents light.
[0019] Figure 6A The invention also includes the amino acid sequence of each of the following enzymes: ATP citrate lyase (SEQ ID NO: 83), acetyl-CoA acetyltransferase (atoB) (SEQ ID NO: 84), 3-hydroxybutyryl-CoA dehydrogenase (SEQ ID NO: 85), enoyl-CoA hydratase (SEQ ID NO: 86), trans-enoyl-CoA reductase (SEQ ID NO: 88), beta-ketothiolase (bktB) (SEQ ID NO: 87), HMG-CoA synthase (SEQ ID NO: 90), truncated HMG-CoA reductase (SEQ ID NO: 91), mevalonate kinase (SEQ ID NO: 92), phosphomevalonate kinase (SEQ ID NO: 93), diphosphomevalonate decarboxylase (SEQ ID NO: 94), isopentenyl-diphosphate delta isomerase (SEQ ID NO: 95), mutant geranyl-diphosphate synthase (ERG20) (SEQ ID NO: 96). WW )(SEQ ID NO:96), olivetol synthase (SEQ ID NO:98), olivetate cyclase (SEQ ID NO:99), CBGA synthase (SEQ ID NO:100), acetyl-CoA carboxylase (SEQ ID NO:100), acetyl-CoA carboxylase (SEQ ID NO:97), CBDA synthase (SEQ ID NO:101), CBCA synthase (SEQ ID NO:102), and hexanoyl-CoA synthetase (SEQ ID NO:89).
[0020] Figure 6BThe amino acid sequence of an engineered enzyme comprising the following formula: enzyme-enzyme linker-cTPR6 spacer-ID linker-ID motif #1-ID motif linker-ID motif #2, wherein the linkers (enzyme linker, ID linker, and ID motif linker) can be the same or different, and ID motif #1 and ID motif #2 can be the same or different. Provided are the amino acid sequences of the following engineered enzymes: ATP citrate lyase (ID1) (SEQ ID NO: 103), acetyl-CoA acetyltransferase (atoB) (ID2) (SEQ ID NO: 104), 3-hydroxybutyryl-CoA dehydrogenase (ID3) (SEQ ID NO: 105), enoyl-CoA hydratase (ID4) (SEQ ID NO: 106), trans-enoyl-CoA reductase (ID5) (SEQ ID NO: 107), beta-ketothiolase (bktB) (ID6) (SEQ ID NO: 108), HMG-CoA synthase (ID7) (SEQ ID NO: 109), truncated HMG-CoA reductase (ID8) (SEQ ID NO: 110), mevalonate kinase (ID9) (SEQ ID NO: 111), phosphomevalonate kinase (ID10) (SEQ ID NO: 112), diphosphomevalonate decarboxylase (ID11) (SEQ ID NO: 113). NO: 113), isopentenyl diphosphate delta isomerase (ID12) (SEQ ID NO: 114), mutant geranyl diphosphate synthase (ERG20 WW )(ID13)(SEQ ID NO:115), olivetol synthase (ID14)(SEQ ID NO:116), olivetate cyclase (ID15)(SEQ ID NO:117), CBGA synthase (ID16)(SEQ ID NO:118) and acetyl-CoA carboxylase (ID17)(SEQ ID NO:211).
[0021] Figure 6C An amino acid sequence comprising a polypeptide scaffold having the formula: N-terminus - [ligand #1 - ID motif #1 ligand - linker - ID motif #2 ligand - scaffolded ID-binding site spacer]n - (Myc)3-tagged C-terminus, wherein n is 16 and the ID motif ligands correspond to the motifs of IDs 1-16, as shown in Table 2. See SEQ ID NO: 119.
[0022] Figure 6DAn amino acid sequence comprising a polypeptide scaffold having the formula: N-terminus - [ligand #1 - ID motif #1 ligand - linker - ID motif #2 ligand - scaffolded ID-binding site spacer]n - (FLAG)3-tagged C-terminus, wherein n is 2 and the ID motif ligands correspond to the motifs of IDs 1 and 17, as shown in Table 2. See SEQ ID NO: 120.
[0023] Figure 7 Schematic diagram of a representative embodiment of a scaffold with minimal requirements for cannabinoid acid synthesis. The scaffold contains enzymes of the upper cannabinoid pathway. In this embodiment, non-scaffolded hexanoyl-CoA synthetase (HCS), non-scaffolded CBDAS, and non-scaffolded CBCAS are also used. ID refers to the enzyme-linked interaction domain; cTPR6 refers to the spacer sequence; and scaffolded ligand refers to a tandem peptide ligand that forms a scaffold binding site specific for each enzyme-linked ID. The target products CBGA, CBG, CBDA, CBD, CBCA, and CBC are highlighted with box marks. CBG can be produced by decarboxylation of CBGA, CBD can be produced by decarboxylation of CBDA, and CBC can be produced by decarboxylation of CBCA. For each decarboxylation, the "Δ" symbol represents heat and the "hν" symbol represents light.
[0024] Figure 8 is a schematic diagram of a representative embodiment of a bidirectional scaffold containing an HCS at the N-terminus of the scaffold, a geranyl pyrophosphate synthase (GPPS) at the C-terminus of the scaffold, and an enzyme of the upper cannabinoid pathway between the HCS and GPPS. In this embodiment, non-scaffolded CBDAS and non-scaffolded CBCAS can also be used. ID refers to the enzyme-linked interaction domain; cTPR6 refers to the spacer sequence; and the scaffolded ligand refers to a tandem peptide ligand that forms a scaffold binding site specific for each enzyme-linked ID. The target products CBGA, CBG, CBDA, CBD, CBCA, and CBC are highlighted with boxes. CBG can be produced by decarboxylation of CBGA, CBD can be produced by decarboxylation of CBDA, and CBC can be produced by decarboxylation of CBCA. For each decarboxylation, the "Δ" symbol represents heat and the "hν" symbol represents light.
[0025] Figure 9Schematic diagram of a representative embodiment of a one-way scaffold containing enzymes of the upper cannabinoid pathway, showing soluble enzymes from precursor pathways (hexanoyl-CoA pathway, mevalonate pathway, and malonyl-CoA pathway), and soluble CBDAS and CBCAS. ID refers to the enzyme-linked interaction domain; cTPR6 refers to the spacer sequence; and scaffolded ligand refers to a tandem peptide ligand that forms a scaffold binding site specific for each enzyme-linked ID. The target products CBGA, CBG, CBDA, CBD, CBCA, and CBC are highlighted with box marks. CBG can be produced by decarboxylation of CBGA, CBD can be produced by decarboxylation of CBDA, and CBC can be produced by decarboxylation of CBCA. For each decarboxylation, the "Δ" symbol represents heat and the "hν" symbol represents light.
[0026] Figure 10 is a schematic diagram of a representative embodiment of a multi-enzyme cannabinoid scaffold in a cell. The multi-enzyme scaffold includes enzymes of the malonyl-CoA (CMA) pathway, enzymes of the upper cannabinoid pathway, and enzymes of the mevalonate pathway. The schematic also depicts separate scaffolds according to an embodiment comprising enzymes of the hexanoyl-CoA pathway, and depicts non-scaffolded CBDAS and non-scaffolded CBCAS. ID refers to the enzyme-linked interaction domain; cTPR6 refers to the spacer sequence; and scaffolded ligand refers to a tandem peptide ligand that forms a scaffold binding site specific for each enzyme-linked ID. The target products CBGA, CBG, CBDA, CBD, CBCA, and CBC are highlighted with box marks. CBG can be produced by decarboxylation of CBGA, CBD can be produced by decarboxylation of CBDA, and CBC can be produced by decarboxylation of CBCA. For each decarboxylation, the "Δ" symbol represents heat and the "hν" symbol represents light.
[0027] Figure 11 is a schematic diagram of one representative embodiment of a multiple enzymatic cannabinoidergic scaffold in the dual compartments of the cell, cytoplasm and mitochondria / plastids.
[0028] Figure 12AThe invention also comprises a nucleotide sequence encoding ATP citrate lyase (SEQ ID NO: 121), acetyl-CoA acetyltransferase (atoB) (SEQ ID NO: 122), 3-hydroxybutyryl-CoA dehydrogenase (SEQ ID NO: 123), enoyl-CoA hydratase (SEQ ID NO: 124), trans-enoyl-CoA reductase (SEQ ID NO: 125), beta-ketothiolase (bktB) (SEQ ID NO: 126), HMG-CoA synthase (SEQ ID NO: 127), truncated HMG-CoA reductase (SEQ ID NO: 128), mevalonate kinase (SEQ ID NO: 129), phosphomevalonate kinase (SEQ ID NO: 130), diphosphomevalonate decarboxylase (SEQ ID NO: 131), isopentenyl-diphosphate delta isomerase (SEQ ID NO: 132), geranyl-diphosphate synthase (ERG20) (SEQ ID NO: 133). WW )(SEQ ID NO: 133), olivetol synthase (SEQ ID NO: 134), olivetate cyclase (SEQ ID NO: 135), CBGA synthase (SEQ ID NO: 136), acetyl-CoA carboxylase (SEQ ID NO: 137), CBDA synthase (SEQ ID NO: 138), CBCA synthase (SEQ ID NO: 139), and hexanoyl-CoA synthetase (SEQ ID NO: 140).
[0029] Figure 12Bcomprising a nucleotide sequence encoding an engineered enzyme having the formula: enzyme-enzyme linker-cTPR6 spacer-ID linker-ID motif #1-ID motif linker-ID motif #2, wherein the enzyme linker, ID linker, and ID motif linker can be the same or different, and wherein ID motif #1 and ID motif #2 can be the same or different. Provided are nucleotide sequences encoding the following engineered enzymes: ATP citrate lyase (ID1) (SEQ ID NO: 141), acetyl-CoA acetyltransferase (atoB) (ID2) (SEQ ID NO: 142), 3-hydroxybutyryl-CoA dehydrogenase (ID3) (SEQ ID NO: 143), enoyl-CoA hydratase (ID4) (SEQ ID NO: 144), trans-enoyl-CoA reductase (ID5) (SEQ ID NO: 145), bktB (ID6) (SEQ ID NO: 146), HMG-CoA synthase (ID7) (SEQ ID NO: 147), truncated HMG-CoA reductase (ID8) (SEQ ID NO: 148), mevalonate kinase (ID9) (SEQ ID NO: 149), phosphomevalonate kinase (ID10) (SEQ ID NO: 150), diphosphomevalonate decarboxylase (ID11) (SEQ ID NO: 151), isopentenyl diphosphate delta isomerase (ID12) (SEQ ID NO: 152). NO:152), mutant geranyl-diphosphate synthase (ERG20 WW )(ID13)(SEQ ID NO:153), olivetol synthase (ID14)(SEQ ID NO:154), olivetate cyclase (ID15)(SEQ ID NO:155), CBGA synthase (ID16)(SEQ ID NO:156) and acetyl-CoA carboxylase (ID17)(SEQ ID NO:157).
[0030] Figure 12C comprising a nucleotide sequence encoding a scaffold polypeptide (SEQ ID NO: 158) comprising a peptide ligand corresponding to IDs 1-16 shown in Table 2 and a three-repeat myc tag at the C-terminus.
[0031] Figure 12D The invention comprises a nucleic acid sequence encoding a scaffold polypeptide (SEQ ID NO: 159) comprising peptide ligands corresponding to ID 1 and 17, and a three-repeat FLAG tag at the C-terminus.
[0032] Figure 13AThe amino acid sequence comprising the scaffold-bound engineered enzyme and soluble hexanoyl-CoA synthetase (HCS) (SEQ ID NO: 209) (encoded by the HCA gene cassette). The scaffold-binding engineered enzymes are ATP citrate lyase (ACL) (ACL-enzyme linker-cTPR6 spacer-ID linker-ID1) (SEQ ID NO: 160); acetyl-CoA acetyltransferase (atoB) (atoB-enzyme linker-cTPR6 spacer-ID linker-ID2) (SEQ ID NO: 161); 3-hydroxybutyryl-CoA dehydrogenase (BHBD) (BHBD-enzyme linker-cTPR6 spacer-ID linker-ID3) (SEQ ID NO: 162); enoyl-CoA hydratase (ECH) (ECH-enzyme linker-cTPR6 spacer-ID linker-ID4) (SEQ ID NO: 163); trans-enoyl-CoA reductase (ECR) (ECR-enzyme linker-cTPR6 spacer-ID linker-ID5) (SEQ ID NO: 164); and β-ketothiolase (bktB) (bktB-enzyme linker-cTPR6 spacer-ID linker-ID6) (SEQ ID NO: 165). NO:165).
[0033] Figure 13B The invention also comprises amino acid sequences of scaffold-binding engineered enzymes encoded by the GPP gene cassette. The scaffold-binding engineered enzymes are HMG-CoA synthase (HMGS) (HMGS-enzyme linker-cTPR6 spacer-ID linker-ID7) (SEQ ID NO: 166); truncated HMG-CoA reductase (tHMGR) (tHMGR-enzyme linker-cTPR6 spacer-ID linker-ID8) (SEQ ID NO: 167); mevalonate kinase (ERG12) (ERG12-enzyme linker-cTPR6 spacer-ID linker-ID9) (SEQ ID NO: 168); phosphomevalonate kinase (ERG8) (ERG8-enzyme linker-cTPR6 spacer-ID linker-ID10) (SEQ ID NO: 169); diphosphomevalonate decarboxylase (MVD1) (MVD1-enzyme linker-cTPR6 spacer-ID linker-ID11) (SEQ ID NO: 170). ID NO: 170); Isopentenyl diphosphate delta-isomerase (IDI1) (IDI1-enzyme linker-cTPR6 spacer-ID linker-ID12) (SEQ ID NO: 171); and Geranyl-diphosphate synthase (ERG20WW) (ERG20WW-enzyme linker-cTPR6 spacer-ID linker-ID13) (SEQ ID NO: 172).
[0034] Figure 13CThe amino acid sequences of the scaffold-bound engineered enzymes, soluble CBDA synthase (SEQ ID NO: 173) and soluble CBCA synthase (SEQ ID NO: 174), encoded by the CAN gene cassette, are included. The scaffold-bound engineered enzymes are olivetol synthase (OS) (OS-enzyme linker-cTPR6 spacer-ID linker-ID14) (SEQ ID NO: 175); olivetate cyclase (OAC) (OAC-enzyme linker-cTPR6 spacer-ID linker-ID15) (SEQ ID NO: 176); CBGA synthase (CBGAS-enzyme linker-cTPR6 spacer-ID linker-ID16) (SEQ ID NO: 177); and acetyl-CoA carboxylase (ACC) (ACC-enzyme linker-cTPR6 spacer-ID linker-ID17) (SEQ ID NO: 178).
[0035] Figure 13D Contains the amino acid sequences of the cannabinoid metabolic scaffold (CBSCFLD)-(Myc)3 (SEQ ID NO: 179) and the malonyl-CoA metabolic scaffold (MCASCFLD)-(FLAG)3 (SEQ ID NO: 180).
[0036] Figure 14A Contains code Figure 13A The codon-optimized nucleotide sequences of the enzymes are shown in FIG.
[0037] Figure 14B Contains code Figure 13B The codon-optimized nucleotide sequences of the enzymes are shown in FIG.
[0038] Figure 14C Contains code Figure 13C The codon-optimized nucleotide sequences of the enzymes are shown in FIG.
[0039] Figure 14D Contains code Figure 13D The codon-optimized nucleotide sequences of the scaffolds of SEQ ID NO: 201 and SEQ ID NO: 202.
[0040] Figure 15A The nucleotide sequence comprising the HCA gene cassette (SEQ ID NO: 203).
[0041] Figure 15B The nucleotide sequence comprising the GPP gene cassette (SEQ ID NO: 204).
[0042] Figure 15CThe nucleotide sequence comprising the CAN gene cassette (SEQ ID NO: 205).
[0043] Figure 15D The nucleotide sequence comprising the SCF gene cassette (SEQ ID NO: 206).
[0044] Figure 15E The nucleotide sequence comprising the SOL gene cassette (SEQ ID NO: 207).
[0045] Figure 16 is a map of the pCCI-Brick plasmid construct.
[0046] Figure 17 Figure 2 is a map of the pESC-TRP ("vHCA") vector construct. In this figure, the vector contains the TRP gene, which allows selection in tryptophan-deficient medium. Similar vectors were also prepared in which the TRP gene was replaced with the LEU gene, which allows selection in leucine-deficient medium, the HIS3 gene, which allows selection in histidine-deficient medium, or the URA3 gene, which allows selection in uracil-deficient medium.
[0047] Figure 18 The line graph depicting the cell proliferation curves is a graph of cell density measurements (OD) recorded at 12-hour intervals for yCBSCF and yCBSOL cultures over a 48-hour incubation period. 600nm ). The initial cell density of all cultures was normalized to OD 600nm = 0.3. For all measurements, n = 3 biological replicates of yCBSCF and yCBSOL cultures. Floating data points represent means with 95% confidence intervals. Dashed lines represent 95% confidence intervals of regression curve fits.
[0048] Figure 19 shows a comparison of cannabinoid and precursor titers from scaffolded and soluble cannabinoid biosynthesis. Representative mass spectra of target analytes isolated from (A) yCBSOL and (B) yCBSCF cultures incubated for 48 hours in basal medium. Bar graphs depict (C) total (polymeric) cannabinoid (CBGA+CBDA+CBCA+CBG+CBD+CBC) titers, (D) cannabinoid precursor (OVA) titers and total parent and decarboxylated derivative (CBGA+CBG, CBDA+CBD, and CBCA+CBC) cannabinoid titers, and (E) isolated parent (COO(H)) cannabinoid (CBGA, CBDA, and CBCA) and decarboxylated derivative (ΔCOOH) cannabinoid (CBG, CBD, and CBC) titers for 48-hour yCBSOL (left) and yCBSCF (right) cultures grown in basal medium. For all measurements, n = 3 biological replicates for yCBSCF and yCBSOL cultures. CB, cannabinoid; cannabigerolic acid, CBGA; cannabigerol, CBG; cannabidiolic acid, CBDA; cannabidiol, CBD; cannabichromenic acid, CBCA; cannabichromene, CBC; olivetolic acid, OVA. Floating asterisks indicate statistically significant differences between strains of yCBSCF versus yCBSOL cultures (determined by Bonferroni's multiple comparison post hoc test; α = 0.05). Bar graphs depict mean values with 95% confidence intervals. *p < 0.05; **p < 0.01; ***p < 0.001; ****p < 0.0001.
[0049] Figure 20 Figure 2 is a bar graph showing the effect of citrate and hexanoate supplementation on scaffolding and soluble cannabinoid biosynthesis. Shown are the total cannabinoid (CBGA+CBDA+CBCA+CBG+CBD+CBC) titers of yCBSOL and yCBSCF cultures incubated for 48 hours in basal, hexanoate (300 mg / L) supplemented, and buffered (pH 6.0) citrate (300 mg / L) supplemented media. Floating asterisks indicate statistically significant inter-strain differences between yCBSCF and yCBSOL cultures (determined by Bonferroni multiple comparison post hoc test; α = 0.05). The asterisked line indicates that the intra-strain difference between the total cannabinoid titer of the basal medium of yCBSCF cultures and the total cannabinoid titer of the medium supplemented with citrate is statistically significant (determined by Bonferroni multiple comparison post hoc test; α = 0.05). The bar graph depicts the mean with 95% confidence intervals. *p<0.05; **p<0.01; ***p<0.001; ****p<0.0001.
[0050] Figure 21Scaffolding and concentration-response parameterization of soluble cannabinoid biosynthesis from citrate are shown. Figure 21 In A, a line graph is shown depicting eight-point concentration ([citrate])-response (total cannabinoid titer) curves fitted by asymmetric sigmoidal (five-parameter) logistic regression, and in Figure 21 In B, a bar graph is shown depicting 48-hour yCB SCF and yCB SOL Concentration-response parameter estimates for cultures (CB 最大 , estimated maximum total cannabinoid titer and citrate EC 50 , estimated citrate concentration producing half-maximal total cannabinoid titer), the cultures were incubated for 48 h in medium supplemented with 0, 10, 30, 100, 300, 1000, 3000, or 10,000 mg / L buffered (pH 6.0) citrate. For all measurements, yCB SCF and yCB SOL n = 3 biological replicates of culture. Floating asterisks indicate yCB SCF With yCB SOL Differences between cultures were statistically significant (determined by Bonferroni multiple comparison post hoc test; α = 0.05). Floating data points and bar graphs depict means with 95% confidence intervals. Dashed lines represent 95% confidence intervals for regression curve fits. *p < 0.05; **p < 0.01; ***p < 0.001; ****p < 0.0001. Detailed Description of the Invention
[0051] This document provides methods and materials for producing cannabinoids in host cells or in vitro using a bidirectional multienzyme scaffold that can control the localization and stoichiometry of enzymes that catalyze the biosynthesis of cannabinoids and cannabinoid precursors. As described herein, the bidirectional multienzyme scaffold and one or more soluble cannabinoid synthases can be used to produce one or more cannabinoids, including cannabigerolic acid (CBGA), cannabidiolic acid (CBDA), cannabicycloleic acid (CBCA), and tetrahydrocannabinolic acid (TECA), as well as the conjugate bases of these cannabinoids, namely cannabigerol, cannabidiolic acid, cannabicycloleate, and tetrahydrocannabinolic acid, respectively, and the decarboxylation products of these cannabinoids, namely cannabigerol (CBG), cannabidiol (CBD), cannabicyclole (CBC), and tetrahydrocannabinol, respectively, as well as the oxidation product of tetrahydrocannabinolic acid, cannabinolic acid, and its decarboxylation product, cannabinol. The bidirectional multi-enzyme scaffolds described herein result in a significant increase in the production of cannabinoids in a recombinant host, including the production of total cannabinoids, CBGA, CBG, CBDA, CBD, CBCA, CBC, and olivetate precursors, compared to the production of cannabinoids in a recombinant host using the same enzymes not bound to the scaffold. As used herein, enzymes not bound to the scaffold are referred to as soluble or non-scaffolded. Although reference may be made herein to a specific form of a cannabinoid or other compound, it is understood that any neutral or ionized form thereof, including any salt form thereof or a decarboxylated derivative thereof (e.g., produced in the presence of heat and light), is also included herein unless otherwise indicated. It will be understood by those skilled in the art that the specific form will depend on factors such as pH and carboxylation state.
[0052] In general, the enzymes described herein can be co-localized on one or more scaffolds and used to produce cannabinoids or cannabinoid precursors that are engineered to contain an interaction domain (ID) that can be separated from the N-terminus or C-terminus of the enzyme by an amino acid spacer sequence. The ID can contain two or more scaffold binding motifs. The engineered enzyme can also include one or more linkers between the enzyme, the spacer, and / or the ID. The engineered enzyme can be bound to a scaffold that is a polypeptide containing a unique ID binding domain, i.e., a tandem peptide ligand, such as Figure 1A and Figure 1BAs shown, the enzyme is co-located on the support. In other words, each enzyme can be engineered to include a protein-protein interaction domain that is specific to one or more ligands (binding sites) on the support, so that the enzyme can be positioned at discrete positions along the support by non-covalent interactions. In some cases, the engineered enzyme can be a chimeric enzyme. Amino acid linkers or spacers can be used to separate the scaffolding ligands. For example, see Horn and Sticht, Frontiers in Bioengineering and Biotechnology, 2015, Vol. 3, P. 191; Whitaker and Dueber, Methods in Enzymology, Chapter 19, " Metabolic Pathway Flux Enhancement by Synthetic Protein Scaffolding ", Vol. 497, 2011, for describing ID, binding domain, linker and spacer. ID can also be referred to as an adaptor domain.
[0053] Typically, each interaction domain consists of two tandem scaffold-binding motifs that extend from the C-terminus of the engineered enzyme and are capable of binding to their corresponding scaffolded peptide ligands, which are constructed in tandem along the scaffold. The dual binding of the enzyme to the scaffold ensures a fixed spatial orientation, increases the binding specificity of each ID-scaffold interaction, and better tethers each enzyme to the scaffold, all of which can improve pathway flux by substrate channeling through each enzymatic step in the scaffolded biosynthetic pathway.
[0054] In some embodiments, there is more than two, for example three, four, five, six, seven, eight, nine or ten or more molecules of each enzyme that are positioned at support.In addition, the ratio of any given enzyme and any other enzyme in biosynthetic pathway is variable.For example, the ratio of a kind of engineered enzyme in approach and the second engineered enzyme in same approach is variable, for example from about 1:5 to about 5:1, for example from about 1:5 to about 2:5, about 2:5 to about 3:5, about 3:5 to about 5:5, about 5:5 to about 5:3, about 5:3 to about 5:2 or about 5:2 to about 5:1.
[0055] Peptide ligands are typically short peptide sequences, 3-50 amino acid residues in length. For example, the peptide ligand can be 3-10, 7-15, 10-20, 15-25, 20-30, 25-35, 30-40, 35-45, or 40-50 amino acids in length. A database of over 200 different motifs is available at elm.eu.org and can be used as described herein. See, for example, Dinkel et al., Nucleic Acids Res. 2014; 42(database content): D259–D266.
[0056] The ID can be a peptide sequence ranging in length from 3 to 200 amino acid residues. For example, the ID can be 3-10, 7-15, 10-20, 15-25, 20-30, 25-35, 30-40, 35-45, 40-50, 45-55, 50-60, 65-75, 70-80, 85-95, 90-100, 100-110, 105-115, 110-120, 115-125, 120-130, 125-135, 130-140, 10-15, 13-150, 135-145, 140-150, 145-155, 150-160, 165-175, 170-180, 175-185, 180-190, 185-195, or 190-200 amino acids in length. For example, the ID can be an SH2 domain, an SH3 domain, a PDZ domain, a GTPase binding domain (GBD), a leucine zipper domain, a PTB domain, an FHA domain, a WW domain, a 14-3-3 domain, a death domain, a caspase recruitment domain, a bromodomain, a chromatin organization modifier, a shadow staining domain, an F-box domain, a HECT domain, a ring finger domain, a sterile alpha motif domain, a glycine-tyrosine-phenylalanine domain, a SNAP domain, a VHS domain, an A NK repeats, armadillo repeats, WD40 repeats, MH2 domains, calmodulin homology domains, Dbl homology domains, gelsolin homology domains, PB1 domains, SOCS box, RGS domains, Toll / IL-1 receptor domains, tetratricopeptide repeats, TRAF domains, Bcl-2 homology domains, coiled-coil domains, bZIP domains, fibronectin receptor domains, FNDC domains, SAMD domains, WBP domains and / or SASH domains. See, e.g., U.S. Patent No. 9,856,460 for a list of domains that can be used as IDs as described herein.
[0057] For example, the ID can be a "Src homology 2" (SH2) or "Src homology 3" (SH3) domain. SH2 domains are highly conserved structures of approximately 100 amino acid residues, consisting of two α-helices and seven β-strands. SH2 domains can have promiscuous or strict specificity for a motif of 3-5 amino acids flanked by phosphorylated tyrosine. See Horn and Sticht, 2015, supra. For example, an SH2 domain that can be used as an ID as described herein can be residues 5-122 of the mouse Ct10 regulon of the kinase adaptor (Crk) protein having GenBank accession number AAH31149.
[0058] SH3 domains are small modules of approximately 60 residues that bind proline-rich ligands, which bind to the domain surface at three shallow grooves formed by conserved aromatic residues and exhibit two different binding orientations. See Horn and Sticht, 2015, supra. In some embodiments, the proline-rich ligand may have a core PXXP motif flanked by positively charged residues. Class I PZP domains recognize ligands that conform to the consensus +XXPXXP (where + is Arg or Lys), while class II domains recognize the PXXPX+ motif and bind to the ligand in the opposite orientation. See Teyra et al., FEBS Lett., 2012 586(17):2631-7. Individual SH3 domains do not interact measurably with other SH3 domain family ligands in vivo, thereby minimizing crosstalk and increasing the number of domain / ligand pairs that can be used simultaneously. See Whitaker and Dueber, 2011, supra. For example, an SH3 domain useful as an ID as described herein may be residues 134-190 of the mouse Crk protein having GenBank Accession No. AAH31149, and its peptide ligand may be PPPALPPKRRR (SEQ ID NO: 1).
[0059] For example, the ID can be a PDZ (PSD-95 / Discs-large / ZO1) domain. PDZ domains are approximately 100 amino acid residues in length and target a specific motif at the C-terminus of a binding partner. Peptide ligands adopt β-strands and, upon binding, extend the existing β-sheet within the PDZ domain. At least four different classes of ligands are known for PDZ domains, which exhibit different binding specificities. See Horn and Sticht, 2015, supra. For example, PDZ domains are divided into two main specificity classes based on different ligand characteristics: Class I PDZ domain recognition (X[T / S] ) motif, class II PDZ domain recognition motif, while the class III PDZ domain recognizes X[ED] motif, where X is any residue, is a hydrophobic amino acid. See, Teyra et al., 2012, supra. PDZ and SH3 domains are found throughout eukaryotic and eubacterial genomes. For example, a PDZ domain that can be used as an ID as described herein can be residues 77-171 of mouse α-syntrophin having GenBank accession number EDL06069, and the peptide ligand can be GVKESLV (SEQ ID NO: 208).
[0060] For example, the ID can be a GBD domain from a protein, such as Wiskott-Aldrich syndrome-like protein (N-WASP). Isolated GBD domains do not adopt a single, discrete structure under physiological conditions, but rather exhibit multiple, loosely coupled conformations in solution. Corresponding peptide ligands are derived from autoinhibited forms of the GBD. See Horn and Sticht, 2015, supra. For example, a GBD domain useful as an ID as described herein can include residues 196-274 of the rat N-WASP protein having GenBank accession number BAA21534, and its peptide ligand, which can be LVGALMHVMQKRSRAIHSSDEGEDQAGDEDED (SEQ ID NO: 2), useful as a peptide ligand as described herein.
[0061] For example, the ID can have a leucine zipper or a synthetic coiled-coil domain. A leucine zipper domain can include multiple interspersed leucine residues separated by approximately seven amino acid residues. Havranek and Harbury ((2003), Nat. Struct. Biol. 10, 45-52) identified new homodimer or heterodimer pairs by varying the residues between leucine zipper pairs based on computational predictions. Reinke et al. ((2010). J. Am. Chem. Soc. 132, 6025–6031) identified three pairs of synthetic coiled coils that did not show measurable self-association. See Whitaker and Dueber, 2011, supra. An example of an ID that may be used as described herein may be ITIRAAFLEKENTALRTEIAELEKEVGRCENIVSKYETRYGPL (SEQ ID NO: 3), and its peptide ligand used as described herein may be LEIRAAFLEKENTALRTRAAELRKRVGRCRNIVSKYETRYGPL (SEQ ID NO: 4).
[0062] For example, the ID can be a cohesion polypeptide that can be localized to a specific cohesion polypeptide on a scaffold described herein. Cohesion-dockerin pairs are particularly useful for in vitro applications because binding is calcium-dependent. See Whitaker and Dueber, 2011, supra.
[0063] Combinations of IDs with high affinity and high specificity for their peptide ligands, i.e., minimal cross-reactivity, can be used as described herein to allow a variety of different enzymes to be bound to the scaffolds provided herein. For example, at least three different enzymes can be positioned on the scaffold. In some embodiments, at least four different enzymes can be positioned on the scaffold. In some embodiments, at least five different enzymes can be positioned on the scaffold. In some embodiments, at least six different enzymes can be positioned on the scaffold. In some embodiments, at least seven different enzymes can be positioned on the scaffold. In some embodiments, at least eight different enzymes can be positioned on the scaffold. In some embodiments, at least nine different enzymes can be positioned on the scaffold. In some embodiments, at least ten different enzymes can be positioned on the scaffold. In some embodiments, at least eleven different enzymes can be positioned on the scaffold. In some embodiments, at least twelve different enzymes can be positioned on the scaffold. In some embodiments, at least fifteen different enzymes can be positioned on the scaffold. In some embodiments, at least seventeen different enzymes can be positioned on the scaffold. In some embodiments, at least eighteen different enzymes can be positioned on the scaffold. In some embodiments, at least twenty different enzymes can be positioned on the scaffold. In some embodiments, at least twenty-one different enzymes may be positioned on the scaffold.
[0064] Table 1 provides exemplary combinations of heterologous IDs, i.e., IDs that are different from one another, that can be used for seventeen different engineered enzymes, and Table 2 provides exemplary combinations of corresponding peptide ligands that can be used to localize seventeen different enzymes to one or more scaffolds. In the embodiments shown in Tables 1 and 2, each ID comprises two tandem peptide motifs, such as corresponding peptide ligands that interact with the tandem peptide motifs. It should be understood that any of the enzymes listed in Tables 1 and 2 can be used in combination with any of the listed combinations of IDs and corresponding peptide ligands.
[0065] Table 1
[0066] Interaction domain motif sequences in engineered enzymes
[0067]
[0068]
[0069]
[0070]
[0071]
[0072]
[0073] Table 2 Sequences of tandem peptide ligands in the scaffold
[0074]
[0075]
[0076]
[0077]
[0078]
[0079] The spacer or joint of ligase and ID, and the binding domain on scaffold, can be a peptide sequence with a length of 6 to 250 amino acid residues. The term "spacer" generally refers to a peptide sequence that is longer and has a more rigid structure, while the term "joint" generally refers to a peptide sequence that is shorter and has a more flexible structure. In an embodiment in which these two terms are used simultaneously, a joint generally refers to a sequence with a length of about 3 to about 50 amino acids, while a spacer generally refers to a longer sequence (for example, a length of about 36 to about 250 amino acids). For example, the length of the joint can be 6-15, 10-20, 15-25, 20-30, 25-35, 30-40, 35-45 or 40-50 amino acids. The spacer length can be, for example, 36-40, 40-50, 45-55, 50-60, 55-65, 60-70, 65-75, 70-80, 75-85, 90-100, 95-105, 100-110, 105-115, 110-120, 115-125, 120-130, 125-135, 130-140, 135-145, 140-150 , 145-155, 150-160, 165-175, 170-180, 175-185, 180-190, 185-195, 190-200, 195-205, 200-210, 205-215, 210-220, 215-225, 220-230, 225-235, 230-240, 235-245, or 240-250 amino acids. See, e.g., Chen et al., Adv Drug Deliv Rev. 2013 65(10): 1357–1369. In either case, the linker / spacer can be a series of small and / or hydrophilic and / or other amino acid residues that can accommodate flexible and / or rigid structures. For example, the linker can be a series of glycine residues, a series of alanine residues, a series of serine residues, or a series of alternating glycine and serine (or threonine) residues, such as (GS)8 (SEQ ID NO: 60), (GS) 10 (SEQ ID NO: 61) or (GS) 15(SEQ ID NO: 62), or primarily comprise glycine residues, such as (GGGGS)3 (SEQ ID NO: 63) or (GGGGS)4 (SEQ ID NO: 64), or comprise any other series of canonical or non-canonical amino acid residues or combinations thereof. In some embodiments, the linker can include glutamic acid, alanine, and lysine residues, such as (EAAAK)2 (SEQ ID NO: 65), (EAAAK)3 (SEQ ID NO: 66), or (EAAAK)4 (SEQ ID NO: 67). See Horn and Sticht, 2015, supra. In some embodiments, the linker can be a combination of glycine, alanine, proline, and methionine residues, such as AAAGGM (SEQ ID NO: 68), AAAGGMPPAAAAGGM (SEQ ID NO: 69), AAAGGM (SEQ ID NO: 70), or PPAAAGGMM (SEQ ID NO: 71). See, for example, U.S. Patent No. 9,856,460.
[0080] Depending on the amino acid composition, the linker or spacer can be structured or unstructured in nature. For example, in some embodiments, the spacer can have a sequence that adopts a more structurally rigid α-helical conformation, while the linker can have a more structurally flexible GS-rich peptide sequence. For example, in some embodiments, the linker can include a flexible GS-rich sequence flanked by one or more rigid α-helical portions, such as a GS-rich sequence flanked by double, triple, or quadruple repeat α-helical portions. For example, in some embodiments, the linker or spacer can have the sequence GSAGSAAGSGEF (SEQ ID NO: 72), KLSGGGGSGGGGSGGGGS (SEQ ID NO: 73), GSAGSAAGSGEFGSAEAAAKEAAAKAGSAGSAAGSGEFGS (SEQ ID NO: 74), GSAGSAAGSGEFAEAAAKEAAAKAGSAGSAAGSGEF (SEQ ID NO: 75), or
[0081] GSAGSAAGSGEFGSAEAAAKEAAAAKEAAAKEAAAKAGSAGSAAGSGEFG S (SEQ ID NO: 76).
[0082] In some embodiments, the ligands on the scaffold can be separated by linkers of 20-50 amino acid residues in length (e.g., 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acid residues in length). In some embodiments, the ID engineered at the C-terminus or N-terminus of each scaffolded enzyme can comprise a linker (e.g., a flexible linker) of 15 to 30 (e.g., 20) amino acid residues in length, flanked by spacers of 15 to 50 (e.g., 36) amino acid residues. In some embodiments, the ID can be separated from the enzyme by a spacer sequence, such as a cTPR6 spacer sequence, which contains a six-repeat rigid α-helical portion and can have the following sequence: AEAWYNLGNAYYKQGDYQKAIEYYQKALELDPNNAEAWYNLGNAYYKQGDYQKAIEYYQKALELDPNNAEAWYNLGNAYYKQGDYQKAIEDYQKALELDPNNLQAEAWKNLGNAYYKQGDYQKAIEYYQKALELDPNNASAWYNLGNAYYKQGDYQKAIEYYQKALELDPNNAKAWYRRGNAYYKQGDYQKAIEDYQKALELDPNNRSRSA (SEQ ID NO: 77).
[0083] In some embodiments, the engineered enzyme may have the formula: enzyme-linker 1- Spacer-Linker 2- motif 1- connector 3- Motif 2, wherein Linkers 1, 2, and 3 may be the same or different, and Motif 1 and Motif 2 may be the same or different. In some embodiments, Linker 1 may be referred to as an enzyme linker, i.e., it connects the enzyme to the spacer, such as the cTPR6 spacer, and may include a GC-rich flexible portion, such as KLSGGGGSGGGGSGGGGS (SEQ ID NO: 73), flanked by rigid α-helical portions. In some embodiments, Linker 2 may be referred to as an ID linker and may include, for example, a GS-rich flexible portion, such as GGGGSGGGGSGGGGAS (SEQ ID NO: 78), flanked by rigid α-helical portions. In some embodiments, Linker 3 may be referred to as a motif linker and may include a GS-rich flexible portion, such as GSAGSAAGSGEFGSAEAAAKEAAAKAGSAGSAAGSGEFGS (SEQ ID NO: 74), flanked by rigid α-helical portions. Table 1 provides non-limiting examples of Motifs 1 and 2 that can be used together to form heterologous IDs. Figure 3 A schematic diagram comprising an exemplary engineered enzyme of this formula complexed with a scaffold. Figure 6B and Figure 13A -C contains the following amino acid sequences: ATP citrate lyase, atoB, 3-hydroxybutyryl-CoA dehydrogenase, enoyl-CoA hydratase, trans-enoyl-CoA reductase, β-ketothiolase (bktB), HMG-CoA synthase, truncated HMG-CoA reductase, mevalonate kinase, phosphomevalonate kinase, diphosphomevalonate decarboxylase, isopentenyl diphosphate delta isomerase, geranyl diphosphate synthase (ERG20 WW ), olivetol synthase, olivetate cyclase, CBGA synthase and acetyl-CoA carboxylase, according to the formula. In some embodiments, linkers 1 and 2 can be (G4S)3, the spacer can be a cTPR6 sequence, and linker 3 can be (GS)8.
[0084] In some embodiments, the scaffold may have the following formula: N-terminus - [ligand #1 - linker - ligand #2 - spacer] n - (optionally labeled) C-terminus, where n is the number of interaction domains. The linker may be referred to as a scaffolded ligand linker and may be used to connect and separate pairs of motif-binding ligands that recruit / localize each enzyme to its scaffold binding site. Such a linker may include a GS-rich flexible portion flanked by a rigid α-helical portion and have a sequence such as GSAGSAAGSGEFAEAAAKEAAAKAGSAGSAAGSGEF (SEQ ID NO: 75). The spacer may be referred to as a scaffolded ID binding site spacer and may be used to connect and separate scaffold binding sites (comprising paired motif-binding ligands) for each enzyme. Such a spacer may include a GS-rich flexible portion flanked by a rigid α-helical portion and have a sequence such as GSAGSAAGSGEFGSAEAAAKEAAAKEAAAKEAAAKAGSAGSAAGSGEFGS (SEQ ID NO: 76). The N-terminus may include a flexible GS-rich sequence to help stabilize and dissolve the scaffold. For example, the N-terminus can have the sequence GSAGSAAGSGEFGSAGSAAGSGEFGSAGSAAGSGEF (SEQ ID NO: 79). The C-terminus can include a flexible GS-rich sequence flanked by rigid α-helical portions to stabilize and dissolve the scaffold, and can optionally be tagged (e.g., with a MYC tag, FLAG tag, or other tags described below) to facilitate purification or detection of the scaffold. For example, a C-terminal sequence with a triple-repeat MYC tag can have the sequence
[0085] GSAGSAAGSGEFGSAEAAAKEAAAKEAAAKEAAAKAGSAGSAAGSGEFGS
[0086] EQKLISEEDLEQKLISEEDLEQKLISEEDLGSAGSAGSAAGSGEFGSAGSAAGSGEFGSAGSAAGSGEF (SEQ ID NO: 80). For example, the C-terminal sequence with a triple repeat FLAG tag can have the sequence GSAGSAAGSGEFGSAEAAAKEAAAKEAAAKEAAAKAGSAGSAAGSGEFGSDYKDDDDKDYKDDDDKDYKDDDDKGSAGSAAGSGEFGSAGSAAGSGEFGSAGSAAGSGEF (SEQ ID NO: 81). Figure 6C and Figure 13D Each includes an example of a scaffold polypeptide having the formula comprising a peptide ligand corresponding to IDS1-16 shown in Table 2, and a three-repeat MYC tag at the C-terminus. For example, Figure 13D Examples of scaffold peptides containing a triple-repeat MYC tag (see Figure 2B of the SCF gene cassette). Figure 6D and Figure 13D Each comprises an example of a scaffold polypeptide comprising peptide ligands corresponding to IDs 1 and 17 shown in Table 2 and a three-repeat FLAG tag at the C-terminus. Thus, the amino acid sequence of the scaffold can depend on the sequence of the peptide ligand that can bind to the selected ID motif of the enzyme.
[0087] In some embodiments, any of the enzymes can be engineered to contain an N-terminal or C-terminal linker motif that allows for covalent (iso-peptide) bonding to the scaffold. See, for example, the SpyTag and SpyCatcher systems described in Zakeri et al., Proc. Natl. Acad. Sci., 2012 109(12) E690-E697.
[0088] In some embodiments relating to the multi-enzyme scaffolds described herein, a first engineered enzyme in a biosynthetic pathway can produce a first product, which can be a substrate for a second engineered enzyme in the biosynthetic pathway, which can produce a second product, which can be a substrate for a third engineered enzyme in the biosynthetic pathway, and so on. In some cases, the second engineered enzyme can be fixed on the scaffold so that it is adjacent to or very close to the first engineered enzyme. The third engineered enzyme can be fixed on the scaffold so that it is adjacent to or very close to the second engineered enzyme. In this way, a high effective concentration of the first product can be obtained, and the second engineered enzyme can efficiently act on the first product, the third engineered enzyme can efficiently act on the second product, and so on.
[0089] exist Figure 1Aand 1B In one example of a multi-enzyme scaffold, an enzyme from the hexanoyl-CoA pathway is included at the N-terminus of the scaffold, an enzyme from the mevalonate pathway is included at the C-terminus of the scaffold, and an enzyme from the cannabinoid pathway is included therebetween. In any pathway, the enzymes can be from a single source, i.e., from a single species or genus, or can be from multiple sources, i.e., from different species or genera. Nucleic acids encoding the enzymes described herein have been identified from various organisms and are readily available in public databases such as GenBank or EMBL (see below).
[0090] The fully assembled multi-enzyme scaffolds provided herein can adopt stoichiometric and spatial arrangements that help maximize pathway flux and minimize the accumulation of pathway intermediates and byproducts. This scaffold can promote substrate passage within and between cannabinoid and cannabinoid precursor pathways. Specifically, this scaffolded system can promote unidirectional flux through each major cannabinoid precursor pathway, converging near the midpoint of the scaffold. The hexanoyl-CoA / olive acid (OVA) pathway can start from the N-terminus of the scaffold, and the mevalonate or MEP pathway can start from the C-terminus of the scaffold. The enzyme that catalyzes the rate-limiting / key step in cannabinoid biosynthesis, CBGA synthase, can be positioned at the intersection of these precursor pathways near the midpoint of the scaffold.
[0091] Through this design, hexanoyl-CoA / olive acid and geranyl pyrophosphate, two major precursors for cannabinoid biosynthesis, can be bidirectionally delivered to CBGA synthase at this junction. CBGA synthase catalyzes the biosynthesis of CBGA, the primary cannabinoid from which all other cannabinoids are derived. According to the law of mass action, substrate channeling within and between scaffolded pathways can accelerate the kinetics of the complex pathways.
[0092] like Figure 1A and 1B As shown, the N-terminal hexanoyl-CoA pathway can include ATP citrate lyase (ACL) (also known as ATP citrate synthase), acetyl-CoA acetyltransferase (atoB), two 3-hydroxy-acyl-CoA dehydrogenases (BHBD), two enoyl-CoA hydratases (ECH), one β-ketothiolase (bktB) and two trans-2-enoyl-CoA-reductases (ECR).
[0093] In such Figure 1A and 1BIn the illustrated hexanoyl-CoA pathway, citric acid from cell metabolism and / or supplemented in the growth medium can serve as a substrate for acetyl-CoA synthesis catalyzed by ACL. ACL is classified under EC 2.3.3.8. Acetyl-CoA can serve as a substrate for acetoacetyl-CoA synthesis catalyzed by atoB. atoB is classified under EC 2.3.1.9. Acetoacetyl-CoA can serve as a substrate for 3-hydroxybutyryl-CoA synthesis catalyzed by BHBD. BHBD is classified under EC 1.1.1.157. 3-Hydroxybutyryl-CoA can serve as a substrate for trans-but-2-enoyl-CoA synthesis catalyzed by ECH. ECH is classified under EC 4.2.1.17. trans-but-2-enoyl-CoA can serve as a substrate for butyryl-CoA synthesis catalyzed by ECR. ECR is classified under EC 1.3.8.1. Butyryl-CoA can serve as a substrate for the synthesis of 3-ketohexanoyl-CoA catalyzed by bktB. bktB is classified under EC 2.3.1.9. The bktB that catalyzes the production of 3-ketohexanoyl-CoA from butyryl-CoA can be the same as or different from the atoB that catalyzes the production of acetoacetyl-CoA from acetyl-CoA. 3-Ketohexanoyl-CoA is a substrate for the synthesis of 3-hydroxyhexanoyl-CoA catalyzed by BHBD. BHBD is classified under EC 1.1.1.157. The BHBD that catalyzes the production of 3-hydroxyhexanoyl-CoA can be the same as or different from the BHBD that catalyzes the production of 3-hydroxybutyryl-CoA. 3-Hydroxyhexanoyl-CoA can be a substrate for the synthesis of trans-hex-2-enoyl-CoA catalyzed by ECH. ECH is classified under 4.2.1.17. The ECH that catalyzes the production of trans-hex-2-enoyl-CoA can be the same or different from the ECH that catalyzes the production of trans-but-2-enoyl-CoA. Trans-hex-2-enoyl-CoA can be the substrate for the synthesis of hexanoyl-CoA catalyzed by ECR. ECRs are classified under EC 1.3.1.38 or EC 1.3.1.44. The ECR that catalyzes the production of hexanoyl-CoA can be the same or different from the ECR that catalyzes the production of butyryl-CoA.
[0094] In some embodiments, a hexanoyl-CoA synthetase (HCS) enzyme can be included in a soluble form in place of or in addition to the scaffolding enzyme of the hexanoyl-CoA pathway, and in some embodiments, hexanoate can be added to the growth medium as a substrate for HCS-catalyzed hexanoyl-CoA production. The HCS can be included on a scaffold positioned at Figure 1A and 1B The N-terminus of the upper cannabinoid pathway is shown, and / or it may be unscaffolded (soluble).
[0095] like Figure 1A and 1BAs shown, the C-terminal mevalonate pathway may include ACL, atoB, hydroxymethylglutaryl-CoA, HMG-CoA synthase (HMGS), HMG-CoA reductase (HMGR), mevalonate kinase (ERG12), phosphomevalonate kinase (ERG8), diphosphomevalonate decarboxylase (MVD1), isopentyl diphosphate isomerase (IDI1) and mutant GPP synthase (mGPPS). Figure 1A and 1B In the illustrated mevalonate pathway, citrate derived from cellular metabolism and / or supplemented in the growth medium can serve as a substrate for acetyl-CoA synthesis catalyzed by ACL. ACL is classified under EC 2.3.3. Acetyl-CoA can serve as a substrate for acetoacetyl-CoA synthesis catalyzed by bktB. bktB is classified under EC 2.3.1.9. Acetoacetyl-CoA can serve as a substrate for HMG-CoA synthesis catalyzed by HMGS. HMG-CoA can serve as a substrate for mevalonate synthesis catalyzed by HMGR. HMGR is classified under EC 1.1.1.88 or 1.1.1.34. Mevalonate can serve as a substrate for mevalonate 5-phosphate synthesis catalyzed by mevalonate kinase. Mevalonate kinase is classified under EC 2.7.1.36. Mevalonate 5-phosphate can serve as a substrate for mevalonate pyrophosphate synthesis catalyzed by phosphomevalonate kinase. Phosphomevalonate kinase is classified under EC 2.7.4.2. Mevalonate pyrophosphate can be a substrate for the synthesis of isopentyl pyrophosphate catalyzed by diphosphomevalonate decarboxylase. Diphosphomevalonate decarboxylase is classified under EC 4.1.1.33. Isopentyl pyrophosphate can be a substrate for the synthesis of dimethylallyl pyrophosphate catalyzed by isopentyl diphosphate isomerase. Isopentyl diphosphate isomerase is classified under EC 5.3.3.2. Dimethylallyl pyrophosphate can be a substrate for the synthesis of geranyl pyrophosphate catalyzed by geranyl pyrophosphate synthase (GPPS). GPPS is classified under EC 2.5.1.1.
[0096] Since acetyl-CoA can be the initial substrate for the biosynthesis of hexanoyl-CoA, mevalonate / geranyl pyrophosphate and malonyl-CoA cannabinoid precursors, Figure 1A and 1B Including ACL at both the N-terminus and C-terminus of the multi-enzyme scaffold allows for direct coupling of the scaffolded pathway to cellular metabolism through ACL-catalyzed production of acetyl-CoA from citric acid cycle-derived citrate. Citrate can also be supplemented to the culture medium (e.g., in a buffered citrate form). In some embodiments, the ACL enzyme is contained only at the N-terminus of the scaffold. In some embodiments, the ACL enzyme is contained only at the C-terminus of the scaffold. In some embodiments, the ACL enzyme is included in a soluble form.
[0097] In some embodiments, the 2-C-methylerythritol 4-phosphate (MEP) pathway, which can also produce geranyl pyrophosphate, can replace the scaffolded mevalonate pathway at the C-terminus of the scaffold, or can be included in soluble form in addition to the scaffolded mevalonate pathway. Figure 5 As shown, the C-terminus of the scaffold can include 1-deoxy-D-xylulose-5-phosphate (DOXP) synthase, DOXP reductoisomerase, MEP cytidine transferase, 4-diphosphocytidine-2-C-methylerythritol (CDPME) kinase, 2-C-methyl-D-erythritol 2,4-cyclodiphosphate (MECDP) synthase, 4-hydroxy-3-methyl-but-2-enyl pyrophosphate (HMBPP) synthase, HMBPP reductase and GPPS. Pyruvate and glyceraldehyde-3-phosphate (G3P) can be used as substrates for DOXP synthesis catalyzed by DOXP synthase. DOXP is classified under EC2.2.1.7. DOXP can be a substrate for MEP synthesis catalyzed by DOXP reductoisomerase (DXR). DXR is classified under EC1.1.1.267. MEP can be a substrate for the synthesis of 4-diphosphocytidyl-2-C-methylerythritol (CDP-ME) catalyzed by 2-C-methyl-D-erythritol 4-phosphate cytidylyltransferase (ISPD). ISPD is classified under EC 2.7.7.60. CDP-ME can be a substrate for the synthesis of 4-diphosphocytidyl-2-C-methyl-D-erythritol 2-phosphate (CDP-MEP) catalyzed by 4-diphosphocytidyl-2-C-methyl-D-erythritol kinase (ISPE). ISPE is classified under EC 2.7.1.148. CDP-MEP can be a substrate for the synthesis of 2-C-methyl-D-erythritol 2,4-cyclodiphosphate (cMEPP) catalyzed by 2-C-methyl-D-erythritol 2,4-cyclodiphosphate synthase (ISPF). ISPF is classified under EC 4.6.1.12. cMEPP can be a substrate for the synthesis of (E)-4-hydroxy-3-methyl-but-2-enyl pyrophosphate (HMBPP) catalyzed by HMB-PP synthase (ISPG). ISPG is classified under EC 1.17.7.1. HMBPP can be a substrate for the synthesis of isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP) catalyzed by 4-hydroxy-3-methylbut-2-enyl diphosphate reductase (ISPH). ISPH is classified under EC 1.17.1.2. IPP and DMAPP can be substrates for the synthesis of geranyl pyrophosphate catalyzed by GPPS. GPPS is classified under EC 2.5.1.1.
[0098] In some embodiments, the mevalonate pathway can replace the scaffolded MEP pathway at the scaffolded C-terminus, or can be included in a soluble form in addition to the scaffolded MEP pathway.
[0099] In such Figure 1A and 1B In the illustrated embodiment, a second multi-enzyme scaffold can be co-expressed to increase the cytosolic titer of malonyl-CoA, another secondary substrate useful for cannabinoid biosynthesis. This scaffold can include ATP citrate lyase (ACL) and acetyl-CoA carboxylase (ACC) in tandem. In some embodiments, ACL and ACC are paired in triplicate or direpeat along the scaffold. If ACL and ACC are paired in triplicate or direpeat, the three or two ACLs on the scaffold can be the same or different, and the three or two ACCs can be the same or different. In any of these embodiments, malonyl-CoA can be supplemented to the growth medium in place of or in addition to the malonyl-CoA pathway provided by the scaffolded malonyl-CoA pathway.
[0100] In any embodiment using ACL enzyme, pyruvate dehydrogenase (E1) and dihydrolipoyl transacetylase (E2) can be substituted for ACL. Figure 4 Shown in, pyruvate dehydrogenase (E1) and dihydrolipoyl transacetylase (E2) can be substituted in the upstream of scaffolding mevalonate, hexanoyl-CoA and malonyl-CoA approach.Use pyruvate dehydrogenase (E1) and dihydrolipoyl transacetylase can allow to use pyruvate rather than citric acid to produce acetyl-CoA as primary substrate.In such embodiment, pyruvate also can be supplemented in the growth medium.Pyruvate dehydrogenase and dihydrolipoyl transacetylase are the integral parts of multienzyme pyruvate dehydrogenase complex, and its energy catalysis produces acetyl-CoA from pyruvate.E1 and E2 are present in bacterium and eukaryote.
[0101] like Figure 1A and Figure 1B As shown, the co-scaffolded upper cannabinoid pathway can include olivetol synthase (OS), olivetate cyclase (OAC) and an aromatic isoprenyltransferase (APT), such as CBGA synthase (CBGAS). The upper cannabinoid pathway can begin by utilizing hexanoyl-CoA and three malonyl-CoAs as substrates for the synthesis of 3,5,7-trioxadodecanoyl-CoA catalyzed by olivetol synthase. Olivetol synthase is classified under EC 2.3.1.206. 3,5,7-trioxadodecanoyl-CoA can serve as a substrate for the synthesis of olivetate catalyzed by OAC. OAC is classified under EC 4.4.1.26.
[0102] At the flux intersection point (near the midpoint of the scaffold) of the converging N-terminal hexanoyl-CoA / cannabinoid and C-terminal mevalonate / MEP pathways, APTs such as CBGAS can utilize olivetate from the hexanoyl-CoA / cannabinoid pathway and geranyl pyrophosphate from the mevalonate or MEP pathways as substrates for cannabinate synthesis. Suitable APTs are classified under EC 2.5.1.102.
[0103] In some embodiments, enzymes in the upper cannabinoid pathway can be scaffolded with hexanoyl-CoA synthetase (HCS) to biosynthesize cannabigerol. In some embodiments, soluble HCS can be used with scaffolded enzymes of the upper cannabinoid pathway to biosynthesize cannabigerol, such as Figure 7 Suitable enzymes for the upper cannabinoid pathway are described above.
[0104] In some embodiments, a minimal bidirectional stent may be used, e.g. Figure 8 The scaffold depicted in FIG, wherein HCS is at the N-terminus of the scaffold, GPPS is at the C-terminus of the scaffold, and enzymes in the upper cannabinoid pathway are scaffolded between HCS and GPPS.
[0105] In some embodiments, for example Figure 9 In the illustrated embodiment, the enzymes in the upper cannabinoid pathway can be scaffolded, while the enzymes in the hexanoyl-CoA pathway, the enzymes in the mevalonate pathway, and the enzymes in the malonyl-CoA pathway can be soluble. In some embodiments, the enzymes in the upper cannabinoid pathway can be scaffolded, while the enzymes in the hexanoyl-CoA pathway, the enzymes in the MEP pathway, and the enzymes in the malonyl-CoA pathway can be soluble. In such embodiments, HCS can replace the soluble forms of the enzymes of the hexanoyl-CoA pathway. Suitable enzymes for each of these pathways are described above.
[0106] In some embodiments, the enzymes in the upper cannabinoid pathway can be scaffolded, while the enzymes in the hexanoyl-CoA synthase, mevalonate or MEP pathways, and the enzymes in the malonyl-CoA pathway can be soluble. Suitable enzymes for each of these pathways are described above.
[0107] In some embodiments, the HCS may be scaffolded at the N-terminus relative to the scaffolded enzymes in the upper cannabinoid pathway, while the enzymes in the mevalonate or MEP pathway and the enzymes in the malonyl-CoA pathway may be soluble. Suitable enzymes for each of these pathways are described above.
[0108] In some embodiments, the enzymes in the upper cannabinoid pathway can be scaffolded, while the enzymes in the hexanoyl-CoA pathway or hexanoyl-CoA synthase and the enzymes in the mevalonate or MEP pathway can be soluble. In some embodiments, the enzymes in the hexanoyl-CoA pathway or hexanoyl-CoA synthase can be scaffolded at the N-terminus relative to the enzymes in the upper cannabinoid pathway, and the enzymes in the mevalonate or MEP pathway can be soluble. In such embodiments, malonyl-CoA can be supplemented. Suitable enzymes for each of these pathways are described above.
[0109] In some embodiments, for example Figure 10In the embodiment shown, a bidirectional scaffold can include enzymes of the malonyl-CoA (MCA) pathway at the N-terminus of the scaffold, enzymes of the mevalonate pathway at the C-terminus of the scaffold, and enzymes of the upper cannabinoid pathway in between. In some embodiments, a bidirectional scaffold can include enzymes of the malonyl-CoA pathway at the N-terminus of the scaffold, enzymes of the MEP pathway at the C-terminus of the scaffold, and enzymes of the upper cannabinoid pathway in between. In such embodiments, enzymes of the hexanoyl-CoA pathway can be on a separate scaffold or can be soluble. In some embodiments, HCS can replace scaffolded or soluble enzymes of the hexanoyl-CoA pathway.
[0110] In some embodiments, each pathway is on a separate scaffold. For example, in one embodiment, enzymes for the upper cannabinoid pathway can be located on one scaffold, enzymes for the mevalonate or MEP pathway can be located on one scaffold, enzymes for the hexanoyl-CoA pathway can be located on one scaffold, and enzymes for the malonyl-CoA pathway can be located on another scaffold.
[0111] In any of the embodiments described herein, the cannabigerolic acids biosynthesized cannabinoids can be isolated and / or used as substrates for the synthesis of other secondary and tertiary cannabinoids using downstream cannabinoid synthases. To produce a wider variety of cannabinoids, the downstream cannabinoid synthases are typically not scaffolded, as scaffolding favors the production of the terminal cannabinoids. However, in some embodiments, one or more downstream cannabinoid synthases may be contained on a scaffold as described herein.
[0112] For example, one or more of cannabidiolic acid synthase (CBDAS), cannabichromenic acid synthase (CBCAS), tetrahydrocannabinolic acid synthase (THCAS) or other cannabinoid synthases can be used to produce other cannabinoid-derived cannabinoids. For example, CBDAS; CBCAS; THCAS; CBDAS and CBCAS; CBDAS and THCAS; CBCAS and THCAS; or CBDAS, CBCAS and THCAS can be used to produce other cannabinoid-derived cannabinoids, such as one or more of cannabidiolic acid, cannabichromenic acid and delta-9 tetrahydrocannabinolic acid. CBDAS is classified under EC 1.21.3.8 and can catalyze the synthesis of cannabidiolic acid from cannabichromenic acid. CBCAS is classified under EC 1.3.3- and can catalyze the synthesis of cannabichromenic acid from cannabichromenic acid. THCAS is classified under EC 1.21.3.7 and can catalyze the synthesis of delta-9 tetrahydrocannabinolic acid from cannabichromenic acid.
[0113] Host cells for cannabinoid production
[0114] Cannabinoids can be produced in host cells or in vitro using the multi-enzyme scaffolds described herein. Suitable host cells include any microorganism, eukaryotic or prokaryotic, such as bacteria (e.g., Escherichia coli, Bacillus, Brevibacterium, Streptomyces, or Pseudomonas), yeast (e.g., Pichia pastoris, Saccharomyces cerevisiae, Yarrowia lipolytica, Kluyveromyces marxianus, or Komagataella phaffii), and other fungi (e.g., Neurospora crassa), and green algae (e.g., Dunaliella sp., Chlorella variabilis), Euglena variabilis, and mutabilis or Chlamydomonas reinhardtii), as well as plant cells that can be maintained in culture (e.g., tobacco, cannabis, or other photosynthetic plant cells), or in the case of plant cells (e.g., cells from tobacco or cannabis plants), that can be engineered in culture and grown as whole transgenic plants (e.g., cells from tobacco or cannabis plants). Such host cells or plants may or may not naturally produce cannabinoids.
[0115] Host cells can be modified to contain one or more exogenous nucleic acids encoding a scaffold as described herein and one or more exogenous nucleic acids encoding an engineered enzyme. As used herein, the term "nucleic acid" encompasses RNA and DNA, including cDNA, genomic DNA, and synthetic (e.g., chemically synthesized) DNA. Nucleic acid can be double-stranded or single-stranded. In the case of a single strand, the nucleic acid can be a sense strand or an antisense strand. In addition, nucleic acid can be circular or linear.
[0116] As used herein, the term "exogenous" with respect to nucleic acids and specific host cells refers to any nucleic acid that is not derived from the specific host cell present in nature. Therefore, once introduced into a host cell, non-naturally produced nucleic acids are considered to be exogenous nucleic acids relative to the host cell. It is important to note that non-naturally produced nucleic acids can include nucleic acid sequences or fragments of nucleic acid sequences present in nature, provided that the nucleic acid does not exist in nature as a whole. For example, a nucleic acid molecule comprising a genomic DNA sequence in an expression vector is a non-naturally produced nucleic acid, and therefore, once introduced into a host cell, it is exogenous relative to the host cell because the nucleic acid molecule does not exist in nature as a whole (genomic DNA plus carrier DNA). Therefore, any vector, autonomously replicating plasmid or virus (such as retrovirus, adenovirus or herpes virus) that does not exist in nature as a whole is considered to be non-naturally produced nucleic acids. Therefore, it is believed that genomic DNA fragments and cDNA produced by PCR or restriction endonuclease treatment are non-naturally produced nucleic acids because they exist in the form of separate molecules that do not exist in nature. It can also be inferred that any nucleic acid comprising a promoter sequence and a polypeptide coding sequence (eg, cDNA or genomic DNA) in an arrangement that does not occur in nature is a non-naturally occurring nucleic acid.
[0117] Naturally occurring nucleic acids can be exogenous to a particular cell. For example, an intact chromosome isolated from a cell of organism X is an exogenous nucleic acid to a cell of organism Y once the chromosome is introduced into the cell of organism Y.
[0118] It should be noted that an exogenous nucleic acid molecule encoding a polypeptide having an enzymatic activity that catalyzes the production of a compound not normally produced by the host cell can be administered to the host cell. Alternatively or additionally, an exogenous nucleic acid molecule encoding a polypeptide having an enzymatic activity that catalyzes the production of a compound normally produced by the host cell can be administered to the host cell. In this case, the recombinant host cell can produce more of the compound, or can produce the compound more efficiently, compared to a similar host cell that has not been genetically modified.
[0119] An enzyme having a specific enzymatic activity can be a naturally occurring or non-naturally occurring polypeptide. A naturally occurring polypeptide is any polypeptide having a naturally occurring amino acid sequence, including wild-type and polymorphic polypeptides. Such naturally occurring polypeptides can be obtained from any species, including but not limited to animals (e.g., mammals), plants, fungi, and bacterial species. A non-naturally occurring polypeptide is any polypeptide having an amino acid sequence that does not exist in nature. Therefore, a non-naturally occurring polypeptide can be a mutant form of a naturally occurring polypeptide, or an engineered polypeptide, such as an engineered enzyme comprising an ID as described herein. For example, a non-naturally occurring polypeptide having geranyl pyrophosphate synthase activity can be a mutant form of a naturally occurring polypeptide having geranyl pyrophosphate synthase activity. For example, the GPPS encoded by Erg20 can include a substitution of phenylalanine at position 96 by tryptophan and a substitution of asparagine at position 127 by tryptophan (referred to as Erg20 WW ). Erg20 WW It is advantageous to generate geranyl pyrophosphate instead of farnesyl pyrophosphate. See Jiang et al., Metab Eng. 2017, 41: 57-66. For example, a truncated HMGR (tHMGR) can be used, such as an N-terminally truncated HMGR comprising a catalytic domain but not comprising the transmembrane or regulatory domain of HMGR. For example, HMGR from Arabidopsis thaliana (A. thaliana) (GenBank accession number J04537) or HMGR from Saccharomyces cerevisiae (containing only residues 646-1025) can be truncated to remove the transmembrane and / or regulatory domain and used in the scaffold to eliminate the bottleneck in the mevalonate pathway. HMGR catalyzes the rate-limiting step in the mevalonate pathway (see, for example, Song et al., 2017, Scientific reports, doi: 10.1038 / s41598-017-15005-4). For example, the nucleic acid encoding atoB from Saccharomyces cerevisiae can be modified to include a synthetic 5'UTR (e.g., a synthetic 5'UTR sequence: 5'-cggcacccctacaaacagaaggaatataaa-3' (SEQ ID NO: 82)), and can be used in a scaffold because it can alter atoB expression to promote flux rebalancing, which favors the generation of acetoacetyl-CoA rather than the reverse reaction product butyryl-CoA (see Kim et al., 2018, Bioresour Technol, doi: 10.1016 / j.biortech.2017.10.014). The polypeptide can be mutated by, for example, sequence addition, deletion, substitution, or a combination thereof.
[0120] Any of the enzymes described herein that can be used to produce one or more cannabinoids may have at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to the amino acid sequence of the corresponding wild-type enzyme. It will be appreciated that sequence identity may be determined based on the mature enzyme (e.g., with any signal sequence removed).
[0121] For example, ACL can be compared to Homo sapiens ACL (see SEQ ID NO: 83, Figure 6A ), or an amino acid sequence having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to an ACL from Rattus norvegicus, Mus musculus, or Ciona intestinalis (e.g., GenBank Accession Nos. AAA74463, AAK56081, and BAB00624, respectively).
[0122] For example, acetyl-CoA acetyltransferase (atoB) can be expressed in the presence of E. coli atoB (see SEQ ID NO: 84, Figure 6A ) or an amino acid sequence having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to atoB from Cupriavidus necator, Clostridium acetobutylicum, or Arabidopsis thaliana (e.g., GenBank Accession Nos. CAJ92573, AAK80816, and AAM67058, respectively). In some embodiments, a malonyl-CoA acyl carrier protein transacylase from Saccharomyces cerevisiae, Homo sapiens, Serratia plymuthica, or Dickeya paradisiaca can be substituted for atoB, e.g., GenBank Accession Nos. DAA10992, AAH30985, AGO55277, and ACS85236, respectively.
[0123] For example, 3-hydroxy-butyryl-CoA dehydrogenase (BHBD) can be expressed as Clostridium acetobutylicum BHBD (see SEQ ID NO: 85, Figure 6A) or an amino acid sequence having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to a BHBD from Escherichia coli, Treponema denticola, or Arabidopsis thaliana (e.g., GenBank Accession Nos. AIZ91493, AAS11105, and AAN17431, respectively).
[0124] For example, the enoyl-CoA hydratase (ECH) can be synthesized with the Clostridium acetobutylicum ECH (see SEQ ID NO: 86, Figure 6A ) or an amino acid sequence having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to an ECH from Acinetobacter oleivorans, Cupribotium necrotum, or Acinetobacter baumannii (e.g., GenBank Accession Nos. ADI91469, CAJ91294, and ACJ57023, respectively).
[0125] For example, β-ketothiolase (bktB) can be combined with Cupriavidus necrotus bktB (see SEQ ID NO: 87, Figure 6A ) or the amino acid sequence of bktB from Escherichia coli, Lactobacillus casei, or Clostridium acetobutylicum (e.g., GenBank Accession Nos. ALI39443, CAQ67083, and AAK80816, respectively) having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%).
[0126] For example, trans-2-enoyl-CoA-reductase (ECR) can be expressed in the presence of Treponema denticola ECR (see SEQ ID NO: 88, Figure 6A ), or an amino acid sequence having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to an ECR from Cupria necrotica, Saccharomyces cerevisiae, or Klebsiella michiganensis (e.g., GenBank Accession Nos. AAP86010, DAA07148, and AIE72439, respectively).
[0127] For example, hexanoyl-CoA synthetase (HCS), an acyl-activating enzyme (AAE), can be synthesized with C. sativa AAE1 (see SEQ ID NO: 89, Figure 6A , GenBank Accession No. AFD33345) or Cannabis sativa L. AAE3 (GenBank Accession No. AFD33347) has at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100%) to the amino acid sequence of Cannabis sativa L. AAE1 and Cannabis sativa L. AAE3 can each utilize hexanoate as a substrate. See Stout et al., Plant J., 71(3):353-365 (2012). In some embodiments, an AAE encoded by CsAAE1 can be used. For the coding sequence, see GenBank Accession No. JN717233. In some embodiments, an AAE encoded by CsAAE3 can be used. For the coding sequence, see GenBank Accession No. JN717233. In some embodiments, both CsAAE1 and CsAAE3 can be used.
[0128] For example, HMG-CoA synthase (HMGS) can be expressed in the presence of Saccharomyces cerevisiae HMGS (see SEQ ID NO: 90, Figure 6A ) or an amino acid sequence having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to an HMGS from Arabidopsis thaliana, Lactobacillus casei, or Homo sapiens (e.g., GenBank Accession Nos. AEE83052, CAQ67081, and AAA62411, respectively).
[0129] For example, an N-terminally truncated or canonical HMG-CoA reductase (HMGR) can be synthesized with Saccharomyces cerevisiae HMGS (see SEQ ID NO: 91, Figure 6A ) or an amino acid sequence having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to an HMGR from Arabidopsis thaliana, Lactobacillus casei, or Homo sapiens (e.g., GenBank Accession Nos. AEE35849, CAQ67082, and AAA52679, respectively).
[0130] For example, mevalonate kinase can be combined with Saccharomyces cerevisiae mevalonate kinase (see SEQ ID NO: 92, Figure 6A ) or the amino acid sequence of mevalonate kinase from Arabidopsis thaliana, Lactobacillus casei, or Homo sapiens (e.g., GenBank Accession Nos. AAD31719, CAQ66794, and AAF82407, respectively).
[0131] For example, phosphomevalonate kinase can be combined with Saccharomyces cerevisiae phosphomevalonate kinase (see SEQ ID NO: 93, Figure 6A ), or an amino acid sequence having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to the mevalonate kinase from Scheffersomyces stipitis, Lactobacillus casei, or Homo sapiens (e.g., GenBank Accession Nos. EAZ63544, CAQ6630609, and AAH06089, respectively).
[0132] For example, the diphosphomevalonate decarboxylase can be combined with the Saccharomyces cerevisiae diphosphomevalonate decarboxylase (see SEQ ID NO: 94, Figure 6A ) or the amino acid sequence of mevalonate diphosphate decarboxylase from Arabidopsis thaliana, Lactobacillus casei, or Homo sapiens (e.g., GenBank Accession Nos. AAC67348, CAQ66795, and AAC50440, respectively).
[0133] For example, isopentyl diphosphate isomerase can be combined with Saccharomyces cerevisiae isopentyl diphosphate isomerase (see SEQ ID NO: 95, Figure 6A ) or an isopentyl diphosphate isomerase from Arabidopsis thaliana, Lactobacillus casei, or Homo sapiens (e.g., GenBank Accession Nos. AAC49920, CAQ66796, and AAP35407, respectively) having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%).
[0134] For example, the geranyl pyrophosphate synthase (GPPS) (also known as geranyl-diphosphate synthase) can have at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to the amino acid sequence of the Saccharomyces cerevisiae GPS or the GPPS from Acinetobacter baumannii, Lactobacillus casei, or Homo sapiens (e.g., GenBank Accession Nos. ACJ56139, CAQ66932, and AAH100, respectively). In some embodiments, a mutant GPPS can be used. For example, the GPPS encoded by Erg20 can include a substitution of phenylalanine at position 96 with tryptophan and a substitution of asparagine at position 127 with tryptophan (referred to as Erg20). WW ) (See SEQ ID NO: 96, Figure 6A ). Erg20 WWFavors the formation of geranyl pyrophosphate rather than farnesyl pyrophosphate. See Jiang et al., Metab Eng. 2017, 41: 57-66. In some cases, glutamate is effective for Erg20 (Erg20 K179E ) can be used to produce GPPS that is beneficial for producing geranyl pyrophosphate. See WO2016010827A1.
[0135] For example, the DOXP synthase can have at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to the amino acid sequence of an Escherichia coli, Clostridium acetobutylicum, Treponema denticola, or Arabidopsis thaliana DOXP synthase (e.g., GenBank Accession Nos. CDH63925, AAK80036, AAS12424, and ANM65835, respectively).
[0136] For example, the DOXP reductoisomerase can have at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to the amino acid sequence of an Escherichia coli, Clostridium acetobutylicum, Treponema denticola, or Arabidopsis thaliana DOXP reductoisomerase (e.g., GenBank Accession Nos. CDH63708, AAK79760, AAS12860, and AAM61343, respectively).
[0137] For example, the MEP cytidine transferase can have at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to the amino acid sequence of an Escherichia coli, Clostridium acetobutylicum, Treponema denticola, or Arabidopsis thaliana MEP cytidine transferase (e.g., GenBank Accession Nos. CDH66380, AAK81121, AAS12810, and BAB21592, respectively).
[0138] For example, the CDPME kinase can have at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to the amino acid sequence of an Escherichia coli, Clostridium acetobutylicum, Treponema denticola, or Arabidopsis thaliana CDPME kinase (e.g., GenBank Accession Nos. CDH64802, AAK80844, AAS11855, and AEC07908, respectively).
[0139] For example, the MECDP synthase can have at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to the amino acid sequence of an Escherichia coli, Nicotiana tabacum, Treponema denticola, or Acinetobacter baumannii MECDP synthase (e.g., GenBank Accession Nos. CDH66379, AHM22925, AAS12811, and ACJ59227, respectively).
[0140] For example, the HMBPP synthase can have at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to the amino acid sequence of an E. coli, Acinetobacter baumannii, Treponema denticola, or Arabidopsis thaliana HMBPP synthase (e.g., GenBank Accession Nos. AAN81487, ACJ58210, AAS11783, and AED97354, respectively).
[0141] For example, the HMBPP reductase can have at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to the amino acid sequence of an E. coli, Acinetobacter baumannii, Treponema denticola, or Arabidopsis thaliana HMBPP reductase (e.g., GenBank Accession Nos. CDH63564, ACJ57384, AAS11585, and AEE86362, respectively).
[0142] For example, acetyl-CoA carboxylase (ACC) can be expressed in the presence of a yeast acetyl-CoA carboxylase (see SEQ ID NO: 97, Figure 6A ) or an acetyl-CoA carboxylase from Homo sapiens, Treponema denticola, or Cupriavida necrotica (e.g., GenBank Accession Nos. AAP94122, AAS11086, and CAQ67359, respectively) having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%).
[0143] For example, the pyruvate dehydrogenase (E1) and the dihydrolipoyl transacetylase (E2) can have at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to the amino acid sequence of Saccharomyces cerevisiae, Escherichia coli, Clostridium acetobutylicum, or Cupriavidus necrotizingus E1 and E2 (e.g., GenBank Accession Nos. DAA07337, AMC97367, CAQ66617, and CAJ92510 for E1, and DAA10474, AUG14916, CAQ66619, and CAJ92511 for E2, respectively).
[0144] For example, olivetol synthase (OS) can be synthesized with OS from cannabis sativa L. (shown in SEQ ID NO: 98 ( Figure 6A )) or the amino acid sequence of OS from Cannabis sativa L. (having GenBank accession number BAG14339) having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100%). See, e.g., Taura et al., FEBS Letters 583 (2009) 2061-2066.
[0145] For example, olivetate cyclase (OAC) can be combined with OAC from cannabis sativa L. ( Figure 6A ) or an amino acid sequence of OAC from Cannabis sativa L. (having GenBank accession number AFN42527) having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100%). See, e.g., Gagne et al., Proc. Natl. Acad. Sci. USA, 2012 109(31): 12811-12816.
[0146] For example, CBGAS can be combined with an aromatic prenyltransferase (APT) from cannabis sativa L. (e.g., as shown in SEQ ID NO: 100 ( Figure 6A ) has at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100%) to the amino acid sequence of CBGAS) shown in ). See, for example, U.S. Patent Publication No. 20120144523A1 and U.S. Patent No. 8,884,100B2. In some embodiments, a soluble APT (e.g., NphB) from Streptomyces can be used. See, for example, Carvalho et al., FEMS Yeast Research, 17, 2017, fox037.
[0147] For example, cannabidiolic acid synthase (CBDAS) can be combined with CBDAS from cannabis sativa L. ( Figure 6A ) or the amino acid sequence of CBDAS from cannabis sativa L. (having GenBank accession number BAF65033) having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100%). See, e.g., Taura et al., FEBS Lett. 581(16), 2929-2934 (2007).
[0148] For example, cannabichromenic acid synthase (CBCAS) can be combined with CBCAS from cannabis sativa L. ( Figure 6A ) or the amino acid sequence of the cannabis CBCAS shown in SEQ ID NO: 2 of WO 2015 / 196275 A1 having at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%). SEQ ID NO: 2 of WO 2015 / 196275 A1 includes a signal peptide of 28 amino acids at the N-terminus. All or part of the signal peptide may be removed from this sequence. CBDAS from Indian hemp (C. indica) or weed hemp (C. ruderalis) may also be used. In some embodiments, nucleic acid sequences encoding cannabis CBCAS as shown in SEQ ID NO: 8 and 9 of WO 2015 / 196275 A1, respectively, which are optimized for E. coli or yeast, may be used.
[0149] For example, the tetrahydrocannabinolic acid synthase (THCAS) can have at least 70% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%) to the amino acid sequence of THCAS from Cannabis sativa L. (having GenBank Accession No. BAC41356). See, e.g., Sirikantaramas et al., J. Biol. Chem. 279(38), 39767-39774 (2004).
[0150] The percent identity (homology) between two amino acid sequences can be determined as follows. First, the amino acid sequences are aligned using the BLAST2 sequence (B12seq) program from a standalone version of BLASTZ containing BLASTP version 2.0.14. This standalone version of BLASTZ can be obtained from the website of Fish & Richardson (e.g., www.fr.com / blast / ) or the website of the National Center for Biotechnology Information of the U.S. government (www.ncbi.nlm.nih.gov). Instructions explaining how to use the B12seq program can be found in the readme file included with BLASTZ. B12seq uses the BLASTP algorithm to compare two amino acid sequences. To compare two amino acid sequences, the options for B12seq are set as follows: -i is set to the file containing the first amino acid sequence to be compared (e.g., C:\seq1.txt); -j is set to the file containing the second amino acid sequence to be compared (e.g., C:\seq2.txt); -p is set to blastp; -o is set to any desired file name (e.g., C:\output.txt); and all other options are set to their default settings. For example, the following command can be used to generate an output file containing a comparison between two amino acid sequences: C:\Bl2seq -i c:\seq1.txt -j c:\seq2.txt -p blastp -o c:\output.txt. If the two compared sequences have homology (identity), the specified output file will display those homologous regions in the form of aligned sequences. If the two compared sequences do not have homology (identity), the specified output file will not display an aligned sequence. A similar procedure can be followed for nucleic acid sequences, except that blastn is used.
[0151] After alignment, the number of matches is determined by counting the number of positions where the same amino acid residue is present in the two sequences. The percent identity (homology) is determined by dividing the number of matches by the length of the full-length polypeptide amino acid sequence and then multiplying the resulting value by 100. It should be noted that the percent identity (homology) values are rounded to the first decimal place. For example, 78.11, 78.12, 78.13, and 78.14 are rounded down to 78.1, while 78.15, 78.16, 78.17, 78.18, and 78.19 are rounded up to 78.2. It should also be noted that the length value is always an integer.
[0152] It is understood that a polypeptide having a particular amino acid sequence can be encoded by many nucleic acids. The degeneracy of the genetic code is well known in the art; that is, for many amino acids, there is more than one nucleotide triplet that serves as a codon for that amino acid. For example, codons in the coding sequence of a given enzyme can be modified using an appropriate codon bias table to obtain optimal expression in a particular species (e.g., bacteria or fungi). For example, Figure 12A The nucleotide sequence shown is encoding ATP citrate lyase, atoB, 3-hydroxybutyryl-CoA dehydrogenase, enoyl-CoA hydratase, β-ketothiolase (bktB), trans-enoyl-CoA reductase, HMG-CoA synthase, HMG-CoA reductase, mevalonate kinase, phosphomevalonate kinase, diphosphomevalonate decarboxylase, isopentenyl diphosphate delta isomerase, geranyl-diphosphate synthase (ERG20 WW ), olivetol synthase, olivetate cyclase, CBGA synthase, CBDA synthase, CBCA synthase, acetyl-CoA carboxylase, and hexanoyl-CoA synthetase. The nucleic acid sequences for ATP citrate lyase, atoB, 3-hydroxybutyryl-CoA dehydrogenase, enoyl-CoA hydratase, trans-enoyl-CoA reductase, bktB, olivetol synthase, olivetate cyclase, CBGA synthase, CBDA synthase, and CBCA synthase were codon-optimized for expression in yeast. Figures 14A-14C Contains code Figures 13A-13C Codon-optimized (for expression in yeast) nucleic acid sequences of engineered enzymes.
[0153] In addition to sequence similarity, it is understood that enzymes and scaffolds having structural and / or functional similarities to the enzymes and scaffolds described herein are also included within the scope herein.
[0154] Provided herein are recombinant host cells that can be used to produce one or more cannabinoids as described herein. For example, individual host cells can contain exogenous nucleic acids to express the scaffold polypeptide and each enzyme fixed to the scaffold. Importantly, it is noted that such host cells can contain any number and / or combination of exogenous nucleic acid molecules. For example, a particular host cell can contain exogenous nucleic acids encoding the scaffold, as well as enzymes encoding the malonyl-CoA pathway, enzymes encoding the hexanoyl-CoA pathway, or other exogenous nucleic acids encoding HCS and mevalonate or MEP pathways. A single exogenous nucleic acid can encode one enzyme or more than one enzyme (e.g., 1 to 10 (or more) enzymes, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, or one or more copies of 2 to 3 enzymes). Therefore, the number of different exogenous nucleic acids required to produce the engineered enzymes positioned on the scaffold will depend on the design and / or specific embodiment of the scaffold. Figure 2A and Figure 2BNon-limiting schematic diagrams of suitable gene cassettes for expression of scaffolds and enzymes are each provided. Figure 12C Provided are nucleic acid sequences encoding scaffold polypeptides comprising peptide ligands corresponding to IDs 1-16 shown in Table 2 and a three-repeat MYC tag. See also Figure 14D , which is the encoding Figure 13D The codon-optimized nucleic acid sequence of the scaffold polypeptide. Figure 12D A nucleic acid sequence encoding a scaffold polypeptide comprising a peptide ligand corresponding to IDs 1 and 17 and a triple repeat FLAG tag is provided. See also Figure 14D .
[0155] In some embodiments, a nucleic acid sequence encoding a self-cleaving peptide can be used to combine multiple nucleic acids encoding polypeptides (e.g., Figure 2A or Figure 2B 2A peptides.
[0156] In addition, the cells described herein can contain a single copy or multiple copies (e.g., about 5, 10, 20, 35, 50, 75, 100, or 150 copies) of a particular exogenous nucleic acid molecule. Likewise, the cells described herein can contain more than one particular exogenous nucleic acid molecule and / or its copies. For example, a particular cell can contain about 50 copies of exogenous nucleic acid molecule X and about 75 copies of exogenous nucleic acid molecule Y.
[0157] Any method can be used to introduce exogenous nucleic acid molecules into host cells. In fact, many methods for introducing nucleic acids into host cells such as bacteria and yeast are well known to those skilled in the art. For example, heat shock, lipofection, electroporation, nucleofection, conjugation, protoplast fusion and gene gun delivery are commonly used methods for introducing nucleic acids into bacteria and yeast cells. See, for example, Ito et al., J. Bacterol. 153: 163-168 (1983); Durrens et al., Curr. Genet. 18: 7-12 (1990); and Becker and Guarente, Methods in Enzymology 194: 182-187 (1991).
[0158] The exogenous nucleic acid molecule that is included in the specific host cell can remain in this host cell in any form.For example, the exogenous nucleic acid molecule can be integrated into the genome of microorganism or remain on free state.In other words, microorganism can be stable or transient transformant.Equally, microorganism as described herein can contain single copy or multiple copies (for example, about 5,10,20,35,50,75,100 or 150 copies) of specific exogenous nucleic acid molecule as described herein.
[0159] Suitable nucleic acid constructs for expressing engineered enzymes and scaffolds include, for example, CRISPR plasmids, baculovirus vectors, phage vectors, plasmids, phagemids, cosmids, Fosmids, bacterial artificial chromosomes, viral vectors (e.g., viral vectors based on vaccinia virus, polio virus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, etc.), artificial chromosomes based on P1, yeast plasmids, yeast artificial chromosomes, and other vectors. Typically, such constructs include regulatory elements that promote the expression of nucleic acid sequences encoding polypeptides. Typically, regulatory elements are DNA sequences that regulate the expression of other DNA sequences at the transcriptional level. Therefore, regulatory elements include, but are not limited to, promoters, enhancers, etc. Any type of promoter can be used to express amino acid sequences from exogenous nucleic acid molecules. Examples of promoters include, but are not limited to, constitutive promoters, tissue-specific promoters, inducible or repressible promoters that respond or do not respond to specific stimuli (e.g., light, oxygen, chemical concentrations, sound, etc.).
[0160] In some embodiments, endogenous yeast promoters with different constitutive activity levels can be used to express the engineered enzymes and / or scaffolds. In order to maintain an excess of enzyme relative to the scaffold molecule, the scaffold can be expressed under the control of the weakest promoter. For example, one or more of the following yeast promoters can be used: a promoter from the gene encoding transcription elongation factor EF-1α (pTEF1), a promoter from the gene encoding phosphoglycerate kinase (PGK1), a promoter from the gene encoding triosephosphate isomerase (pTPI1), a promoter from the gene encoding hexose transporter (pHXT7), HXT7, a promoter from the gene encoding pyruvate kinase 1 (pPYK1), a promoter from the gene encoding alcohol dehydrogenase 1 (pADH1), or a promoter from the gene encoding triphosphate dehydrogenase (pTDH3). For example, in Figure 2A In the embodiment shown, the pTPI1 promoter can be used to express enzymes of the upper hexanoyl-CoA (HCA) pathway, enzymes of the lower HCA pathway, enzymes of the upper mevalonate (MVA) pathway, enzymes of the lower MVA pathway, and enzymes of the lower cannabinoid (CB) pathway, while the pTEF1 promoter can be used to express enzymes of the upper CB pathway, atoB enzymes, and enzymes of the malonyl-CoA pathway, and the pADH1 promoter can be used to express the scaffold. Among these promoters, the pADH1 promoter has the weakest activity ( Figure 2A The pTEF1 promoter has the strongest activity ( Figure 2A +++ in), the activity of pTPI1 promoter was between the other two ( Figure 2A In some embodiments, the Gal 1-10 promoter (e.g., from Saccharomyces cerevisiae) can be used. See, e.g. Figure 17 .
[0161] The nucleic acid construct can also include a selective marker, such as a selective marker for antibiotics such as neomycin resistance, ampicillin resistance, tetracycline resistance, chloramphenicol resistance, or kanamycin resistance. In some embodiments, a nutritional marker gene for the prototrophy of essential nutrients such as tryptophan (TRP1), uracil (URA3), histidine (HIS3), leucine (LEU2), lysine (LYS2), or methionine can be included on the nucleic acid construct. See, for example, Figure 17. As shown in Example 3, four different auxotrophic markers are used to sequentially select transformed cells containing the desired combination of nucleic acids encoding enzymes and scaffolds. For example, yeast cells transformed with a vector containing the TRP gene and nucleic acids encoding enzymes of the hexanoyl-CoA pathway are grown in a culture medium lacking tryptophan. Transformed cells grown in a culture medium lacking tryptophan are selected and further transformed with a vector containing the LEU gene and nucleic acids encoding mevalonate pathway enzymes. The resulting transformed cells are grown on a culture medium lacking tryptophan and leucine, and cells grown in a culture medium lacking tryptophan and leucine are transformed with a vector containing the HIS gene and nucleic acids encoding cannabinoid pathway enzymes. The resulting transformed cells are grown on a culture medium lacking tryptophan, leucine, and histidine, while cells grown in a culture medium lacking tryptophan, leucine, and histidine are transformed with a vector containing the URA3 gene and nucleic acids encoding a scaffold. The resulting transformed cells are grown on a culture medium lacking tryptophan, leucine, histidine, and uracil. Cells grown in a medium lacking tryptophan, leucine, histidine, and uracil contain the desired combination of enzymes and scaffolds, e.g. Figure 1B shown.
[0162] In some embodiments, the encoded enzymes (e.g., one or more enzymes from the cannabinoid biosynthetic pathway, the mevalonate pathway, the MEP pathway, the hexanoyl-CoA pathway, or hexanoyl-CoA synthetase) and / or scaffolds can include targeting sequences that can be used to direct the enzymes or scaffolds to one of several different intracellular compartments, including, for example, the endoplasmic reticulum (ER), mitochondria, plastids (e.g., chloroplasts), vacuoles, Golgi apparatus, or protein storage vesicles (PSVs). For example, mitochondrial or plastid targeting sequences can be used to promote mitochondrial or plastid compartmentalization of cannabinoid / cannabinoid precursor biosynthesis, such that the encoded enzymes and scaffolds are expressed in the mitochondria or plastids of the host cell.
[0163] In some embodiments, cannabinoid / cannabinoid precursor biosynthesis can be performed by co-expressing one or more engineered enzymes and scaffolds in both the cytosolic compartment and the plastids or mitochondria of the host cell. See, e.g. Figure 11 It should be understood that although Figure 11Depicted are scaffolds comprising enzymes of the hexanoyl-CoA pathway, enzymes of the upper cannabinoid pathway, and enzymes of the mevalonate pathway. Dual compartment engineering can be performed with any of the scaffolds and enzymes described herein. For example, dual compartment engineering can be performed in two compartments by co-expressing the scaffold and enzymes of the hexanoyl-CoA pathway, enzymes of the upper cannabinoid pathway, and enzymes of the MEP pathway in both the cytosolic compartment and the mitochondrial plastid of the host cell. Dual compartment engineering can also be performed by engineering separate haploid yeast strains for cytosolic and mitochondrial / plastid cannabinoid biosynthesis, and then crossing these two haploid strains to generate a diploid lineage that is heterozygous for both cytosolic and mitochondrial / plastid cannabinoid biosynthesis.
[0164] In some embodiments, the engineered enzyme and / or scaffold further comprises a tag (e.g., c-myc, FLAG, polyhistidine (e.g., hexa-histidine), hemagglutinin (HA), glutathione-S-transferase (GST), or maltose binding protein (MBP)) that can be used to purify the recombinant protein, or a detectable marker (e.g., luciferase, green fluorescent protein (GFP), or chloramphenicol acetyltransferase (CAT)). For example, in Figure 6C and 6D In the embodiment shown, the scaffold can include a myc tag (eg, (Myc)3 tag) or a FLAG tag (FLAG)3 tag at the C-terminus.
[0165] In some embodiments, the host cell can be engineered to increase the availability of acetyl-CoA for the biosynthesis of cannabinoids and cannabinoid precursors. For example, the mitochondrial enzyme isocitrate dehydrogenase-1 (IDH1) can be placed under transient microRNA-mediated induction inhibition. Since mitochondrial IDH1 is primarily responsible for the consumption of the cellular citrate pool, microRNA-mediated IDH1 inhibition can increase the availability and cytoplasmic shuttling of citrate to produce acetyl-CoA by ATP citrate lyase. The resulting increase in acetyl-CoA bioavailability can further increase the titer of downstream hexanoyl-CoA and geranyl pyrophosphate by improving the initial substrate availability of the hexanoyl-CoA and mevalonate pathways. The combined metabolic engineering of acetyl-CoA can alleviate the problems associated with siphoning acetyl-CoA from the endogenous metabolism of the host cell.
[0166] In some embodiments, one or more conventional and / or modern gene editing technologies can be used to produce a recombinant host. For example, clustered, regularly interspaced, short palindromic repeats (CRISPR) technology can be used to modify the expression of endogenous nucleic acids. The CRISPR / Cas system includes components of the prokaryotic adaptive immune system, which is functionally similar to eukaryotic RNA interference, using RNA base pairing to guide DNA or RNA cutting. The function of the Cas9 protein is an endonuclease, and the CRISPR RNA (crRNA) and transactivating RNA (tracerRNA) sequences are complexed with the Cas9 enzyme to target the target DNA sequence (Makarova et al., Nat Rev Microbiol 9 (6): 467-477, 2011). Modification of a single targeting RNA can be sufficient to change the nucleotide target of the Cas protein. In some cases, crRNA and tracrRNA can be engineered into a cr / tracrRNA hybrid (also referred to as "guide RNA" or "gRNA") to guide Cas9 cutting activity (Jinek et al., Science, 337 (6096): 816-821, 2012). The CRISPR / Cas system can be used in various prokaryotic and eukaryotic organisms (see, e.g., Jiang et al., Nat Biotechnol, 31(3):233-239, 2013; Dicarlo et al., Nucleic Acids Res, doi:10.1093 / nar / gkt135, 2013; Cong et al., Science, 339(6121):819-823, 2013; Mali et al., Science, 339(6121):823-826, 2013; Cho et al., Nat Biotechnol, 31(3):230-232, 2013; and Hwang et al., Nat Biotechnol, 31(3):227-229, 2013).
[0167] Another gene editing technology can include sequence-specific nucleases created by fusing transcription activator-like effectors (TALEs) to the catalytic domains of, for example, folded endonucleases. Original and customized TALE nucleases ("TALENs") are fused to guide DNA double-strand breaks at specific target sites. See, for example, Christian et al., Genetics 186:757-761 (2010) and U.S. Patent Publication No. 20110145940.
[0168] Other suitable gene insertion techniques include the use of retroviral vectors and gene gun particle gene delivery systems (commonly known as "gene guns").
[0169] Methods for identifying and / or selecting host cells containing exogenous nucleic acids or modified endogenous nucleic acids are well known to those skilled in the art. Such methods include, but are not limited to, introducing and expressing negative selection markers such as antibiotic resistance genes, PCR and nucleic acid hybridization techniques such as Northern and Southern analysis. In some cases, immunohistochemistry and biochemical techniques can be used to determine whether a microorganism contains a specific nucleic acid by detecting the expression of an enzymatic polypeptide encoded by a specific nucleic acid molecule. For example, antibodies specific to the encoded enzyme can be used to determine whether a specific cell contains the encoded enzyme. In addition, biochemical techniques can be used to determine whether a cell contains a specific nucleic acid molecule encoding an enzymatic polypeptide by detecting the organic products produced by the expression of the enzymatic polypeptide.
[0170] Also provided herein are isolated nucleic acid molecules. The term "isolated" used herein with respect to nucleic acid refers to naturally occurring nucleic acids that are not adjacent to the two sequences (one at the 5' end and one at the 3' end) immediately adjacent to the naturally occurring genome of the organism from which they originate. For example, an isolated nucleic acid can be, but is not limited to, a recombinant DNA molecule of any length, provided that one of the nucleic acid sequences that are typically closely flanking the recombinant DNA molecule in the naturally occurring genome is removed or does not exist. Therefore, isolated nucleic acids include, but are not limited to, recombinant DNA that exists as an isolated molecule independently of other sequences (cDNA or genomic DNA fragments produced by PCR or restriction endonuclease treatment), and recombinant DNA that is incorporated into a vector, an autonomously replicating plasmid, a virus (such as a retrovirus, adenovirus, or herpes virus), or the genomic DNA of a prokaryotic or eukaryotic organism. In addition, an isolated nucleic acid can include a recombinant DNA molecule that is a part of a hybrid or fusion nucleic acid sequence.
[0171] The term "isolated" used herein with respect to nucleic acids also includes any non-naturally produced nucleic acids, because such non-naturally produced nucleic acid sequences do not exist in nature and do not have the sequence immediately following in the naturally produced genome. For example, non-naturally produced nucleic acids such as engineered nucleic acids are considered to be isolated nucleic acids. Engineered nucleic acids can be manufactured using general molecular cloning or chemical nucleic acid synthesis techniques. The isolated non-naturally produced nucleic acids can be independent of other sequences or incorporated into vectors, self-replicating plasmids, viruses (such as retroviruses, adenoviruses or herpes viruses) or genomic DNA of prokaryotes or eukaryotes. In addition, non-naturally produced nucleic acids can include nucleic acid molecules that are part of a hybrid or fusion nucleic acid sequence.
[0172] It will be apparent to one skilled in the art that a nucleic acid present among hundreds to millions of other nucleic acid molecules, for example, in a cDNA or genomic library or in a gel slice containing restriction enzyme digests of genomic DNA, is not considered an isolated nucleic acid.
[0173] In some embodiments, one or more cannabinoids can be produced in vitro using the scaffolds and immobilized enzymes described herein, using lysates from recombinant host cells (e.g., buffered cell lysates) as a source of scaffold and enzyme, using multiple lysates from different host cells as a source of scaffold and enzyme, or using a cell-free reaction buffer such as a synthesis reaction buffer. For example, following immunoprecipitation of a C-terminally Myc / Flag-spiked enzyme-bound scaffold, the scaffold-enzyme complex can be maintained in a reaction buffer supplemented with citric acid and / or glucose (or other carbon sources), which allows for in vitro scaffolded cannabinoid biosynthesis.
[0174] Production of cannabinoids using recombinant hosts
[0175] Typically, one or more cannabinoids can be produced by providing a recombinant host such as a recombinant microorganism and culturing the microorganism with a culture medium. Generally speaking, the culture medium and / or culture conditions can allow the microorganism to grow to a sufficient density and efficiently produce cannabinoids. For example, aerobic batch fermentation can be performed on the microorganism. In some embodiments, one or more precursors (e.g., citrate, glucose, hexanoic acid and / or other carbon sources and / or malonyl-CoA) are supplemented in the culture medium. In some embodiments, a buffered citrate of about 30 mg / L to about 10,000 mg / L (e.g., about 100 mg / L to about 5,000 mg / L, about 200 mg / L to about 4,000 mg / L, about 300 mg / L to about 3,000 mg / L, or about 350 mg / L to about 1,000 mg / L), pH 6.0 can be added to the culture medium.
[0176] For large-scale production processes, any method can be used, such as those described elsewhere (Manual of Industrial Microbiology and Biotechnology, 2nd edition, edited by A.L. Demain and J.E. Davies, ASM Publishing Company; and Principles of Fermentation Technology, P.F. Stanbury and A. Whitaker, Pergamon). Briefly, a large container (e.g., a 100-gallon, 200-gallon, 500-gallon, or larger volume container) containing a suitable culture medium is inoculated with a specific microorganism. After inoculation, the microorganism is incubated to produce biomass. Once the desired biomass or cell confluence is reached, part or all of the broth containing the microorganism can be transferred to a second container. The second container can be of any size. For example, the second container can be larger, smaller, or of the same size as the first container. Typically, the second container is larger than the first container so that additional culture medium can be added to the broth from the first container. In addition, the culture medium in the second container can be the same or different from the culture medium used in the first container. The system can be expanded to include an array consisting of any number of individual containers.
[0177] After transfer, the microorganisms can be incubated to produce one or more cannabinoids. Once produced, the cannabinoids can be isolated by any method. For example, the biomass can be removed from the broth using conventional separation techniques, and the cannabinoids can be obtained from the biomass using conventional separation procedures (e.g., extraction, such as non-polar extraction with hexane followed by ethyl acetate), high performance liquid chromatography (e.g., HPLC with diode array detector (HPLC-DAD)), gas chromatography with flame ionization detection (GC-FID), or ion exchange procedures).
[0178] The host cells described herein can produce one or more cannabinoids at a concentration of at least about 10 mg / L (e.g., at least about 15 mg / L, 25 mg / L, 50 mg / L, 75 mg / L, 100 mg / L, 150 mg / L, 200 mg / L, 250 mg / L, or more). For example, in some embodiments, total cannabinoids (the total amount of CBG, CBGA, CBD, CBDA, CBC, and CBCA) can be produced at a concentration of at least about 10 mg / L, 15 mg / L, 20 mg / L, 40 mg / L, 60 mg / L, 80 mg / L, or 100 mg / L or more. For example, in some embodiments, the total cannabinoids (the sum of CBG, CBGA, CBD, CBDA, CBC, and CBCA) can be produced at concentrations of about 10 mg / L to about 500 mg / L (e.g., 20 mg / L to 450 mg / L, 40 mg / L to 380 mg / L, 60 mg / L to 280 mg / L, 60 mg / L to 250 mg / L, 60 mg / L to 150 mg / L, 80 mg / L to 400 mg / L, 80 mg / L to 300 mg / L, 80 mg / L to 250 mg / L, 80 mg / L to 200 mg / L, 80 mg / L to 175 mg / L, 90 mg / L to 400 mg / L, 90 mg / L to 300 mg / L, 90 mg / L to 250 mg / L, or 90 mg / L to 150 mg / L). In some embodiments, one or more individual cannabinoids (e.g., one or more of CBG, CBGA, CBD, CBDA, CBC, and CBCA) may be produced at a concentration of at least about 1 mg / L, 2 mg / L, 5 mg / L, 10 mg / L, 15 mg / L, 20 mg / L, 25 mg / L, 30 mg / L, 35 mg / L, 40 mg / L, 45 mg / L, 50 mg / L, 55 mg / L, 60 mg / L, 65 mg / L, 70 mg / L, 75 mg / L, 80 mg / L, 85 mg / L, 90 mg / L, 95 mg / L, 100 mg / L, or more.For example, in some embodiments, one or more individual cannabinoids may be produced at a concentration of about 1 mg / L to about 100 mg / L (e.g., 2 to 90 mg / L, 2 to 80 mg / L, 2 to 70 mg / L, 2 to 60 mg / L, 2 to 50 mg / L, 2 to 40 mg / L, 2 to 30 mg / L, 2 to 20 mg / L, 2 to 15 mg / L, 3 to 90 mg / L, 3 to 80 mg / L, 3 to 70 mg / L, 3 to 60 mg / L, 3 to 50 mg / L, 3 to 40 mg / L, 3 to 30 mg / L, 3 to 20 mg / L, 3 to 15 mg / L, 4 to 90 mg / L, 4 to 80 mg / L, 4 to 70 mg / L, 4 to 60 mg / L, 4 to 50 mg / L, 4 to 40 mg / L, 4 to 30 mg / L, 4 to 20 mg / L, or 4 to 15 mg / L).
[0179] The present invention is further illustrated by the following examples, which do not limit the scope of the present invention described in the claims. Example
[0180] Example 1 - General Methods
[0181] Enzymatic constructs
[0182] Each enzyme construct is designed to include an interaction domain (ID) comprising two tandem N-terminal or C-terminal ligand binding motifs separated from the given enzyme and from each other by an amino acid sequence containing a flexible, GS-rich linker flanked by rigid α-helical spacer sequences. The motifs comprising each enzyme ID are capable of specifically binding to tandem peptide ligands, which form ID binding sites at discrete positions along the synthetic intracellular polypeptide scaffold. Expression of each enzyme is controlled by a constitutive or inducible promoter. The nucleic acids encoding the enzymes can be codon-optimized, for example, for expression in yeast.
[0183] Scaffolded constructs
[0184] ID binding sites comprising tandem peptide ligands specific for the tandem scaffold binding motifs, comprising the ID of each enzyme, are inserted at discrete locations along the intracellular polypeptide scaffold.
[0185] The tandem ligands comprising each scaffolded ID binding site are separated from each other by a 36-amino acid residue sequence comprising a flexible, GS-rich linker flanked by a rigid α-helical spacer sequence, while the scaffolded ID binding sites themselves are separated from each other by a 50-amino acid residue (or any other number of amino acid residues) sequence comprising a flexible, GS-rich linker flanked by a rigid α-helical spacer sequence. Specifically, the scaffold binding site for each enzyme in the hexanoyl-CoA pathway is located proximal to the ATP citrate lyase and acetyl-CoA acetyltransferase at the N-terminus of the primary scaffold (in catalytic order). The scaffold binding site for each enzyme in the upper cannabinoid pathway is located proximal to (immediately downstream of) the binding site for the hexanoyl-CoA pathway enzyme. The scaffold binding site for each enzyme in the mevalonate (or MEP) pathway is located proximal to the ATP citrate lyase and acetyl-CoA acetyltransferase at the C-terminus of the primary scaffold (in catalytic order). The enzyme catalyzing the rate-limiting / critical step in cannabinoid biosynthesis (CBGA synthase, the final enzymatic step in the upper cannabinoid pathway) is located at the intersection of converging cannabinoid precursor pathways near the midpoint of the scaffold.
[0186] Assessment of cannabinoidergic potential by transient transfection
[0187] Competent yeast and / or green algae cells were transiently transfected with plasmids encoding various arrangements of scaffolds and enzymes. To establish baseline cannabinoid capacity, cells were first transiently transfected with the enzymes required for cannabinoid biosynthesis (but not the scaffolds), and the biosynthesized cannabinoids were then extracted, isolated, and quantified as described below (see "Cannabinoid Extraction, Isolation, and Analytical Characterization"). To measure the improvement in cannabinoid capacity conferred by multi-enzymatic scaffolding, a subset of the above cells were co-transfected with plasmids encoding one or more multi-enzymatic scaffolds described herein, and the biosynthesized cannabinoids were extracted, isolated, and quantified. The presence of plasmid DNA was confirmed by PCR, functional gene expression was confirmed by qRT-PCR, protein / polypeptide production was confirmed by Western blotting, and the scaffold of each enzyme was confirmed by co-immunoprecipitation of a C-terminal myc / flag-spiked scaffold followed by Western blot analysis of each co-immunoprecipitated enzyme.
[0188] Engineering of stable cannabinoid-producing cell lines
[0189] The construct can be integrated into the genome of a host cell (e.g., yeast, green algae, or other suitable host) by stable transfection. Gene integration is confirmed by PCR, functional gene expression is confirmed by qRT-PCR, and protein / polypeptide production is confirmed by Western blotting. Gene expression / protein synthesis is confirmed by comparing qRT-PCR and Western blotting results between samples with and without genetic engineering. To evaluate the improvement in cannabinoid capacity brought about by the multi-enzyme scaffold to stably engineered cannabinoid-producing cell lines, cannabinoid biosynthesis is compared between cells stimulated for enzyme expression without the scaffold and cells stimulated for enzyme and scaffold expression.
[0190] Validation of multi-enzyme scaffolding
[0191] To verify the success of the multi-enzyme scaffold in transient transfection and stable engineered cells, a myc tag (or other immunoprecipitation tag) was inserted at the N-terminus or C-terminus of the polypeptide scaffold. The scaffolded enzymes were selectively co-immunoprecipitated by affinity chromatography using anti-myc affinity beads. Western blotting was performed to detect and quantify each co-immunoprecipitated enzyme.
[0192] Aerobic batch fermentation
[0193] Stably engineered cannabinoid-producing yeast, green algae or other host cells are grown in a bioreactor (or any other vessel) by aerobic batch fermentation (or any other culture technique).
[0194] Extraction, separation and analytical characterization of cannabinoids
[0195] After fully initiating cannabinoid biosynthesis, the engineered yeast / green algae cells are precipitated by centrifugation and washed with TBS. The supernatant (liquid culture medium) is poured out and collected. After washing with TBS, the precipitated cells are resuspended in ethanol adjusted with NaOH and lysed by repeated freezing and thawing and ultrasonic treatment. The biosynthesized cannabinoid fermentation product is then harvested from the lysate and supernatant by performing a triple non-polar extraction using hexane and then ethyl acetate. The resulting organic fractions are combined and rotary evaporated. The biosynthesized cannabinoids are then quantitatively and qualitatively measured using high performance liquid chromatography with a diode array detector (HPLC-DAD) or gas chromatography-flame ionization detector (GC-FID).
[0196] In the following examples, each 48-hour culture was lysed / homogenized by ultrasonication. The ultrasonicated samples were then subjected to triple liquid-liquid extraction with ethyl acetate (one volume equivalent of ethyl acetate was extracted each time). After separation, the ethyl acetate fractions collected from each sample were combined and the combined samples were centrifuged. The ethyl acetate was then removed from each sample in a vacuum oven and the residual sample was resuspended in 10 mL of methanol for analytical characterization. The analytical characterization of all samples was performed by a licensed independent third-party analytical testing agency (Precision Plant Molecules, Denver, Colorado). HPLC-DAD was used to quantitatively and qualitatively measure each parent and derivative cannabinoid and the cannabinoid precursor OVA.
[0197] Example 2- Synthetic gene cassette assembly / synthesis, plasmid preparation, and polycistronic vector construction
[0198] Five synthetic gene cassettes (named HCA, GPP, CAN, SCF, and SOL) were constructed for the biosynthesis of cannabinoids in heterologous cells or cell-free reaction buffers. Figure 2B These cassettes collectively encode all scaffold-binding engineered enzymes and the polypeptide scaffolds to which the engineered enzymes can bind.
[0199] The HCA gene cassette encodes scaffold-bound engineered enzymes for scaffolded hexanoyl-CoA biosynthesis, namely ACL, atoB, BHBD, ECH, ECR, and bktB, and a soluble HCS for the production of additional hexanoyl-CoA from hexanoate-supplemented culture medium or cell-free reaction buffer. Figure 13A The GPP gene cassette encodes scaffold-binding engineered enzymes for scaffolded geranyl pyrophosphate (GPP) biosynthesis, namely HMGS, tHMGR, ERG12, ERG8, MVD1, IDI1, and ERG20 WW See also Figure 13B The CAN gene cassette encodes the scaffold-binding engineered enzymes for the biosynthesis of scaffolded OAC, malonyl-CoA, and CBGA, i.e., OS and OAC, ACC, and CBGAS, respectively, and all enzymes for the biosynthesis of soluble (non-scaffolded) CBDA and CBCA, i.e., CBDAS and CBCAS, respectively. Figure 13C The SCF gene cassette encodes a polypeptide scaffold for bidirectional scaffolded cannabinoid biosynthesis and scaffolded malonyl-CoA biosynthesis, namely, the cannabinoid metabolic scaffold (CBSCF) and the malonyl-CoA metabolic scaffold (MCASCF), respectively, as well as additional copies of ACL and atoB to enhance acetyl-CoA biosynthesis from supplemented and / or endogenous citrate and acetoacetyl-CoA biosynthesis from acetyl-CoA, respectively. Figure 13DThe SOL gene cassette lacks the polypeptide scaffold for bidirectional scaffolded cannabinoid biosynthesis and scaffolded malonyl-CoA biosynthesis (i.e., it is used for soluble cannabinoid biosynthesis), but, like the SCF gene cassette, encodes additional copies of ACL and atoB to enhance acetyl-CoA biosynthesis from supplemented and / or endogenous citrate and acetoacetyl-CoA biosynthesis from acetyl-CoA. The amino acid sequences of the engineered enzymes ACL and atoB are shown in Figure 13A .
[0200] Gene cassettes were assembled / synthesized using self-cleaving 2A peptides (P2A) to link multiple codon-optimized (for S. cerevisiae) gene sequences assigned to each cassette. To improve P2A cleavage, a GSG linker (comprising a single serine residue flanked by a single glycine residue) was inserted at the interface between each constituent gene sequence and the P2A linker sequence fused thereto (having the format: gene cassette sequence 1–SG–P2A linker–gene cassette sequence 2–GSG–P2A linker–gene cassette sequence 3–GSG–P2A linker-), and so on. Codon-optimized nucleic acid sequences encoding engineered enzymes and scaffolds are described in [ 1 ]. Figures 14A-14D After assembly, each synthetic gene cassette was inserted into the pCCI-Brick plasmid, resulting in plasmids named pHCA, pGPP, pCAN, pSCF, and pSOL as described in Table 3. The complete gene cassettes inserted into the plasmids are shown in Table 3. Figures 15A-15E Each of these plasmids was then used to amplify each synthetic gene cassette by standard plasmid preparation. Plasmid DNA encoding each complete synthetic gene cassette was cloned into the SpeI / XhoI cloning sites of a polycistronic yeast auxotrophic selection vector, generating vectors designated vHCA, vGPP, vCAN, vSCF, and vSOL, as described in Table 3, to allow iterative antibiotic / auxotrophic selection of only those cells that were transformants of one or more of these polycistronic vectors.
[0201] Table 3
[0202]
[0203]
[0204]
[0205] The genes assigned to each synthetic gene cassette and the plasmids and vectors into which each synthetic gene cassette was inserted are listed in Table 3 , and the amino acid sequences encoded by each synthetic gene cassette are provided in Figures 13A-13D The codon-optimized nucleotide sequence fragments comprising each synthetic gene cassette are detailed in Figures 14A-14D The complete nucleotide sequence of each fully assembled synthetic gene cassette (the complete insert sequence of each plasmid and expression vector) is provided in Figures 15A-15EThe general map of the pCCI-Brick plasmid is shown in Figure 16 In, and a general map of the polycistronic yeast auxotrophic selection vector is shown in Figure 17 middle.
[0206] Example 3 - Engineering of cannabinoid-producing cells
[0207] To engineer a novel heterologous pathway for the biosynthesis of cannabinoids from citrate, and to evaluate the impact of bidirectional multi-enzyme scaffolding thereon, competent Saccharomyces cerevisiae cells were sequentially / iteratively transformed with, and auxotrophically selected for, expression of vHCA, vGPP, vCAN, and either vSCF (for scaffolded cannabinoid biosynthesis) or vSOL (for non-scaffolded / soluble cannabinoid biosynthesis) constructs.
[0208] All vector transformations and auxotrophic selection procedures were performed as follows: An aliquot of an overnight S. cerevisiae culture was inoculated into 100 mL of YPD medium (10 g / L yeast nitrogen base, 20 g / L peptone, and 20 g / L D-(+)-glucose) to an OD of 600nm = 0.3 (stationary phase), and grown in an orbital shaker at 30°C and 225 RPM to OD 600nm =1.6. Then by centrifuging 3 minutes with 3000xg and then aspirating culture medium to harvest cells. The cell pellet of harvest is washed 2 times with 50mL cooling nuclease-free water subsequently, and with 50mL cooling electroporation buffer (1M sorbitol / 1mMCaCl ) wash 1 time. The washed cell is conditioned by hatching 30 minutes in 20Ml 0.1MLiAc / 10mM DTT in the orbital shaker of 30 DEG C and 225RPM, harvest, wash 1 time with 50mL electroporation buffer, harvest, and be resuspended in 100 μ L electroporation buffer. By electroporation at 2.5kV and 25μF, resuspended cells are transformed with a certain amount of carrier containing 3 μ g target DNA insert (using the carrier-insert ratio calculation of each carrier). 8 mL of YPD medium containing 1 M sorbitol was then added to the electroporated cell suspension and the resulting suspension was incubated in an orbital shaker at 30° C. and 225 RPM for 1 hour. To isolate the desired transformants by auxotrophic selection, cells were harvested, resuspended in the appropriate yeast nitrogen base (YNB) withdrawal (selection) medium, transferred to baffled culture bottles, and incubated overnight in an orbital shaker at 30° C. and 225 RPM as described subsequently for each iterative transformation step. Transformation and selection protocols were used sequentially for each assigned vector.
[0209] Using the above method, the initial culture cells of electrocompetent Saccharomyces cerevisiae are first transformed with vHCA, which encodes the scaffold-binding engineered enzyme required for the biosynthesis of HCA from citric acid. Cells transformed with vHCA (named yHCA) are selected by resuspending and incubating in tryptophan-deficient YNB medium. The selected yHCA cells (i.e., cells grown in tryptophan-deficient YNB medium) are then transformed with vGPP, which encodes the scaffold-binding engineered enzyme required for the biosynthesis of GPP from citric acid. Cells co-transformed with vHCA and vGPP (called yHCAGPP) are selected by resuspending and incubating in tryptophan- and leucine-deficient YNB medium. Selected yHCAGPP cells (i.e., cells grown in tryptophan- and leucine-deficient YNB medium) were then transformed with vCAN, which encodes the scaffold-bound engineered enzymes required for the biosynthesis of malonyl-CoA from citrate, the biosynthesis of olivate from HCA and malonyl-CoA, the biosynthesis of OVA (olivate) from olivate, and the biosynthesis of CBGA from OVA and GPP, as well as the soluble enzymes required for the biosynthesis of CBDA and CBCA from CBGA. Cells co-transformed with vHCA, vGPP, and vCAN (designated yCB 亲本 ) were selected by resuspension and incubation in tryptophan-, leucine-, and histidine-deficient YNB medium.
[0210] Then yCB containing cells grown in tryptophan, leucine and histidine deficient YNB medium 亲本 The culture was split into two separate cultures. The first split yCB 亲代 Cultures were transformed with vSCF, which encodes CBSCF (cannabinoid metabolic scaffold) and MCASCF (malonyl-CoA metabolic scaffold) as well as additional copies of ACL and atoB. Cells co-transformed with vHCA, vGPP, vCAN, and vSCF (designated yCB) were selected by resuspension and incubation in YNB medium deficient in tryptophan, leucine, histidine, and uracil. SCF ). The second separated yCB 亲本 Cultures were transformed with vSOL, which encodes additional copies of ACL and atoB but lacks CBSCF and MCASCF. Cells co-transformed with vHCA, vGPP, vCAN, and vSOL (designated yCB) were also selected by resuspending and incubating in YNB medium deficient in tryptophan, leucine, histidine, and uracil. SOL ).
[0211] To quantify the improvement in cannabinoid-producing capacity conferred by multi-enzyme scaffolding, triplicate yCBs were cultured in 100 mL YPD medium at 30 °C and 400 RPM in an incubator shaker for 48 h. SOLand yCB SCF Cannabinoid titers were compared between cultures. SOL and yCB SCF The proliferation rate of each culture was initially diluted to OD 600nm =0.3, and OD was recorded at 12-hour intervals thereafter 600nm Measurement results. Figure 18 Proliferation curves are depicted in Figure 2. Additional sum-of-squares F tests indicated that there were no significant differences in the proliferation curves of yCBSCF and yCBSOL cultures for any parameter over the 48-hour incubation period, indicating that the scaffold did not affect cell proliferation.
[0212] Total cannabinoid titers, parent (carboxylated) cannabinoids (CBGA, CBDA, and CBCA) titers, derivative (decarboxylated) cannabinoids (CBG, CBD, and CBC) titers, and procannabinoid (OVA) titers were measured. As shown in Figure 19, mixed ANOVA was performed to test the difference between strains (F 1,4 =943.8; p<0.0001) and analyte (cannabinoid and procannabinoid) titers (F 10,40 =216.4; p<0.0001), and a significant strain x analyte interaction (F 10,40 =131.4; p<0.0001). yCBSCF cultures showed increased titers of total cannabinoids (p<0.0001), OVA precursor (p<0.0001), CBG(A) (p<0.0001), CBD(A) (p<0.0001), CBC(A) (p<0.0001), CBGA (p<0.0001), CBDA (p<0.0001), CBCA (p<0.0001), CBG (p<0.0001), CBD (p<0.01), and CBC (p<0.001) relative to yCBSOL cultures.
[0213] Example 4 - Effects of Citrate and Hexanoate Supplementation on Scaffolding and Soluble Cannabinoid Biosynthesis
[0214] To evaluate the effect of supplementation of the medium with citrate and hexanoate precursors, triplicate yCBs were grown in 100 mL YPD medium containing 300 mg / L buffered citrate (pH 6.0) or hexanoate for 48 h at 30°C and 400 RPM on an orbital shaker. SOL and yCB SCF Cannabinoid titers were compared in the cultures. All cultures were initially diluted to an OD 600nm = 0.3. The cannabinoid titers of cultures grown in YPD medium, YPD medium supplemented with citrate, and YPD medium supplemented with hexanoic acid were evaluated and analyzed by ANOVA. Figure 20As shown, mixed ANOVA tested the differences between strains (F 1,4 =457.5; p<0.0001) and medium supplements (F 2,8 = 312.5; p < 0.0001), and a significant strain x medium supplement interaction (F 2,8 =289.6; p<0.0001). yCBSCF, but not yCBSOL, cultures showed increased total cannabinoid titers when cultured in medium supplemented with 300 mg / L citrate compared to basal medium cultures (p<0.0001). When cultured in medium supplemented with 300 mg / L hexanoate, the total cannabinoid titers of both yCBSCF and yCBSOL cultures were not different relative to basal medium. For all measurements, n=3 biological replicates for yCBSCF and yCBSOL cultures. In addition, yCBSCF cultures showed increased total cannabinoid titers relative to yCBSOL cultures when cultured in basal medium (p<0.0001, data also reported in Figure 19) and medium supplemented with 300 mg / L citrate (p<0.0001) and hexanoate (p<0.0001).
[0215] To characterize the concentration-response relationship of citrate-supplemented media, triplicate yCBs grown in 100 mL YPD medium containing 0, 10, 30, 100, 300, 1000, 3000, and 10,000 mg / L buffered citrate (pH 6.0) at 30°C and a 400 RPM orbital shaker for 48 h were compared. SOL and yCB SCF Cannabinoid titers between cultures. All cultures were initially diluted to OD 600nm = 0.3. After quantification, an asymmetric sigmoidal (five-parameter) logistic regression was calculated to fit the concentration-response curve, from which yCB was derived. SOL With yCB SCF The maximum cannabinoid titer in culture (CB 最大 ) estimated values and citrate EC for cannabinoid biosynthesis 50 Concentration-response curve, CB 最大 Estimates and citrate EC50 estimates are as follows Figure 21 As shown. Mixed ANOVA detected strains (F 1,8 =69.9; p<0.0001) and parameters (F 1,8 = 66.7; p < 0.0001), and a significant strain x parameter interaction (F 1,8=5.3; p<0.05) for concentration-response parameter estimates (CBmax and citrate EC50). yCBSCF cultures showed significantly increased CBmax (p<0.0001) and citrate EC50 (p<0.001) estimates compared to yCBSOL cultures.
[0216] Other Implementations
[0217] It should be understood that although the invention has been described in conjunction with the detailed description, the above description is intended to illustrate rather than limit the scope of the invention, which is defined by the appended claims. Other aspects, advantages and improvements are also within the scope of the following claims.
Claims
1. A host cell capable of producing one or more cannabinoids, the host cell comprising: (a) a first exogenous nucleic acid encoding a first polypeptide having CBGA synthase activity and comprising a first heterologous interaction domain, (b) a second exogenous nucleic acid encoding a second polypeptide having oliverate cyclase activity and comprising a second heterologous interaction domain, (c) a third exogenous nucleic acid encoding a third polypeptide having olivetol synthase activity and comprising a third heterologous interaction domain, (d) a fourth exogenous nucleic acid encoding a fourth polypeptide having trans-2-enoyl-CoA reductase activity and comprising a fourth heterologous interaction domain, (e) a fifth exogenous nucleic acid encoding a fifth polypeptide having enoyl-CoA hydratase activity and comprising a fifth heterologous interaction domain, (f) a sixth exogenous nucleic acid encoding a sixth polypeptide having 3-hydroxybutyryl-CoA dehydrogenase activity and comprising a sixth heterologous interaction domain, (g) a seventh exogenous nucleic acid encoding a seventh polypeptide having acetyl-CoA acetyltransferase activity and comprising a seventh heterologous interaction domain, (h) an eighth exogenous nucleic acid encoding an eighth polypeptide having ATP citrate lyase activity and comprising an eighth heterologous interaction domain, (i) a ninth exogenous nucleic acid encoding a ninth polypeptide having geranyl diphosphate synthase activity and comprising a ninth heterologous interaction domain, (j) a tenth exogenous nucleic acid encoding a tenth polypeptide having isopentenyl diphosphate isomerase activity and comprising a tenth heterologous interaction domain, (k) an eleventh exogenous nucleic acid encoding an eleventh polypeptide having diphosphomevalonate decarboxylase activity and comprising an eleventh heterologous interaction domain, (1) a twelfth exogenous nucleic acid encoding a twelfth polypeptide having phosphomevalonate kinase activity and comprising a twelfth heterologous interaction domain, (m) a thirteenth exogenous nucleic acid encoding a thirteenth polypeptide having mevalonate kinase activity and comprising a thirteenth heterologous interaction domain, (n) a fourteenth exogenous nucleic acid encoding a fourteenth polypeptide having HMG-CoA reductase activity and comprising a fourteenth heterologous interaction domain, (o) a fifteenth exogenous nucleic acid encoding a fifteenth polypeptide having HMG-CoA synthase activity and comprising a fifteenth heterologous interaction domain, and (p) a sixteenth exogenous nucleic acid encoding a polypeptide scaffold comprising a peptide ligand for each of the first to fifteenth heterologous interaction domains, wherein each of the first to fifteenth heterologous interaction domains are different, wherein the peptide ligand for each of the first to fifteenth heterologous interaction domains is different, wherein, in the order of extending toward the first direction from the peptide ligand for the first heterologous interaction domain, the polypeptide scaffold comprises (1) the peptide ligand for the second heterologous interaction domain, (2) the peptide ligand for the third heterologous interaction domain, (3) the peptide ligand for the fourth heterologous interaction domain, (4) the peptide ligand for the fifth heterologous interaction domain, (5) the peptide ligand for the sixth heterologous interaction domain, (6) the peptide ligand for the seventh heterologous interaction domain, and (7) the peptide ligand for the eighth heterologous interaction domain, Wherein, in the order starting from the peptide ligand for the first heterologous interaction domain and extending toward the other direction, the polypeptide scaffold comprises (1) the peptide ligand for the ninth heterologous interaction domain, (2) the peptide ligand for the tenth heterologous interaction domain, (3) the peptide ligand for the eleventh heterologous interaction domain, (4) the peptide ligand for the twelfth heterologous interaction domain, (5) the peptide ligand for the thirteenth heterologous interaction domain, (6) the peptide ligand for the fourteenth heterologous interaction domain, (7) the peptide ligand for the fifteenth heterologous interaction domain, (8) the peptide ligand for the seventh heterologous interaction domain, and (9) the peptide ligand for the eighth heterologous interaction domain.
2. The host cell of claim 1 , wherein the host cell further comprises (q) a seventeenth exogenous nucleic acid encoding an acetyl-CoA carboxylase and comprising a seventeenth heterologous interaction domain, and (r) an eighteenth exogenous nucleic acid encoding a polypeptide scaffold comprising a peptide ligand to each of the eighth and seventeenth heterologous interaction domains.
3. The host cell of claim 1 , wherein the host cell further comprises exogenous nucleic acids encoding cannabidiolic acid synthase and cannabichromenic acid synthase.
4. The host cell of claim 1, wherein the host cell further comprises exogenous cannabidiolic acid synthase.
5. The host cell of claim 1, wherein the host cell further comprises exogenous cannabichromenic acid synthase.
6. The host cell of claim 1, wherein the host cell is a bacterial or yeast host cell.
7. The host cell of claim 1, wherein the bacterial cell is selected from the group consisting of Escherichia coli, Bacillus, Brevibacterium, Streptomyces and Pseudomonas cells.
8. The host cell of claim 1, wherein the yeast cell is selected from the group consisting of Pichia pastoris, Saccharomyces cerevisiae, Yarrowia lipolytica, Kluyveromyces marxianus, and Komagataella phaffii cells.
9. The host cell of claim 1, wherein the host cell is an algae or plant cell.
10. The host cell of claim 9, wherein the algae is a Dunaliella sp., Chlorella variabilis, Euglena mutabilis, or Chlamydomonas reinhardtii cell.
11. The host cell of claim 9, wherein the plant cell is a Cannabis or tobacco cell.
12. wherein each of said polypeptides has the following formula: enzyme-linker 1-spacer-linker 2-motif 1-linker 3-motif 2, wherein linker 1, linker 2 and linker 3 are the same or different, wherein motif 1 and motif 2 are the same or different, and wherein motif 1 and motif 2 form said heterologous interaction domain.
13. The host cell of claim 12, wherein the scaffold polypeptide comprises a linker between adjacent peptide ligands. The host cell of claim 13 , wherein the scaffold polypeptide is tagged with a MYC tag, a FLAG tag, or a HA tag.
15. The host cell of claim 12, wherein the linker is a flexible GS-rich sequence flanked by rigid α-helical portions.
16. The host cell of claim 12, wherein the spacer is a cTPR6 spacer.
17. The host cell of claim 1, wherein a constitutive promoter is operably linked to one or more of the exogenous nucleic acids encoding the polypeptide, or is operably linked to the sixteenth exogenous nucleic acid encoding the polypeptide scaffold.
18. The host cell of claim 1, wherein a first constitutive promoter is operably linked to one or more of the exogenous nucleic acids encoding the polypeptide, and a second constitutive promoter is operably linked to the sixteenth exogenous nucleic acid encoding the polypeptide scaffold.
19. The host cell of claim 18, wherein the constitutive activity level of the constitutive promoter used to express the polypeptide scaffold is weaker than that of the constitutive promoter used to express the polypeptide.
20. The host cell of claim 1, wherein each of the exogenous nucleic acids comprises an inducible promoter operably linked to a sequence encoding the polypeptide or the polypeptide scaffold.
21. The host cell of claim 20, wherein the promoter is the GAL1-10 promoter.
22. A method of producing one or more cannabinoids, the method comprising culturing a host cell under conditions wherein the host cell produces the one or more cannabinoids, wherein the host cell comprises: (a) a first exogenous nucleic acid encoding a first polypeptide having CBGA synthase activity and comprising a first heterologous interaction domain, (b) a second exogenous nucleic acid encoding a second polypeptide having oliverate cyclase activity and comprising a second heterologous interaction domain, (c) a third exogenous nucleic acid encoding a third polypeptide having olivetol synthase activity and comprising a third heterologous interaction domain, (d) a fourth exogenous nucleic acid encoding a fourth polypeptide having trans-2-enoyl-CoA reductase activity and comprising a fourth heterologous interaction domain, (e) a fifth exogenous nucleic acid encoding a fifth polypeptide having enoyl-CoA hydratase activity and comprising a fifth heterologous interaction domain, (f) a sixth exogenous nucleic acid encoding a sixth polypeptide having 3-hydroxybutyryl-CoA dehydrogenase activity and comprising a sixth heterologous interaction domain, (g) a seventh exogenous nucleic acid encoding a seventh polypeptide having acetyl-CoA acetyltransferase activity and comprising a seventh heterologous interaction domain, (h) an eighth exogenous nucleic acid encoding an eighth polypeptide having ATP citrate lyase activity and comprising an eighth heterologous interaction domain, (i) a ninth exogenous nucleic acid encoding a ninth polypeptide having geranyl diphosphate synthase activity and comprising a ninth heterologous interaction domain, (j) a tenth exogenous nucleic acid encoding a tenth polypeptide having isopentenyl diphosphate isomerase activity and comprising a tenth heterologous interaction domain, (k) an eleventh exogenous nucleic acid encoding an eleventh polypeptide having diphosphomevalonate decarboxylase activity and comprising an eleventh heterologous interaction domain, (1) a twelfth exogenous nucleic acid encoding a twelfth polypeptide having phosphomevalonate kinase activity and comprising a twelfth heterologous interaction domain, (m) a thirteenth exogenous nucleic acid encoding a thirteenth polypeptide having mevalonate kinase activity and comprising a thirteenth heterologous interaction domain, (n) a fourteenth exogenous nucleic acid encoding a fourteenth polypeptide having HMG-CoA reductase activity and comprising a fourteenth heterologous interaction domain, (o) a fifteenth exogenous nucleic acid encoding a fifteenth polypeptide having HMG-CoA synthase activity and comprising a fifteenth heterologous interaction domain, and (p) a sixteenth exogenous nucleic acid encoding a polypeptide scaffold comprising a peptide ligand for each of the first to fifteenth heterologous interaction domains, wherein each of the first to fifteenth heterologous interaction domains are different, wherein the peptide ligand for each of the first to fifteenth heterologous interaction domains is different, wherein, in the order of extending toward the first direction from the peptide ligand for the first heterologous interaction domain, the polypeptide scaffold comprises (1) the peptide ligand for the second heterologous interaction domain, (2) the peptide ligand for the third heterologous interaction domain, (3) the peptide ligand for the fourth heterologous interaction domain, (4) the peptide ligand for the fifth heterologous interaction domain, (5) the peptide ligand for the sixth heterologous interaction domain, (6) the peptide ligand for the seventh heterologous interaction domain, and (7) the peptide ligand for the eighth heterologous interaction domain, Wherein, in the order starting from the peptide ligand for the first heterologous interaction domain and extending toward the other direction, the polypeptide scaffold comprises (1) the peptide ligand for the ninth heterologous interaction domain, (2) the peptide ligand for the tenth heterologous interaction domain, (3) the peptide ligand for the eleventh heterologous interaction domain, (4) the peptide ligand for the twelfth heterologous interaction domain, (5) the peptide ligand for the thirteenth heterologous interaction domain, (6) the peptide ligand for the fourteenth heterologous interaction domain, (7) the peptide ligand for the fifteenth heterologous interaction domain, (8) the peptide ligand for the seventh heterologous interaction domain, and (9) the peptide ligand for the eighth heterologous interaction domain.
23. The method of claim 22, wherein the host is cultured in a medium supplemented with buffered citrate, glucose, hexanoic acid and / or other carbon sources.
24. The method of claim 22, wherein the host is cultured in a medium supplemented with malonyl-CoA.
25. The method of claim 22, wherein the host is cultured in a medium supplemented with buffered citrate.
26. The method of claim 22, further comprising extracting the one or more cannabinoids from the host cell.
Citation Information
Patent Citations
TAL effector-mediated DNA modification
US20110145940A1
Aromatic Prenyltransferase from Cannabis
US20120144523A1
Aromatic prenyltransferase from Cannabis
US8884100B2
Nucleic acids encoding synthetic scaffolds and host cells genetically modified with the nucleic acids
US9856460B2
Cannabichromenic acid synthase from cannabis sativa
WO2015196275A1