Combination of nucleic acid sequences encoding proteins derived from Helichrysum umbraculigerum, and transgenic cells, tissues and organisms containing the same
By employing nucleic acid sequences from Helichrysum umbraculigerum to produce cannabinoids in transgenic cells and plants, the challenges of high costs and yield limitations in traditional plant cultivation are addressed, enabling efficient and stable cannabinoid production.
Patent Information
- Application Number
- JP2025514322
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-19
- Filing Date
- 2023-09-07
- Publication Date
- 2025-09-11
AI Technical Summary
The high costs and challenges associated with growing and maintaining large plants, as well as the difficulty in obtaining high yields of cannabinoids, hinder their therapeutic use due to their sensitivity to oxidation and light.
The use of nucleic acid sequences encoding enzymes from Helichrysum umbraculigerum, such as acyl-activating enzymes, polyketide synthases, polyketide cyclases, prenyltransferases, and cannabichromenic acid synthases, to produce cannabinoids in transgenic cells and plants, enabling large-scale production.
This approach allows for the efficient synthesis of cannabinoids, overcoming the limitations of traditional plant-based methods by providing a cost-effective and stable production process.
Smart Images

Figure 2025530214000001_ABST
Abstract
Description
[Technical Field]
[0001] Reference to electronic sequence list The contents of the electronic sequence listing (YEDA-P-010-PCT ST26.xml; size: 251,312 bytes; created on August 20, 2023) are incorporated herein by reference in their entirety.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 404,645, filed September 8, 2022, entitled "Combination of Nucleic Acid Sequences Encoding Proteins Derived from Helichrysum umbraculigerum, and Transgenic Cells, Tissues, and Organisms Comprising the Same," and U.S. Provisional Patent Application No. 63 / 453,112, filed March 19, 2023, entitled "Combination of Nucleic Acid Sequences Encoding Proteins Derived from Helichrysum umbraculigerum, and Transgenic Cells, Tissues, and Organisms Comprising the Same," the contents of both applications are incorporated herein by reference in their entireties.
[0003] FIELD OF THE INVENTION The present invention relates to a combination of enzymes derived from Helichrysum umbraculigerum, including polynucleotides encoding the enzymes from Helichrysum umbraculigerum, and methods of using same, such as to produce cannabinoids. [Background technology]
[0004] Cannabinoids are terpenophenolic compounds found in Cannabis sativa, an annual plant belonging to the Cannabaceae family. This plant contains over 400 chemical compounds and approximately 70 cannabinoids, the latter of which accumulate primarily in glandular trichomes. Naturally occurring cannabinoids, such as tetrahydrocannabinol (THC), have been used to treat a wide range of medical conditions, including glaucoma, AIDS wasting, neuropathic pain, the treatment of spasticity associated with multiple sclerosis, fibromyalgia, and chemotherapy-induced nausea. THC is also effective in treating allergies, inflammation, infections, epilepsy, depression, migraines, bipolar disorder, anxiety disorders, drug dependence, and drug withdrawal syndrome.
[0005] Additional active cannabinoids include cannabidiol (CBD), an isomer of THC and a potent antioxidant and anti-inflammatory compound known to protect against acute and chronic neurodegeneration; cannabigerol (CBG), found in high concentrations in hemp and acting as a high-affinity α2-adrenergic receptor agonist, a medium-affinity 5-HT1A receptor antagonist, and a low-affinity CB1 receptor antagonist, potentially possessing antidepressant properties; and cannabichromene (CBC), which possesses anti-inflammatory, antifungal, and antiviral properties. Many plant cannabinoids have therapeutic potential in various diseases and may play important roles in plant defense and pharmacology. Therefore, biotechnological production of cannabinoids and cannabinoid-like compounds with therapeutic properties is of paramount importance. Therefore, cannabinoids are considered promising pharmaceutical agents with beneficial effects in the treatment of various diseases.
[0006] Despite the known beneficial effects of cannabinoids, their therapeutic use is hindered by the high costs associated with growing and maintaining large plants and the difficulty of obtaining high yields of cannabinoids. Extraction, isolation, and purification of cannabinoids from plant tissues are particularly challenging because cannabinoids are easily oxidized and sensitive to light and heat.
[0007] Therefore, there is a need to develop methodologies that allow for the large-scale production of cannabinoids for therapeutic use. Summary of the Invention
[0008] According to a first aspect, there is provided an isolated DNA molecule comprising at least a first nucleic acid sequence encoding a first protein and at least a second nucleic acid sequence encoding a second protein, wherein the first protein and the second protein are derived from Helichrysum umbraculigerum and belong to an enzyme family selected from the group consisting of acyl-activating enzymes (AAEs), polyketide synthases (PKSs), polyketide cyclases (PKCs), prenyltransferases (PTs), and cannabichromenic acid synthases (CBCASs), and the first protein and the second protein belong to different enzyme families.
[0009] According to another aspect, there is provided an artificial nucleic acid molecule comprising an isolated DNA molecule disclosed herein.
[0010] According to another aspect, there is provided a plasmid or agrobacterium comprising an artificial nucleic acid molecule disclosed herein.
[0011] According to another aspect, there is provided a transgenic cell comprising: (a) an isolated DNA molecule of the present invention; (b) an artificial nucleic acid molecule disclosed herein; (c) a plasmid or Agrobacterium disclosed herein; or (d) any combination of (a)-(c).
[0012] According to another aspect, there is provided an extract derived from the transgenic cells disclosed herein, or any fraction thereof.
[0013] According to another aspect, there is provided a transgenic plant, transgenic plant tissue, or plant part comprising: (a) an isolated DNA molecule of the invention; (b) an artificial nucleic acid molecule disclosed herein; (c) a plasmid or Agrobacterium disclosed herein; (d) a transgenic cell disclosed herein; or (e) any combination of (a)-(d).
[0014] According to another aspect, there is provided a composition comprising: (a) an isolated DNA molecule of the invention; (b) an artificial nucleic acid disclosed herein; (c) a plasmid or Agrobacterium disclosed herein; (d) a transgenic cell disclosed herein; (e) an extract disclosed herein; (f) a transgenic plant tissue or plant part disclosed herein; or (g) any combination of (a)-(f) and an acceptable carrier.
[0015] According to another aspect, there is provided a method for synthesizing a cannabinoid, a precursor thereof, or a combination thereof, comprising the steps of: (a) providing a transgenic cell or a cell transfected with an isolated DNA molecule of the invention or an artificial nucleic acid molecule disclosed herein; and (b) culturing the transgenic cell or transfected cell of step (a) so that at least a first protein and a second protein encoded by the artificial nucleic acid molecule are expressed, thereby synthesizing the cannabinoid, a precursor thereof, or any combination thereof.
[0016] According to another aspect, there is provided an extract of a transgenic or transfected cell obtained according to the methods disclosed herein.
[0017] According to another aspect, there is provided a composition comprising the extract disclosed herein and an acceptable carrier.
[0018] In some embodiments, the isolated DNA molecule is derived from H. umbraculigerum and further comprises at least a third nucleic acid sequence encoding a third protein belonging to an enzyme family selected from the group consisting of AAE, PKS, PKC, PT, and CBCAS, wherein the first protein, the second protein, and the third protein belong to different enzyme families.
[0019] In some embodiments, the isolated DNA molecule is derived from H. umbraculigerum and further comprises at least a fourth nucleic acid sequence encoding a fourth protein belonging to an enzyme family selected from the group consisting of AAE, PKS, PKC, PT, and CBCAS, wherein the first protein, the second protein, the third protein, and the fourth protein belong to different enzyme families.
[0020] In some embodiments, the isolated DNA molecule is derived from H. umbraculigerum and further comprises at least a fifth nucleic acid sequence encoding a fifth protein belonging to an enzyme family selected from the group consisting of AAE, PKS, PKC, PT, and CBCAS, wherein the first protein, the second protein, the third protein, the fourth protein, and the fifth protein belong to different enzyme families.
[0021] In some embodiments, the isolated DNA is derived from H. umbraculigerum and further comprises a nucleic acid sequence encoding a protein belonging to an enzyme family selected from the group consisting of uridine diphosphate (UDP)-glycosyltransferases (UGTs), alcohol acyltransferases (AATs), and both.
[0022] In some embodiments, (a) the AAE is encoded by a nucleic acid sequence having at least 89% homology to any one of SEQ ID NOs: 1-11, and any combination thereof; (b) the PKS is encoded by a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 23-26, and any combination thereof; (c) the PKC is encoded by a nucleic acid sequence having at least 88% homology to any one of SEQ ID NOs: 31-38, and any combination thereof; (d) the PT is encoded by a nucleic acid sequence having at least 91% homology to any one of SEQ ID NOs: 47-58, and any combination thereof; (e) the CBCAS is encoded by a nucleic acid sequence having at least 82% homology to any one of SEQ ID NOs: 71-79, and any combination thereof; or (f) any combination of (a)-(e).
[0023] In some embodiments, (a) the UGT is encoded by a nucleic acid sequence having at least 87% homology to any one of SEQ ID NOs: 89-101, and any combination thereof; (b) the AAT is encoded by a nucleic acid sequence having at least 87% homology to any one of SEQ ID NOs: 115-129, and any combination thereof; or (c) both (a) and (b).
[0024] In some embodiments, (a) the AAE comprises an amino acid sequence at least 93% identical to any one of SEQ ID NOs: 12-22; (b) the PKS comprises an amino acid sequence at least 93% identical to any one of SEQ ID NOs: 27-30; (c) the PKC comprises an amino acid sequence at least 87% identical to any one of SEQ ID NOs: 39-46; (d) the PT comprises an amino acid sequence at least 92% identical to any one of SEQ ID NOs: 59-70; (e) the CBCAS comprises an amino acid sequence at least 86% identical to any one of SEQ ID NOs: 80-88; or (f) any combination of (a)-(e).
[0025] In some embodiments, (a) the UGT comprises an amino acid sequence having at least 90% homology to any one of SEQ ID NOs: 102-114; (b) the AAT comprises an amino acid sequence having at least 91% homology to any one of SEQ ID NOs: 130-144; or (c) both (a) and (b).
[0026] In some embodiments, (a) AAE consists of any one of the amino acid sequences of SEQ ID NOs: 12-22; (b) PKS consists of any one of the amino acid sequences of SEQ ID NOs: 27-30; (c) PKC consists of any one of the amino acid sequences of SEQ ID NOs: 39-46; (d) PT consists of any one of the amino acid sequences of SEQ ID NOs: 59-70; (e) CBCAS consists of any one of the amino acid sequences of SEQ ID NOs: 80-88; or (f) any combination of (a) to (e).
[0027] In some embodiments, (a) the UGT consists of the amino acid sequence of any one of SEQ ID NOs: 102-114; (b) the AAT consists of the amino acid sequence of any one of SEQ ID NOs: 130-144; or (c) both (a) and (b).
[0028] In some embodiments, the isolated DNA molecule comprises a plurality of isolated DNA molecule types.
[0029] In some embodiments, each type of the plurality of isolated DNA molecule types encodes one or more proteins belonging to different enzyme families.
[0030] In some embodiments, the transgenic cell is any one of a unicellular organism, a cell of a multicellular organism, and a cell in culture.
[0031] In some embodiments, the unicellular organism comprises a fungus or a bacterium.
[0032] In some embodiments, the fungus is a yeast cell.
[0033] In some embodiments, the transgenic cell is a transgenic Cannabis sativa cell.
[0034] In some embodiments, the extract comprises cannabinoids, precursors thereof, or combinations thereof.
[0035] In some embodiments, the precursor is selected from the group consisting of acyl-coenzyme A (CoA), polyketides, resorcinoid precursors, and any combination thereof.
[0036] In some embodiments, the acyl is a C1-C8 alkyl.
[0037] In some embodiments, the acyl-CoA is hexanoyl-CoA.
[0038] In some embodiments, the polyketide is a tetraketide.
[0039] In some embodiments, the tetraketide is a linear tetraketide.
[0040] In some embodiments, the resorcinoid precursor is olivetolic acid.
[0041] In some embodiments, the cannabinoid is cannabigerolic acid (CBGA), CBCA, or both.
[0042] In some embodiments, the artificial nucleic acid molecule is an expression vector.
[0043] In some embodiments, the transgenic or transfected cell is a prokaryotic or eukaryotic cell.
[0044] In some embodiments, the transgenic or transfected cell is a C. sativa cell.
[0045] In some embodiments, the method further comprises, prior to step (a), introducing or transfecting an artificial nucleic acid molecule into a cell, thereby obtaining a transgenic or transfected cell.
[0046] In some embodiments, the method further comprises extracting the transgenic or transfected cell, thereby obtaining an extract from the transgenic or transfected cell.
[0047] In some embodiments, the extract comprises cannabinoids, precursors thereof, or any combination thereof.
[0048] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in practicing or testing embodiments of the present invention, exemplary methods and / or materials are described below. In case of conflict, the patent specification, including definitions, will control. Furthermore, the materials, methods, and examples are illustrative only and are not intended to be necessarily limiting.
[0049] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description set forth hereinafter. It should be understood, however, that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description. [Brief explanation of the drawings]
[0050] [Figure 1A]Figures 1A-1I contain chemical structures, images, chromatograms, tables, and micrographs demonstrating that H. umbraculigerum biosynthesizes CBGA 1 and other terpenoids in all aerial parts of the plant. Figure 1A illustrates the proposed biosynthetic pathways for CBGA 1 and heliCBGA 2. [Figure 1B] Photographs of the inflorescence (top) and shoot (bottom) of an H. umbraculigerum plant. [Figure 1C] Total ion chromatogram of the ethanol extract of fresh H. umbraculigerum leaves. The most abundant peaks of identified metabolites are indicated and color-coded according to their terpenoid class. CBGA 1 and heliCBGA 2 are highlighted in red and blue, respectively. [Figure 1D] Figure 1 shows absolute quantification of CBGA1 in various plant tissues [% w / w per fresh weight, n=3; % w / w per dry weight (DW) for freeze-dried leaves, n=5]. Reported values for Cannabis have been added for comparison. [Figure 1E] FIG. 1 shows the chemical structures and names of selected terpenophenols with similar chemical formulas to 1-3. [Figure 1F] Representative cryo-SEM of the adaxial upper surface region of a leaf showing stalked glandular trichomes (indicated by arrows). [Figure 1G] FIG. 1 shows representative confocal micrographs of the adaxial upper surface region of a leaf showing stalked glandular trichomes (indicated by arrows). [Figure 1H] TEM micrograph showing the multicellular structure of various cell types in stalked glandular trichomes at the secretory stage. BC, basal cells; SC, stalk cells; NC, neck cells; DC, disc cells; SCv, secretory cavity. The dashed line indicates the surface of the SCv. [Figure 1I]High-magnification images show the ultrastructure of DCs. CW, cell wall; M, mitochondria; N, nucleus; P, plastid; PSP, periplasmic space; V, vacuole; Vs, vesicles. Arrows indicate active exocytotic secretion from vesicles into the periplasmic space.
[0051] [Figure 2A] Figures 2A-2E include fluorescence micrographs, graphs, and schematic diagrams showing that cannabinoid-related gene expression correlates with the accumulation of cannabinoid metabolites in the glandular trichomes of H. umbraculigerum. Figure 2A shows an optical image of a leaf cross-section showing that CBGA accumulates in the stalked glandular trichomes of the leaf. [Figure 2B] Figure 2B shows MALDI-MSI of m / z 361.23 ± 0.01 Da, showing that CBGA 1 accumulates in the stalked glandular trichomes of the leaves. The glandular trichomes in Figure 2A are marked for improved interpretation. The signal in Figure 2B corresponds to the protonated m / z of CBGA 1 and geranylphlorocaprophenone 4. [Figure 2C] Figure 1 shows the normalized enrichment score (NES) of each co-expressed module in each tissue. Module M4 is highlighted because it is highly expressed in trichomes and leaves. [Figure 2D] Figure 1 shows a spaghetti chart depicting the expression profile of module M4. The expression levels of individual genes are shown as gray lines. Colored lines highlight the expression of candidate genes from the pathway. [Figure 2E]This figure shows the genomic landscape of the eight longest scaffolds of the H. umbraculigerum assembly. Track i represents gene density; ii represents repetitive element density; iii represents 3' Tran-Seq coverage; and iv represents TrueSeq coverage. These metrics are calculated over non-overlapping windows of 0.1 Mb. Zooming in on the marked region of scaffold 1 reveals a tandem gene cluster containing seven PKSs. In this study, the enzymes HuPKS1-3 and HuTKS4 were cloned and functionally characterized.
[0052] [Figure 3A] Figures 3A-3F include heat maps, graphs, and tables illustrating the discovery of core cannabinoid biosynthetic pathway enzymes. Figure 3A shows gene expression in young leaves, roots, and trichomes of the putative enzymes characterized in this study [log(cpm+1), n=3]. The most active enzymes in this study are highlighted in pink. AAE, acyl-activating enzyme; PKS, type III polyketide synthase; PKC, polyketide cyclase; PT, prenyl-transferase. [Figure 3B] Figure 1 shows the products of recombinant enzyme assays of purified HuAAE proteins using various alkyl (single- and medium-chain FAs) and aromatic (cinnamic and coumaric) substrates. Peak areas were used for comparison (mean ± standard deviation; n = 3). CoAT, acyl-CoA-transferase; EV, empty. [Figure 3C] Figure 1 shows the products of coupled recombinant enzyme assays of either EV or Cannabis olivetolic acid cyclase (CsOAC) with HuPKS in the presence of hexanoyl-CoA and malonyl-CoA. PDAL, pentyl diacetic acid lactone; HTAL, hexanoyl triacetic acid lactone; OA92, olivetolic acid; PCP95, fluorocaprophenone. Peak areas were used for comparison (mean ± standard deviation; n = 3). OA92 and PCP95 were identified using analytical standards ([MH] = 223.097 Da). [Figure 3D]Activity assay of microsomal fractions expressing prenyltransferases (PTs) using a series of aromatic substrates and either geranylprophosphate (GPP) or isopentenyl pyrophosphate (IPP) as the isoprenoid donor. Circles represent monoprenylated or isoprenylated products observed in H. umbraculigerum or in vitro assays. VA, divarinolic acid; DHSA 93, dihyrostilbenic acid; ND, not detected; CBGAS, cannabigerolic acid synthase. [Figure 3E] Steady-state kinetic analysis of HuPT1, HuPT3, and HuCBGAS4 using OA 92 and GPP. Michaelis-Menten K values were calculated using various concentrations (0.5 µM–3 mM) and a constant concentration (1 mM) of each substrate (n=3). Literature K values for Cannabis sativa CsGOT4 were included for comparison. [Figure 3F] Phylogenetic analysis of PT proteins from H. umbraculigerum and other plants. Protein selection was based on functionally characterized enzymes, as described by de Bruijn et al. (2020). Phylogenetic groups with different substrates are indicated by colored circles. HuPT proteins are highlighted in red, while Cannabis and Rhododendron dauricum PTs that prenylate cannabinoids are highlighted in blue. H. umbraculigerum flowers and Cannabis leaves highlight active HuCBGA4 and CsGOT4, respectively. A complete list of protein IDs is available in Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023).
[0053] [Figure 4A]Figures 4A-4F include phylogenetic trees, heat maps, tables, chromatograms, and compound structures illustrating the functional characteristics of cannabinoid-modulating enzymes. Figure 4A shows a phylogenetic analysis of selected uridine diphosphate glycosyltransferase (UGT) proteins from H. umbraculigerum, Arabidopsis thaliana, rice (Oryza sativa), and stevia (Stevia rebaudiana). Phylogenetic groups are annotated according to the Arabidopsis thaliana UGT family classification (numbers in colored circles). HuUGT proteins are highlighted in red, while other proteins from non-cannabinoid-producing plant species previously shown to be capable of glycosylation of cannabinoids are highlighted in blue. A complete list of protein IDs is available in Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023). H. umbraculigerum flowers exhibit active HuCBGT1, HuCBGT6, and HuOAGT11. 4-Hydroxybenzoic acid (4-HBA) and 2,4-dihydroxybenzoic acid (2,4-DHBA), which are structurally similar to OA92 and CBGA1, are located next to the UGT enzymes that glycosylate them. The glycosylated hydroxyls are highlighted. [Figure 4B] Gene expression in young leaves, roots, and trichomes of putative UGT and alcohol acyltransferase (AAT) enzymes characterized in this study [log(cpm+1), n=3]. The most active enzymes in this study are highlighted in pink. [Figure 4C] Figure 1 shows a comparison of steady-state kinetics of HuOAGT11 and HuUGT13 with OsUGT and SrUGT using OA92 and uridine diphosphate glucose (UDP-Glc). Assays were performed using various concentrations (0.5 µM - 3 mM) and a constant concentration (1 mM) of each substrate (n = 3). [Figure 4D] Figure 12B shows extracted ion chromatograms of monoglucosides with theoretical m / z values after enzymatic assays using purified enzyme in the presence of UDP-Glc and a series of aromatic substrates (additional assays are shown in Figure 12B). One to three glucosylated compounds were observed for each substrate. Peaks were putatively assigned by MS / MS fragmentation patterns (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). Compounds naturally observed in H. umbraculigerum are marked with a green asterisk. Chromatograms were normalized to the highest value. [Figure 4E] Figure 1 shows extracted ion chromatograms of O-acylated cannabinoids after enzymatic assays using purified HuCoAT5 in the presence of various aromatic substrates as acyl donors and acceptors. The major ion products in each LC-MS / MS chromatogram were selected. A single peak was observed for each substrate pair. The detected analog peaks shifted in retention time depending on the change in hydrophobicity of the acyl group. Identification was performed according to MS / MS fragmentation and retention time (Figure 13 and Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). Compounds naturally observed in H. umbraculigerum are marked with a purple asterisk. Chromatograms were normalized to the highest value. [Figure 4F] Potential glycosylation sites and observed O-acylation sites were highlighted in blue and / or purple, respectively, in each chemical structure.
[0054] [Figure 5A]Figures 5A-5D include combination diagrams and graphs showing in vivo reconstitution of the core cannabinoid pathway in heterologous systems: Figure 5A shows the co-expression of various combinations of Cannabis-derived HuCoAT6, HuTKS4, and HuCBGAS4 with CsOAC and CsOLS in N. benthamiana leaves; [Figure 5B] FIG. 1 shows co-expression of various combinations of HuCoAT6, HuTKS4, and HuCBGAS4 from Cannabis with CsOAC and CsOLS in N. benthamiana leaves. [Figure 5C] FIG. 1 shows the co-expression of various combinations of Cannabis-derived HuCoAT6, HuTKS4, and HuCBGAS4 with CsOAC and CsOLS in S. cerevisiae yeast. [Figure 5D] Coexpression of various combinations of Cannabis-derived HuCoAT6, HuTKS4, and HuCBGAS4 with CsOAC and CsOLS in S. cerevisiae yeast. The gray, yellow, and green boxes on the left side of the graph indicate the biosynthetic genes included in the coexpression experiment, while the blue boxes indicate supplementation with geranyl pyrophosphate (GPP) and either (5A and 5C) sodium hexanoate (HexNa) or (5B and 5D) OA 92. Peak areas were used for comparison (mean ± standard deviation; n = 3–6). N. benthamiana primarily produced glycosylated products identified according to a previously performed in vitro UGT enzyme assay (Figure 4D and Figure 12B). All metabolites were identified by accurate mass, retention time and MS / MS spectra (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). EV, empty vector.
[0055] [Figure 6]This figure shows a schematic diagram illustrating the parallel and divergent evolution of the cannabinoid biosynthetic pathway. This diagram provides a side-by-side comparison of the cannabinoid biosynthetic pathways in H. umbraculigerum and cannabis. At the top, the phylogenetic relationships of Arabidopsis thaliana, tomato (Solanum lycopersicum), sunflower (Helianthus annuus), lettuce (Letuca sativa), hemp (Cannabis sativa), and Helicrysum umbraculigerum are shown, along with the evolutionary distance between Cannabis and Helicrysum umbraculigerum. This tree was constructed based on the entire proteome of each species using the Word-based software Prot-SpaM. In this study, hybrid but previously undescribed metabolites were produced by reacting naturally biosynthesized cannabinoids (green) with uridine diphosphate glucose (UDP-Glc) or acyl-CoA in the presence of HuCoAT5, HuCBGT1, or HuCBGT6 enzymes (blue) from H. umbraculigerum. AAE, acyl-activating enzyme; OLS, olivetol synthase; OAC, olivetolate cyclase; GOT, geranyl pyrophosphate:olivetolate geranyltransferase; CBDAS, cannabidiolic acid synthase; CBCAS, cannabichromenic acid synthase; THCAS, (-)-Δ9-trans-tetrahydrocannabinolic acid synthase; AAE, acyl-activating enzyme; PT, prenyltransferase; UGT, uridine diphosphate glycosyltransferase; AAT, alcohol acyltransferase. The active enzymes identified in this study have been given their names.CoAT, acyl-CoA-transferase; TKS, tetraketide synthase; PKC, polyketide cyclase; CBGAS, cannabigerolic acid synthase; OAGT, olivetolic acid UGT; CBGT, cannabinoid UGT; CBAT, cannabinoid acyl-transferase; BBE-like, berberine bridge enzyme-like; Cyc, cyclase; CYP, cytochrome P450.
[0056] [Figure 7A] Figures 7A-7B include chromatograms and compound structures showing LC-MS / MS fingerprinting of CBGA 1, heliCBGA 2, and APHA 3 in H. umbraculigerum. Figure 7A shows extracted ion chromatograms and MS / MS spectral matching of H. umbraculigerum leaf extract with cannabigerolic acid (CBGA 1 [MH] = 359.222 Da), heli-cannabigerolic acid (heliCBGA 2 [MH] = 393.206 Da), and pre-amorphous stilbole (APHA 3 [MH] = 391.191 Da) standards or authentic metabolites. To confirm the assignments, CBGA 1 and heliCBGA 2 were purified and analyzed by NMR. [Figure 7B] Stable isotope labeling of CBGA1, heliCBGA2, and APHA3 by feeding H. umbraculigerum leaves with hexane-D11 acid, phenylalanine-D5, or phenylalanine-C9. MS / MS spectra of the unlabeled and labeled forms show similar fragmentation patterns with mass shifts corresponding to the labeled portion of the molecule.
[0057] [Figure 8]Figures 8A-8J show photomicrographs and images depicting stalked glandular trichomes in H. umbraculigerum leaves and flowers. Figure 8A shows a representative cryo-SEM micrograph of a side view of a flower specimen showing stalked glandular trichomes (indicated by arrows). Figure 8B shows a representative cryo-SEM micrograph of a side view of a flower specimen showing stalked glandular trichomes (indicated by arrows). Figure 8C shows an optical micrograph depicting the biserial structure of stalked glandular trichomes in H. umbraculigerum leaves. Figure 8D shows selected TEM micrographs of trichomes in H. umbraculigerum leaves at various stages of secretion. Figure 8E shows selected TEM micrographs of trichomes in H. umbraculigerum leaves at various stages of secretion. Figure 8F shows selected TEM micrographs of H. umbraculigerum leaf trichomes at various stages of secretion. High-magnification images show the ultrastructure of disc cells (DCs). CW, cell wall; M, mitochondria; N, nucleus; P, plastid; PSP, periplasmic space; SCv, secretory cavity; V, vacuole; Vs, vesicles. Arrows indicate active secretion from vesicles to PSPs by exocytosis. (Figure 8D) During the presecretory stage, DCs contained highly dense cytoplasm lined with ER and multiple ribosomes. There was no SCv or PSP, and plastids were large and resembled proplastids. (Figure 8E) During the secretory stage, detachment of the apical DC wall led to the formation of SCvs. Electron-lucent secretory products exuded from plastids within vesicles separated by an electron-dense layer. The vesicles released their contents into PSPs by exocytosis, where the secretory products accumulated before being secreted into SCvs. (Figure 8F) The DC of mature trichomes at the post-secretory stage was largely vacuolated, with the cytoplasm restricted to a small remaining region. Plastids at this stage were degenerated, and no vesicles were observed. The cell wall was largely cutinized, and the SCv was large. MALDI-MSI of the m / z 361.23 ± 0.01 Da signal from the abaxial leaf region (Figure 8G) and the adaxial leaf region (Figure 8H) after partial removal of the trichomes with duct tape (the detached region is outlined in green).Areas where the trichomes were partially or completely removed show less or no signal compared to untouched areas. (Figure 8I) Optical image of a cross-section of a receptacle and (Figure 8J) MALDI-MSI of m / z 361.23 ± 0.01 Da. The glandular trichomes in Figures 8G-8H and 8J correspond to the protonated m / z of CBGA 1 and geranylfluorocaprofenone 4. The white dashed lines in Figures 8G-8J indicate the analyzed areas.
[0058] [Figure 9] A schematic diagram showing predicted parallel metabolic pathways for the biosynthesis of cannabinoids and other terpenophenols present in H. umbraculigerum is included. The enzyme types predicted to catalyze each reaction are indicated by 1–8. Further functional groups and rearrangements include hydroxylation, double bond isomerization or reduction, cyclization, etc. Alkyl chains can be linear or branched, with lengths ranging from 1 to 7 carbons. AAE, acyl-activating enzyme; PKS, type III polyketide synthase; PKC, polyketide cyclase; PT, prenyltransferase; UGT, uridine diphosphate glycosyltransferase; AAT, alcohol acyltransferase; DBR, double bond reductase; and CHI, chalcone isomerase. The active enzymes identified in this study are labeled accordingly. CoAT, acyl-CoA-transferase; TKS, tetraketide synthase; CBGAS, cannabigerolic acid synthase; OAGT, olivetolic acid UGT; CBGT, cannabinoid UGT; CBAT, cannabinoid acyl-transferase.
[0059] [Figure 10A] Figures 10A-10E include chromatograms, schematic diagrams, chemical structures, and curves showing the functional characteristics of HuAAEs, HuPKSs, and HuPTs. Figure 10A shows the ion abundances of triple quadrupole analysis of acyl-CoA produced in vitro by HuAAEs compared to analytical standards (Std). [Figure 10B] FIG. 1 is a schematic diagram showing the steps and types of products and by-products synthesized in vitro by recombinant HuPKS with or without Cannabis olivetolate cyclase (CsOAC). [Figure 10C] Figure 1 shows triple quadrupole ion abundances of OA 92 and olivetol products from coupled recombinant enzyme assays of HuPKS with either empty vector (EV) or Cannabis olivetol acid cyclase (CsOAC) in the presence of hexanoyl-CoA and malonyl-CoA. [Figure 10D] MS / MS spectra of the prenylated OA 92 products from cannabigerolic acid synthase (HuCBGAS4) using isopentenyl pyrophosphate (IPP), geranyl prophosphate (GPP), or farnesyl pyrophosphate (FPP) as the prenyl donor: CBPA 19, cannabiprenilic acid; CBGA 1, cannabigerolic acid; sesquiCBGA, sesquicannabigerolic acid (MS / MS spectra correspond to published data for Cannabis 15). [Figure 10E] Steady-state kinetic analysis of H. umbraculigerum prenyltransferases HuPT1, HuPT3, and HuCBGAS4 using OA 92 and GPP. Michaelis-Menten K values for each enzyme were calculated using various concentrations (0.5 µM - 3 mM) and a constant concentration (1 mM) of each substrate (n = 3 technically independent samples; measurements were plotted separately).
[0060] [Figure 11A] 11A-11D include phylogenetic trees showing the phylogenetic analysis of enzymes and whole proteomes from H. umbraculigerum and different plant species. Phylogenetic analysis of AAEs from H. umbraculigerum and other plants. [Figure 11B]FIG. 1 shows a phylogenetic analysis of PKSs from H. umbraculigerum and other plants. [Figure 11C] Phylogenetic analysis of PT proteins from H. umbraculigerum and other plants. H. umbraculigerum and Cannabis proteins are highlighted in red and blue, respectively, with active enzymes shown in flowers and leaves, respectively. A complete list of protein IDs is available in Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023). Bootstrap values are shown at each branch node. (Figure 11A) Protein selection was based on either (Figure 11A) Arabidopsis thaliana enzymes or (Figures 11B-11C) functionally tested enzymes. Phylogenetic groups by substrate or function are shown in different colors. None of the active H. umbraculigerum enzymes clustered with any known Cannabis proteins. [Figure 11D] Phylogenetic relationships of Arabidopsis (Arabidopsis thaliana), tomato (Solanum lycopersicum), sunflower (Helianthus annuus), lettuce (Letuca sativa), hemp (Cannabis sativa), and Helicrysum umbraculigerum, showing the evolutionary distance of the last two species (shown by their flowers and leaves, respectively). The tree was constructed based on the whole proteome of each species using the Word-based software Prot-SpaM.
[0061] [Figure 12A]Figures 12A-12C include graphs, chromatograms, chemical structures, and curves showing the functional characteristics of HuUGT. Figure 12A shows the activity (n=1) of a lysate containing HuUGT using olivetolic acid (OA92), cannabigerolic acid (CBGA1), and helicambigerolic acid (heliCBGA2) as substrates and uridine diphosphate glucose (UDP-Glc) as the glycosyl donor. The reactions demonstrate different substrate specificities and product types. Representative peaks correspond to the chromatogram obtained for HuCBUGT1. The most abundant product in each assay is indicated by an asterisk. EV, empty vector. [Figure 12B] Figure 1 shows the in vitro production of monoglucosides using purified UGTs and additional substrates. Extracted ion chromatograms of the observed monoglucosides using UDP-Glc and either DHSA 93, olivetol, CBG, CBD, Δ9-THC, PCP 95, naringenin chalcone 97, or pinocrine chalcone 100. The substrates naringenin chalcone 97 and pinocrine chalcone 100 contained a mixture of chalcone and the respective flavanone. All LC-MS chromatograms were selected based on the theoretical m / z values of the respective metabolites of interest. [Figure 12C] Figure 1 shows a comparison of the steady-state kinetics of UGTs with OA 92 and UDP-Glc. HuOAUGT11 and HuUGT13 were compared with UGTs derived from rice (OsUGT) and stevia (SrUGT). Kinetic values were calculated using varying concentrations (0.5 mM - 3 mM) and a constant concentration (1 mM) of each substrate (n = 3 technically independent samples; measurements were plotted separately). Because no analytical standard was available for Glc-OA 92, a calibration curve for OA 102 was used to calculate V0 and Vmax.
[0062] [Figure 13A]Figures 13A-13C include chemicals, chromatograms, and a phylogenetic tree showing the functional characteristics of HuAAT. Figure 13A shows the stable dual-isotope labeling of O-MeButCBGA120 by feeding H. umbraculigerum leaves with either 2-methylbutyric-D9 acid or hexane-D11 acid. MS / MS spectra of the unlabeled and dual-labeled forms show fragmentation patterns with mass shifts corresponding to the labeled portion of the molecule. The red- or purple-colored fragments correspond to the m / z of specific fragments bearing labeled alkyl chains or acyl groups, respectively. [Figure 13B] Figure 1 shows the activity of lysates containing HuAAT with different acyl donors and cannabinoid acceptors. Extracted ion chromatograms were selected based on the theoretical m / z values of each metabolite. Only HuCBAT5 and HuAAT14 (red and blue, respectively) acylated CBGA1 and heliCBGA2 with both acyl-CoAs. EV, empty vector; Std, standard; ButCoA, butyryl-CoA; HexCoA, hexanoyl-CoA. [Figure 13C] Phylogenetic analysis of HuAAT proteins and BAHD AATs identified from other plants. A maximum likelihood tree was constructed using MEGA11 software with 100 bootstrap tests based on the MUSCLE multiple alignment. Evolutionary distances were calculated using a JTTmatrix-based method. Bootstrap values are shown at each branch node. Phylogenetic groups of different AAT types are indicated by circles based on Tuominen et al. (2011). Active HuCBAT5 and HuAAT14 clustered in phylogenetic group IIIa, with BAHDs of diverse catalytic functions. A complete list of protein IDs is available in Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, 817-831 (2023).
[0063] [Figure 14] Figure 1 shows chromatograms and chemical structures of the acylated cannabinoids observed after enzymatic assays using purified HuCBAT5. OA 92, olivetolic acid; CBGA 1, cannabigerolic acid; HeliCBGA 2, helicannabigerolic acid; CBDA, cannabidiolic acid. Complete MS / MS product data are available in Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, pp. 817-831 (2023). MS / MS fragmentation and retention times correspond to O-acylated cannabinoids found in plants.
[0064] [Figure 15]Figures 15A-15F include schematics, chromatograms, and tables showing the reconstruction of the core cannabinoid pathway in heterologous systems. Figure 15A shows a schematic representation of the products observed in N. benthamiana leaves after co-expression of Cannabis-derived CsOAC with various combinations of HuCoAT6, HuTKS4, and HuCBGAS4. Figure 15D shows a schematic representation of the products observed in S. cerevisiae after co-expression of Cannabis-derived CsOAC with various combinations of HuCoAT6, HuTKS4, and HuCBGAS4. NbUGT, N. benthamiana uridine diphosphate-glycosyltransferase; HexNa, sodium hexanoate; GPP, geranylproline; OA 92, olivetolic acid. Figure 15B shows extracted ion chromatograms and MS / MS spectra showing glycosylated OA (Glc-OA 102), glycosylated polycaprophenone (Glc-PCP1 / 2), and glycosylated naringenin chalcone (Glc-naringenin chalcone 1 / 2) after feeding with HexNa and GPP(I). Figure 15C shows extracted ion chromatograms and MS / MS spectra showing (15C) glycosylated cannabigerolic acid (Glc-CBGA 109) after feeding with OA 92 and GPP(II). Glycosylated metabolites synthesized by recombinant stevia (SrUGT) or rice (OsUGT) enzymes were used as references for identification of N. benthamiana products by accurate mass, retention time, and MS / MS spectra. EV, empty vector; UDP-Glc, uridine diphosphate glucose. Figure 15E shows extracted ion chromatograms of OA 92, PCP 95, and CBGA 1 products observed in unfed yeast. Identification was according to analytical standards. Figure 15F shows a summary of the products observed in each assay. PDAL, pentyl acyl diacetic acid lactone; HTAL, hexanoyl acyl triacetic acid lactone. DETAILED DESCRIPTION OF THE INVENTION
[0065] The present invention, in some embodiments, relates to and methods of using DNA molecules comprising at least one first nucleic acid sequence encoding a first protein and at least one second nucleic acid sequence encoding a second protein, wherein the first protein and the second protein are derived from Helichrysum umbraculigerum.
[0066] In some embodiments, any one of the first protein and the second protein belongs to an enzyme family selected from acyl-activating enzymes (AAEs), polyketide synthases (PKSs), polyketide cyclases (PKCs), prenyltransferases (PTs), cannabichromene acid synthases (CBCASs), uridine diphosphate (UDP)-glycosyltransferases (UGTs), and alcohol acyltransferases (AATs).
[0067] In some embodiments, the DNA molecule further comprises at least a third nucleic acid sequence derived from H. umbraculigerum and encoding a third protein belonging to an enzyme family selected from AAE, PKS, PKC, PT, CBCAS, UGT, and AAT.
[0068] In some embodiments, the DNA molecule further comprises at least a fourth nucleic acid sequence derived from H. umbraculigerum and encoding a third protein belonging to an enzyme family selected from AAE, PKS, PKC, PT, CBCAS, UGT, and AAT.
[0069] In some embodiments, the DNA molecule further comprises at least a fifth nucleic acid sequence derived from H. umbraculigerum and encoding a third protein belonging to an enzyme family selected from AAE, PKS, PKC, PT, CBCAS, UGT, and AAT.
[0070] In some embodiments, the DNA molecule further comprises at least a sixth nucleic acid sequence derived from H. umbraculigerum and encoding a third protein belonging to an enzyme family selected from AAE, PKS, PKC, PT, CBCAS, UGT, and AAT.
[0071] In some embodiments, the DNA molecule further comprises at least a seventh nucleic acid sequence derived from H. umbraculigerum and encoding a third protein belonging to an enzyme family selected from AAE, PKS, PKC, PT, CBCAS, UGT, and AAT.
[0072] In some embodiments, the first protein and the second protein belong to different enzyme families.
[0073] In some embodiments, the first protein, the second protein, and the third protein belong to different enzyme families.
[0074] In some embodiments, the first protein, the second protein, the third protein, and the fourth protein belong to different enzyme families.
[0075] In some embodiments, the first protein, the second protein, the third protein, the fourth protein, and the fifth protein belong to different enzyme families.
[0076] In some embodiments, the first protein, the second protein, the third protein, the fourth protein, the fifth protein, and the sixth protein belong to different enzyme families.
[0077] In some embodiments, the first protein, the second protein, the third protein, the fourth protein, the fifth protein, the sixth protein, and the seventh protein belong to different enzyme families.
[0078] According to some embodiments, (a) the AAE protein is encoded by a nucleic acid sequence having at least 89% homology to any one of SEQ ID NOs: 1-11; (b) the PKS is encoded by a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 23-26; (c) the PKC is encoded by a nucleic acid sequence having at least 88% homology to any one of SEQ ID NOs: 31-38; (d) the PT is encoded by a nucleic acid sequence having at least 91% homology to any one of SEQ ID NOs: 47-58; (e) the CBCAS is encoded by a nucleic acid sequence having at least 82% homology to any one of SEQ ID NOs: 71-79; or (f) any combination of (a)-(e).
[0079] In some embodiments, the DNA molecule is derived from Helichrysum umbraculigerum and further comprises a nucleic acid sequence encoding one or more proteins or enzymes belonging to the uridine diphosphate (UDP)-glycosyltransferase (UGT) family; the alcohol acyltransferase (AAT) family; or both.
[0080] In some embodiments, (a) the UGT is encoded by a nucleic acid sequence having at least 87% homology to any one of SEQ ID NOs: 89-101, and any combination thereof; (b) the AAT is encoded by a nucleic acid sequence having at least 87% homology to any one of SEQ ID NOs: 115-129, and any combination thereof; or (c) both (a) and (b).
[0081] In some embodiments, the DNA molecule comprises at least two nucleic acid sequences encoding at least two enzymes, each enzyme belonging to a different family, and the at least two families are selected from AAE, PKS, PKC, PT, CBCAS, UGT, and AAT.
[0082] In some embodiments, the DNA molecule is an isolated DNA molecule. In some embodiments, the DNA molecule is a complementary DNA (cDNA) molecule.
[0083] As used herein, the term "DNA molecule" refers to a polynucleotide that comprises or consists of deoxyribonucleotides.
[0084] As used herein, the terms "isolated polynucleotide" and "isolated DNA molecule" refer to a nucleic acid molecule that is essentially free from contaminating cellular components, such as carbohydrates, lipids, or other proteinaceous impurities that naturally accompany the nucleic acid. Typically, an isolated DNA or RNA preparation contains the nucleic acid in a highly purified form, for example, at least about 80% pure, at least about 90% pure, at least about 95% pure, greater than 95% pure, or greater than 99% pure. In some embodiments, the isolated polynucleotide is any one of DNA, RNA, and cDNA. In some embodiments, the isolated polynucleotide is a synthetic polynucleotide. Polynucleotide synthesis is well known in the art and can be performed, for example, by linking or covalently linking multiple nucleic acid molecules together using primer linkers.
[0085] The term "nucleic acid" is well known in the field of molecular biology. As used herein, "nucleic acid" generally refers to any molecule (e.g., a chain) of DNA, RNA, or derivatives or analogs thereof that contain nucleotides. Nucleotides are composed of a nucleoside and a phosphate group. The nitrogenous base of a nucleoside includes naturally occurring purine or pyrimidine nucleosides, such as those found in DNA (e.g., adenine "A," guanine "G," thymine "T," or cytosine "C") or RNA (e.g., A, G, uracil "U," or C).
[0086] The term "nucleic acid molecule" includes, but is not limited to, single-stranded RNA (ssRNA), double-stranded RNA (dsRNA), single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), small RNAs, circular nucleic acids, fragments of genomic DNA or RNA, degraded nucleic acids, amplification products, modified nucleic acids, plasmids or organelle nucleic acids, and artificial nucleic acids such as oligonucleotides.
[0087] Includes.
[0088] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 89%, at least 92%, at least 95%, or at least 97% homology or identity to SEQ ID NO:1, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 89%-95%, 90%-97%, 95%-99%, or 90%-100% homology or identity to SEQ ID NO:1. Each possibility represents a separate embodiment of the present invention.
[0089] Includes:
[0090] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 79%, at least 83%, at least 87%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:2, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 79%-85%, 80%-92%, 82%-99%, or 80%-100% homology or identity to SEQ ID NO:2. Each possibility represents a separate embodiment of the present invention.
[0091] Includes.
[0092] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 86%, at least 88%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:3, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 86%-94%, 88%-97%, 86%-100%, or 92%-99% homology or identity to SEQ ID NO:3. Each possibility represents a separate embodiment of the present invention.
[0093] Includes.
[0094] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:4, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 88%-95%, 89%-99%, 91-98%, or 88%-100% homology or identity to SEQ ID NO:4. Each possibility represents a separate embodiment of the present invention.
[0095] Includes.
[0096] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:5, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 88%-95%, 89%-99%, 91-98%, or 88%-100% homology or identity to SEQ ID NO:5. Each possibility represents a separate embodiment of the present invention.
[0097] Includes.
[0098] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 89%, at least 90%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:6, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 89%-95%, 90%-99%, 91-98%, or 89%-100% homology or identity to SEQ ID NO:6. Each possibility represents a separate embodiment of the present invention.
[0099] Includes.
[0100] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 85%, at least 87%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:7, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 85%-94%, 88%-97%, 85%-100%, or 92%-99% homology or identity to SEQ ID NO:7. Each possibility represents a separate embodiment of the present invention.
[0101]
[0102] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 84%, at least 87%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:8, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 84%-94%, 88%-97%, 84%-100%, or 92%-99% homology or identity to SEQ ID NO:8. Each possibility represents a separate embodiment of the present invention.
[0103]
[0104] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:9, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 88%-95%, 89%-99%, 91-98%, or 88%-100% homology or identity to SEQ ID NO:9. Each possibility represents a separate embodiment of the present invention.
[0105]
[0106] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 89%, at least 92%, at least 95%, or at least 97% homology or identity to SEQ ID NO: 10, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 89%-95%, 90%-97%, 95%-99%, or 90%-100% homology or identity to SEQ ID NO: 10. Each possibility represents a separate embodiment of the present invention.
[0107]
[0108] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 89%, at least 92%, at least 95%, or at least 97% homology or identity to SEQ ID NO: 11, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 89%-95%, 90%-97%, 95%-99%, or 90%-100% homology or identity to SEQ ID NO: 11. Each possibility represents a separate embodiment of the present invention.
[0109] Comprises or consists of.
[0110] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 83%, at least 85%, at least 87%, at least 95%, or at least 99% homology or identity to SEQ ID NO:23, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 83%-100%, 88%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO:23. Each possibility represents a separate embodiment of the present invention.
[0111] Comprises or consists of.
[0112] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 83%, at least 85%, at least 87%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:24, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 83%-100%, 87%-100%, 90%-100%, or 93%-100% homology or identity to SEQ ID NO:24. Each possibility represents a separate embodiment of the present invention.
[0113] Comprises or consists of.
[0114] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 83%, at least 87%, at least 89%, at least 90%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:25, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 83%-100%, 88%-100%, 93%-100%, or 95%-100% homology or identity to SEQ ID NO:25. Each possibility represents a separate embodiment of the present invention.
[0115] Comprises or consists of.
[0116] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 82%, at least 85%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:26, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 82%-100%, 86%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO:26. Each possibility represents a separate embodiment of the present invention.
[0117] In some embodiments, the DNA molecule has the nucleic acid sequence: ATGGCGGAGTTCACACATTTAGTGGTGGTTAAGTTCAAAGAAGAGGTGGTTGTGGAGGATATTATGAAAGGGTTGGAGAAACTTGTATCTCAACTTGATAGTGTCAAGTCCTTTGTTTGGGGAAAGGATATTGAAAGCATGGAGATGTTAAGGCAAGGATTCACCCATGCAATCATGATGACATTTGGTTCTAAAGAAGATTTTACTGCATTTCAATCCCACCCAAACCATGTTGAATTCTCGGCTACGTTTTCAGCAGCAATCGAAAAGATCGTTCTTCTTGATTTCCCAGTTGTTGCTGTCAAGACTGCAACTGCTTGA (SEQ ID NO: 31) Comprises or consists of.
[0118] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 72%, at least 75%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 31, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 72%-95%, 72%-100%, 75%-99%, or 80%-100% homology or identity to SEQ ID NO: 31. Each possibility represents a separate embodiment of the present invention.
[0119] In some embodiments, the DNA molecule comprises the nucleic acid sequence: (SEQ ID NO: 32) Comprises or consists of.
[0120] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 32, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 50%-95%, 55%-98%, 60%-99%, or 50%-100% homology or identity to SEQ ID NO: 32. Each possibility represents a separate embodiment of the present invention.
[0121] In some embodiments, the DNA molecule comprises the nucleic acid sequence: (SEQ ID NO: 33) Comprises or consists of.
[0122] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 67%, at least 72%, at least 78%, at least 85%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 33, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 67%-95%, 70%-98%, 75%-99%, or 67%-100% homology or identity to SEQ ID NO: 33. Each possibility represents a separate embodiment of the present invention.
[0123] In some embodiments, the DNA molecule has the nucleic acid sequence: ATGGGAGAAGTGAAGCACATACTTTTAGCGAAGTTTAAGGATGGAATCTCGGAACAACAGATCCAGCATCTCATCACAGGTTATGCTAACCTCGTCAATCTCGTTGAACCCATGAAGTCTTTTCGATGGGGAAAAGATGTGAGCATTGAGAATCTGCACCAAGGCTTTACTCATGTGTTCGAGTCAACCTTTGAAACCACTGAAGGCATTGCAACTTATATATCTCATCCTGCTCATGTCGAGTTCGCCACTGGTTTCCTGGATCAACTGGAAAAAGTCATAGTCATCGACTACAAACCTACATCAGTTGACCCGTGA (SEQ ID NO: 34) Comprises or consists of.
[0124] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 74%, at least 78%, at least 85%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 34, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 74%-95%, 78%-98%, 80%-99%, or 75%-100% homology or identity to SEQ ID NO: 34. Each possibility represents a separate embodiment of the present invention.
[0125] In some embodiments, the DNA molecule comprises the nucleic acid sequence: (SEQ ID NO: 35) Comprises or consists of.
[0126] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 69%, at least 75%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 35, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 69%-95%, 70%-100%, 80%-99%, or 68%-100% homology or identity to SEQ ID NO: 35. Each possibility represents a separate embodiment of the present invention.
[0127] In some embodiments, the DNA molecule has the nucleic acid sequence: ATGGCAGTTGCTCAACTTTCTTCCTCCCTCTGTATCTCCACACCCGCTAGAATCTCTACTGGTTCTGGGTTTTCGTCATCAGGTTTGCCTCGGATTGGGACAACGTTTGTATGCGGTTCAGGTTCGCCTCTTGTGATATCTGGAACATATCATCAGAAGGCTCGAGTACATAAGCCTGCAGCATTATCTGTGAGATGTGAACAAAGTAGTAAGGATGGAAATGGTTTAAATGTGTGGCTTGGTCGAACAGCAATGGTTGGCTTTGCAGTGGCAATTAGTGTTGAAGTATCAACTGGGAAGGGGCTTCTTGAGAACTTTGGGCTCACATCACCCTTGCCAACAGTGGCCTTGGCACTGACTGCACTTGGGGGCGTTCTTACAGCACTTTTCATCTTCCAGTCTGCTTCTGAGAGTTGA (SEQ ID NO: 36) Comprises or consists of.
[0128] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 73%, at least 75%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 36, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 73%-95%, 73%-100%, 80%-99%, or 80%-100% homology or identity to SEQ ID NO: 36. Each possibility represents a separate embodiment of the present invention.
[0129] In some embodiments, the DNA molecule comprises the nucleic acid sequence: (SEQ ID NO: 37) Comprises or consists of.
[0130] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 69%, at least 78%, at least 85%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 37, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 69%-95%, 70%-98%, 71%-99%, or 69%-100% homology or identity to SEQ ID NO: 37. Each possibility represents a separate embodiment of the present invention.
[0131] In some embodiments, the DNA molecule has the nucleic acid sequence: ATGGCGGAGTTCACACATTTAGTGGTGGTTAAGTTCAAAGAAGAGGTGGTTGTAGAGGATATTATGAAAGGGTTGGAGAAACTTGCATCTCAACTTGATAGTGTCAAGTCCTTTGTTTGGGGAAAGGATATTGAAAGCATGGAGATGTTAAGGCAAGGATTCACCCATGCAATCATGATGACATTTGGTTCTAAAGAAGATTTTACTGCATTTCAATCCCACCCAAACCATGTTGAATTCTCGGCTACGTTTTCAGCAGCAATCGAAAAGATCGTTCTTCTTGATTTCCCAGTTGTTGCAGTCAAGACTGCAACTGCTTGA (SEQ ID NO: 38) Comprises or consists of.
[0132] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 88%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homology or identity to SEQ ID NO: 38, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 88%-95%, 88%-98%, 89%-99%, or 88%-100% homology or identity to SEQ ID NO: 38. Each possibility represents a separate embodiment of the present invention.
[0133] Comprises or consists of.
[0134] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 75%, at least 79%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 47, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 75%-100%, 80%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 47. Each possibility represents a separate embodiment of the present invention.
[0135] Comprises or consists of.
[0136] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 48, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 80%-100%, 85%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 48. Each possibility represents a separate embodiment of the present invention.
[0137] Comprises or consists of.
[0138] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 75%, at least 79%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 49, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 75%-100%, 80%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 49. Each possibility represents a separate embodiment of the present invention.
[0139] Comprises or consists of.
[0140] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 91%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 50, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 91%-100%, 93%-100%, 95%-100%, or 97%-100% homology or identity to SEQ ID NO: 50. Each possibility represents a separate embodiment of the present invention.
[0141] Comprises or consists of.
[0142] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 91%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:51, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 91%-100%, 93%-100%, 95%-100%, or 97%-100% homology or identity to SEQ ID NO:51. Each possibility represents a separate embodiment of the present invention.
[0143] Comprises or consists of.
[0144] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 90%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 52, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 90%-100%, 92%-100%, 95%-100%, or 97%-100% homology or identity to SEQ ID NO: 52. Each possibility represents a separate embodiment of the present invention.
[0145] Comprises or consists of.
[0146] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 77%, at least 79%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 53, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 77%-100%, 85%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 53. Each possibility represents a separate embodiment of the present invention.
[0147] Comprises or consists of.
[0148] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 89%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homology or identity to SEQ ID NO: 54, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 89%-100%, 92%-100%, 94%-100%, or 97%-100% homology or identity to SEQ ID NO: 54. Each possibility represents a separate embodiment of the present invention.
[0149] Comprises or consists of.
[0150] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 76%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% homology or identity to SEQ ID NO: 55, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 76%-100%, 83%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 55. Each possibility represents a separate embodiment of the present invention.
[0151] Comprises or consists of.
[0152] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% homology or identity to SEQ ID NO: 56, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 75%-100%, 80%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 56. Each possibility represents a separate embodiment of the present invention.
[0153] Comprises or consists of.
[0154] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 76%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% homology or identity to SEQ ID NO: 57, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 76%-100%, 85%-100%, 90%-100%, or 96%-100% homology or identity to SEQ ID NO: 57. Each possibility represents a separate embodiment of the present invention.
[0155] Comprises or consists of.
[0156] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 77%, at least 79%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 58, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 77%-100%, 85%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 58. Each possibility represents a separate embodiment of the present invention.
[0157] Comprises or consists of.
[0158] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 68%, at least 75%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 71, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 68%-95%, 75%-100%, 72%-99%, or 68%-100% homology or identity to SEQ ID NO: 71. Each possibility represents a separate embodiment of the present invention.
[0159] Comprises or consists of.
[0160] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 71%, at least 77%, at least 85%, at least 93%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 72, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 71%-95%, 75%-98%, 80%-99%, or 71%-100% homology or identity to SEQ ID NO: 72. Each possibility represents a separate embodiment of the present invention.
[0161] Comprises or consists of.
[0162] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 69%, at least 75%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 73, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 69%-95%, 75%-100%, 72%-99%, or 69%-100% homology or identity to SEQ ID NO: 73. Each possibility represents a separate embodiment of the present invention.
[0163] Comprises or consists of.
[0164] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 79%, at least 85%, at least 92%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 74, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 79%-98%, 80%-99%, 82%-99%, or 79%-100% homology or identity to SEQ ID NO: 74. Each possibility represents a separate embodiment of the present invention.
[0165] In some embodiments, the DNA molecule comprises the nucleic acid sequence: (SEQ ID NO: 75) Comprises or consists of.
[0166] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 82%, at least 87%, at least 92%, at least 96%, or at least 99% homology or identity to SEQ ID NO: 75, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 82%-98%, 83%-99%, 85%-99%, or 82%-100% homology or identity to SEQ ID NO: 75. Each possibility represents a separate embodiment of the present invention.
[0167] Comprises or consists of.
[0168] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 80%, at least 87%, at least 93%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 76, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 80%-98%, 81%-99%, 85%-99%, or 80%-100% homology or identity to SEQ ID NO: 76. Each possibility represents a separate embodiment of the present invention.
[0169] Comprises or consists of.
[0170] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 79%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 77, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 79%-95%, 82%-97%, 81%-98%, or 79%-100% homology or identity to SEQ ID NO: 77. Each possibility represents a separate embodiment of the present invention.
[0171] In some embodiments, the DNA molecule comprises the nucleic acid sequence: (SEQ ID NO: 78) Comprises or consists of.
[0172] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 80%, at least 85%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 78, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 80%-95%, 85%-98%, 89%-99%, or 80%-100% homology or identity to SEQ ID NO: 78. Each possibility represents a separate embodiment of the present invention.
[0173] Comprises or consists of.
[0174] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 79%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 79, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 79%-95%, 82%-97%, 81%-98%, or 79%-100% homology or identity to SEQ ID NO: 79. Each possibility represents a separate embodiment of the present invention.
[0175] Includes.
[0176] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 77%, at least 79%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 89, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 77%-95%, 78%-100%, 79%-99%, or 77%-100% homology or identity to SEQ ID NO: 89. Each possibility represents a separate embodiment of the present invention.
[0177] Comprises or consists of.
[0178] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 76%, at least 77%, at least 85%, at least 93%, at least 97%, or at least 99% homology or identity to SEQ ID NO:99, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 76%-95%, 77%-98%, 80%-99%, or 76%-100% homology or identity to SEQ ID NO:90. Each possibility represents a separate embodiment of the present invention.
[0179] In some embodiments, the DNA molecule comprises the nucleic acid sequence: (SEQ ID NO: 91) Comprises or consists of.
[0180] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 78%, at least 80%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO:91, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 79%-95%, 78%-100%, 80%-99%, or 79%-100% homology or identity to SEQ ID NO:91. Each possibility represents a separate embodiment of the present invention.
[0181] Comprises or consists of.
[0182] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 87%, at least 92%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 92, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 87%-100%, 88%-99%, 89%-99%, or 87%-100% homology or identity to SEQ ID NO: 92. Each possibility represents a separate embodiment of the present invention.
[0183] Comprises or consists of.
[0184] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 87%, at least 92%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 93, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 87%-100%, 88%-99%, 89%-99%, or 87%-100% homology or identity to SEQ ID NO: 93. Each possibility represents a separate embodiment of the present invention.
[0185] Comprises or consists of.
[0186] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 80%, at least 87%, at least 93%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 94, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 80%-98%, 81%-99%, 85%-99%, or 80%-100% homology or identity to SEQ ID NO: 94. Each possibility represents a separate embodiment of the present invention.
[0187] In some embodiments, the DNA molecule comprises the nucleic acid sequence: (SEQ ID NO: 95) Comprises or consists of.
[0188] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 77%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 95, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 77%-95%, 82%-97%, 81%-98%, or 77%-100% homology or identity to SEQ ID NO: 95. Each possibility represents a separate embodiment of the present invention.
[0189] Comprises or consists of.
[0190] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 82%, at least 85%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 96, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 82%-95%, 83%-98%, 82%-99%, or 82%-100% homology or identity to SEQ ID NO: 96. Each possibility represents a separate embodiment of the present invention.
[0191] Comprises or consists of.
[0192] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 79%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 97, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 79%-95%, 82%-97%, 81%-98%, or 79%-100% homology or identity to SEQ ID NO: 97. Each possibility represents a separate embodiment of the present invention.
[0193] Comprises or consists of.
[0194] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 78%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 98, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 78%-95%, 82%-97%, 81%-98%, or 78%-100% homology or identity to SEQ ID NO: 98. Each possibility represents a separate embodiment of the present invention.
[0195] Comprises or consists of.
[0196] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 82%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO:99, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 82%-95%, 82%-97%, 83%-98%, or 82%-100% homology or identity to SEQ ID NO:99. Each possibility represents a separate embodiment of the present invention.
[0197] Comprises or consists of.
[0198] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 74%, at least 80%, at least 85%, at least 87%, at least 93%, or at least 99% homology or identity to SEQ ID NO: 100, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 74%-95%, 75%-97%, 76%-98%, or 74%-100% homology or identity to SEQ ID NO: 100. Each possibility represents a separate embodiment of the present invention.
[0199] Comprises or consists of.
[0200] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 80%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 101, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 80%-95%, 82%-97%, 81%-98%, or 80%-100% homology or identity to SEQ ID NO: 101. Each possibility represents a separate embodiment of the present invention.
[0201] Comprises or consists of.
[0202] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 84%, at least 87%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 115, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 84%-100%, 88%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 115. Each possibility represents a separate embodiment of the present invention.
[0203] Comprises or consists of.
[0204] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 77%, at least 85%, at least 93%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 116, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 77%-100%, 80%-100%, 85%-100%, or 93%-100% homology or identity to SEQ ID NO: 116. Each possibility represents a separate embodiment of the present invention.
[0205] Comprises or consists of.
[0206] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 87%, at least 90%, at least 93%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 117, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 87%-100%, 90%-100%, 93%-100%, or 95%-100% homology or identity to SEQ ID NO: 117. Each possibility represents a separate embodiment of the present invention.
[0207] Comprises or consists of.
[0208] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 82%, at least 90%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 118, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 82%-100%, 85%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 118. Each possibility represents a separate embodiment of the present invention.
[0209] Comprises or consists of.
[0210] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 74%, at least 80%, at least 85%, or at least 95% homology or identity to SEQ ID NO: 119, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 74%-100%, 80%-100%, 87%-100%, or 95%-100% homology or identity to SEQ ID NO: 119. Each possibility represents a separate embodiment of the present invention.
[0211] Comprises or consists of.
[0212] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 79%, at least 87%, at least 93%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 120, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 79%-100%, 85%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 120. Each possibility represents a separate embodiment of the present invention.
[0213] Comprises or consists of.
[0214] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 82%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 121, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 82%-100%, 85%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 121. Each possibility represents a separate embodiment of the present invention.
[0215] Comprises or consists of.
[0216] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 83%, at least 85%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 122, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 83%-100%, 88%-100%, 92%-100%, or 95%-100% homology or identity to SEQ ID NO: 122. Each possibility represents a separate embodiment of the present invention.
[0217] Comprises or consists of.
[0218] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 77%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 123, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 77%-100%, 82%-100%, 87%-100%, or 95%-100% homology or identity to SEQ ID NO: 123. Each possibility represents a separate embodiment of the present invention.
[0219] Comprises or consists of.
[0220] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 84%, at least 89%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 124, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 84%-100%, 88%-100%, 93%-100%, or 97%-100% homology or identity to SEQ ID NO: 124. Each possibility represents a separate embodiment of the present invention.
[0221] Comprises or consists of.
[0222] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 82%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 125, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 82%-100%, 85%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 125. Each possibility represents a separate embodiment of the present invention.
[0223] Comprises or consists of.
[0224] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 72%, at least 80%, at least 85%, at least 87%, at least 93%, or at least 99% homology or identity to SEQ ID NO: 126, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 72%-100%, 79%-100%, 86%-100%, or 91%-100% homology or identity to SEQ ID NO: 126. Each possibility represents a separate embodiment of the present invention.
[0225] Comprises or consists of.
[0226] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 79%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 127, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 79%-100%, 85%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 127. Each possibility represents a separate embodiment of the present invention.
[0227] Comprises or consists of.
[0228] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 82%, at least 85%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 128, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 82%-100%, 88%-100%, 93%-100%, or 97%-100% homology or identity to SEQ ID NO: 128. Each possibility represents a separate embodiment of the present invention.
[0229] Comprises or consists of.
[0230] In some embodiments, the DNA molecule comprises a nucleic acid sequence having at least 87%, at least 91%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 129, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the DNA molecule comprises a nucleic acid sequence having 87%-100%, 90%-100%, 94%-100%, or 97%-100% homology or identity to SEQ ID NO: 129. Each possibility represents a separate embodiment of the present invention.
[0231] In some embodiments, the DNA molecule comprises more than one nucleic acid sequence. In some embodiments, the polynucleotide comprises more than one type of polynucleotide.
[0232] As used herein, the term "plurality" includes any integer greater than or equal to two.
[0233] In some embodiments, the plurality of nucleic acid sequences encodes proteins of different enzymatic functions or families as described herein. In some embodiments, the plurality of nucleic acid sequences encodes at least two proteins of the same enzymatic function or family as described herein. In some embodiments, the plurality of nucleic acid sequences encodes multiple proteins of multiple different enzymatic functions or families as described herein.
[0234] In some embodiments, the DNA molecule encodes a protein characterized by acyl-activating enzyme (AAE) activity. In some embodiments, the DNA molecule encodes an AAE protein. In some embodiments, the AAE is an AAE from Helichrysum umbraculigerum. In some embodiments, the DNA molecule encoding the protein characterized by acyl-activating enzyme (AAE) activity comprises a nucleic acid sequence set forth in SEQ ID NOs: 1-11.
[0235] As used herein, the terms "acyl-activating enzyme" and "AAE" are interchangeable and refer to any peptide, polypeptide, or protein capable of catalyzing the activation of a carboxylic acid. In some embodiments, the AAE activity comprises forming a thioester bond or forming a thioester bond. In some embodiments, the AAE activity comprises conjugating a carboxyl group to an amine group. In some embodiments, the AAE activity comprises conjugating a carboxyl group to an alcohol. In some embodiments, the AAE is an acid-thiol ligase.
[0236] In some embodiments, the DNA molecule encodes a protein characterized by polyketide synthesis activity. In some embodiments, the DNA molecule encodes a protein that is a polyketide synthase (PKS). In some embodiments, the PKS is a PKS derived from Helichrysum umbraculigerum. As used herein, the terms "polyketide synthase" and "PKS" encompass any enzyme derived from H. umbraculigerum that has or is a functional analog of the "olivetol synthase" or "OLS" of Cannabis sativa. In some embodiments, the DNA molecule encoding a protein characterized by polyketide synthesis activity comprises the nucleic acid sequence set forth in SEQ ID NOs: 23-26.
[0237] As used herein, the terms "polyketide synthase" and "PKS" are interchangeable and refer to any peptide, polypeptide, or protein capable of catalyzing the elongation of a ketide or polyketide chain. In some embodiments, the PKS activity transacylation. In some embodiments, the PKS activity includes Claisen condensation. In some embodiments, the PKS activity includes reducing a β-keto group to a β-hydroxy group. In some embodiments, the PKS activity includes HO decomposition, thereby yielding an α-β-unsaturated alkene, providing an α-β-unsaturated alkene, or generating an α-β-unsaturated alkene. In some embodiments, the PKS activity includes reducing an α-β-double bond to a single bond. In some embodiments, the PKS activity includes hydrolyzing a polyketide chain or a completed polyketide chain from the acyl carrier protein domain of the PKS. In some embodiments, the PKS activity includes polymerizing and / or ligating a diketide substrate into a polyketide chain. In some embodiments, the PKS activity comprises elongating a diketide into a polyketide chain. In some embodiments, the PKS activity comprises elongating a polyketide chain.
[0238] In some embodiments, the DNA molecule encodes a protein characterized by polyketide cyclization activity. In some embodiments, the DNA molecule encodes a protein that is a polyketide cyclase (PKC). In some embodiments, the PKC is a PKC derived from Helichrysum umbraculigerum. As used herein, the terms "polyketide cyclase" and "PKC" encompass any enzyme derived from H. umbraculigerum that possesses or is a functional analog of "olivetolic acid cyclase" or "OAC" from Cannabis sativa. In some embodiments, the DNA molecule encoding a protein characterized by polyketide cyclization activity comprises a nucleic acid sequence set forth in SEQ ID NOs: 31-38.
[0239] As used herein, the terms "polyketide cyclase" and "PKC" are interchangeable and refer to any peptide, polypeptide, or protein capable of folding and / or cyclizing polyketides. In some embodiments, PKC activity comprises the action of a cyclase subunit. In some embodiments, PKC activity comprises site-specific keto-reductase activity.
[0240] In some embodiments, the DNA molecule encodes a protein characterized by prenyl transfer activity. In some embodiments, the DNA molecule encodes a protein that is a prenyltransferase (PT). In some embodiments, the PT is a PT derived from Helichrysum umbraculigerum. As used herein, the terms "prenyltransferase" and "PT" encompass any enzyme derived from H. umbraculigerum that is functionally similar to or is a functional analog of "geranyl pyrophosphate:olivetolic acid geranyltransferase" or "GOT" from Cannabis sativa. In some embodiments, the GOT is GOT4 or CsGOT4. In some embodiments, the DNA molecule encoding a protein characterized by prenyl transfer activity comprises a nucleic acid sequence set forth in SEQ ID NOs: 47-58.
[0241] As used herein, the terms "prenyltransferase" and "PT" are interchangeable and refer to any peptide, polypeptide, or protein capable of transferring an aryl prenyl group to an acceptor molecule. In some embodiments, PT activity involves cyclization. In some embodiments, PT activity involves transferring an aryl prenyl group to an acceptor molecule.
[0242] In some embodiments, the DNA molecule encodes a protein characterized by cannabigerolic acid (CBGA) cyclization or cyclization activity. In some embodiments, the cyclization activity includes cyclization of CBGA to CBCA. In some embodiments, the polynucleotide encodes a protein capable of cyclizing or cyclizing CBGA to CBCA. In some embodiments, the DNA molecule encodes a protein capable of synthesizing CBCA or characterized as a CBCA synthase (CBCAS). In some embodiments, the CBCAS is a CBCAS derived from Helichrysum umbraculigerum. As used herein, the terms "CBCA synthase" and "CBCSA" encompass any enzyme derived from H. umbraculigerum that has or is a functional analog of a CBCA synthase (e.g., CsCBCAS) in Cannabis sativa. In some embodiments, the DNA molecule encoding a protein characterized by CBGA cyclization or cyclization activity comprises the nucleic acid sequence set forth in SEQ ID NOs: 71-79.
[0243] In some embodiments, the polynucleotide encodes a protein characterized by catalytic activity for transferring the glucuronic acid moiety of UDP-glucuronic acid to a small hydrophobic molecule (e.g., a UGT). In some embodiments, the polynucleotide encodes a protein characterized by glycosyltransferase catalytic activity. In some embodiments, the polynucleotide encodes a protein characterized by being able to transfer the glucuronic acid moiety of UDP-glucuronic acid to a cannabinoid or precursor thereof. In some embodiments, the polynucleotide encodes a protein characterized by having catalytic activity for glycosylating a cannabinoid or precursor thereof. In some embodiments, the polynucleotide encodes a UGT enzyme.
[0244] In some embodiments, the UGT is a UGT derived from Helichrysum umbraculigerum. As used herein, the term "UGT" encompasses any enzyme derived from H. umbraculigerum and having or characterized as having an activity described herein.
[0245] In some embodiments, the UGT protein is encoded by a DNA molecule comprising SEQ ID NOs: 89-101.
[0246] In some embodiments, the DNA molecule encodes a protein characterized by its ability to act on acyl groups. In some embodiments, the DNA molecule encodes a protein characterized by catalytic activity for transferring an acyl group from a donor molecule to an acceptor molecule. In some embodiments, the acceptor molecule is a hydrophobic molecule, a small molecule, or both. In some embodiments, the donor molecule comprises an acyl group, CoA, or both. In some embodiments, the DNA molecule encodes a protein characterized by acyltransferase catalytic activity. In some embodiments, the DNA molecule encodes a protein characterized by its ability to transfer an acyl group to a cannabinoid. In some embodiments, the DNA molecule encodes a protein characterized by having catalytic activity for acylation of cannabinoids. In some embodiments, the acyltransferase (AT) is an alcohol acyltransferase (AAT). In some embodiments, the DNA molecule encodes an AT enzyme. In some embodiments, the polynucleotide encodes an AAT enzyme.
[0247] In some embodiments, the AAT is AAT derived from Helichrysum umbraculigerum. As used herein, the term "AAT" encompasses any enzyme derived from H. umbraculigerum and having or characterized as having an activity described herein.
[0248] In some embodiments, the AAT protein is encoded by a DNA molecule comprising or consisting of SEQ ID NOs: 115-129.
[0249] In some embodiments, the artificial vector comprises a plasmid. In some embodiments, the artificial vector comprises an Agrobacterium containing the artificial nucleic acid molecule or is an Agrobacterium containing the artificial nucleic acid molecule. In some embodiments, the artificial vector is an expression vector. In some embodiments, the artificial vector is a plant expression vector. In some embodiments, the artificial vector is for use in expressing any one or any combination of AAE, PKS, PKC, PT, or CBCAS encoding nucleic acid sequences disclosed herein. In some embodiments, the artificial vector is further for use in expressing a UGT, an AAT, or both. In some embodiments, the artificial vector is for use in heterologous expression in a cell, tissue, or organism of any one or any combination of AAE, PKS, PKC, PT, or CBCAS encoding nucleic acid sequences disclosed herein. In some embodiments, the artificial vector is further for use in heterologous expression in a cell, tissue, or organism of a UGT, an AAT, or both. In some embodiments, the artificial vectors are for use in producing or producing acyl-coenzyme A (acyl-CoA), polyketides, cannabinoids, such as CBGA, CBCA, any precursors thereof, or any combination thereof, in a cell, tissue, or organism. In some embodiments, the artificial vectors are further used in producing or producing modified acyl-coenzyme A (acyl-CoA), polyketides, cannabinoids, such as CBGA, CBCA, any precursors thereof, or any combination thereof, in a cell, tissue, or organism, wherein the modifications further comprise an acyl group, a glycan (e.g., glycosylation), or both.
[0250] It is well known to those skilled in the art to express polynucleotides in cells. This can be achieved by transfection, viral infection, or direct modification of the cell's genome, among many other methods. In some embodiments, the DNA molecule is in an expression vector, such as a plasmid or viral vector. The vector nucleic acid sequence generally contains at least a replication origin for propagation in cells, and optionally further contains elements such as heterologous polynucleotide sequences, expression control elements (e.g., promoters, enhancers), selection markers (e.g., antibiotic resistance), polyadenine sequences, etc.
[0251] The vector may be a DNA plasmid delivered by non-viral or viral methods. The viral vector may be a retroviral vector, a herpesvirus vector, an adenovirus vector, an adeno-associated virus vector, a Virucaviridae virus vector, or a poxvirus vector. Barley stripe mosaic virus (BSMV), tobacco rattle virus, and cabbage leaf curl geminivirus (CbLCV) may also be used. The promoter may be active in plant cells. The promoter may be a viral promoter.
[0252] In some embodiments, the DNA molecule disclosed herein is operably linked to a promoter. The term "operably linked" is intended to mean that the nucleotide sequence of interest is linked to one or more regulatory elements in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). In some embodiments, a promoter is operably linked to the polynucleotide of the present invention. In some embodiments, the promoter is a heterologous promoter. In some embodiments, the promoter is an endogenous promoter.
[0253] In some embodiments, vectors are introduced into cells by standard methods including electroporation (e.g., as described in From et al., Proc. Natl. Acad. Sci. USA 82, 5824 (1985)), heat shock, infection with viral vectors, high velocity ballistic penetration by small particles carrying nucleic acid either within the matrix of small beads or particles or on the surface (Klein et al., Nature 327:70-73 (1987)), use of gene guns, e.g., coated particles, and needle-shaped particles, Agrobacterium Ti plasmids, etc.
[0254] As used herein, the term "promoter" refers to a group of transcriptional control modules clustered around the initiation site of RNA polymerase, i.e., RNA polymerase II. Promoters consist of individual functional modules, each consisting of approximately 7-20 bp of DNA, containing one or more recognition sites for transcriptional activator or repressor proteins. Promoters may extend upstream or downstream from the transcription start site and may be of any size, ranging from a few base pairs to several kilobases.
[0255] In some embodiments, the DNA molecule is transcribed by RNA polymerase II (RNAP II and Pol II). RNAP II is an enzyme found in eukaryotic cells that is known to catalyze the transcription of DNA to synthesize precursors of mRNA and most snRNAs and microRNAs.
[0256] In some embodiments, a plant expression vector is used. In one embodiment, the expression of the polypeptide coding sequence is driven by several promoters. In some embodiments, a viral promoter, such as the 35S RNA and 19S RNA promoters of CaMV [Brisson et al., Nature 310:511-514 (1984)] or the coat protein promoter for TMV [Takamatsu et al., EMBO J. 6:307-311 (1987)] is used. In another embodiment, a plant promoter, such as the small subunit promoter of RUBISCO [Coruzzi et al., EMBO J. 3:1671-1680 (1984); and Brogli et al., Science 224:838-843 (1984)] or a heat shock promoter, such as soybean hspl7.5-E or hspl7.3-B [Gurley et al., Mol. Cell. Biol. 6:559-565 (1986)] is used. In one embodiment, the construct is introduced into plant cells using Ti plasmid, Ri plasmid, plant virus vector, direct DNA transformation, microinjection, electroporation and other techniques well known to those skilled in the art.See, for example, Weissbach & Weissbach [Methods for Plant Molecular Biology, Academic Press, NY, Section VIII, pp421-463 (1988)].Other expression systems, such as insect and mammalian host cell systems well known in the art, can also be used according to the present invention.
[0257] In some embodiments, expression vectors containing regulatory elements derived from eukaryotic viruses, such as retroviruses, are used in accordance with the present invention. SV40 vectors include pSVT7 and pMT2. In some embodiments, vectors derived from bovine papillomavirus include pBV-1 MTHA, and vectors derived from Epstein-Barr virus include pHEBO and p205. Other exemplary vectors include pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo-5, baculovirus pDSVE, and any other vector that allows protein expression under the direction of the SV-40 early promoter, SV-40 late promoter, metallothionein promoter, mouse mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown to be effective for expression in eukaryotic cells.
[0258] In some embodiments, recombinant viral vectors are used for in vivo expression, offering advantages such as systemic infection and target specificity. In one embodiment, systemic infection is inherent in the life cycle of, for example, retroviruses, a process in which a single infected cell produces many progeny viral particles, which then infect nearby cells. In one embodiment, this results in the rapid infection of a large area, most of which were not initially infected by the original viral particle. In one embodiment, a viral vector is produced that cannot spread systemically. In one embodiment, this feature can be useful when the desired goal is to introduce a specific gene into only a localized number of target cells.
[0259] In some embodiments, plant viral vectors are used. In some embodiments, wild-type viruses are used. In some embodiments, degraded viruses are used as known in the art. In some embodiments, Agrobacterium is used to introduce the vectors of the invention into the viruses.
[0260] A variety of methods can be used to introduce the expression vector of the present invention into cells. Such methods are generally described in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Springs Harbor Laboratory, New York (1989, 1992), Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989), Chang et al., Somatic Gene Therapy, CRC Press, Ann Arbor, Mich. (1995), Vega et al., Gene Targeting, CRC Press, Ann Arbor, Mich. (1995), Vectors: A Study of Molecular Cloning Vectors and Their Uses, Butterworths, Boston Mass. (1988), and Gilboa et al., Biotechniques 4(6):504-512, 1986, and include, for example, stable or transient transfection, lipofection, electroporation, infection with Agrobacterium Ti plasmids and recombinant viral vectors. Additionally, see US Pat. Nos. 5,464,764 and 5,487,992 for positive-negative selection methods.
[0261] It will be understood that, in addition to containing the necessary elements for the transcription and translation of the inserted coding sequence (encoding a polypeptide), the expression constructs of the invention may also contain sequences designed to optimize the stability, production, purification, yield, or activity of the expressed polypeptide.
[0262] In some embodiments, the artificial vector comprises a polynucleotide encoding a protein comprising an amino acid sequence described herein.
[0263] According to some embodiments, there is provided a protein encoded by (a) a DNA molecule disclosed herein; (b) an artificial vector disclosed herein; or a plasmid or Agrobacterium disclosed herein.
[0264] In some embodiments, the protein is an isolated protein.
[0265] As used herein, the terms "peptide," "polypeptide," and "protein" are used interchangeably and refer to polymers of amino acid residues. In another embodiment, the terms "peptide," "polypeptide," and "protein" as used herein encompass natural peptides, peptidomimetics (usually containing non-peptide bonds or other synthetic modifications), and peptide analogs, peptoids, and semipeptoids, or any combination thereof. In another embodiment, the described peptides, polypeptides, and proteins have been modified to be more stable in an organism or to enhance their ability to penetrate cells. In one embodiment, the terms "peptide," "polypeptide," and "protein" apply to naturally occurring amino acid polymers. In another embodiment, the terms "peptide," "polypeptide," and "protein" apply to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids.
[0266] As used herein, the term "isolated protein" refers to a protein that is essentially free from contaminating cellular components, such as carbohydrates, lipids, or other proteinaceous impurities associated with naturally occurring nucleic acids. Typically, an isolated protein contains a highly purified form of the protein, for example, at least about 80% pure, at least about 90% pure, at least about 95% pure, greater than 95% pure, or greater than 99% pure. In some embodiments, the isolated protein is a synthesized protein. Protein synthesis is well known in the art and can be achieved, for example, by heterologous expression in transformed cells, as exemplified herein.
[0267] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 12) Comprises or consists of.
[0268] In some embodiments, the protein comprises an amino acid sequence having at least 91%, at least 93%, at least 95%, or at least 97% homology or identity to SEQ ID NO: 12, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 91%-97%, 92%-99%, 93%-98%, or 90%-100% homology or identity to SEQ ID NO: 12. Each possibility represents a separate embodiment of the present invention.
[0269] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 13) Comprises or consists of.
[0270] In some embodiments, the protein comprises an amino acid sequence having at least 83%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 13, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 83%-95%, 85%-99%, 83%-100%, or 84%-97% homology or identity to SEQ ID NO: 13. Each possibility represents a separate embodiment of the present invention.
[0271] In some embodiments, the protein has the amino acid sequence: MTEEEKNKAESMGIKTYAWSDFLHLGSKNPSELQTPKATDICTIMYTSGTSGDPKGVILTHENATTNIRGVDLFMEQFEDKMTVDDVYISFLPLAHILDRMIEEYFFRSGASVGFYHGDINALKEDLAELKPTFLAGVPRVLEKIHEGVLKGLEEVNPRRRKIFSILYNHKLKYMKAGYKHKYASPLADLLAFRKVKNRLGGRIRLMVSGGAPLSTEIEEFMRVTSCAFV AQGYGLTETCGLATLGFPDEMCMIGTVGSPFVYTELRLEEVSDMGYDPLANPPRGEICVKGKTPFAGYYKNPELTNEVMKDGWFHTGDIGEMQPNGVLKIIDRKKHLIKLSQGEYIALEYLEKVYCITPILEDIWVYGDSFKSSLVAVAVPNKENAEKWADQKGLKVSYSELCTLTQFRDYIQSELKSTAERNKLRGFEHIKAIIVEPRTFEGDQELLTATMKKRRNKLLNRYKEGIDNLYKNLAANKR (SEQ ID NO: 14) Comprises or consists of.
[0272] In some embodiments, the protein comprises an amino acid sequence having at least 86%, at least 88%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 14, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 86%-93%, 86%-95%, 88%-97%, or 86%-100% homology or identity to SEQ ID NO: 14. Each possibility represents a separate embodiment of the present invention.
[0273] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 15) Comprises or consists of.
[0274] In some embodiments, the protein comprises an amino acid sequence having at least 86%, at least 88%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 15, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 86%-93%, 86%-95%, 88%-97%, or 86%-100% homology or identity to SEQ ID NO: 15. Each possibility represents a separate embodiment of the present invention.
[0275] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 16) Comprises or consists of.
[0276] In some embodiments, the protein comprises an amino acid sequence having at least 89%, at least 92%, at least 94%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 16, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 89%-95%, 89%-98%, 90%-99%, or 89%-100% homology or identity to SEQ ID NO: 16. Each possibility represents a separate embodiment of the present invention.
[0277] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 17) Comprises or consists of.
[0278] In some embodiments, the protein comprises an amino acid sequence having at least 93%, at least 94%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 17, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 93%-98%, 93%-99%, 93%-100%, or 95%-100% homology or identity to SEQ ID NO: 17. Each possibility represents a separate embodiment of the present invention.
[0279] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 18) Comprises or consists of.
[0280] In some embodiments, the protein comprises an amino acid sequence having at least 84%, at least 87%, at least 91%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 18, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 84%-99%, 85%-99%, 84%-100%, or 90%-100% homology or identity to SEQ ID NO: 18. Each possibility represents a separate embodiment of the present invention.
[0281] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 19) Comprises or consists of.
[0282] In some embodiments, the protein comprises an amino acid sequence having at least 82%, at least 87%, at least 91%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 19, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 82%-99%, 83%-99%, 82%-100%, or 85%-100% homology or identity to SEQ ID NO: 19. Each possibility represents a separate embodiment of the present invention.
[0283] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO:20) Comprises or consists of.
[0284] In some embodiments, the protein comprises an amino acid sequence having at least 86%, at least 88%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO:20, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 86%-93%, 86%-95%, 88%-97%, or 86%-100% homology or identity to SEQ ID NO:20. Each possibility represents a separate embodiment of the present invention.
[0285] In some embodiments, the protein has the amino acid sequence: MTFQQLRSEVWLVAYALDTLGVEKGSAIAIDMPMDVKSVVIYLAIVLAGYVVVSIADSFAAGEISTRLVLSKAKAIFTQDLIIRGDRSHPLYSRVVDAQSPLAIVIPTRGSSFSIKLRDGDISWHDFLERANTYRNVEFVAVERPVEAFSNILFSSGTTGEPKAIPWTLATPFKAGADAWCHMDVHKGDVVAWPTNLGWMMGPWLIYASLLNGGSLALYNGSPLTSGFAKFVQDAKVTLLGVIPSIVRAW RTNNSTAGFDWSTIRCFGSTGEASNTDECLWLMGRAHYKPVIEYCGGTEIGGGFITGSLLQPQCLSAFSTPSLGCKLLILGEDGIPIPQNAPGIGELALNPLMFGASSTLLNANHYDVYFKGMPSWNGKVLRRHGDVFERTSKGYYRAHGRADDTMNLGGIKVSSVEIERVCNSIDDRILETAAIGVTPSGGGPERLVIVVAFKDGSGSKPDLIKLKVTLNSALQKNLNPLFKVSDVVPFPSLPRTATNKVMRRVLRQQLTQIGQNSKL (SEQ ID NO: 21) Comprises or consists of.
[0286] In some embodiments, the protein comprises an amino acid sequence having at least 89%, at least 92%, at least 94%, at least 97%, or at least 99% homology or identity to SEQ ID NO:21, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 89%-95%, 89%-98%, 90%-99%, or 89%-100% homology or identity to SEQ ID NO:21. Each possibility represents a separate embodiment of the present invention.
[0287] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 22) Comprises or consists of.
[0288] In some embodiments, the protein comprises an amino acid sequence having at least 88%, at least 92%, at least 94%, at least 97%, or at least 99% homology or identity to SEQ ID NO:22, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 88%-95%, 89%-98%, 90%-99%, or 88%-100% homology or identity to SEQ ID NO:22. Each possibility represents a separate embodiment of the present invention.
[0289] In some embodiments, the protein has the amino acid sequence: MASSINISKIREAQRAQGPASILAVGTANPSNCVYQADYPDYYFRITKSEHMVDLKRKFKRMCDQSMIRKRYMQITEEYLKENPNICEYMAPSLDARQDVVVVEVPKLGKEAATKAIKEWGQPKSKITHLIFCTTSGVDMPGADYQLTKLLGLCPSVKRFMMYQQGCFAGGTVLRLAKDIAENNKGARVLVVCSEITAVIFRGPNDTHLDSLIGQALFGDGASSVIVGSDPDLTTERPLFEIISAAQTILPDSEGAIDGHLREAGLTFHLLKDVPRLISKNIEKALTQAFSPLGISDWNSIFWVTHPGGPAILDQVELKLGLKEEKMRTTRHVLSEYGNMSSACVFFVLDEMRKRSAKGGARTTGEGLDWGVLFGFGPGLTVETVVLHSLPTTMSIAT (SEQ ID NO: 27) Comprises or consists of.
[0290] In some embodiments, the protein comprises an amino acid sequence having at least 92%, at least 96%, at least 98%, or at least 99% homology or identity to SEQ ID NO: 27, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 92%-100%, 95%-100%, 96%-100%, or 98%-100% homology or identity to SEQ ID NO: 27. Each possibility represents a separate embodiment of the present invention.
[0291] In some embodiments, the protein has the amino acid sequence: MASSINISKIREAQRAQGPASILAVGTANPSNCVYQADYPDYYFRITKSEHMVDLKEKFQRMCDKSMIRKRHIHITEEFLKENPNLCEYMAPSLDTRQDVVVVEVPKLGKEAATKAIKEWGQPKSKITHLIFCTTSGVDMPGADYQLTKLLGLHPSVKRFMMYQQGCFAGGTVLRLAKDLAENNKGARVLAVCSEITAVTFRGPNDTHIDSLVGQALFGDGAAAVIVGSDPDLTTERPLFEIISAAQTILPNSEGAIDGHVREVGVTIHILKDVPVLISKNIEKALTQAFSPLGISDWNSIFWVVHPGGPAILDQVELKLGLKEEKMRTTRHVLSEYGNMSSACVFFVLDEMRKRSAKGGARTTGEGLDWGVLFGFGPGLTVETVVLHSLPTTMSIAT (SEQ ID NO: 28) Comprises or consists of.
[0292] In some embodiments, the protein comprises an amino acid sequence having at least 91%, at least 94%, at least 95%, or at least 97% homology or identity to SEQ ID NO: 28, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 91%-100%, 94%-100%, 97%-100%, or 97%-100% homology or identity to SEQ ID NO: 28. Each possibility represents a separate embodiment of the present invention.
[0293] In some embodiments, the protein has the amino acid sequence: MASSINISKIREAQRAQGPASILAVGTANPSNCVYQADYPNYYFRITKSEHMVDLKRKFKRMCDQSMIRKRYMQITEEYLKENPNICEYMAPSLDARQDVVVVEVPKLGKEAATKAIKEWGQPKSKITHLIFCTTSGVDMPGADYQLTKLLGLCPSVKRFMMYQQGCFAGGTVLRLAKDIAENNKGARVLVVCSEITAVIFRGPNDTHLDSLIGQALFGDGASSVIVGSDPDLTTERPLFEIISAAQTILPDSEGAIDGHLREAGLTFHLLKDVPGLISKNIEKALTQAFSPLGISDWNSIFWVTHPGGPAILDQVELKLGLKEEKMRASRHVLSEYGNMSSACVFFILDEMRKKSDEDGAPTTGEGLDWGVLFGFGPGLTVETVVLHSLPTTMSIAT (SEQ ID NO: 29) Comprises or consists of.
[0294] In some embodiments, the protein comprises an amino acid sequence having at least 93%, at least 95%, or at least 97% homology or identity to SEQ ID NO: 29, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 93%-100%, 94%-100%, 96%-100%, or 98%-100% homology or identity to SEQ ID NO: 29. Each possibility represents a separate embodiment of the present invention.
[0295] In some embodiments, the protein has the amino acid sequence: MASSINISKIREAQRAQGPASILAVGTANPSNYEIQADFPDYYFRVTKSEHMADMKGTFQRMCDKSMIRKRHMLITEEFLKENPNLCEYMAPSLDTRQDVVVVEVPKLGKEAATKAIKEWGQPKSKITHLIFCTTTGVDMPGADYQLTKLLGLAPSVKRFMIYQQGCFAGGTVLRLAKDIAENNKGARVLAVCSEITAMSFRGPNDTHVDSLVGQALFGDGAAAVIVGSDPDLTTERPLFEIISAAQTILPNSEGAIDGHVREVGLTIHILKDVPVLISKNIEKALTQAFSPLGISDWNSIFWIVHPGGPAILDQVELKVGLKKEKMATSRHVLSEYGNMSSACVFFIMDEMRKRSAKGGARTTGEGLDWGVLFGFGPGLTVETVVLHSLPTTM (SEQ ID NO: 30) Comprises or consists of.
[0296] In some embodiments, the protein comprises an amino acid sequence having at least 88%, at least 92%, at least 95%, or at least 97% homology or identity to SEQ ID NO: 30, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 88%-100%, 91%-100%, 93%-100%, or 95%-100% homology or identity to SEQ ID NO: 30. Each possibility represents a separate embodiment of the present invention.
[0297] In some embodiments, the protein has the amino acid sequence: MAEFTHLVVVKFKEEVVVEDIMKGLEKLVSQLDSVKSFVWGKDIESMEMLRQGFTHAIMMTFGSKEDFTAFQSHPNHVEFSATFSAAIEKIVLLDFPVVAVKTATA (SEQ ID NO: 39) Comprises or consists of.
[0298] In some embodiments, the protein comprises an amino acid sequence having at least 86%, at least 91%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 39, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 86%-99%, 88%-98%, 90%-99%, or 89%-100% homology or identity to SEQ ID NO: 39. Each possibility represents a separate embodiment of the present invention.
[0299] In some embodiments, the protein has the amino acid sequence: MSSLQNKFIEHIALIKIKPGVESTTLIDKLNGLSSIEVLLHFSAGELLGSSHGFTHIVHCRVRSKDDLQIYLTHPIHLHLADDTLPLLDDVTVVDWFSSNSDIVDPPKPGSAMRVTLLKLKHDSTESNKLVVIEGIKNQFKGIEDVIVTTTFGENLFHEMHENFSIEIDKGYSIGSIAFVPGSADFQVLNSKVDNNKLNDLTESEVVVDYVFPSAN (SEQ ID NO: 40) Comprises or consists of.
[0300] In some embodiments, the protein comprises an amino acid sequence having at least 45%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 40, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 45%-90%, 50%-99%, 65%-98%, or 55%-100% homology or identity to SEQ ID NO: 40. Each possibility represents a separate embodiment of the present invention.
[0301] In some embodiments, the protein has the amino acid sequence: MSSEEQIVEHVVLFKVKPDADPSKVAAWVNGLNGLTSLQLALHLSAGQLIRCRSSSLTFTHMLHSRYRSKEHLRQYTVHPEHVRVVTEGKSIIDDVMALDWMISNGAASSVCPKPGSAVRVGFYKLMESLGEIEKARVLEVMGGIEELSVGESFCDDRAKGYTIASTAVFPNGNPAADLDLYHSGDQLLLKEEVMKDSIQSVVVVDYVIPSP (SEQ ID NO: 41) Comprises or consists of.
[0302] In some embodiments, the protein comprises an amino acid sequence having at least 71%, at least 80%, at least 90%, or at least 99% homology or identity to SEQ ID NO: 41, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 71%-97%, 75%-99%, 80%-98%, or 71%-100% homology or identity to SEQ ID NO: 41. Each possibility represents a separate embodiment of the present invention.
[0303] In some embodiments, the protein has the amino acid sequence: MGEVKHILLAKFKDGISEQQIQHLITGYANLVNLVEPMKSFRWGKDVSIENLHQGFTHVFESTFETTEGIATYISHPAHVEFATGFLDQLEKVIVIDYKPTSVDP (SEQ ID NO: 42) Comprises or consists of.
[0304] In some embodiments, the protein comprises an amino acid sequence having at least 87%, at least 92%, at least 96%, or at least 97% homology or identity to SEQ ID NO: 42, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 87%-97%, 88%-99%, 90%-98%, or 87%-100% homology or identity to SEQ ID NO: 42. Each possibility represents a separate embodiment of the present invention.
[0305] In some embodiments, the protein has the amino acid sequence: MLCAPARTRLLPSISLLPSQHNIFRRLNCLIHRRNHHQTPITMSAQQQIVEHVVLFKVKPDVDSSKVAAMVNGLNGLTSLDLTLHLSAGQLLRSRSSSLTFTHMLHSRYRSKDDLREYAAHPDHVRVVTENIKPVIDDIMAVDWISNDASVSPKPGSAMRVTFLKLKENLGENEKSRVLEVIGGIKNQFKSIEELSVGENFSHDRAKGYTIASIAVLPGPSELEALDSNTELVKLEKEKVKDLLESVVVVDYVIPSLQSASL (SEQ ID NO: 43) Comprises or consists of.
[0306] In some embodiments, the protein comprises an amino acid sequence having at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 43, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 85%-97%, 87%-99%, 89%-98%, or 85%-100% homology or identity to SEQ ID NO: 43. Each possibility represents a separate embodiment of the present invention.
[0307] In some embodiments, the protein has the amino acid sequence: MAVAQLSSSLCISTPARISTGSGFSSSGLPRIGTTFVCGSGSPLVISGTYHQKARVHKPAALSVRCEQSSKDGNGLNVWLGRTAMVGFAVAISVEVSTGKGLLENFGLTSPLPTVALALTALGGVLTALFIFQSASES (SEQ ID NO: 44) Comprises or consists of.
[0308] In some embodiments, the protein comprises an amino acid sequence having at least 79%, at least 82%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 44, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 79%-95%, 79%-99%, 80%-98%, or 79%-100% homology or identity to SEQ ID NO: 44. Each possibility represents a separate embodiment of the present invention.
[0309] In some embodiments, the protein has the amino acid sequence: MIEHIVLLKFKSDVDSTKVESMINELNGLASLDVALDVSAGKILRVSSTSSSSLTFTHLFRCCFRSADDQQVFSTHPDHLRVAIEVRPVIEDMVVVDLVSKTTIDSPNPGSAMKVRIFKLKDDLIEDSKLVVMEGIKNELKAVEHIRFGDNINVMAKGYSIAMIAFFPDLESSVAGAEIVKDYIESELVVDFVFPPPNVTSHS (SEQ ID NO: 45) Comprises or consists of.
[0310] In some embodiments, the protein comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 45, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 50%-90%, 55%-99%, 60%-97%, or 50%-100% homology or identity to SEQ ID NO: 45. Each possibility represents a separate embodiment of the present invention.
[0311] In some embodiments, the protein has the amino acid sequence: MAEFTHLVVVKFKEEVVVEDIMKGLEKLASQLDSVKSFVWGKDIESMEMLRQGFTHAIMMTFGSKEDFTAFQSHPNHVEFSATFSAAIEKIVLLDFPVVAVKTATA (SEQ ID NO: 46) Comprises or consists of.
[0312] In some embodiments, the protein comprises an amino acid sequence having at least 87%, at least 93%, at least 95%, or at least 97% homology or identity to SEQ ID NO: 46, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 87%-97%, 88%-99%, 89%-98%, or 87%-100% homology or identity to SEQ ID NO: 46. Each possibility represents a separate embodiment of the present invention.
[0313] In some embodiments, the protein has the amino acid sequence: MELSLSSSSSSSLPQLHTHPSSSSSSSHYIKKSPFFINKFNNHTKCKFHNSSALRTNFFYTTITKTSSSRFVLNKNPNQFSVKACSQVGSAGSDPALNKVADFKDAFWRFLRPHTIRGTALGSVSLVTRALLENPNLIRWSLLLKAFSGLVALICGNGYIVGINQIYDIGIDKVNKPYLPIAAGDLSVQSAWFLVLAFAMVGVIIVGMNFGPFITSLYSLGLFLGTIYSVPPLRMKRFPVVAFLIIATVRGFLLNFGVYYAVRAALGLTFQWSSAVAFITTFVTLFALVIAITKDLPDVEGDRKFQISTFATKLGVRNIALLGSGLLLINYIGSIVAALYMPQAFRSSLMIPLHTILASCLIYQAWILERANYTQEAIAGYYRFVWNLFYSEYIIFPFI (SEQ ID NO: 59) Comprises or consists of.
[0314] In some embodiments, the protein comprises an amino acid sequence having at least 82%, at least 85%, at least 90%, or at least 99% homology or identity to SEQ ID NO: 59, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 82%-99%, 85%-98%, 84%-99%, or 82%-100% homology or identity to SEQ ID NO: 59. Each possibility represents a separate embodiment of the present invention.
[0315] In some embodiments, the protein has the amino acid sequence: MATMASSLLNPLSCSIKPNSNRLPLPTPISLSRSCRRLTIKATETDANEVKPKAPEKAPAASGSGFNQILGIKGAKQETNKWKIRVQLTKPVTWPPLIWGVVCGAAASGNFQWTVEDVAKSIVCMLMSGPFLTGYTQTINDWYDRDIDAINEPYRPIPSGAISENEVITQIWVLLLGGIGLAGILDVWAGHKSPTIFYLALGGSLLSYIYSAPPLKLKQNGWIGNFALGASYISLPWWAGQALFGTLTPDIVVLTLLYSIAGLGIAIVNDFKSVEGDRKMGLQSLPVAFGEETAKWICVGAIDITQLSIAGYLLGSGKPYYALALVGLIVPQIFFQFKYFLKDPVKYDVKYQASAQPFLILGLLVTALATSH (SEQ ID NO: 60). Comprises or consists of.
[0316] In some embodiments, the protein comprises an amino acid sequence having at least 92%, at least 93%, at least 95%, at least 97%, at least 98%, or at least 99% homology or identity to SEQ ID NO: 60, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 92%-98%, 93%-99%, 94%-98%, or 92%-100% homology or identity to SEQ ID NO: 60. Each possibility represents a separate embodiment of the present invention.
[0317] In some embodiments, the protein has the amino acid sequence: MKSLIIGSFSNKVSCYSPSLPDSSSSLIPTGCYHVSLRTFQRNRAIQAQSSLVRCNIGKFNETLLLSRKRSTKHVACAVSEQPIEPDATNPQSSLPNALDAFYRFSRPHTVIGTALSIVSVSLLAVQKLSDFSPLFFIGVFEAIVAAFFMNIYIVGLNQLSDIEIDKVNKPYLPLASGEYSVQTGIIIVSSFAVMSFWLGWIVGSWPLFWALFISFLLGTAYSINIPMLRWKRFALVAAMCILAVRAIIVQVAFYLHIQTFVYGRLAVFPKPVIFATGFMSFFSVVIALFKDIPDIVGDKIFGIQSFTVRMGQKRVFWICILLLEIAYGVAILVGASSPFLWSRYITVLGHAILGLILWGRAKSTDLESKSAITSFYMFIWQLFYAEYLLIPLVR (SEQ ID NO: 61) Comprises or consists of.
[0318] In some embodiments, the protein comprises an amino acid sequence having at least 89%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 61, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 89%-97%, 89%-99%, 90%-98%, or 89%-100% homology or identity to SEQ ID NO: 61. Each possibility represents a separate embodiment of the present invention.
[0319] In some embodiments, the protein has the amino acid sequence: MELSLSSSSSSSLPQLHTHPSSSSSSSHYIKKSPFFINKFNNHTKCKFHNSSALRTNFFYTTITKTSSSRFVLNKNPNQFSVKACSQVGSAGSDPALNKVADFKDAFWRFLRPHTIRGTALGSVSLVTRALLENPNLIRWSLLLKAFSGLVALICGNGYIVGINQIYDIGIDKVNKPYLPIAAGDLSVQSAWFLVLAFAMVGVIIVGMNFGPFITSLYSLGLFLGTIYSVPPLRMKRFPVVAFLIIATVRGFLLNFGVYYAVRAALGLTFQWSSAVAFITTFVTLFALVIAITKDLPDVEGDRKFQISTFATKLGVRNIALLGSGLLLINYIGSIVAALYMPQAFRSSLMIPLHTILASCLIYQAWILERANYTQRSQYFDMSSCRRR (SEQ ID NO: 62). Comprises or consists of.
[0320] In some embodiments, the protein comprises an amino acid sequence having at least 81%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 62, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 81%-97%, 83%-99%, 84%-98%, or 81%-100% homology or identity to SEQ ID NO: 62. Each possibility represents a separate embodiment of the present invention.
[0321] In some embodiments, the protein has the amino acid sequence: MELSLSSSSSSSLPQLHTHPSSSSSSSHYIKKSPFFINKFNNHTKCKFHNSSALRTNFFYTTITKTSSSRFVLNKNPNQFSVKACSQVGSAGSDPALNKVADFKDAFWRFLRPHTIRGTALGSVSLVTRALLENPNLIRWSLLLKAFSGLVALICGNGYIVGINQIYDIGIDKVNKPYLPIAAGDLSVQSAWFLVLAFAMVGVIIVGMNFGPFITSLYSLGLFLGTIYSVPPLRMKRFPVVAFLIIATVRGFLLNFGVYYAVRAALGLTFQWSSAVAFITTFVTLFALVIAITKDLPDVEGDRKFQISTFATKLGVRNIALLGSGLLLINYIGSIVAALYMPQVKTTSIDHYRPYSFLVDLPGQNGITLAA (SEQ ID NO: 63) Comprises or consists of.
[0322] In some embodiments, the protein comprises an amino acid sequence having at least 81%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 63, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 81%-97%, 83%-99%, 84%-98%, or 81%-100% homology or identity to SEQ ID NO: 63. Each possibility represents a separate embodiment of the present invention.
[0323] In some embodiments, the protein has the amino acid sequence: MATMASSLLNPLSCSIKPNSNRLPLPLPIPISLSRSCRRLTIKATETDANEVKPKAPEKAPAASGSGFNQILGIKGAKQETNKWKIRVQLTKPVTWPPLIWGVVCGAAASGNFQWTVEDVAKSIVCMLMSGPFLTGYTQTINDWYDRDIDAINEPYRPIPSGAISENEVITQIWVLLLGGIGLAGILDVWAGHKSPTIFYLALGGSLLSYIYSAPPLKLKQNGWIGNFALGASYISLPWWAGQALFGTLTPDIVVLTLLYSIAGLGIAIVNDFKSVEGDRKMGLQSLPVAFGEETAKWICVGAIDITQLSIAGYLLGSGKPYYALALVGLIVPQIFFQFKYFLKDPVKYDVKYQASAQPFLILGLLVTALATSH (SEQ ID NO: 64) Comprises or consists of.
[0324] In some embodiments, the protein comprises an amino acid sequence having at least 92%, at least 93%, at least 95%, at least 97%, at least 98%, or at least 99% homology or identity to SEQ ID NO:64, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 92%-98%, 93%-99%, 94%-98%, or 92%-100% homology or identity to SEQ ID NO:64. Each possibility represents a separate embodiment of the present invention.
[0325] In some embodiments, the protein has the amino acid sequence: MASLAIGSLGSPSSRQCSSPVASSSSFAIGSQIASKFLRISKFDKTKNSPLTLQQKHINKSIDQSFFEPLPLPLHKINKDKFKLYATSTNNPQFDATHDLKTPEVSIINFVDALYRLIRPYTAVVTIVSVVAMSLLTVNSLSDFSPLFFIKVVQALIGGIFMQMYVSGFNQICDIELDKVNKQSLPLAAGELSMKTAIVIASLSAIMSLSIGWFVGSPPLLWCLVWWFIVGTAYSANVLPYLRWKRFPFTAAFCAMTSRALVLPIGYYLHMQNSIPGVSALLSRPILFAVAMLSAFSLSAMFFKDIPDIKGDRMHGIKSLAIKLGEKRVYWISISIIEIAYIAAAFIGATSPISWSKYVTIIGHLGMGLLLWVRARSVDPTNTVAVQSMYMFLIKLVYAEYGLISLVR (SEQ ID NO: 65) Comprises or consists of.
[0326] In some embodiments, the protein comprises an amino acid sequence having at least 71%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 65, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 71%-90%, 75%-99%, 73%-97%, or 71%-100% homology or identity to SEQ ID NO: 65. Each possibility represents a separate embodiment of the present invention.
[0327] In some embodiments, the protein has the amino acid sequence: MKSLIIGSFSNKVSCYSPSLPDSSSSLIPTGCYHVSLRTFQRNRAIQAQSSLVRCNIGKFNETLLLSRKRSTKHVACAVSEQPIEPDATNPQSSLPNALDAFYRFSRPHTVIGTALSIVSVSLLAVQKLSDFSPLFFIGVFEAIVAAFFMNIYIVGLNQLSDIEIDKVNKPYLPLASGEYSVQTGIIIVSSFAVMSFWLGWIVGSWPLFWALFISFLLGTAYSINIPMLRWKRFALVAAMCILAVRAIIVQVAFYLHIQTFVYGRLAVFPKPVIFATGFMSFFSVVIALFKDIPDIVGDKIFGIQSFTVRMGQKRVFWICILLLEIAYGVAILVGASSPFLWSRYITVLGHAILGLILWGRAKSTDLESKSAITSFYMFIWQLFYAEYLLIPLVR (SEQ ID NO: 66) Comprises or consists of.
[0328] In some embodiments, the protein comprises an amino acid sequence having at least 89%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 66, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 89%-97%, 89%-99%, 90%-98%, or 89%-100% homology or identity to SEQ ID NO: 66. Each possibility represents a separate embodiment of the present invention.
[0329] In some embodiments, the protein has the amino acid sequence: MLIHHEHFLTTGFESSNDRAAYSINFSKQHHLHMASIATGSLCRPTSHQFSIPVASSSSFATGSQFASKFLHISISAKKSSLTLQQRHIHKNIDQSFLKPLALQKLNKDKFKLNGTSPDNPQFDATHDLKTQIESTINFVDVLYRLLRPYALLQMGLCVVTMSLLTVESLSDFSPLFFVKVAQALIGGIFMQMYVNGFNQICDIELDK VNKPSLPLASGELSKTTTIVVSSLSAITSLSIGWFVGSPPLLWSLVVWFIAGTTYSANLPYLRWKRFPFTNMFCNLTMALVVPIGTYLHMENSIHGVSTLLSRPLLFTVAMCTVFPVSIILFKDIPDIKGDRMHGMKSLAIILGEKRTYWICIWILEITYIAAAFFGATSPISWSKYVTIISHLGMGFLLWLRSKSVDVKNTVAVQSMYMFLWKLLYAEYGLILLVR (SEQ ID NO: 67) Comprises or consists of.
[0330] In some embodiments, the protein comprises an amino acid sequence having at least 68%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 67, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 68%-97%, 69%-99%, 70%-98%, or 68%-100% homology or identity to SEQ ID NO: 67. Each possibility represents a separate embodiment of the present invention.
[0331] In some embodiments, the protein has the amino acid sequence: MFIHHEQFLTTGFESSNDRAAYSINFLKQHHLHMVSIATGSLCRPTSHRFSIPVASSSSFATGSQFASISAKKSSLTLKQRHTHKNIDQSFFKPLALQKMNKGKFKLNATSPDNSQLDATHDLKTQIESIINFVDVLYRLIRPYVVLGMGVTIVTMCLLTVDSLSDFSPLFFVKVAQALIGSIFMAMYVNSFNEICDIELDKVNKP SLPLASGELSMTTAIVVSSLSAIMSLSIGWFVGSPPLLWSLVVWFILGTAYSANLPYLRWKRFPLTTLSSALTMGALVIPIGNYMHMENSIRGVTTLLSRPLLFAVAMCAAFHVSTILFKDIPDIKGDRMHGMKSLAIKLGEKRMYWICIWILEIAYIAAAFFGATSPISWSKYVTIISHLGMGFLLWLRSKSVDVKNTVAVQSMYMFLWKLFYVEHGLILLVR (SEQ ID NO: 68) Comprises or consists of.
[0332] In some embodiments, the protein comprises an amino acid sequence having at least 66%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 68, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 66%-97%, 67%-99%, 70%-98%, or 66%-100% homology or identity to SEQ ID NO: 68. Each possibility represents a separate embodiment of the present invention.
[0333] In some embodiments, the protein has the amino acid sequence: MASIATGSLCRPTSHRFSIHVASSSSFATGSQFASKILQISISAKKSSLTLQQRHIHKNIDQSFFKPLALQKMNKDKFKLNATSPDNPQFDATRDLKTQIESIIKFVDVLYRLLRPYAILEMGLSVVTMSLLTVESLSDFSPLFFVKVAQALIGGIFMQMYVNGFNQICDIELDKVNKPSLPLASGELSTTT TIVVSSLSAIMSLSIGWFVGSPPLLWSLVVWFIVGTTYSTNLPYLRWKRFPFTAMFCNLTRALVVPIGTYLHMKNSIHEVSTLLSRPLLFAVAMCTVFPISIILFKDIPDIKGDRMHGMKSLAIILGEERTYWICIWILEIAYIAAAFFGATSPISWSKYVMIISHLGMGFLLWLRSKSVDVKNTVAVQSMYMFLWKLLYAEYGLILLVR (SEQ ID NO: 69) Comprises or consists of.
[0334] In some embodiments, the protein comprises an amino acid sequence having at least 68%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 69, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 68%-97%, 69%-99%, 70%-98%, or 68%-100% homology or identity to SEQ ID NO: 69. Each possibility represents a separate embodiment of the present invention.
[0335] In some embodiments, the protein has the amino acid sequence: MASLAIGSLGSPSSRQCSSPVASSSSFAIGSQIASKFLRISKFDKTKNSPLALQQKHINKSIDQSFFEPLPLPLHKINKDKFKLYATSTNNPQFDATHDLKTPEVSIINFVDALYRLIRPYTAVVTIVSVVAMSLLTVNSLSDFSPLFFIKVVQALIGGIFMQMYVSGFNQICDIELDKVNKQSLPLAAGELSMKTAIVIASLSAIMSLSIGWFVGSPPLLWCLVWWFIVGTAYSANVLPYLRWKRFPFTAAFCAMTSRALVLPIGYYLHMQNSIPGVSALLSRPILFAVAMLSAFSLSAMFFKDIPDIKGDRMHGIKSLAIKLGEKRVYWISISIIEIAYIAAAFIGATSPISWSKYVTIIGHLGMGLLLWVRARSVDPTNTVAVQSMYMFLIKLVYAEYGLISLVR (SEQ ID NO: 70) Comprises or consists of.
[0336] In some embodiments, the protein comprises an amino acid sequence having at least 71%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 70, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 71%-100%, 80%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 70. Each possibility represents a separate embodiment of the present invention.
[0337] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 80) Comprises or consists of.
[0338] In some embodiments, the protein comprises an amino acid sequence having at least 69%, at least 75%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 80, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 69%-99%, 70%-98%, 75%-99%, or 69%-100% homology or identity to SEQ ID NO: 80. Each possibility represents a separate embodiment of the present invention.
[0339] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 81) Comprises or consists of.
[0340] In some embodiments, the protein comprises an amino acid sequence having at least 81%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 81, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 92%-98%, 93%-99%, 94%-98%, or 92%-100% homology or identity to SEQ ID NO: 81. Each possibility represents a separate embodiment of the present invention.
[0341] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 82) Comprises or consists of.
[0342] In some embodiments, the protein comprises an amino acid sequence having at least 86%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 82, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 86%-97%, 87%-99%, 88%-98%, or 86%-100% homology or identity to SEQ ID NO: 82. Each possibility represents a separate embodiment of the present invention.
[0343] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 83) Comprises or consists of.
[0344] In some embodiments, the protein comprises an amino acid sequence having at least 69%, at least 75%, at least 80%, at least 85%, at least 92%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 83, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 69%-97%, 70%-99%, 75%-98%, or 69%-100% homology or identity to SEQ ID NO: 83. Each possibility represents a separate embodiment of the present invention.
[0345] In some embodiments, the protein has the amino acid sequence: MDQYVITKFISYLLAVFMALFCSDPTADKFLQCFTKDSNATDSNFVFTQENTQYSSVLESTIINLRFATSITPKPIAVITPLSYSHVQSAILCSKKIGYRIRIRSGGHDYAGVSYTSYDHDHTPFVVLDLKELRTITIDSGENTSWVESGATVGELYYWVSQKSRNLGFPAGICPTVGVGGHLSGGGVGTMVRKYGLAADNVIDARIIDVNGRILDRKSMGEDLFWAIRGGGGASFGVIVAWKVNLVYVPEKSFGF (SEQ ID NO: 84) Comprises or consists of.
[0346] In some embodiments, the protein comprises an amino acid sequence having at least 84%, at least 87%, at least 90%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 84, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 84%-97%, 86%-99%, 85%-98%, or 84%-100% homology or identity to SEQ ID NO: 84. Each possibility represents a separate embodiment of the present invention.
[0347] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 85) Comprises or consists of.
[0348] In some embodiments, the protein comprises an amino acid sequence having at least 72%, at least 75%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 85, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 72%-99%, 74%-98%, 78%-99%, or 72%-100% homology or identity to SEQ ID NO: 85. Each possibility represents a separate embodiment of the present invention.
[0349] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 86) Comprises or consists of.
[0350] In some embodiments, the protein comprises an amino acid sequence having at least 69%, at least 75%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 86, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 69%-99%, 70%-98%, 75%-99%, or 69%-100% homology or identity to SEQ ID NO: 86. Each possibility represents a separate embodiment of the present invention.
[0351] In some embodiments, the protein has the amino acid sequence: MGEDLFWAIRGGGGGSFGVVVAWMVNLVHVPEKVTAFTIVRTLEQGGSDLFNKWQHVGPKLTKDLFISVIIQPISVWNGNGTVQVIFNSMYLGTVDKLMKTVNSSFPELGLQAKDCTEMSWIQSVLYFAGYPIEGSMDVLKDRKPQTRRYFNNKSDHVKEPIPKERLEDLWKWCMEGDFPILLMDPLGGKMNEIDTTRIPYPYRNGYSYMIQYVETWENIGDSEKRISWMRQMYENMTPYVSKNPRSAYVNYRDLDLGKNDNAKNTSYLEAMKWGSKYFGDNFKRLAMVKGVVDPDNFFFHEQSIPPLKV (SEQ ID NO: 87) Comprises or consists of.
[0352] In some embodiments, the protein comprises an amino acid sequence having at least 71%, at least 75%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 87, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 75%-99%, 74%-98%, 78%-99%, or 71%-100% homology or identity to SEQ ID NO: 87. Each possibility represents a separate embodiment of the present invention.
[0353] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 88) Comprises or consists of.
[0354] In some embodiments, the protein comprises an amino acid sequence having at least 74%, at least 79%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 88, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 74%-99%, 78%-98%, 81%-99%, or 74%-100% homology or identity to SEQ ID NO: 88. Each possibility represents a separate embodiment of the present invention.
[0355] In some embodiments, the protein has the amino acid sequence: MTNSELVFIPSPGAGHLPPTVELAKLLLHREPQLSVTIIIMNLPHETKPTTETRMSTPRLRFIDIPKDESTKDLISRHTFISAFLEHQKPHVRNIVRSITESDSVRLVGFVVDMFCIAMMDVANELGAPTYLYFTSSAASLGLMFCLQAKRDDEEFDVTELKDKDSELSIPCYTNPLPAKLLPSVLFDKRGGSKTFIDLARKYRESRGIVVNTFQELESYAIEY LASSNANVPPVFPVGAILNQEKKVNDDKTEEIMTWLNEQPESSVVFLCFGSMGSFGEDQIKEIALAIEESGQRFLWSLRRPPSNENKYPKEYENFGEVLPEGFLERTSSVGKVIGWAPQMAVLSHSSVGGFVSHCGWNSTLESIWCGVPVAAWPLYAEQQLNAFKLVVELGLAVEIKIDYRSENEIILTSKEIESGIRRLMNDEELRMKVKEMKGNSRFAVSEGGSSYVSIRRFIDLVMTKE (SEQ ID NO: 102) Comprises or consists of.
[0356] In some embodiments, the protein comprises an amino acid sequence having at least 75%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 102, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 75%-99%, 76%-98%, or 75%-100% homology or identity to SEQ ID NO: 102. Each possibility represents a separate embodiment of the present invention.
[0357] In some embodiments, the protein has the amino acid sequence: MPTSELVFIPSPGVGHLSPTIELVNQLLHRDQRLSVTIIVMKFSLESKHDTETPTSTPRLRFIDIPYDESAMALINPNTFLSAFVEHNKPHVRNIVRDISESNSVRLAGFVVDMFCVAMTDVVNEFEIPTYIYFTSTANLLGLMFYLQAKRDDEGFDVTVLKDSESEFLSVPSYVNPVPAKVLPDAVLDKNGGSQMCLDLAKGFRESKGIIVNTFQELERRGIEHLLSS NMNLPPVFPVGPILNLRNAPNDGKTADIMTWLNDHPENSVVFLCFGSMGSFEKEQVKEIAIAIEQSGQRFLWSLRRPTSLEKFEFPKDYENPEEVLPKGFLERTKGVGKVIGWAPQMAVLSHPSVGGFVSHCGWNSTLESIWCGVPIAAWPLYAEQKINAFQLVVEMGMAAEIRIDYRTNTRPGGGKEMMVMAEEIESGIRKLMSDDEMRKKVKGMKDKSRAAVLEGGSSHTSIGILIENLVSITI (SEQ ID NO: 103) Comprises or consists of.
[0358] In some embodiments, the protein comprises an amino acid sequence having at least 76%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 103, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 76%-99%, 80%-98%, or 76%-100% homology or identity to SEQ ID NO: 103. Each possibility represents a separate embodiment of the present invention.
[0359] In some embodiments, the protein has the amino acid sequence: MVGLKCFWILQKGFRESKGIIVNTFQELERRGIEHLLSSNMDLPPVFPVGPILNLRNARNDGKMADIMTWLNDQPENSVVFLCFGSRGSFKEEQVKEIAIAIEQSGQRFLWSLRRPTSIETFEFPKYYENPEEVLPKGFLERTKSVGKVIGWAPQMAVLSHPSVGGFVSHCGWNSTLESIWCGVPIAAWPLYAEQQTNAFQLVVEMGMAAEIRIDYRTNTPLVGGKDMMVTAEEIERGIRKLMSDDEMRKKVKDMKDKSRGAVLEGGSSHTSIGNLIDVLVSITI (SEQ ID NO: 104) Comprises or consists of.
[0360] In some embodiments, the protein comprises an amino acid sequence having at least 77%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 104, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 77%-99%, 79%-98%, or 77%-100% homology or identity to SEQ ID NO: 104. Each possibility represents a separate embodiment of the present invention.
[0361] In some embodiments, the protein has the amino acid sequence: MATNNLHFLLIPHIGPGHTIPMIDMAKLLAKQPNVMVTIATTPLNITRYGHTLADAINSFRFFEVPFPAVEAGLPEGCESTDKIPSMDLVPNFLTAIGMLEQKLEEHFHLLEPRPNCIISDKYMSWTGDFADKYRIPRIMFDGMSCFNELCYNNLYENKVFEGMHETEPFVVPGLPDKIELTRKQLPPEFNPSSIDTSEFRQRARDAEVRAYGVVINSFEELEQEYVNEYKKLRK GKVWCIGPLSLCNSDNSDKAQRGNIASVDEEKCLKWLDSHEADSVVYACFGSLVRVNTPQLIELGLGLEASNRPFIWVVRSVHREKEVEEWLVESGFEERIKDRGLIIRGWAPQVLILSHPSIGGFLTHCGWNSTLESVCAGVPMITWPQFAEQFINEKLIVQVLGIGVGVGVDSVVHVGEEDRSGVKVKRESVTKAIEKVMDDEIDGNERRRRSKEFGKIANNAIKEGGSSYLNLTLLIQDIMRYANADASS (SEQ ID NO: 105) Comprises or consists of.
[0362] In some embodiments, the protein comprises an amino acid sequence having at least 88%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 105, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 88%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 105. Each possibility represents a separate embodiment of the present invention.
[0363] In some embodiments, the protein has the amino acid sequence: MEKTPHIAIVPSPGMGHLIPLVEFAKKLKNHHNIHATFIIPNDGPLSISQKVFLDSLPNGLNYLILPPVNFDDLPQDTQIETRISLMVTRSLDSLREVFKSLVVEKNMVALFIDLFGTDAFDVAIEFGVSPYVFFPSTAMALSLFLYLPKLDQMVSCEYRELPEPVQIPGCIPVRGQDLVDPVQDRKNDAYKWVLHNAKKYSMAKGIAVNSFKELEGGALNALLE DEPGKPKVYPVGPLVQTGFSCDVDSIECLKWLDGQPCGSVLYISFGSGGTLSSSQLNELAMGLELSEQRFIWVVRSPNDQPNATYFDSHGHKDPLGFLPKGFLERTKGIGFVIPSWAPQAQILSHSATGGFLTHCGWNSILETVVHGVPVIAWPLYAEQKMNAVSLTEGIKMALRPTVGENGIVGRLEVARVVKSLLEGEEGKAIRSRVRDLKDAAANVLSKDGSSTKTLDQLAVQLKKQELS (SEQ ID NO: 106) Comprises or consists of.
[0364] In some embodiments, the protein comprises an amino acid sequence having at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 106, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 90%-100%, 93%-100%, 95%-100%, or 97%-100% homology or identity to SEQ ID NO: 106. Each possibility represents a separate embodiment of the present invention.
[0365] In some embodiments, the protein has the amino acid sequence: MTQKQMQMQPHFLLVTYPAQGHINPSLQFAERLIRLGVKVTFTTTVSAYRRMSKAGNISEFLNFAAFSDGFDDGFNFETDDHGLFLTQLRSRGKDSLKETILSNAKNGTPISCLVYTLLLPWAPEVARGLNVPSAFLWIQPASVLRLYYYYFNGYNELIGDDCNEPSWSIQLPGLPLLKS (SEQ ID NO: 107) Comprises or consists of.
[0366] In some embodiments, the protein comprises an amino acid sequence having at least 77%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 107, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 77%-100%, 79%-100%, 80%-100%, or 90%-100% homology or identity to SEQ ID NO: 107. Each possibility represents a separate embodiment of the present invention.
[0367] In some embodiments, the protein has the amino acid sequence: MTKIQQQPHFLLVTYPAQGHINPSLRFAERLIRLGVKVTFTITVSAYRRMSKAGHISEFLNFAVFSDGFDDGFNSKTDDYGLFLTQFRSRGKDSLKETILSNAKNGTPVSCLVYTLLLPWAPEVARGLNVPSAFLWIQPASVLRLYYYYFNGYNELIGDDCNEPSWSIQLPGLPLLKSRDLPSFCLPSNPYADVLTLVKEHLDVLDLEEKPKILVNSFDELEREALNEIDGKLKMVAVGPLIPSAFFGWTGCI (SEQ ID NO: 108) Comprises or consists of.
[0368] In some embodiments, the protein comprises an amino acid sequence having at least 73%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 108, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 73%-100%, 77%-100%, 85%-100%, or 90%-100% homology or identity to SEQ ID NO: 108. Each possibility represents a separate embodiment of the present invention.
[0369] In some embodiments, the protein has the amino acid sequence: MGSWRNSRTTSTKFLWLILPLMVVTVIIGVKKSNYGSKYNYPWVWSSVINSYSSSAVKEDVTVVAEGPVESFGLRSTVVNGGGVVAEGPSEDFGFNSSYPPLAMEDEMDVELPAIAKEDDLNATLSGPDLFVSANQTGGLHVDIGINSKYTSLDKLEARLGQVRAAIKEAESGNRTYDPDYVPEGPMYWHAASFHRSYLEMEKQFKVFVYEEGEPPIFHNGPCKNIYAMEGNFIYHMETTKFRTKNPEKAHTFFLPMSAAMMVRFIFERDPNVDHWRPMKQTIKDYVDLVGGKYPFWNRSLGADHFTVACHDWVSKVFYPIIFMLLLVFIFRMSTGC (SEQ ID NO: 109) Comprises or consists of.
[0370] In some embodiments, the protein comprises an amino acid sequence having at least 81%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 109, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 81%-100%, 85%-100%, 87%-100%, or 91%-100% homology or identity to SEQ ID NO: 109. Each possibility represents a separate embodiment of the present invention.
[0371] In some embodiments, the protein has the amino acid sequence: MSTVEVAKLLVNRDHRLFITFLIIQPPSSGSGSAITTYIESLAEKAMDRISFIELPQDKIPPPRYPKSLPTAESKAHPLIFMIEFIKCHCKYVRNIVSDMISQPSSGRVAGLVIDMLCFSMMDVANEFNIPTYVFVTSNAAFLGFYLYVQILSNDQNQDVVELSKSDTEISVPGFVKPVPTKVFWTVVRTKEGLDFVLSSAQKLRQAKAIMVNTFLELETHAIKSL SDDTSIPPVYPVGPILNLEGGAGKTFDNDISRWLDSQPPSSVVFLCFGSHGCFDEIQVKEIAHALEQSGHRFLWSLRRPPSDQTLKVPGDYEDPGVVLPEGFLERTAGRGKVIGWAPQVMVLAHRAVGGFVSHCGWNSLLESLWFGVPTATWPIYAEQQMNAFEMVVELGLAVEITLDYRNDMDMFIVTAQEIESGIRKVMEDNEVRTKVKERSEKSRAAVAEGGSSYASVGHLIKEFTGNIS (SEQ ID NO: 110) Comprises or consists of.
[0372] In some embodiments, the protein comprises an amino acid sequence having at least 74%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 110, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 74%-100%, 79%-100%, 85%-100%, or 90%-100% homology or identity to SEQ ID NO: 110. Each possibility represents a separate embodiment of the present invention.
[0373] In some embodiments, the protein has the amino acid sequence: MSSFINFVESTTQLQPQFEQLIQTLLPITAIISDGFLMWTQDSAEKFNIPRLVFYGTNIFFMTMCNIMAQFKPHAAVNSDDEAFDVPGFTRFKLTANDFEPPFNEVEPKGSMLDFLLEQQKAMVRSHGLVVNSFYEIEHEFNVYWNQNYGPKAWLMGPFCVAKPYASNVMDSEISTKVVKKSAWIQWLDRKLAANEPVLYISFGTQAEASMEHLHEVAIGLERSNVSFIWVVKAKQMQLIGAGFEERVKGRGKVVTEWVDQMEILKHEIVSGFLSHCGWNSLLESMCVGVPVLAMPLMADQLLNARLVVEEIGMGLRLWPRGMVARGIVGAEEVEKMVVELMEGEGGRRVRKRVIEVREMAYGAMKEGGSSSRTLDSLIDHVCEAFHKTV (SEQ ID NO: 111) Comprises or consists of.
[0374] In some embodiments, the protein comprises an amino acid sequence having at least 76%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 111, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 76%-100%, 80%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 111. Each possibility represents a separate embodiment of the present invention.
[0375] In some embodiments, the protein has the amino acid sequence: MGSLKKGAHILIFPFPAQGHMLPLLDLTHHLATNGLTITILVTPKNLPILNPLLSSSPNIQPLVFPFPPHPRLPPHVENVKDIGNHANVPITNSLAKLQDQIIQWFNSHHNPPVAIISDFFLGWTQHLANKLGIPRVGFFSSGAYLTAVLDYVCHNIKTVRSQEETVFHDLPNSPCFKFEHLPGLAQIYKESDPEWELVLDGHIANGLSWGWIVNTFDGLE SRYMEYLTKKMGVGRVFGVGPVNLLNGSDPMTRGKSESGSDSGVLNWLDGKPDGSVLYVCFGSQKFLTNDQMEGLSIGLEQSGVHYVWVVKDEQGDAIRSGSGRGLVVTGWAPQVSILGHGAVGGFLSHCGWNSVLEAIVNGVMILAWPMEADQFVNAKLLVDDHGIGVWVCEGPNTVPDSTELARKIGESMSTDKSEKVKAKEMKNKANEAVKEGGSSSMELSRLVKELSNFETNGP (SEQ ID NO: 112) Comprises or consists of.
[0376] In some embodiments, the protein comprises an amino acid sequence having at least 81%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 112, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 81%-100%, 85%-100%, 90%-100%, or 93%-100% homology or identity to SEQ ID NO: 112. Each possibility represents a separate embodiment of the present invention.
[0377] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 113) Comprises or consists of.
[0378] In some embodiments, the protein comprises an amino acid sequence having at least 71%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 113, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 71%-100%, 77%-100%, 85%-100%, or 90%-100% homology or identity to SEQ ID NO: 113. Each possibility represents a separate embodiment of the present invention.
[0379] In some embodiments, the protein has the amino acid sequence: MSLVTNNPHLLVYPLPTSGHIIPLLDLTDLLLRRGLTITVVISTTDLTLLDTLLSSHPTSLHKLYFPDPEIGPSSHPVIARIIATQKLFDPIVKWFESHPSPPVAIISDFFLGWTNELASRLGIRRVVFSPSGALGHSILQSLWRDVAEINAKNVDGNGNYSISFTDIPNSPEFHWWQLSQLLRVHREGDPDFEFFRNGMLANTKSWGIVYNTFERIEKV YIDHVKKQIGHDRVWAIGPLLPEEHGPVGSTARGGSSVVPPHDLLTWLDKKPHDSVVYICFGSRLTLSEKQMSALASALELSNVDFILCVKASGSSFIPSGFEDRVVGRGFVIKGWAPQLAIL RHRAVGSFVTHCGWNSTLEGVSSGVMMLTWPMGADQYANAKLLVDQLGVGKRVCEGGPESVPDSTELARLLEESLSGDTSERVKVKELSREANTAVKEGTSIRDLNMFVNLLSEL (SEQ ID NO: 114) Comprises or consists of.
[0380] In some embodiments, the protein comprises an amino acid sequence having at least 78%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 114, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 78%-100%, 85%-100%, 90%-100%, or 93%-100% homology or identity to SEQ ID NO: 114. Each possibility represents a separate embodiment of the present invention.
[0381] In some embodiments, the protein has the amino acid sequence: MATQVKTEEKHLKVEIINKTYVKPETPLGRKECQLVTFDLPYIAFYYNQKLIIYKGGVEEFEDTVEKLKDGLKVVLGEFHQLAGKLDKDDDGVFKVVYDDDMDGVEVLSAVAEDTATADLMDEEGTIKLKELVPYNSVLNIEGLHRPLLSIQITKLKDGLVLGCAFNHAILDGTSTWHFMSSWAQICSGSKSISAAPFLDRTQARNTRVKLDLTPPAQT NGNSNGDTNGDASATKPPAPAPLREKIFKFSESAIDKIKAKINANPPEGSTKPFSTFQSLSTHIWHAVTRARNLKPEDYTVFTVFADCRKRVDPPMPDSYFGNLIQAIFTVTAAGLLQANPPEFAASMIQKAIDMHDAKAIEARNKEWESNPIIFQYKDAGVNCVAVGSSPRFKVYDVDFGFGKPESVRSGANNRFDGMVYLYQGKSGGRSIDVEISLDASAMGNLEKDKEFLIQE (SEQ ID NO: 130) Comprises or consists of.
[0382] In some embodiments, the protein comprises an amino acid sequence having at least 87%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 130, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 87%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 130. Each possibility represents a separate embodiment of the present invention.
[0383] In some embodiments, the protein has the amino acid sequence: MASLPLLTVLEQSHVSPPPATVVDKSLSLTFFDFLWLTQPPIHNLFFYEFSIDETQFVETIVPSLKNSLSITLQHFYPFAGNLILFPDNKRPEIRYVEGDYVMVTFAKSSLDFNELVGNHPRDCDQFYDLIPPLGESVKTSEFRKIPLFSVQVTFFPQKGVSIGMTNHHSLGDASTRFCFLNAWTSISRSSSDESFLANGTKPFYDRVISNPKLDQ SYLKFSKIDTLYEKYQPLSLSRPSNKLRGTFILTRKILNELKKSVSIKLPTLSYVSSFTVACGYIWSCIAKSRNDDLQLFGFTIDCRARLDPPVPSTYFGNCVGGCMAMAKTTLLTEDDGFITAAKLLGESLHKTLTESGGIVKDIEVFEDLFKDGLPTTMIGVAGTPKLKFYETDFGWGNPKKVETISIDYNMSISMNACRESKDDLEIGVCLMNTEMEAFVRLFDEGLESYV (SEQ ID NO: 131) Comprises or consists of.
[0384] In some embodiments, the protein comprises an amino acid sequence having at least 72%, at least 80%, at least 89%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 131, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 72%-100%, 80%-100%, or 90%-100% homology or identity to SEQ ID NO: 131. Each possibility represents a separate embodiment of the present invention.
[0385] In some embodiments, the protein has the amino acid sequence:MGSENVHKIMKINITKSSFVQPSKPTVLPTNHIWTSNLDLVVGRIHILTVYFYRPNGASNFFDPIVMKKALADVLVSFYPMAGRISKDDNGRVVINCNDEGVLFVEAESDSTLDDFGEFTPSPELRQLTPTIDYSGDISTYPLFFAQVTHFKCGGVGFGCGVFHTLADGLSSIHFINTWSDMARGLSIAIPPFTDRTLLRAREPPTPTFDHV EYHLPPSMKTTSQTNKSRKPSTAMLKLTLDQLNALKAAAKNEGGNTNYSTYEILAAHLWRCACKARGLPDDQLTKLYVATDGRSRLSPQLPPGYLGNVVFTATPVAKSADLTTQPLSNAASLIRTTLTKMDNDYLRSAIDYLEVQPDLSALIRGPSYFASPNLNINTWTRLPVHDADFGWGRPVFMGPAVILYEGTIYVLPSPNNDRSMSLAVCLDADEQPSFEKFLYDF (SEQ ID NO: 132) Comprises or consists of.
[0386] In some embodiments, the protein comprises an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO: 132, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 90%-100%, 95%-100%, or 97%-100% homology or identity to SEQ ID NO: 132. Each possibility represents a separate embodiment of the present invention.
[0387] In some embodiments, the protein has the amino acid sequence: MPSSSSSPSSTADSVTIISKCTVYPHMKNSTPESLQLSVSDLPMLSCQYIQKGVLLSQPPPNHTNNIISHLKLSLSKTLSHFPPLAGRLSTDSHGHVSIICNDSGVEFVHSTANHLHTHQILPLNSDVHPCFKTFFAFDKTLSYAGHHQPIAAVQVTELADGLFIGCTVNHAVVDGTSFWNFFNTFAEITKGCQKVTNLPDFSRENVFISPVVLPLPSGGPSATFSGD EPLRERIIHFSRDAILKMKFRANNPLWRQPQNSDLDDTEIYGKVCNDINGKVNGAFKPKSEISSFQSLCGQLWRAVTRARKFNDPIKTTTFRMAVNCRHRLDPKVDKLYFGNLIQSIPTVASVGELLSHDLSWAANELHQNVVAHDNATVRRGVKDWENNPKLFPLGNFDGAMITMGSSPRFPMYNNDFGWGRPMAVRSGKANKFDGKISAFPGRDGDGSVDLEVVLAPETMACLERDHEFMQYVS (SEQ ID NO: 133) Comprises or consists of.
[0388] In some embodiments, the protein comprises an amino acid sequence having at least 86%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 133, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 86%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 133. Each possibility represents a separate embodiment of the present invention.
[0389] In some embodiments, the protein has the amino acid sequence: (SEQ ID NO: 134) Comprises or consists of.
[0390] In some embodiments, the protein comprises an amino acid sequence having at least 59%, at least 65%, at least 75%, at least 85%, at least 90%, or at least 99% homology or identity to SEQ ID NO: 134, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 59%-100%, 70%-100%, 80%-100%, 90%-100%, 95%-100%, or 97%-100% homology or identity to SEQ ID NO: 134. Each possibility represents a separate embodiment of the present invention.
[0391] In some embodiments, the protein has the amino acid sequence: MEVPDQFHLNILEQCHVSPSPNSIIPSFSLPLTFLDIPWLFYPSNQTLFFFPEPPPKTTIITTLKQSLSLTLHHFHPLAGNLSLPSPPAEPHIVYTKNDSIALTIAQTNTNIHHLSCNHPRSVKNLYSLLPKLPSPSMSRETHVGLVIPLLTIQITVFADLGYSIGVTMQHAAVDERTFDQFMKCWASVCTSLLKNDSLFTFKSTPWYDRSVIIDPKSLKTTF LKQWWNRSNSLNESHDQENDDHDLVLATFVLSSLDINMIKNHILAKCKMINEDPPLHLSPYVSACAYLWKCLIKIQETHDSIKGGPLYLGFNAGGITRLGYDIPSTYFGNCIAFGRCKAFESELLGDNGIVFAAKSIGKEIKRLDKDVLGGANKWISDWDELTIRLLGSPKVDSYGMDFGWGKVEKVEKISSISNHGRVNVISLSGCKDFKGGIEIGVVLSVAKMNVFTSLFHGGLMEFAY (SEQ ID NO: 135) Comprises or consists of.
[0392] In some embodiments, the protein comprises an amino acid sequence having at least 71%, at least 80%, at least 90%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 135, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 71%-100%, 80%-100%, 87%-100%, or 95%-100% homology or identity to SEQ ID NO: 135. Each possibility represents a separate embodiment of the present invention.
[0393] In some embodiments, the protein has the amino acid sequence: MKNKNPTSVIREALAKVLVFYYPFAGRLKEGPARKLMVDCSGEGVLFIEAEADVTLKQFGDALQPPFPCLEELLYDVPGSTGILDTPLLLIQVTRLLCGGFIFALRLNHTMSDAAGLVQFMTGLGEMAQGASRPSTLPVWQRELLFARDPPRVTCTHHEYTEVEDTNGTIIPLDDMAHKSFFFGPSEISALRRFVPSYLKKCSTFEVLTACLWRCRTIALQPDPEEEMRMICIVNARGKFNPPLLPKGYYGNGFAIPVAISTAGDLSSKPLGHALELVMKAKSNVTEEYMRSVADLMVIKGRPHYTVVRSYLVSDVTHAGFDVVDFGWGKASYGGPAKGGVGAIPGVVTFFIPFTNHKGESGIVLPICLPSAAMDKFVEELNKMLVPDNNEQVLREHKLLVLARL (SEQ ID NO: 136) Comprises or consists of.
[0394] In some embodiments, the protein comprises an amino acid sequence having at least 88%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 136, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 88%-100%, 92%-100%, 97%-100%, or 99%-100% homology or identity to SEQ ID NO: 136. Each possibility represents a separate embodiment of the present invention.
[0395] In some embodiments, the protein has the amino acid sequence: MAQIDTPLTFKVRRHAPELIAPAKPTPRELKPLSDIDDQEGLRFHIPVIQFYRSDPKMKNKNPASVIREALAKVLVFYYPFAGRLKEGPARKLMVDCSGEGVLFIEAEADVTLKQFGDALQPPFPCLEELLYDVPGSTGVLDTPLLLIQVTRLLCGGFIFALRLNHTMSDAPGLVQFMTGLGEMAQGASRPSTLPVWQRELLLARDPPRVTCTHHEYTEVED TKGTIIPLDDMAHKSFFFGPSEISALRRFVPSYLKKCSTFEVLTACLWRCRTIALQPDPEEEMRIICIVNARGKFNPPLPKGYYGNGFAFPVAISTAGDLSSKPLGHALELVMKAKSDVTEEYMRSIADLMVIKGRPHFTVVRSYLVSDVTHAGFDVVDFGWGKAAYGGPAKGGVGAIPGVASFYIPFTNHKGESGIVLPICLPSAAMDKFVEELNKMLVPDNNEQVLREHKLLVLARL (SEQ ID NO: 137) Comprises or consists of.
[0396] In some embodiments, the protein comprises an amino acid sequence having at least 91%, at least 93%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 137, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 91%-100%, 93%-100%, 95%-100%, or 97%-100% homology or identity to SEQ ID NO: 137. Each possibility represents a separate embodiment of the present invention.
[0397] In some embodiments, the protein has the amino acid sequence: MEIQVINYSSKLVKPLTPTPTANRYYNISFTDELVPTIYVPLILYYATPKNPNGDHFENICDRLEESLSKTLSDFYPLAARFIRKLSLIDCNDQGVLFVLGNVNIRLSDVTGLGLTFKTSVLNDFLPCEIGGADEVDDPMLCVKVTTFECGGFAIGMCFSHRLSDMGTMCNFINNWAARTIGEYDNEKHTPIFNSPLYFPQRGLPELDLKVPRSSIGVKN AARMFHFNGKAISSMREVFGVDENGSRRLSKVQLVVALLWKAFVRIDDVNDGQSKASFLIQPVGLRDKVVPPLPSNSFGNFWGLATSQLGPGEGHKIGFQEYFYILRESIKKRARDCAKILTHGEEGYGVVIDPYLESNQKIADNGTNFYLFTCWCKFSFYEADFGCGKPIWASTGKFPVQNLVIMMDDNEGDGVEAWVHLDDKRMNELEQDPDVKLYACNLA (SEQ ID NO: 138).
[0398] In some embodiments, the protein comprises an amino acid sequence having at least 73%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 138, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 73%-100%, 80%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 138. Each possibility represents a separate embodiment of the present invention.
[0399] In some embodiments, the protein has the amino acid sequence: MKLAVKESVIVKPSKTTPCQQIWTSNLDLVVGRIHILTVYLYRPNGSSNFFDSMVLKKALADVLVSFFPVAGRLDKDGDGRVVIDCNGEGVLFVEAEADCCIDDFGEITPSPELRRLVPTVDYSGDMSSYPLFITQVTRFKCGGVSLGCGLHHTLSDGLSALHFINTWSDVARGLSVAIPPFIDRSLLRARDPPSPVFDHIEYHPPPSLITPLQNQ KNASHSRSASTLILRLTLHQINNLKSKAKGDGSMYHSTYEILAAHLWRCACKARGLANDQPTKLYVATDGRSRLIPPLPPGYLGNVVFTATPVAKSGDFESESLAETARRIRSELGKMNDEYLRSAIDYLESVSDISTLVRGPTYFASPNLNVNSWTRLPIYESDFGWGRPIFMGPASILYEGTIYIIPSPSGDRSVSLAVCLDPDHMALFKECLYVF (SEQ ID NO: 139).
[0400] In some embodiments, the protein comprises an amino acid sequence having at least 83%, at least 88%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 139, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 83%-100%, 88%-100%, 94%-100%, or 97%-100% homology or identity to SEQ ID NO: 139. Each possibility represents a separate embodiment of the present invention.
[0401] In some embodiments, the protein has the amino acid sequence: MKLAVKESVIVKPSKTTPCQQIRTSNLDLVAGRIHILVVFFYRPNGSSNFFDSLVLKKALADVLVPFFPVAGRFSEDGDGRVVIDCNGEGVLFVESEADCCIDDFGEITLSPELQQLVPTVDYSGDMSSYPLFIAQVTRFKCGGVSLGWGLHHTLLDGLSALHFVNTWGDVARGLSVAIQPFIDRSLLRARDPPTPVFDHIEYHPPPS LITPLQNQKNASHSRSASTLILQLTPDQIKNLKSKAKGDGSMYHSTYEILAAHLWRCACKARGLANDQPTKLYVAANGRSRLIPPLPPGYLGNVVFNATHVAKSGDFESESLAETARRIHCELGKMNDEYFRSAIDYLESVDDISTLVKGPTYFASPNLNVYSWIGIPIYACDFGWGQPIFMRPASFLYDGSIYIIPSPSGDRSVLLAVCLDPDHMDLFKECLYAF (SEQ ID NO: 140) Comprises or consists of.
[0402] In some embodiments, the protein comprises an amino acid sequence having at least 76%, at least 84%, at least 92%, or at least 99% homology or identity to SEQ ID NO: 140, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 76%-100%, 83%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 140. Each possibility represents a separate embodiment of the present invention.
[0403] In some embodiments, the protein has the amino acid sequence:MVMISKLLRLGRRKLHTIVSRDTIRPSSPTPSHSKTYNLSLLDQIAVNSYVPIVAFYPSSNVCRSSDDKTLELKNSLSKILTHYYPFAGRMKKNRPTVVDCNDEGVEFVEARNTNSLSDFLQQSEHEDLDQLFPDDCVWFKQNLKGSINDANNSSVCPLSIQVNHFACGGVAVATSLRHKIGDGSSALNFIKHWAAVTSHSRAGNHQIDATSPIINPHFIS YPTRTFKLPDRSPYIPPSDVVSKSFVFPNTNIKDLQAKVVTMTMGSRQPIVNPTRADVVSWLLHKCVVAAATKRISGNFKESCVISPLNLRNKLEEPLPETSIGNIFYLITFPISNNHGDLMPDDFISQLRLGIRKFQNIRNLETALRTVEEMISETFILGTAESMDTSYVYSSIRGFPMYDIDFGWGKPVKVTVGGALKNLSILMDTPDVNGIEALVSLDKQDMKILLNDPELLAFCL (SEQ ID NO: 141) Comprises or consists of.
[0404] In some embodiments, the protein comprises an amino acid sequence having at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% homology or identity to SEQ ID NO: 141, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 60%-100%, 70%-100%, 80%-100%, or 90%-100% homology or identity to SEQ ID NO: 141. Each possibility represents a separate embodiment of the present invention.
[0405] In some embodiments, the protein has the amino acid sequence: MSTSDKMKITIRESSMIKPSKPTPDQRIWNSNLDLVVGRIHILTLYFFRPNGSSDFFDSEVLKQSLADVLVSFFPMAGRLGLDGDGRVEINCNGEGVLFVEAEADCSIDDFGEITPSPELRRLAPTVDYSGDISSYPLVITQVTHFKCGGVSLGCGLHHTLSDGLSSLHFINTWSDVTRGLPVAIPPFVDRTVLRARDPPTVVFDHVEYHTPPSMTSSL DKDKPQSEDVHVSTSMLRLTLDQINALKAKGKGDGIVYHSTYEILAAHLWRCACKARGLLNDQMTKLYVATDGRSRLIPPLPPGYLGNVVFTATPIAKSGELQQEPLATTARKIHTELAKMDDKYLRSALDYLESQQDLSALIRGPAYFACPNLNINSWTRLPIYDADFGWGRPIFMGPASILYEGTIYIIPSPSGDRSVSLAVCLDPSHMPLFQKYLYEL (SEQ ID NO: 142).
[0406] In some embodiments, the protein comprises an amino acid sequence having at least 85%, at least 89%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 142, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 85%-100%, 90%-100%, 93%-100%, or 96%-100% homology or identity to SEQ ID NO: 142. Each possibility represents a separate embodiment of the present invention.
[0407] In some embodiments, the protein has the amino acid sequence: MVNVEIISNEYIKPSSPTPPHLKIYNLSILDQLIPAPYAPIILYYPNQDHINDFEVHERLKLLKDSLSKTLTRFYPLAGTIKGDLSIDCNDIGAYFAVAHVNTRLDVFLNHPDLDLINCFLPRGPYLNGSSEGSCVSNVQVNIFECCGIAISLCISHKILDGAALSTFLKAWAGTSYGSKEVVYPNMSAPSLFPAKDLWLKDSSMVMFGSLF KMGKCSTKRFVFDSSKLSFLKAKASLNGLKDPTRVEVVSALLWKCIMAASEENTGSWKPSLLSHVVNLRKRLVSTLSEDSIGNLIWLASAECRTNAQSRLSDLVEKVRDSVSKINSEFVKKIQGDKGTKVMEESLKSMKDCADYIGFTSWCKMGFYDVDFGWGKPVWVCGSVCEGSPVFMNFVILMDTKYGDGIEAWVSLDEHEMHILKHNPELLEYASIDPSPLQMNK (SEQ ID NO: 143) Comprises or consists of.
[0408] In some embodiments, the protein comprises an amino acid sequence having at least 82%, at least 85%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 143, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 82%-100%, 85%-100%, 90%-100%, or 93%-100% homology or identity to SEQ ID NO: 143. Each possibility represents a separate embodiment of the present invention.
[0409] In some embodiments, the protein has the amino acid sequence: MGTIYQSPMIKSSTPKIIEDLKVIIHDTFTIFPPHETEKRSMFLSNIDQVLTFNVETVHFFAANPDFPPQVVAEKLKLALSKALVPYDFLAGRLKLNHESQRFEFDCNGAGARFVVGSSEFELGEIGDLVYPNPGFRQLVQKSYDNLELHEKPLCILQLTSFKCGGFALGVATNHATFDGLSFKTFLQNLGSLAADQPLAVDPCNDRHLLAARSPPKV QFDHPELLKIPTGTDIPNPTVFDCPESQLDFKIFNLTSDDIAHLKTKAKDGPGSTNAKITGFNVVAAHVWRCKALSSGSEYDPERVSTVLYAVDIRSRLNLPLSLAGNAVLSAYASAKCKEIEEGPLSRLVEMVTEGTNRMTGEYARSVIDWGEVNKGFPNGEFLISSWWRLGFADVEYPWGKPRYSCPVVYHRKDIILLFPDIVGADNNNEVNVLVALPGKEMEKFETLFHKFLA (SEQ ID NO: 144) Comprises or consists of.
[0410] In some embodiments, the protein comprises an amino acid sequence having at least 88%, at least 92%, at least 95%, or at least 99% homology or identity to SEQ ID NO: 144, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 88%-100%, 90%-100%, 93%-100%, or 95%-100% homology or identity to SEQ ID NO: 144. Each possibility represents a separate embodiment of the present invention.
[0411] In some embodiments, the protein comprising the amino acid sequence set forth in SEQ ID NOs: 12-22 is an AAE.
[0412] In some embodiments, the protein comprising the amino acid sequence set forth in SEQ ID NOs: 27-30 is a PKS.
[0413] In some embodiments, the protein comprising the amino acid sequence set forth in SEQ ID NOs: 39 to 46 is PKC.
[0414] In some embodiments, the protein comprising the amino acid sequence set forth in SEQ ID NOs: 59-70 is a PT.
[0415] In some embodiments, the protein comprising the amino acid sequence set forth in SEQ ID NOs: 80 to 88 is CBCAS.
[0416] In some embodiments, the protein comprising the amino acid sequence set forth in SEQ ID NOs: 102-114 is a UGT.
[0417] In some embodiments, the protein comprising the amino acid sequence set forth in SEQ ID NOs: 130-144 is AAT.
[0418] The terms "homology" and "identity," when used interchangeably herein, refer to the sequence identity between two amino acid sequences or two nucleic acid sequences, with identity being the more strict comparison. The phrases "percent identity or homology" and "% identity or homology" refer to the percentage of sequence identity found in a comparison of two or more amino acid sequences or nucleic acid sequences. Two or more sequences can be identical from 0 to 100%, or any value therebetween. Identity can be determined by comparing positions in each sequence that can be aligned for comparison purposes with a reference sequence. If a position in the compared sequence is occupied by the same nucleotide base or amino acid, the molecules are identical at that position. The degree of identity between amino acid sequences is a function of the number of identical amino acids at positions shared by the amino acid sequences. The degree of identity between nucleic acid sequences is a function of the number of identical or matching nucleotides at positions shared by the nucleic acid sequences. The degree of homology between amino acid sequences is a function of the number of amino acids at positions shared by the polypeptide sequences.
[0419] The following are non-limiting examples for calculating homology or sequence identity between two sequences (these terms are used interchangeably herein): The sequences are aligned for optimal comparison purposes (e.g., gaps can be introduced into one or both of the first and second amino acid or nucleic acid sequences for optimal alignment, and non-homologous sequences can be ignored for comparison purposes). Optimal alignment is determined using the GAP program in the GCG software package as the highest score using the Blossum 62 scoring matrix with a gap penalty of 12, a gap extension penalty of 4, and a frameshift gap penalty of 5. The amino acid residues or nucleotides at corresponding amino acid or nucleotide positions are then compared. If a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, the molecules are identical at that position. The percent identity between two sequences is a function of the number of identical positions shared by the sequences.
[0420] In some embodiments, the % homology or identity described herein is calculated or determined using the Basic Local Alignment Search Tool (BLAST). In some embodiments, the % homology or identity described herein is calculated or determined using the Blossum 62 scoring matrix.
[0421] In some embodiments, the protein comprises or is characterized by acyl-activating enzyme activity.
[0422] In some embodiments, the acyl is selected from a C1-C8 alkyl chain and an α-unsaturated phenylalkylcarboxylic acid.
[0423] In some embodiments, the acyl is a C1 alkyl chain. In some embodiments, the acyl is a C2 alkyl chain. In some embodiments, the acyl is a C3 alkyl chain. In some embodiments, the acyl is a C4 alkyl chain. In some embodiments, the acyl is a C5 alkyl chain. In some embodiments, the acyl is a C6 alkyl chain. In some embodiments, the acyl is a C7 alkyl chain. In some embodiments, the acyl is a C8 alkyl chain.
[0424] In some embodiments, the C1-C8 alkyl chain is hexanoic acid. In some embodiments, the acyl is hexanoic acid.
[0425] In some embodiments, the α-unsaturated phenylalkyl carboxylic acid comprises cinnamic acid or a derivative thereof.
[0426] In some embodiments, cinnamic acid derivatives include hydroxylated derivatives of cinnamic acid.
[0427] In some embodiments, the hydroxylated derivative of cinnamic acid comprises or is coumaric acid.
[0428] In some embodiments, the protein comprises or is characterized by a polyketide synthesis activity, as described herein, hi some embodiments, the protein is characterized by having the activity of polymerizing a diketide substrate into a polyketide.
[0429] In some embodiments, the diketide substrate is obtained by coupling an acyl-CoA starter unit.
[0430] In some embodiments, the acyl-CoA starter unit is selected from acetyl-CoA, butyryl-CoA, hexanoyl-CoA, octanoyl-CoA, cinnamoyl-CoA, coumaroyl-CoA, or any combination thereof.
[0431] In some embodiments, the acyl-CoA is or comprises hexanoyl-CoA, cinnamoyl-CoA, or both.
[0432] In some embodiments, the acyl-CoA is hexanoyl-CoA.
[0433] In some embodiments, the polyketide comprises a tetraketide. In some embodiments, the polyketide comprises a linear polyketide. In some embodiments, the polyketide comprises a straight-chain tetraketide.
[0434] In some embodiments, the protein comprises or is characterized by a polyketide cyclization or cyclization activity, as described herein, hi some embodiments, the protein is characterized by having the activity to cyclize polyketides.
[0435] In some embodiments, the polyketide cyclization comprises aldol cyclization, Claisen cyclization, or both.
[0436] In some embodiments, the polyketide comprises an acyl group as described herein.
[0437] In some embodiments, the protein comprises or is characterized by prenyl transfer activity as described herein. In some embodiments, the protein is characterized by being able to transfer a prenyl group to a substrate molecule. In some embodiments, the protein is characterized by being able to transfer an aryl prenyl group to an acceptor molecule. In some embodiments, the protein is a prenyl diphosphate synthase. In some embodiments, the protein is a trans-prenyl transferase. In some embodiments, the protein is a cis-prenyl transferase.
[0438] In some embodiments, the prenyl group is selected from dimethylallyl diphosphate, geranyl diphosphate, farnesyl diphosphate, or geranylgeranyl diphosphate.
[0439] In some embodiments, the protein has Formula I: [ka] (wherein (i) R1 is selected from C1-C8 alkyl, α-unsaturated phenylalkylcarboxylic acid, or α-saturated phenylalkylcarboxylic acid; and R2 is OH; or (ii) R1 is OH and R2 is selected from C1-C8 alkyl, α-unsaturated phenylalkylcarboxylic acid, or α-saturated phenylalkylcarboxylic acid).
[0440] In some embodiments, the compound has a formula selected from the following: [ka] (wherein R3 is C1-C8 alkyl and R4 is an α-unsaturated phenylalkyl carboxylic acid).
[0441] In some embodiments, the compound is selected from the following group: [ka]
[0442] In some embodiments, the compound is: [ka]
[0443] In some embodiments, the protein is characterized by cannabigerolic acid (CBGA) cyclization or cyclization activity. In some embodiments, the cyclization activity includes cyclizing CBGA to CBCA. In some embodiments, the protein is characterized by being capable of cyclizing or cyclizing CBGA to CBCA. In some embodiments, the protein is characterized by being capable of synthesizing CBCA or being a CBCA synthase (CBCAS).
[0444] In some embodiments, the protein is characterized in that it is capable of transferring the glucuronic acid moiety of UDP-glucuronic acid to a cannabinoid or a precursor thereof.
[0445] In some embodiments, the protein is characterized in that it is capable of transferring an acyl group from a donor molecule to a cannabinoid.
[0446] According to some embodiments, there is provided a transgenic cell comprising: (a) a DNA molecule disclosed herein; (b) an artificial nucleic acid molecule disclosed herein; (c) a plasmid or agrobacterium disclosed herein; (d) a protein disclosed herein; or any combination thereof.
[0447] In some embodiments, the cell further comprises a nucleic acid sequence encoding at least one enzyme associated with cannabinoid production from Cannabis sativa, hi some embodiments, the at least one enzyme associated with cannabinoid production from Cannabis sativa is selected from olivetol synthase (OLS), olivetolic acid cyclase (OAC), prenyltransferase 1 (PT1 / GOT1), PT4 / GOT4, or any combination thereof.
[0448] In some embodiments, the at least one enzyme associated with cannabinoid production from Cannabis sativa is selected from an OLS, an OAC, or both.
[0449] As used herein, the term "transgenic cell" refers to any cell that has undergone human manipulation at the genomic or genetic level. In some embodiments, a transgenic cell has been introduced with an exogenous polynucleotide, such as a DNA molecule disclosed herein. In some embodiments, a transgenic cell includes a cell into which an artificial vector has been introduced. In some embodiments, a transgenic cell is a cell that has undergone a mutation or modification of its genome. In some embodiments, a transgenic cell is a cell that has undergone CRISPR genome editing. In some embodiments, a transgenic cell is a cell that has undergone a targeted mutation of at least one base pair in its genome. In some embodiments, an exogenous polynucleotide (e.g., a DNA molecule disclosed herein) or vector is stably integrated into the cell. In some embodiments, a transgenic cell expresses a polynucleotide of the invention. In some embodiments, a transgenic cell expresses a vector of the invention. In some embodiments, a transgenic cell expresses a protein of the invention. In some embodiments, a transgenic cell is a cell that does not contain a polynucleotide of the invention, but has been transformed or genetically modified to contain a polynucleotide of the invention. In some embodiments, CRISPR technology is used to modify the genome of a cell, as described herein.
[0450] In some embodiments, the cells are cells of unicellular organisms, cells of multicellular organisms, and cells in culture.
[0451] In some embodiments, the unicellular organism comprises a fungus or a bacterium.
[0452] In some embodiments, the fungus is a yeast cell.
[0453] In some embodiments, the cell is an insect cell. In some embodiments, the cell comprises an insect cell line.
[0454] The variety of insect cell lines suitable for transformation and / or heterologous expression is common and will be apparent to those skilled in the art. Non-limiting examples of such insect cell lines include, but are not limited to, Sf-9 cells, SR+ Schneider cells, S2 cells, etc.
[0455] According to some embodiments, there is provided an extract derived from the transgenic cells disclosed herein, or any fraction thereof.
[0456] In some embodiments, the extract comprises a DNA molecule disclosed herein, a protein disclosed herein, or any combination thereof.
[0457] According to some embodiments, there is provided a homogenate, lysate, extract, any combination thereof, or any fraction thereof derived from the transgenic cells disclosed herein.
[0458] Methods and / or means for extracting, lysing, homogenizing, fractionating, or any combination thereof, cells or cultures thereof are common and will be apparent to those skilled in the art of cell biology and biochemistry. Non-limiting examples include, but are not limited to, pressure lysis (e.g., using a French press), enzymatic lysis, soluble-insoluble phase separation (e.g., to obtain a supernatant and pellet), detergent-based lysis, solvents (e.g., polar or nonpolar solvents), liquid chromatography mass spectrometry, etc.
[0459] In some embodiments, a transgenic plant, a transgenic plant tissue, or a plant part is provided. In some embodiments, a transgenic plant, or any part, seed, tissue, or organ thereof, is provided, comprising at least one transgenic plant cell of the present invention. In some embodiments, the transgenic plant, transgenic plant tissue, or plant part comprises (a) a DNA molecule disclosed herein; (b) an artifact disclosed herein; (c) a plasmid or Agrobacterium disclosed herein; (d) a protein of the present invention; (e) a transgenic cell disclosed herein; or any combination thereof.
[0460] In some embodiments, the transgenic plant, transgenic plant tissue, or plant part consists of transgenic plant cells of the present invention. In some embodiments, the transgenic plant, transgenic plant tissue, or plant part comprises at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% transgenic cells of the present invention, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the transgenic plant, transgenic plant tissue, or plant part comprises 20%-50%, 20%-60%, 20%-70%, 20%-80%, 20%-90%, or 20%-100% transgenic cells of the present invention. Each possibility represents a separate embodiment of the present invention.
[0461] In some embodiments, the transgenic plant, transgenic plant tissue, or plant part is or is derived from a Cannabis sativa plant. In some embodiments, the transgenic plant is a C. sativa plant.
[0462] In some embodiments, the transgenic plant, transgenic plant tissue, or plant part is or is derived from hemp. In some embodiments, C. sativa includes or is hemp.
[0463] According to some embodiments, there is provided a composition comprising any one of (a) the DNA molecules of the invention; (b) artificial vectors; (c) plasmids or Agrobacterium; (d) proteins of the invention; (e) transgenic cells; (f) extracts; (g) transgenic plant tissues or plant parts; and (h) any combination of (a)-(g), and an acceptable carrier, as disclosed herein.
[0464] As used herein, the term "carrier," "excipient," or "adjuvant" refers to any component of a composition, such as a pharmaceutical or nutritional supplement, that is not an active ingredient. As used herein, the term "pharmaceutically acceptable carrier" refers to a non-toxic, inert solid, semi-solid liquid filler, diluent, encapsulating agent, any type of formulation auxiliary, or simply a sterile aqueous medium, such as physiological saline. Some examples of materials which can serve as pharmaceutically acceptable carriers are sugars such as lactose, glucose, and sucrose, starches such as corn starch and potato starch, cellulose and its derivatives such as sodium carboxymethylcellulose, ethylcellulose, and cellulose acetate; powdered tragacanth; malt, gelatin, talc; excipients such as cocoa butter and suppository wax; oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; glycols such as propylene glycol, polyols such as glycerin, sorbitol, mannitol, and polyethylene glycol; esters such as ethyl oleate and ethyl laurate, agar; buffers such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline, Ringer's solution; ethyl alcohol and phosphate buffers, and other non-toxic compatible substances used in pharmaceutical formulations. Some non-limiting examples of materials that can function as carriers herein include sugar, starch, cellulose and its derivatives, powered tragacanth, malt, gelatin, talc, stearic acid, magnesium stearate, calcium sulfate, vegetable oil, polyol, alginic acid, pyrogen-free water, isotonic saline, phosphate buffer, cocoa butter (suppository base), emulsifiers (e.g., carbomer, hydroxypropyl cellulose, sodium lauryl sulfate) and other non-toxic, pharmaceutically compatible materials used in other pharmaceutical preparations. Wetting agents and lubricants such as sodium lauryl sulfate, as well as colorants, flavorings, excipients, stabilizers, antioxidants, and preservatives may also be present. Any non-toxic, inert, and effective carrier can be used to formulate the compositions contemplated herein.In this regard, suitable pharmaceutically acceptable carriers, excipients and diluents are well known to those skilled in the art, and can be found in, for example, The Merck Index, Thirteenth Edition, edited by Budavari et al., Merck & Co., Inc., Rahway, NJ (2001); the CTFA (Cosmetic, Toiletry, and Fragrance Association) International Cosmetic Ingredient Dictionary and Handbook, Tenth Edition (2004); and "Inactive Ingredient Guide," US Food and Drug Administration (FDA) Center for Drug Evaluation and Research (CDER) Office of Management, all of whose contents are incorporated herein by reference in their entirety.The examples of pharmaceutically acceptable carriers, carriers and diluents useful in the compositions of the present invention include distilled water, physiological saline, Ringer's solution, dextrose solution, Hank's solution and DMSO. These additional inactive ingredients, as well as effective formulation and administration procedures, are well known in the art and are described in standard textbooks, such as Goodman and Gillman's: The Pharmacological Bases of Therapeutics, 8th Ed., edited by Gilman et al., Pergamon Press (1990); Remington's Pharmaceutical Sciences, 18th Ed., Mack Publishing Co., Easton, Pa. (1990); and Remington: The Science and Practice of Pharmacy, 21st Ed., Lippincott Williams & Wilkins, Philadelphia, Pa., (2005), the contents of each of which are incorporated herein by reference in their entirety. The compositions described herein may also be contained in artificially engineered structures such as liposomes, ISCOMS, sustained-release particles, and other vehicles that increase the half-life of peptides or polypeptides in serum.Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and the like. Liposomes for use with the peptides described herein are generally formed from standard vesicle-forming lipids, including neutral and negatively charged phospholipids and sterols, such as cholesterol. The choice of lipid is generally determined by considerations such as liposome size and stability in the blood. Various methods for preparing liposomes are available, as reviewed, for example, in Coligan, JE et al., Current Protocols in Protein Science, 1999, John Wiley & Sons, Inc., New York. See also U.S. Patent Nos. 4,235,871, 4,501,728, 4,837,028, and 5,019,369.
[0465] The carriers may, in total, constitute from about 0.1% to about 99.99999% by weight of the pharmaceutical compositions presented herein.
[0466] Synthesis method According to some embodiments, a method for synthesizing a cannabinoid, a precursor thereof, or any combination thereof is provided.
[0467] According to some embodiments, there is provided a method of synthesizing acyl-coenzyme A (CoA), a polyketide, a compound represented by Formula I, a compound represented by Formula II, a cannabinoid, or any combination thereof.
[0468] In some embodiments, the method further comprises glycosylating the compound of Formula I, the compound of Formula II, the cannabinoid, or any combination thereof. In some embodiments, the method further comprises transferring an acyl group to the compound of Formula I, the compound of Formula II, the cannabinoid, or any combination thereof.
[0469] As used herein, the term "cannabinoid" or "cannabinoids" generally refers to a heterogeneous family of molecules that exhibit pharmacological properties by interacting with specific receptors. To date, two membrane receptors for cannabinoids, designated CB1 and CB2, both of which are G protein-coupled, have been identified. CB1 receptors are primarily expressed in the central and peripheral nervous systems, while CB2 receptors have been reported to be more abundant in cells of the immune system.
[0470] In some embodiments, the cannabinoid includes any compound shown in FIG.
[0471] According to some embodiments, the method comprises the steps of: (a) providing a transgenic cell or a cell transfected with a DNA molecule of the invention or an artificial nucleic acid molecule disclosed herein; and (b) culturing the transgenic cell or transfected cell of step (a) so that at least a first protein and a second protein encoded by the DNA molecule or artificial nucleic acid molecule are expressed, thereby synthesizing a cannabinoid, a precursor thereof, or any combination thereof.
[0472] In some embodiments, the precursor is selected from acyl-coenzyme A (CoA), a polyketide, a resorcinoid precursor, or any combination thereof.
[0473] In some embodiments, the resorcinoid precursor is olivetolic acid.
[0474] In some embodiments, the cannabinoid includes or is CBGA, CBCA, or both.
[0475] According to some embodiments, a method for obtaining an extract from a transgenic or transfected cell is provided.
[0476] In some embodiments, the method comprises culturing the transgenic or transfected cells in a medium and extracting the transgenic or transfected cells.
[0477] In some embodiments, the method comprises the steps of: (a) culturing the transgenic or transfected cells in a medium; and (b) extracting the transgenic or transfected cells, thereby obtaining an extract from the transgenic or transfected cells.
[0478] In some embodiments, the transgenic or transfected cell comprises a DNA molecule or a plurality of DNA molecules of the invention disclosed herein.
[0479] In some embodiments, the transgenic or transfected cells comprise an artificial nucleic acid molecule or vector disclosed herein.
[0480] In some embodiments, the cell is a transgenic cell or a cell transfected with a DNA molecule disclosed herein.
[0481] In some embodiments, the method further comprises a step prior to step (a) comprising introducing or transfecting a cell with an artificial nucleic acid molecule or vector disclosed herein.
[0482] Methods for introducing or transfecting an artificial nucleic acid molecule or vector into a cell are common and will be apparent to those skilled in the art.
[0483] In some embodiments, the introducing or transfecting step comprises transferring an artificial nucleic acid molecule or vector comprising a DNA molecule disclosed herein into the cell; or modifying the genome of the cell to include a polynucleotide disclosed herein. In some embodiments, the transferring comprises transfection. In some embodiments, the transferring comprises transformation. In some embodiments, the transferring comprises lipofection. In some embodiments, the transferring comprises nucleofection. In some embodiments, the transferring comprises viral infection.
[0484] As used herein, the terms "transfect" and "introduce" are interchangeable.
[0485] In some embodiments, the contacting is performed in a cell-free system.
[0486] The type of suitable cell-free system for expression and / or synthesis utilizing any one of the DNA molecules or molecules of the invention and the protein or proteins of the invention disclosed herein will be apparent to those skilled in the art.
[0487] In some embodiments, the method further comprises a step prior to step (b) comprising separating the cultured transgenic or transfected cells from the culture medium.
[0488] Methods for separating cells from the medium are common and include, but are not limited to, centrifugation, ultracentrifugation, or others, as will be apparent to one of skill in the art.
[0489] According to some embodiments, there is provided an extract of a transgenic or transfected cell obtained according to the methods disclosed herein.
[0490] In some embodiments, the extract comprises cannabinoids, precursors thereof, or any combination thereof.
[0491] In some embodiments, the extract includes CBGA, CBCA, or both.
[0492] According to some embodiments, there is provided a medium or portion thereof separated from cultured transgenic or transfected cells obtained according to the methods disclosed herein.
[0493] According to some embodiments, there is provided a composition comprising: (a) an extract as disclosed herein; (b) a medium or portion thereof as disclosed herein; or (c) any combination of (a) and (b), and an acceptable carrier as described herein.
[0494] In some embodiments, a portion includes a fraction or a plurality thereof.
[0495] Where a range of values is provided, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limit of that range, and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, provided there is no specific excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0496] As used herein, the term "about," when used in conjunction with a value, refers to plus or minus 10% of the reference value. For example, a length of about 1,000 nanometers (nm) refers to a length of 1,000 nm ± 100 nm.
[0497] It should be noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to a "polynucleotide" includes a plurality of such polynucleotides; a reference to "the polypeptide" includes a reference to one or more polypeptides and equivalents thereof known to those of skill in the art, and so forth. It should be further noted that the claims may be drafted to exclude any element. Accordingly, this statement is intended to serve as a prerequisite for using exclusive terminology, such as "only," "only," and the like, in connection with the recitation of claim elements or the use of "negative" limitations.
[0498] When a convention similar to "such as at least one of A, B, and C, etc." is used, such configuration is generally intended in the sense that one of ordinary skill in the art would understand that convention (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having A only, B only, C only, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Furthermore, one of ordinary skill in the art will appreciate that virtually any disjunctive word and / or term presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either term, or both terms. For example, the phrase "A or B" is understood to include the possibilities of "A" or "B" or "A and B."
[0499] It is understood that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination. All combinations of embodiments related to the present invention are specifically embraced by the present invention and are disclosed herein as if each and every combination were individually and explicitly disclosed. Furthermore, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein as if each and every such subcombination were individually and explicitly disclosed herein.
[0500] Additional objects, advantages, and novel features of the present invention will become apparent to those skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as delineated hereinabove and claimed in the claims below finds experimental support in the following examples.
[0501] Various embodiments and aspects of the present invention as delineated hereinabove and claimed in the claims section below find experimental support in the following examples. [Example]
[0502] Generally, the nomenclature used herein and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological, and recombinant DNA techniques. Such techniques are fully explained in the literature. See, for example, "Molecular Cloning: A Laboratory Manual," Sambrook et al. (1989); "Current Protocols in Molecular Biology," Volumes I-III, edited by Ausubel, R.M. (1994); Ausubel et al., "Current Protocols in Molecular Biology," John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning," John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA," Scientific American Books, New York; Birren et al. (eds.), "Genome Analysis: A Laboratory Manual Series," Vols. 1-4, Cold Spring Harbor Laboratory Press, New York. New York (1998); U.S. Patent Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659; and 5,272,057; "Cell Biology: A Laboratory Handbook," Volumes I-III, edited by Cellis, JE (1994); "Culture of Animal Cells—A Manual of Basic Technique" by Freshney, Wiley-Liss, NY (1994), Third Edition; "Current Protocols in Immunology," Volumes I-III, Coligan JE(eds.), Basic and Clinical Immunology (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds.), Stategies for Protein Purification and Characterization - A Laboratory Course Manual, CSHL Press (1996); all of which are incorporated by reference. Other general references are provided throughout this document.
[0503] material and method material Unless otherwise stated, all analyzed metabolites were >95% pure. CBGA 1, CBCA 15, CBDA, acetic acid, propionic acid, butyric acid, pentanoic acid, hexanoic acid, heptanoic acid, octanoic acid, ±2-methylbutyric acid, phenylalanine, hexane-D 11 Acid (D > 98%), GPP, IPP, FPP, phloretin 98, naringenin 96, malonyl-CoA (≥ 90%), acetyl-CoA (≥ 93%), butyryl-CoA (≥ 90%), hexanoyl-CoA (≥ 85%), octanoyl-CoA, iso-valeryl-CoA (≥ 90%), olivetol, and sodium hexanoate were purchased from Sigma-Aldrich (Rehovot, Israel). 9 -THCA was purchased from Silicon Scientific Equipment Ltd. (Or Yehuda, Israel). Acetic acid-D3 acid (D>99%), propionic acid-D5 acid (D>99%), butyric acid-D5 acid (D>98%), pentanoic acid-D9 acid (D>98%), heptanoic acid-D5 acid (D>99%), octanoic acid-D5 acid (D>99%), isobutyric acid-D7 acid (D>98%), 2-methylbutyric acid-D9 acid (D>99%), isovaleric acid-D9 acid (D>98%), isocaproic acid-D 11 Acids (D>98%) were purchased from C / D / N isotopes (Quebec, Canada). Phenylanine-D5 (D>98%) and phenylalanine- 13 C9,15 N1( 13 C, 15 N>99%) was synthesized at Cambridge Isotope Laboratories (Andover, MA). HeliCBGA 2 (NP009525, 90%) was purchased from Analyticon Discovery GmbH (Potsdam, Germany). APHA 3 was reported as an impurity (NP015136, 5%) in the heliCBGA analytical metabolites. OA 92 (>90%), VA (>90%), and iso-butyryl-CoA were purchased from Cayman Chemical (Ann Arbor, MI, USA). PCP 95, naringenin chalcone 97, and pinocembrin chalcone 100 were purchased from Wuhan ChemFaces Biochemical Co. Ltd. (Hubei, China). Cinnamoyl-CoA and coumaroyl-CoA were purchased from TransMIT GmbH (Hesse, Germany).
[0504] H. umbraculigerum seeds (Silverhill seeds, Cape Town, South Africa) were germinated and grown in a greenhouse under a long-day photoperiod. Plants were propagated by cuttings.
[0505] Supply experiment All feeding solutions were 0.5 mg ml -1The FA solution was prepared as an aqueous solution of the precursor. The pH of the FA solution was adjusted to 5.5–6.0. Phenylalanine feeding experiments were performed on young mother plant leaves, which were excised by cutting the proximal end of the pedicel with scissors in water, leaving 1–2 cm of the pedicel. For the FA feeding experiments, 10 cm young cuttings were obtained from the mother plants. The lower leaves were removed, leaving 4–5 leaves on each stem, and the stems were peeled to increase the uptake of the labeling solution. Three to four leaves or young seedlings were immersed in the aqueous solution [DDW (control), unlabeled, or labeled precursor; each group consisted of a minimum of three biological replicates]. All feeding experiments were conducted in a controlled environment at 25°C under constant fluorescent lighting and humidity for 48–96 hours, with the tubes periodically refilled. After completion, the fresh leaves were rinsed with a small amount of water, gently dried, flash-frozen, and stored at -80°C for extraction.
[0506] LC-MS chemical analysis Unless otherwise stated, 100 mg of frozen, powdered plant tissue was extracted with 300 μl of ethanol, sonicated for 15 min, stirred for 30 min, and centrifuged at 14,000 g for 10 min. The supernatant was filtered through a 0.22 μm syringe filter and analyzed at the resulting concentration. Detection was performed using an ultra-high-performance liquid chromatography-tandem quadrupole time-of-flight (UPLC-qTOF) system consisting of a UPLC (Waters Acquity) with a diode array detector connected to either a XEVO G2-S QTof (Waters) or a Synapt HDMS (Waters) using both targeted and untargeted approaches, as described in Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, pp. 817-831 (2023). Chromatographic separation was performed on a 100 mm x 2.1 mm i.d. (inner diameter), 1.7 μm UPLC BEH C18 column (Waters Acquity). The mobile phase consisted of 0.1% formic acid in acetonitrile:water (5:95, v / v; phase A) and 0.1% formic acid in acetonitrile (phase B). Terpenophenols were analyzed using UPLC method 1 as follows: initial conditions were 40% B for 1 min, increased to 100% B by 23 min, held at 100% B for 3.8 min, decreased to 40% B by 27 min, and held at 40% B by 29 min for system re-equilibration. The flow rate was 0.3 ml min -1 The column temperature was maintained at 35°C. Intermediates and glycosylated metabolites were analyzed using UPLC Method 2 as follows: initial conditions were 0% to 28% B over 22 min, increased to 100% B by 36 min, held at 100% B for 2 min, decreased to 0% B by 38.5 min, and held at 40% B until 40 min for system re-equilibration. The flow rate was 0.3 ml min -1 The column temperature was maintained at 35 °C. Electrospray ionization (ESI) was used in positive or negative ionization mode with an m / z range of 50–1,000 Da. Mass spectrometry was performed at a capillary voltage of 1 kV, a source temperature of 140 °C, a desolvation temperature of 450 °C, and a desolvation gas flow rate of 800 l h. -1Argon was used as the collision gas. The MS system was calibrated with sodium formate, and Leu-enkephalin was used as the lock mass. Data acquisition for non-targeted analysis was performed using MS E MS / MS experiments were performed in negative ionization mode using the NMR spectroscopy (MS / MS) algorithm. Collision energy was set to 4 eV for the low-energy function and a 15-50 eV ramp for the high-energy function. The R package Miso was implemented as previously described. Differential metabolites were selected if the fold change was 10 or greater and the p-value was less than 0.05. MS / MS experiments were performed in positive or negative ionization mode according to the specific protonated or deprotonated masses with the following settings: 1 kV capillary spray; cone voltage 30 eV; collision energy ramp 10-45 eV for positive mode and 15-50 eV for negative mode.
[0507] Absolute quantification of CBGA 1 Fresh samples of leaves (dark and light), flowers, stems, and roots were collected from plants at the flowering stage. Florets and receptacles were separated using a scalpel and analyzed separately. All tissues were flash-frozen in liquid N2 and ground to a fine powder. To measure CBGA1 content in dried tissues, fresh leaves were flash-frozen, ground, and freeze-dried. For extraction, 100 mg of frozen powder was accurately weighed in triplicate and extracted with 1 ml ethanol, prepared as previously described. Samples were diluted several-fold and injected to fit the linear range of the calibration curve. Injections were performed in multiple reaction monitoring (MRM) mode on a UPLC (Waters) coupled to a triple quadrupole detector (TQ-S, Waters). The system was operated as follows, using the same column and mobile phase as for UPLC-qTOF analysis. The initial conditions were 57% B, increased to 85% B by 4 min, increased to 100% B by 4.2 min, held at 100% B until 6 min, decreased to 67% B by 6.2 min, and held at 67% B until 7 min to allow the system to re-equilibrate. The flow rate was 0.6 ml min -1The column temperature was maintained at 40 °C. The instrument was operated in negative mode with a capillary voltage of 1.5 kV and a cone voltage of 40 V. Absolute quantification of CBGA 1 was performed by external calibration using two different transitions (359.3 > 191.2, 32 V for quantification and 359.3 > 315.4, 21 V for qualification).
[0508] Purification of metabolites for NMR analysis A total of 86 g of fresh leaves were flash-frozen in liquid N2, ground to a fine powder using an electric grinder, extracted with 600 ml ethanol, sonicated for 20 min, and stirred for 30 min. The supernatant was filtered, evaporated using a rotary evaporator at 40 °C, and lyophilized. The extract was reconstituted in 25 ml acetonitrile and used either for direct purification (after 10-fold dilution) or prefractionation by medium-pressure liquid chromatography (MPLC). The Buchi Sepacore MPLC system was equipped with two C-605 pump modules, a C-620 control unit, a C-660 fraction collector, a C-640 UV photometer (Buchi Labortechnik AG, Switzerland), and a C18 manually packed column. The mobile phase consisted of acetonitrile:water (5:95, v / v; phase A) and acetonitrile (phase B) using the following multi-step gradient method: Initial conditions were 0% B for 10 min, increased to 99% B by 530 min, and slowly increased to 100% B by 660 min. The flow rate was 15 ml min -1The injection volume was 15 ml, and the wavelengths were 210, 224, 270, and 350 nm. 100 ml fractions were collected throughout the run and analyzed by UPLC-qTOF to select specific metabolites for purification. Selected fractions were evaporated at 40 °C using a rotary evaporator, lyophilized, reconstituted in ethanol or methanol (only fractions containing Glc-OA 102 and Glc-DHSA 103) and filtered through a 0.22 μm syringe filter. Metabolite purification was performed on an Agilent 1290 Infinity II UPLC system (System 1, general instrument setup according to Jozwiak et al., 2020) or a UPLC system (Waters Acquity) equipped with a binary pump, autosampler, fraction manager, and diode array detector (System 2), using the same mobile phase as UPLC-qTOF. Triggering was performed using specific UV wavelengths depending on the metabolite.
[0509] Method development was performed on System 1 by acquiring both MS and UV signals. MS spectra were acquired in negative full scan mode from m / z 50 to 1,700. The HPLC column was either XBridge (BEH C18, 250 × 4.6 mm, 5 μm i.d.; Waters) or Luna (C18, 250 × 4.6 mm, 5 μm i.d.; Phenomenex), and conditions were adjusted and optimized for each metabolite. With this system, the eluate containing the metabolite of interest was collected at 1.8 ml min. -1The metabolites were mixed with a water makeup stream and then captured on solid-phase extraction (SPE) cartridges (10 × 2 mm Hysphere resin GP cartridges). Each cartridge was loaded four times with the same metabolite, and 36–72 cartridges were used to capture a single metabolite, depending on the concentration of the injected sample. After collection, the SPE cartridges were dried with a stream of N2 and eluted with 150 μl of methanol. The eluates containing the same metabolites were pooled, dried under a stream of N2, and stored at −20 °C until NMR analysis. A UPLC BEH C18 column (100 mm × 2.1 mm, 1.7 μm; Waters) was used in System 2, separate from the metabolites Glc-OA 102 and Glc-DHSA 103, which were fractionated on a Luna Phenyl-Hexyl column (150 mm × 2 mm, 3 μm; Phenomenex). The flow rate was 0.3 ml min -1 The column temperature was maintained at 35°C. All other conditions were adjusted and optimized according to the sample. The eluates containing the metabolites of interest were collected in 2 ml HPLC vials. The eluates containing the same metabolites were pooled, dried under a stream of N2, lyophilized, and stored at -20°C until NMR analysis.
[0510] NMR spectroscopy The purified metabolites were resuspended in 300 μl of methanol-d4, dried under a stream of N2, and resuspended in 0.01% 3-(trimethylsilyl)propionic acid-2,2,3,3-d4 sodium salt (TMSP, 1 H and 13 The metabolites were reconstituted in 70 μl of methanol-d4 containing 1.7 mm NMR microtubes for structure elucidation. NMR spectra were collected on a Bruker AVANCE NEO-600 NMR spectrometer equipped with a 5 mm TCI-xyz CryoProbe. All spectra were acquired at 298 K. The structures of the various metabolites were analyzed in one-dimensional (1D) 1 H NMR spectrum, as well as various two-dimensional (2D) NMR spectra: 1 H- 1 H correlation spectroscopy (COSY), 1 H- 1 H total correlation spectroscopy (TOCSY), 1 H-1 H rotating frame nuclear Overhauser spectroscopy (ROESY), 1 H- 13 C heteronuclear single quantum coherence (HSQC), and 1 H- 13 The C heteronuclear multiple bond correlation (HMBC) spectra were determined.
[0511] one dimensional 1 H NMR spectra were collected using 16,384 data points and a 2.5-second recycle delay. Two-dimensional COSY, TOCSY, and ROESY spectra were acquired using 16,384–8,192 (t2) × 400–512 (t1) data points. 2D TOCSY spectra were acquired using an isotropic mixing time of 100–300 ms. In this study, we used a TOCSY-less ROESY experiment, which effectively suppresses the TOCSY transition in ROESY experiments. T-ROESY spectra were recorded using a spin-lock pulse of 100–400 ms. 2D HSQC and 2D HMBC spectra were collected using 4,096 (t2) × 400–512 (t1) data points. Multiplicity-edited HSQC allows us to distinguish between methyl and methine groups, which produce positive correlations, and methylene groups, which appear as negative peaks. Setting the HMBC delay for the evolution of long-range coupling, J H,C Long-range coupling of .gtoreq.8 Hz was observed. All data were processed and analyzed using TopSpin 4.1.1 software (Bruker).
[0512] MALDI Imaging For peeling experiments, fresh whole leaves from young plants were attached to glass slides using tape on either the abaxial or adaxial side, slowly peeled off from above / below the midrib using duct tape, and dried overnight under moderate vacuum. Images were captured using a digital camera. To localize metabolites to individual trichomes, fresh leaves and flowers were dissected and sprayed with matrix as described previously. Sections were imaged with a Nikon DS-Ri2 microscope. MALDI imaging was performed using a 7 T Solarix FT-ICR (Fourier transform ion cyclotron resonance) mass spectrometer (Bruker Daltonics). Data sets were analyzed using lock mass calibration (DHB matrix peak: [3DHB + H - 3HO]). + Mass spectra were collected using a laser beam (m / z 409.055408 Da) under positive ionization at a frequency of 1 kHz, 40% laser power, 200 laser shots per pixel, and pixel sizes of 50, 15, or 25 µm for detached trichomes and sectioned leaves and flowers, respectively. Each mass spectrum was recorded in broadband mode over the m / z range of 150–3,000 with an acquisition time window of 1 MHz, resulting in an estimated resolution of 115,000 at m / z 400. Spectra were normalized to root-mean-square intensity, and pixel interpolation was turned on to plot MALDI images at the theoretical m / z ±0.005%.
[0513] Cryo-SEM, TEM, and confocal microscopy For cryo-scanning electron microscopy (cryo-SEM) analysis, frozen samples were attached to a holder using either mechanical clamps or glue made from concentrated PVP solution. The holder containing the sample was then flash-frozen in liquid N2 and transferred to a BAF 60 cryo-fracture apparatus (Leica Microsystems, Vienna, Austria) using a VCT 100 Vacuum Cryo Transfer Apparatus (Leica) and sublimated at -95 °C for 30 min. The sample was then transferred to an Ultra 55 cryo-SEM (Zeiss, Germany) using a VCT 100 shuttle and observed at -95 °C without coating, primarily using an InLens+SE detector in mixed mode at 1–1.3 kV. For transmission electron microscopy (TEM) analysis, H. umbraculigerum leaves were fixed in 4% paraformaldehyde and 2% glutaraldehyde in 0.1 M cacodylate buffer (pH 7.4) containing 5 mM CaCl, followed by postfixation in 1% osmium tetroxide supplemented with 0.5% potassium hexacyanoferrate trihydrate and potassium dichromate in 0.1 M cacodylate (1 h), stained with 2% uranyl acetate in water (1 h), dehydrated in graded ethanol solutions, and embedded in Agar 100 epoxy resin (Agar Scientific Ltd., Stansted, UK). Ultrathin sections (70–90 nm) were observed and photographed with an FEI Tecnai SPIRIT (FEI, Eidhoven, The Netherlands) transmission electron microscope equipped with a OneView Gatan Camera and operating at 120 kV. Confocal microscopy of trichomes was performed on a Nikon Eclipse A1 microscope. Because trichomes do not fluoresce, they were imaged using transmitted light. To better visualize the trichomes, chlorophyll autofluorescence was used as contrast. A far-red laser was used to detect chlorophyll autofluorescence (excitation: 640 nm; emission: 663–738 nm).
[0514] Trichome enrichment Trichomes were enriched according to modified guidelines from Bergau et al. Briefly, young leaves were harvested, immersed in ice-cold distilled water, and then polished using a BeadBeater machine (Biospec Products, Bartlesville, Oklahoma). A polycarbonate chamber was filled with 15 g of plant material and half-filled with glass beads (0.5 mm diameter), XAD-4 resin (1 g / g plant material), and 80% ethanol. The leaves were beaten for 2–4 pulses of 1 min each. This procedure was carried out at 4°C, and the chamber was cooled on ice after each pulse. After polishing, the contents of the chamber were first filtered through a kitchen mesh strainer and then through a 100 μm nylon mesh to remove the plant material, glass beads, and XAD-4 resin. Residual plant material and beads were scraped off the mesh and rinsed twice with 80% ethanol passed through a 100 μm mesh. The presence of concentrated glandular trichome secretory cells was confirmed by visualization with an inverted light microscope.
[0515] Genome assembly High-molecular-weight DNA was extracted from young frozen leaves and sequenced at the UC Davis Genome Center. Sequencing was performed on the Pacbio Sequel II platform using approximately 12 kilobases of DNA SMRT library preparation according to the manufacturer's protocol. Three different SMRT 8M cells were used to obtain a total of 57.8 Gb of HiFi data (approximately 44× haploid coverage). In addition to the Pacbio HiFi data, PE 2×150 Illumina Hi-C data with 200 M reads was obtained by Phase Genomics. Both the Pacbio HiFi and HiC data were integrated using Hifiasm software to generate a chromosome-scale haplotype-resolved assembly. Further scaffolding was performed using the Hi-C data, and read mapping was performed according to the Arima Genomics pipeline and SALSA software. Hi-C heatmap visualization was performed using Juicer, and quality metrics were obtained using the Assemblathon 2 script. Finally, to incorporate CDS sequences from the transcriptome data, the assembly was soft-masked for repetitive elements using EDTA with the -cds flag. Details of the parameters for each command are provided in github.com / Luisitox / Helichrysum_paper.
[0516] RNA sequencing and genome annotation RNA was extracted from seven tissues: young leaves, old leaves, florets and receptacles, stems, roots, and trichomes (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). RNA integrity was confirmed using a TapeStation instrument. Paired-end Illumina libraries were prepared for five tissues and sequenced on an Illumina HiSeq 3000 instrument (PE 2 × 150, approximately 40M reads per sample) and processed according to the Freedman and Weeks guidelines. Briefly, random sequencing errors were corrected using Rcorrector, and uncorrectable reads were removed. Adapter and quality trimming were performed using TrimGalore!. Ribosomal RNA was filtered by discarding reads mapping to the SILVA_132_LSURef and SILVA_138_SSURef nonredundant databases using bowtie2. Fastq quality checks at each step were performed using MultiQC. The remaining reads were pooled and used for genome-guided and genome-independent de novo transcriptome assembly using Trinity.
[0517] Iso-Seq data were obtained from four tissues (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9:817-831 (2023)) and processed with isoseq3. Fusion and unspliced transcripts were removed, and only poly(A)-positive transcripts were retained as a unique set of high-quality isoforms. Iso-Seq and Trinity transcripts were aligned to an assembly using minimap2, and the BAM files were incorporated into the PASA pipeline to generate RNA-based gene model structures. Furthermore, de novo gene structures were obtained using software Breaker2 and alignment of long and short read BAM files as external training evidence. Ab initio and RNA-based gene models were integrated using EvidenceModeler, followed by the final round of the PASA pipeline. Gene functional annotation was performed on predicted mature transcripts using TransDecoder. TransDecoder considers HMMER hits against PFAM and BLASTP hits against the UniProt database as similarity retention criteria. Further annotation of protein-coding transcripts was performed by obtaining the best hits from BLASTP searches against other plant protein databases (sunflower id UP000215914_4232, Arabidopsis id UP000006548_3702, tomato id UP000004994_4081, rice id UP000059680_39947, and Uniprot protein fasta files with NCBI id GCF_900626175.1_cs10). Signal peptides were predicted with SignalP, transmembrane domains were predicted with TMHMM, and GO and KEGG terms were obtained with Trinotate. The complete scripts used for protein functional annotation can be found at github.com / Luisitox / Helichrysum_paper. BUSCO was used at multiple stages of the analysis to assess the completeness of various versions of both the transcriptome and genome.
[0518] 3' RNA sequencing and gene co-expression network analysis UMI-based 3' RNAseq for three replicates of seven tissues was obtained as described. Adapter and quality trimming was performed in two steps using TrimGalore!, including PolyA trimming mode. Reads were mapped to the genome using STAR UMI deduplication with the UMI tool, and counts were obtained using featureCounts. Normalization was performed using the varianceStabilizingTransformation algorithm in DESeq2, and the CEMItools package was used for coexpression analysis (dissimilarity threshold 0.6, p-value 0.1). Circos and gene cluster plots
[0519] Gene and TE densities were calculated by intersecting the corresponding gff files with 0.1 Mb non-overlapping windows using bedtools makewindows and bedtools intersect. True-seq and Tran-seq coverage were calculated in BedGraph format using bedtools genomecov. Circos (circus) plots were created using the R circlize package, and gene cluster plots were created using the gggenes package. The complete R script can be found at github.com / Luisitox / Helichrysum_paper.
[0520] Phylogenetic analysis of functionally tested enzymes The selection of proteins for each family analyzed in this study was based on functionally tested enzymes according to the studies referenced in each figure. A complete list of IDs can be found in Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, pp. 817–831 (2023). Maximum likelihood trees were constructed with 100 bootstrap tests based on MUSCLE multiple alignments using MEGA11 software. Evolutionary distances were calculated using a JTTmatrix-based method.
[0521] Orthology and synteny analysis The proteome was analyzed for all available annotated Asteraceae genomes present in NCBI: GCA_3112345.1 (Artemisia annua), GCA_9363875.1 (Mikania micrantha), GCA_23376185.1 (Endive), GCA_23525715.1 (Cichorium intybus), GCA_23525745.1 (Burdock root), GCA_23525975.1 (Yacon), GCA_24762085.1 (Pigweed), GCF_1531365.2 (Artichoke), GCF_1531365.2 (Cynara cardunculus). var. scolymus), GCF_2127325.2 (sunflower (Helianthus annuus)), GCF_2870075.4 (lettuce (Lactuca sativa)), GCF_10389155.1 (artemisonia canadensis)), and GCA_900626175.1 (hemp (Cannabis sativa)). Orthogroups and their phylogenetic relationships were inferred using Orthofinder. The genomic locations and putative functions of all genes belonging to the orthogroups of HuCoAT6 (OG0014461), HuOLS4 (OG0000313), and HuCBGAS4 (OG0002538) were determined using the corresponding GFF files, and plots were created using the gggenomes package. The phylogenetic gene tree generated by Orthofinder was plotted using MEGA11.
[0522] β-Glucosidase assay for preparation of DHSA 93 The MPLC fractions (50 ml each) containing Glc-DHSA 103 were evaporated at 40°C using a rotary evaporator, lyophilized, and reconstituted in 15 ml of McIlvaine buffer (20 mM, pH 5.0). Reactions were carried out in separate 20 ml vials and incubated at 45°C for 24 hours. Each reaction contained 6 ml of McIlvaine buffer (pH 5.0), 3 ml of 0.1 mg / ml McIlvaine buffer, and 103 mg of 0.1 mg / ml McIlvaine buffer. -1 of almond β-glucosidase solution (6U mg -1 The reaction consisted of 1.5 ml of Glc-DHSA 103-containing fractions (Sigma-Aldrich) and 3 ml of ethyl acetate:diethyl ether (1:1). Metabolites were extracted using three volumes of ethyl acetate:diethyl ether (1:1), evaporated using a rotary evaporator, and reconstituted in 5 ml of methanol. The product from the reaction contained a mixture of both glucosylated and non-glucosylated metabolites. Therefore, DHSA 93 was purified using System 2 and reconstituted in 100 μl of methanol for enzymatic assays. The purified DHSA 93 was analyzed by UPLC-qTOF to confirm that the purified fractions did not contain Glc-DHSA 103.
[0523] Expression of AAE, PKS, PKC, UGT and AAT in Escherichia coli and protein purification The HuAAE1-6, HuUGT1-13, and HuAAT1-15 coding sequences from H. umbraculigerum and previously characterized sequences from rice (OsUGT) and stevia (SrUGT) were individually cloned into EcoRI-digested pET28b vector using the ClonExpress II One-Step Cloning Kit (Vazyme, Germany). HuPKS1-4, HuPKC1-5, CsOLS, and CsOAC were ligated into the pOPINF vector (digested with HindIII and KpnI) using the ClonExpress II One-Step Cloning Kit (Vazyme, Germany). Due to the high sequence similarity of the coding sequences, HuPKS2-4 were synthesized by Twist Biosciences. All constructs were expressed in Escherichia coli BL21(DE3) cells (a complete list of primers can be found in Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). Bacterial starters were grown overnight in LB medium at 37 °C, diluted 1:100 in fresh LB, and re-incubated at 37 °C. Once the cultures reached an A600 of 0.6, protein expression was induced with 400 μM isopropyl-1-thio-β-d-galactopyranoside (IPTG) overnight at 15 °C. Bacterial cells were maintained in 50 mM Tris-HCl pH 8.0, 0.5 mM phenylmethylsulfonyl fluoride (PMSF, Sigma-Aldrich) solution in isopropanol, 10% glycerol and protease inhibitor cocktail (Sigma-Aldrich), and 1 mg ml -1Cells were lysed by sonication in lysozyme (Sigma-Aldrich). Whole-cell extracts were either retained for functional activity or used for protein purification. Purification of hexahistidine-tagged proteins was performed with Ni-NTA agarose beads (Adar Biotech). Proteins were eluted with 200 mM imidazole (Fluka) in a buffer containing 50 mM NaH2PO4, pH 8.0, and 0.5 M NaCl. Protein concentrations of the eluted fractions were measured using the Pierce™ 660 nm Protein Assay Reagent (Thermo Scientific).
[0524] AAE enzyme assay Recombinant AAE assays were performed at 40°C for 10 min in a 20 μl reaction mixture containing 0.1 μg of recombinant AAE, 50 mM HEPES pH 9.0, 8 mM ATP, 10 mM MgCl2, 0.5 mM CoA, and 4 mM of the sodium salt of each acid (acetic, butyric, hexanoic, octanoic, cinnamic, and coumaric). The reaction was terminated with 2 μl of 1 M HCl and stored on ice until analysis. After centrifugation at 15,000 g for 5 min at 4°C, the sample was diluted 1:100 in water and analyzed in MRM mode on a TQ-S system using the same column as described above. The system was operated with aqueous buffer pH 7.0 (10 mM ammonium acetate, 5 mM NH4HCO2, phase A) and acetonitrile (phase B). The flow rate was 0.3 ml min -1The column temperature was maintained at 25°C. Metabolites were analyzed using a 15-minute multistep gradient: initial conditions were 1% B increasing to 35% B by 10.5 minutes, then increasing to 100% B by 11 minutes, holding at 100% B for 1 minute, decreasing to 1% B by 12.5 minutes, and holding at 1% B for 15 minutes for system re-equilibration. The instrument was operated in positive mode with a capillary voltage of 3.0 kV and a cone voltage of 50 V. The identity of the metabolites was confirmed with authentic standards. Two different transitions were used for the following analyses: acetyl-CoA (810.52 > 303.30, 27.0 V; 810.52 > 428.25, 24.0 V); butyryl-CoA (838.58 > 331.30, 28.0 V; 838.58 > 331.30, 25.0 V); hexanoyl-CoA (866.65 > 359.40, 28.0 V; 866.65 > 428.2 V) 5, 26.0V); octanoyl-CoA (894.65 > 387.55, 30.0V; 894.65 > 428.25, 28.0V); coumaroyl-CoA (914.59 > 407.37, 30.0V; 914.59 > 428.25, 28.0V); cinnamoyl-CoA (898.59 > 391.37, 30.0V; 898.59 > 428.25, 28.0V).
[0525] PKS and PKC enzyme assays Individual and combined HuPKS and PKC (HuOAC or CsOAC) assays were performed as described by Gagne et al. (2012) with some modifications. Enzyme assays were performed in 50 μL of 20 mM HEPES pH 7.2, 5 mM DTT, 1.8 mM malonyl-CoA, and 0.6 mM hexanoyl-CoA. HuPKS (5 μg) and PKC (10 μg) were added individually or in combination. The reaction mixture was incubated at 30°C for 3 h. The reaction was terminated by extraction with 100 μL of methanol, vortexing, and centrifugation at 15,000 g for 10 min. The supernatant was filtered and analyzed on both a UPLC-qTOF and a triple quadrupole system. The column and mobile phase were the same as for metabolic profiling. The initial conditions were 10% B, increased to 70% by 6 min, increased to 100% B by 6.2 min, held at 100% B until 8 min, decreased to 10% B by 8.5 min, and held at 10% B until 11 min for re-equilibration of the system. The flow rate was 0.3 ml min -1 The column temperature was maintained at 35 °C. UPLC-qTOF was run in both polarities in MS or MS / MS mode using parameters similar to those described above. The TQ-S system was operated in MRM mode in both positive (for Olivetol) and negative modes, with a capillary voltage of 3.5 or 1.5 kV, respectively, and a cone voltage of 40 or 20 V, respectively. Two different transitions were used for the following analyses: OA 92 (223.1 > 179.1, 15.0 V; 223.1 > 137.1, 20.0 V); PDAL (181.2 > 137.1, 10.0 V; 181.2 > 97.1, 20.0 V); HTAL (223.1 > 179.1, 10.0 V; 223.1 > 125.1, 10.0 V); PCP 95 (223.1 > 179.1, 20.0 V; 223.1 > 81.0, 25.0 V); and olivetol (181.1 > 111.0, 10.0 V; 181.1 > 71.2, 10.0 V). The identity of olivetol, OA 92, and PCP 95 was confirmed with authentic standards.
[0526] PT enzyme assay The HuPT1-4 genes from H. umbraculigerum were separately cloned into the pESC-TRP vector. Microsome preparation from yeast cells transformed with the pESC-TRP vector was performed as described by Jozwiak et al. (2020). PT enzyme assays were performed using CsPT4 with some modifications. 8 The assay was performed as previously described. Microsomes from yeast expressing HuPT were resuspended in 3.3 ml of buffer (10 mM Tris-HCl, 10 mM MgCl2, pH 8.0, 10% glycerol) and homogenized with a tissue grinder. Enzyme assays were performed in 50 μL volumes using 2 μl of each membrane preparation dissolved in reaction buffer (50 mM Tris-HCl, 10 mM MgCl2, pH 8.0), 500 μM of an aromatic acceptor (OA 92, VA, DHSA 93, PCP 95, naringenin chalcone 97, or pinocrine chalcone 100), and 500 μM of an isoprenoid (IPP, GPP, or FPP). Samples were incubated at 30°C for 1 h. Kinetic assays were similarly performed using 1 mM GPP and various concentrations of OA 92 (0.5 μM to 1.5 mM) incubated at 30 °C for 15 min. Samples were extracted with 100 μl of ethanol, vortexed, and centrifuged. The supernatants were filtered and analyzed by UPLC-qTOF as with terpenoid phenolics (UPLC Method 1).
[0527] UGT enzyme assay UGT enzyme assays were performed as described by Cai et al. (2021) with some modifications. UGT assays using different aromatic substrates were performed by mixing 1.5 μl of UDP-Glc solution (80 mM, final concentration: 2.5 mM), 27.5 μl of Tris buffer (100 mM, pH 8.0), 1 μl of each substrate (50 mM, final concentration: 1 mM), and 20 μl of lysate enzyme solution. The reactions were incubated at 30°C for 1 hour. The reactions were terminated by extraction with 100 μl of methanol, vortexing, and centrifugation at 15,000 g for 10 minutes. The supernatant was filtered and analyzed by UPLC-qTOF using UPLC method 2. Assays using purified UGTs were performed in the presence of 1.5 μl of 80 mM UDP-Glc, 46.5 μl of Tris buffer (100 mM, pH 8.0), and 1 μl of each enzyme in 2 μl of cannabinoid acceptor (OA 92, DHSA 93, CBGA 1, heliCBGA 2, CBDA, Δ 9 -THCA, CBCA 15, Olivetol, CBG, CBD or Δ 9 The assay was performed by mixing purified enzyme (1.5 μg μl) dissolved in 45 μl of Tris buffer (100 mM, pH 8.0). -1 Kinetic assays were performed using various (0.5 µM–3 mM) and constant (1 mM) concentrations of OA 92 and UDP-Glc. The total reaction volume was 50 µl. To stop the reaction, 100 µl of methanol was added to each tube, and metabolites were extracted and analyzed as described above.
[0528] AAT enzyme assay Recombinant AAT assays using different donor and acceptor substrates were performed using 7 μl of cannabinoid acceptor (OA, CBGA, or heliCBGA, 1 mg ml -1) with 58 μl of potassium phosphate buffer (100 mM, pH 7.4), 5 μl of acyl-CoA donor (butyryl-CoA, hexanoyl-CoA, iso-valeryl-CoA, or acetyl-CoA, 10 mM), and 30 μl of enzyme solution. The reaction was incubated at 30°C for 3 h. The sample was extracted with 100 μl of ethanol, then vortexed and centrifuged. The supernatant was filtered and used for UPLC-qTOF analysis using the same column, mobile phase, and MS parameters as those described above for terpenophenols. Initial conditions were 40% B for 1 min, increased to 100% B until 14 min, held at 100% B for 3.8 min, decreased to 40% B until 18 min, and held at 40% B until 20 min for system re-equilibration. The flow rate was 0.3 ml min -1 The column temperature was maintained at 35°C.
[0529] Assays using purified HuCBAT5 enzyme were performed using 2 μl of cannabinoid acceptors (OA, CBGA, heliCBGA, CBDA, Δ 9 The reaction was performed by mixing 15 mM hydroxybenzoate (HuCBAT5) with 2 μl of acyl-CoA donor (butyryl-CoA, isobutyryl-CoA, hexanoyl-CoA, isovaleryl-CoA, or acetyl-CoA, 10 mM), 44 μl of potassium phosphate buffer (100 mM, pH 7.4), and 2 μl of purified HuCBAT5 enzyme solution. The reaction was incubated at 30 °C for 3 h. To stop the reaction, 50 μl of ethanol was added to each tube, and the acylated metabolites were extracted and analyzed by UPLC-qTOF in both MS and MS / MS modes, similar to terpenoids (UPLC Method 1). From the LC-MS / MS analysis, extracted ion chromatograms with the major products were selected as follows: CoA-free cannabinoid acceptors: OA 92 > 179.107, CBGA 1, CBCA 15 > 191.107, heliCBGA2 > 225.092, CBDA, Δ 9-THCA>245.154; acylated cannabinoids: OA 92>179.107, CBGA 1>231.102, heliCBGA 2>265.086, CBDA>245.154, Δ 9 -THCA>245.154, CBCA 15>191.107).
[0530] Transient expression of selected genes in Nicotiana benthamiana (N. benthamiana) Overexpression constructs for GFP (as a negative control), CsOLS, and CsOAC were generated using GoldenBraid cloning as described in Jozwiak et al. (2020) into the final vector pAlpha2-Ubq10-CCD-Ter10. HuCoAT6, HuTKS4, and HuCBGAS were amplified using the ClonExpress II One Step Cloning Kit (Vazyme) and cloned into the BsaI-digested pAlpha2-NPT II-Ubq10-CCD-Ter10 vector. A complete list of oligonucleotides used for cloning can be found in Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, 817–831 (2023). All plasmids were sequenced and transformed into Agrobacterium tumefaciens strain GV3101 by electroporation. A. tumefaciens carrying the overexpression constructs was grown overnight at 28°C in Luria-Bertani (LB) medium in the presence of kanamycin and gentamicin. Bacterial cells were harvested by centrifugation, washed, and resuspended in infiltration buffer (10 mM MES, 2 mM MgCl2, 2 mM Na3PO4, 0.5% glucose, and 100 mM acetosyringone) to an OD600 of 0.3. Equal volumes of A. tumefaciens suspensions carrying different expression vectors were combined to obtain the desired gene combinations and incubated at room temperature for 2 hours. Using a 1 ml needleless syringe, the solution was infiltrated into 4- or 5-week-old N. benthamiana leaves from the abaxial side. Two days after the initial infiltration, substrates (0.5 mM each) were infiltrated into the same leaf area, and leaves were collected 24 hours later for metabolite analysis. Leaf samples were flash frozen, extracted with 300 μl of methanol as described above, and analyzed on a similar UPLC system coupled to an Orbitrap IQ-X Tribrid MS (Thermo Scientific, Bremen, Germany) using UPLC method 2 in negative mode.The source parameters were: sheath gas flow rate, auxiliary gas flow rate, and sweep gas flow rate: 45, 10, and 1 arbitrary unit, respectively; vaporizer temperature: 300 °C; ion transfer tube temperature: 275 °C; spray voltage: 2.3 kV. The instrument was a data-dependent MS / MS (MS-dd-MS). 2 ) using full MS 1 It worked. Full MS 1 Data acquisition in dd-MS mode was with a resolution of 60,000, a scan range of 100–1000 m / z, a normalized automatic gain control (AGC) target of 25%, and a maximum injection time (IT) of 50 ms. 2 Data acquisition in mode was at a resolution of 15,000, a normalized AGC target of 20%, a maximum IT of 150 ms, an isolation window of 1.5 m / z, and a normalized collision energy of 40. Metabolite identification was performed using analytical standards and / or products from in vitro UGT enzyme assays (Figure 4D and Figure S12B).
[0531] Heterologous expression in S. cerevisiae For expression of HuCoAT6, HuTKS4, CsOAC, and HuCBGAS in S. cerevisiae, the CDS was amplified and the purified amplicons were transfected into a series of pESC (Amp R) plasmid to allow simultaneous expression of two genes from a single plasmid. HuCoAT6 and HuTKS4 were inserted into the pESC-HIS plasmid, linearized with SalI and SacI restriction enzymes, respectively, using the ClonExpress II One Step Cloning Kit (Vazyme). HuCBGAS and CsOAC were similarly cloned into the pESC-TRP plasmid, linearized with SalI / SacI restriction enzymes, respectively. A complete list of primers used for cloning can be found in Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, 817-831 (2023). The pESC constructs were transformed into S. cerevisiae WAT11 using the Yeastmaker yeast transformation system (Clontech). We transformed yeast cells with a combination of pESC vectors that allowed for simultaneous expression of all four genes. Transformed yeast were grown on SD minimal medium supplemented with the appropriate amino acids and 2% glucose. Colonies were screened, and the presence of the transgene was confirmed by colony PCR. To induce gene expression, transformed cells were grown in 2 ml of minimal medium containing 2% glucose. After 24 h, they were transferred to minimal medium containing 2% galactose without additional supplementation or supplemented with either GPP (0.21 mM) and sodium hexanoate (1 mM) or OA 92 (0.2 mM) and grown for an additional 24 h at 30°C. The cultures were transferred to 2 ml Eppendorf tubes and centrifuged at 8,000 g for 1 min. The cell pellets were weighed and lysed using a bead beater at 22 Hz for 6 min with the addition of two volumes of glass beads (500 μm diameter) and 500 μl of MeOH. The lysed cells were centrifuged at 14,000 rpm for 5 min, and the clear supernatant was collected and dried using a SpeedVac. The dried residue was dissolved in 100 μl of methanol, filtered through a 0.22 μm filter, and analyzed by LC-MS as detailed for the N. benthamiana samples.
[0532] Example 1 H. umbraculigerum produces CBGA Because two previous reports on the presence of cannabinoids, specifically CBGA1, in H. umbraculigerum were contradictory, we decided to perform comprehensive chemical profiling of cannabinoids in various H. umbraculigerum tissues. We confirmed that CBGA1 is a major component of H. umbraculigerum, accumulating in leaves at up to 4.3% dry weight (Figure 1C-1D), comparable to the maximum concentration typically measured in the inflorescences of Cannabis species (Figure 1D). CBGA1, its phenethyl analog heliCBGA2, and preamorphous stilbole (APHA, 3), the stilbene form of heliCBGA2, represent three major peaks in the total ion chromatogram of ethanol extracts of fresh leaves (Figure 1C-1E, and Figure 7A).
[0533] We predicted that the biosynthesis of CBGA1 and heliCBGA2 is derived from hexanoic acid and phenylalanine, respectively (Figure 1A). Therefore, we administered unlabeled and stable isotope-labeled hexanoic acid (hexane-D) to the leaves of H. umbraculigerum. 11 acid) or phenylalanine (phenylalanine-D5 or phenylalanine- 13C9), and the labeled and unlabeled mass spectra and their respective tandem mass spectrometry (MS / MS) spectra were compared (Figure 7B). As a result, the newly derived isotopologues were detected as coeluting chromatographic peaks (unlabeled and labeled forms) with mass shifts and MS / MS fragmentation patterns corresponding to the isotopically labeled portions of the molecules. These findings validated the presence of alkyl and aralkyl cannabinoids in H. umbraculigerum and confirmed that their biosynthesis originates from the polyketide pathway and the phenylpropanoid pathway, respectively. Feeding experiments revealed the presence of several additional major prenyl-acyl-phloroglucinoids, prenylchalcones and prenylflavanones, with chemical formulas similar to 1–3. Based on the previously identified core structure and the MS / MS fragmentation spectra of each metabolite, we assigned these peaks to the structure shown in Figure 1E [see also Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)].
[0534] Example 2 Cannabinoids accumulate in glandular trichomes We used various high-resolution imaging techniques to investigate whether H. umbraculigerum, like Cannabis, develops and accumulates cannabinoids in glandular trichomes. We found that in flowers, the involucre bracts of the flower head contained numerous nonglandular and glandular trichomes. In individual florets, glandular trichomes were particularly abundant at the tips of the corolla lobes (Figures 8A-8B). In leaves, both the adaxial and abaxial surfaces were densely covered with both nonglandular and glandular trichomes (Figure 1F). The glandular trichomes were slightly elevated from the epidermis and consisted of two rows of stalks and a globular head (Figure 8C). Two disc cells (DCs) were observed in the subcutaneous space of the globular head (Figure 1G). In Cannabis, cannabinoid biosynthesis occurs in these cells. The multicellular bilayer structure of the trichome further consisted of two basal cells (BC, not always observed), stalk cells (SC), neck cells (NC), and a secretory cavity (SCv) (Figure 1H). The DC of the secretory trichome exhibited exudation of electron-lucent secretions from plastids into vesicles, followed by exocytosis of their contents into the periplasmic space (PSP), where they were accumulated and then secreted into the SCv (Figure 1I and Figures 2D-2F).
[0535] Next, we applied matrix-assisted laser desorption / ionization mass spectrometry imaging (MALDI-MSI) to spatially localize cannabinoids in H. umbraculigerum. We first analyzed the abaxial and adaxial leaf surfaces after partially removing the trichomes (Figures 8G-8H, imaging corresponding to CBGA 1 and geranylfluorocaprophenone 4, m / z [M+H]). +=361.237 Da). As shown, the metabolite was detected in intact sections, whereas areas where the trichomes had been partially or completely removed showed little or no signal, respectively. We further analyzed cross-sections of H. umbraculigerum leaves and flowers. We transversely cut the leaves so that the adaxial and abaxial trichomes were exposed on both sides (Figure 2A). In flowers, we cut the receptacle to expose the trichomes on the outer surface of the involucre (Figure 8I). As shown in Figures 2B and 8J for leaf and flower samples, respectively, CBGA1 was found exclusively in the glandular trichomes.
[0536] Example 3 H. umbraculigerum produces both classical and novel cannabinoids Cannabis produces a variety of CBGA-type analogs with aliphatic chain lengths (1–7 carbon atoms) derived from various linear short- and medium-chain fatty acids (FAs). We observed several of these analogs in H. umbraculigerum leaves, including cannabigerovaric acid (CBGVA 9), cannabigerol butyrate (CBGBA 10), cannabigerohexolic acid (CBGHA 11), and cannabigerophorolic acid (CBGPA 12), corresponding to chains of 3, 4, 6, and 7 carbon atoms, respectively (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, 817–831 (2023)). We also observed two metabolites with similar masses and fragmentation patterns to CBGA 1 and CBGHA 11, which we attribute to cannabinoids derived from these branched FAs (13 and 14, respectively; Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). These branched cannabinoids have not been identified in cannabis. We also found small amounts of CBCA 15 and its aromatic analog helichromenic acid (heliCBCA 16) and their hydroxylated forms (17 and 18, respectively; Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)), as well as isoprenyl forms of CBGA 1 and heliCBGA 2 by MS / MS fragmentation (CBPA 19 and heliCBPA 20, respectively; Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)).The present inventors found that Δ in both tissues. 9 -No THCA or CBDA type cannabinoids were detected.
[0537] Several additional peaks showed MS / MS fragments and chemical formulas corresponding to one or two hydroxylations of metabolites with five carbon atom chains. 11 The hydroxylated amorfurthine was labeled after obtaining the acid (21-33, Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). Interestingly, the hydroxylated amorfurthine was observed to have a similar fragmentation pattern to that of cannabinoids (m / z difference 33.984 Da), suggesting that similar chemical structures and enzymes are involved in its metabolism (34-46, Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). We purified this metabolite, 26, and identified a new tetrahydroxanthan-type cannabinoid (12-OH-cyclocannabigerolic acid 26) by NMR spectroscopy. According to its MS / MS fragmentation pattern, we also putatively identified cyclocannabigerolic acid (cycloCBGA47) and the similar amorfurtine type [12-OH-heli-cyclocannabigerolic acid (12-OH-helicycloCBGA39) and helicyclocannabigerolic acid (helicycloCBGA48), respectively; Berman et al., "Parallel evolution of cannabinoid biosynthesis", Nature Plants 9 817-831 (2023)].
[0538] Current feeding experiments indicate that prenyl-acyl-phloroglucinoids, prenylchalcones, and prenylflavanones are derived from similar precursors to cannabinoids and amorfurthins (49-91, Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). A summary of identified metabolites 1-91 is provided in Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)).
[0539] Example 4 A proposed cannabinoid biosynthetic pathway in H. umbraculigerum We hypothesize that the core cannabinoid pathway leading to CBGA 1 in H. umbraculigerum is composed of enzymes and reactions similar to those in cannabis (Figure 9). These include acyl-activating enzymes (AAEs) that activate hexanoic acid to hexanoyl-CoA; type III polyketide synthases (PKSs) and polyketide cyclases (PKCs) that produce olivetolic acid (OA 92); and membrane-bound aromatic prenyltransferases (PTs) that geranylate OA 92 to CBGA 1. In addition to CBGA 1 and other cannabinoids, we propose that all identified terpenoids are produced via five parallel pathways (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, pp. 817-831 (2023)). According to this scheme, cannabinoids and phloroglucinoids originate from a common linear or branched FA precursor activated via the same AAE enzyme. Amorfurthines and chalcones are derived from phenylalanine to form cinnamic or coumaric acids and activated via AAE enzymes (similar or different from polyketide enzymes). These activated intermediates can be further reduced by double bond reductases (DBRs) to form dihydrogen intermediates. The activated precursors are then elongated by one or more PKS-type enzymes using three malonyl-CoA units and further cyclized by PKS in a Claisen reaction to form the phloroglucinoid or chalcone skeleton, or in a PKC-assisted aldol reaction to form cannabinoids and amorfurthines. A fifth pathway uses the chalcone isomerase (CHI) enzyme, which cyclizes the chalcone to a flavanone. All of these intermediates are further prenylated by one or more PTs to form various types of terpenophenols. While the majority of the encapsulated molecules are monoprenylated, other prenylated forms have also been observed.Terpenophenols can be further cyclized by berberine bridge-like enzymes (BBEs) to produce cyclized metabolites such as CBCA 15, cyclocannabinoids, and cycloamorphurtines (26, 47, 39, 48), as well as the cyclophloroglucinoids previously identified by Pollastro et al. (2017). Further functionalization and rearrangements include hydroxylation, double bond isomerization, or reduction. To support these five pathways, we identified primary intermediates (before prenylation) from all corresponding metabolic pathways in H. umbraculigerum (92–101, Berman et al., “Parallel evolution of cannabinoid biosynthesis,” Nature Plants 9, 817–831 (2023)).
[0540] Example 5 Elucidating the core cannabinoid pathway To identify the enzymes responsible for cannabinoid biosynthesis in H. umbraculigerum, we obtained a haplotype-resolved dual genome assembly using 44x Pacbio HiFi reads and 200M reads from Illumina HiC chromatin interaction data (haploid size approximately 1.3 Gb; Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). After scaffolding, the N50 of the primary assembly was 174 Mb, with eight scaffolds exceeding 10 Mb (Figure 2C and Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). We also obtained RNA-seq data for various tissues using PacBio Iso-Seq, Illumina True-Seq, and Illumina UM-aware 3'Transeq (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). We soft-masked the genome (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). We obtained gene models with a BUSCO completeness of 98.7% for the primary assembly and 99.3% for all transcripts, including those missing from the genome (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). Based on Figure 2B, we predicted that biosynthetic genes would be highly expressed in trichomes.Weighted gene coexpression network analysis of transcriptome data from H. umbraculigerum tissues (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)) revealed a transcriptional module enriched for FA and terpenoid biosynthetic genes induced in trichomes and leaves (Figure 2D-2E, and Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). This module contained two AAEs, three PKSs, one stress-related protein (potential PKC), and one PT (Figure 2E and Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). Notably, three of these PKSs were also located in a tandem gene cluster consisting of seven enzymes of the same type (Figure 2C). This region showed a strong signature of long terminal repeat (LTR) transposition activity, which may explain the observed gene duplication pattern (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). Overall, we selected six HuAAEs, four HuPKSs, five HuPKCs, and four HuPTs for further characterization (Figure 3A and Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). The four selected PKSs exhibited subtle amino acid differences that would have been overlooked without the genomic sequences. Because we were unable to amplify the different variants from cDNA, we synthetically generated the genes.
[0541] The first step in cannabinoid biosynthesis involves the formation of acyl-CoA thioesters by members of the AAE superfamily. Because different acyl moieties serve as substrates for these enzymes, we tested acetate, butyrate, hexanoate, octanoate, cinnamate, and coumarate. In vitro assays using purified recombinant proteins showed that HuAAE2 and HuAAE4 efficiently produced butyryl-CoA, while HuAAE2 showed greater activity toward acetate and formed acetyl-CoA (Figure 3B and Figure 10A). HuAAE6 (HuCoAT6) was the only enzyme with activity toward both medium-chain alkyl (e.g., hexanoate and octanoate) and aralkyl (e.g., cinnamate and coumarate) precursors required for the five terpenoids observed in H. umbraculigerum. Interestingly, HuAAE4 belongs to the same family as the most active Cannabis enzymes, while HuCoAT6 is located within the family of long-chain acyl-CoA synthetases (LACS, Figure 11A and Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)).
[0542] In Cannabis, the next step is carried out by a coupled enzymatic reaction involving CsOLS and the accessory protein CsOAC, resulting in the condensation of hexanoyl-CoA with three malonyl-CoA molecules to produce OA92. In vitro assays show that decarboxylation of unstable intermediates occurs, generating additional byproducts not naturally identified in plant extracts [olivetol, pentyl acyl diacetic acid lactone (PDAL), and hexanoyl acyl triacetic acid lactone (HTAL); Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, 817-831 (2023)]. PDAL and HTAL are produced by spontaneous lactonization of unstable triketide and tetraketide intermediates, whereas olivetol is produced by CsOLS in the absence of CsOAC in an aldol decarboxylation cyclization reaction similar to the production of resveratrol by stilbene synthase (STS). When CsOAC is also present in the reaction, OA92 is produced at the expense of olivetol. Here, we cloned and expressed the E. coli HuPKS1-4, HuPKC1-5, CsOLS, and CsOAC enzymes and tested their ability to form OA92 in a coupled in vitro assay using all possible combinations of hexanoyl-CoA and malonyl-CoA (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, pp. 817-831 (2023)). In the absence of PKC, all HuPKSs produced PDAL and HTAL byproducts, but HuPKS1, HuPKS2, and HuPKS4 also produced olivetol (Figures 3C and 10C). When reactions were performed in combination with CsOAC, particularly with HuPKS4 (HuTKS4), olivetol decreased and OA 92 increased (Figure 3C and Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)).However, significantly lower amounts of olivetol and OA 92 were observed in all reactions compared to HTAL and PDAL (Figure 3C). Interestingly, regardless of CsOAC, all HuPKSs produced the phloroglucinoid precursor phlorocaprophenone 95 (PCP), present in H. umbraculigerum (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023) and Figure 3C). This suggested that the same HuPKS enzymes could perform both aldol and Claisen cyclization reactions. This phenomenon has previously been observed for CHS and STS enzymes that produce different amounts of both naringenin and resveratrol, as well as PKSs that produce different ratios of resorcinolic acid and phloroglucinoid products. Interestingly, the HuPKS protein sequence did not cluster with known resorcinoleic acid or phloroglucinoid-producing PKSs, such as CsOLS, Rhododendron dauricum orcinol synthase (RdOS), or hop (Humulus lupulus) valerophenone synthase (HlVPS) (Figure 11B). Neither HuPKS1-HuPKS4 nor combinations containing CsOLS and HuPKC enzymes (selected based on their expression profiles and sequence homology to CsOAC) resulted in the formation of OA92 (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, pp. 817-831 (2023)). This suggests that the cyclization and possibly stabilization of the tetraketide intermediate is mediated by a different class of enzymes than in cannabis. This was previously suggested to occur in Rhododendron dauricum in the production of orsellinic acid by RdOS and an as yet unidentified PKC enzyme. 20Alternatively, H. umbraculigerum may contain another CsOAC homolog that we did not characterize in this study.
[0543] In the next step, OA92 or an OA derivative is prenylated by aromatic PTs to form CBGA1 and its derivatives. We expressed four enzymes in yeast and purified microsomal fractions for use in enzyme assays (HuPT1-4, Figure 3D). We examined a series of aromatic substrates and either geranyl pyrophosphate (GPP) or isopentenyl pyrophosphate (IPP) as the isoprenoid donor. All HuPTs geranylated OA92 and divalinolic acid (VA) to produce CBGA1 and CBGVA9, respectively. HuPT4 also geranylated aromatic dihydrostilbenic acid (DHSA93) and was the only enzyme that isoprenylated OA92 and DHSA93 (Figure 3D). HuPT4 was also active with farnesyl pyrophosphate (FPP) to produce sesquicannabigerolic acid (SesquiCBGA, Figure 10D). Kinetic assays of HuPTs using GPP and OA 92 revealed that HuPT4 (HuCBGAS4) exhibited a Michaelis-Menten Km value smaller than that reported for Cannabis sativa CsPT4 (Figure 3e and Figure S10E). Interestingly, none of the HuPTs prenylated phloroglucinoid or chalcone intermediates, and none of their sequences clustered with previously known terpenophenolic PTs (Figure 3F).
[0544] To gain more insight into the pathway's evolution, we searched for orthologous enzymes in Cannabis and all other Asteraceae species with annotated genomes. To our knowledge, these species do not accumulate terpenoids. Similarly, similar to the phylogenetic relationships observed for the functionally tested enzymes (i.e., AAEs, PKSs, and PTs, Figure 11), the enzymes that enable H. umbraculigerum to produce cannabinoids evolved independently in this lineage. In particular, for PKS-type enzymes, gene duplication and subsequent specialization likely occurred multiple times within this family. Interestingly, the PTs from Cannabis and H. umbraculigerum did not cluster in the same orthogroup, suggesting that they are derived from a distant evolutionary ancestor.
[0545] Example 6 Decorated cannabinoids are formed by UGT and BAHD enzymes Glycosylated cannabinoids have not been reported to occur naturally in plants. Here, we identified glucosylated OA (Glc-OA 102) and glucosylated DHSA (Glc-DHSA 103), as well as glucosylated C3-C6 alkyl chain intermediates (104-108), glucosylated CBGA (Glc-CBGA 109) and heliCBGA (Glc-heliCBGA 110), and their isoprenylated forms (Glc-CBPA 111 and Glc-heliCBPA 112) (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, 817-831 (2023)). All of these metabolites showed a neutral loss of 162.053 Da, corresponding to a fragment similar to the hexose and non-glucosylated compounds. No diglucosylated metabolites were identified in the extract. In Arabidopsis thaliana, uridine 5'-diphospho-glucuronosyltransferases (AtUGT89B1, AtUGT71B1, AtUGT75B1, and AtUGT71B2) catalyze the glycosylation of several hydroxybenzoic acids (HBA and DHBA) structurally similar to OA 92 (Figure 4A). We selected 13 gene candidates in H. umbraculigerum based on sequence similarity to these proteins and a positive correlation between gene expression and accumulation of glucosylated metabolites (HuUGT1-13, Figure 4A-B; see Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, pp. 817-831 (2023)).
[0546] Eleven of the 13 UGTs from H. umbraculigerum were expressed in E. coli, and their enzymatic activities were examined using OA 92, CBGA 1, and heliCBGA 2 in reactions containing uridine diphosphate glucose (UDP-Glc) as the sugar donor. Eight of the 11 enzymes showed activity toward various substrates, including HuUGT1-2, HuUGT4-7, HuUGT11, and HuUGT13 (Figure 12A). The production of Glc-OA 102 in the enzymatic assay using UDP-Glc was supported by the NMR assignment of the glucose moiety. We next purified the four most active enzymes (HuUGT1, HuUGT6, HuUGT11, and HuUGT13) and performed in vitro assays with a range of natural and unnatural cannabinoid substrates in H. umbraculigerum (Figure 4D and Figure 12B; Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). We also included enzymes from stevia and rice (SrUGT and OsUGT, respectively), which have been reported to have cannabinoid glycosylation activity, even though these plants do not produce cannabinoids. All enzymes were active with a variety of substrate specificities and products. For example, HuUGT1 and HuUGT6 were most active toward cannabinoids (HuCBUGT1 and HuCBUGT6), whereas HuUGT11 (HuOAUGT11) and HuUGT13 were highly active toward cannabinoid intermediates but largely inactive toward prenylated metabolites. Diglucosylation of acidic metabolites was observed only with HuCBUGT6, whereas olivetol, cannabidiol (CBD), and cannabigerol (CBG) were diglucosylated by different HuUGTs depending on the metabolite (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, pp. 817-831 (2023)).Interestingly, UGTs from H. umbraculigerum also glucosylated phloroglucinoid and flavonoid precursors naturally occurring in plants (Figure S4B), yet no glucosylated forms were observed in plant extracts. Kinetic assays of HuOAUGT11, HuUGT13, OsUGT, and SrUGT showed highly significant catalytic activity of HuOAUGT11 with OA92 and UDP-Glc compared with all other enzymes (Figures 4C and S4C). HuOAGT11 is also co-expressed with other cannabinoid-related enzymes (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)) and is most likely the enzyme responsible for generating the large amounts of Glc-OA 102 and Glc-DHSA 103 produced in H. umbraculigerum (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)).
[0547] Previous reports have identified isoprenylated O-acylated amorfurtines in H. umbraculigerum, but not geranylated or alkylated forms, which are also found in Cannabis. Here, we identified a diverse group of O-acylated cannabinoids and amorfurtines, including O-acylated alkyl (113-130) and aralkyl (131-141) metabolites (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). We hypothesized that the acyl group was derived from short- or medium-chain FAs (Figure 9) and verified this using precursor isotope labeling (Figure 13A). The majority of alkylcannabinoids in this group possess a five-carbon tail (hexane-D). 11Both alkyl and aralkyl metabolites contained isoprenyl or monoprenyl and linear or branched short-chain O-acyl groups, as indicated by specific labeling (Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9 817-831 (2023)). To confirm the identity of this group of metabolites, we purified O-methylbutyryl-cannabigerolic acid (O-MeButCBGA 120) and O-methylbutyryl-helicannabigerolic acid (O-MeButheliCBGA 138) and confirmed their structures by NMR.
[0548] O-acylation of specialized metabolites in plants is often catalyzed by BAHD-type alcohol acyltransferase (AAT) enzymes. Therefore, we selected 15 H. umbraculigerum BAHD homologs, four of which were coexpressed with other cannabinoid-related enzymes (Figures 2E and 4B and Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, 817-831 (2023)). Twelve of the 15 AATs were expressed in Escherichia coli and their activity was examined using butyryl- and hexanoyl-CoA as acyl donors and CBGA1 and helicobacter CBGA2 as acceptors. Only HuAAT5 and HuAAT14 showed activity with these substrates (Figure 13B). Phylogenetic analysis showed that these two enzymes clustered in phylogenetic group IIIa, which represents BAHDs with diverse catalytic functions (Figure 13C). HuAAT5 (HuCBAT5) was selected for detailed characterization with a series of acyl donors and acceptors because it produced larger amounts of product. It accepted all tested acyl donors and acylated OA 92, CBGA 1, heliCBGA 2, and CBDA, yielding a single O-acyl-cannabinoid from each substrate pair (Figures 4E-F, and Figure 14; see Berman et al., "Parallel evolution of cannabinoid biosynthesis," Nature Plants 9, 817-831 (2023)). Many of the cannabinoids produced were those naturally observed in plants (indicated by an asterisk in Figure 4E). Meanwhile, this enzyme exhibited a Δ 9 It was inactive with -THCA and CBCA 15. Therefore, it appears that only the hydroxyl at C5 is acylated. Furthermore, O-acyl esterification in H. umbraculigerum was observed only with prenylated cannabinoids and amorfurthin, but not with their intermediates.
[0549] Example 7 In vivo reconstitution of the core cannabinoid pathway in a heterologous system We verified the in planta activity of the enzymes for CBGA 1 by transiently co-expressing different combinations of HuCoAT6, HuTKS4, and HuCBGAS4 with the cannabis (CsOAC) and CsOLS in N. benthamiana leaves. After infiltrating the leaves with sodium hexanoate and GPP, we observed the production of glycosylated forms of OA 92 (HuTKS4 + CsOAC or CsOLS + CsOAC) and PCP 95 (HuTKS4 only, Figures 5A and 15A-B). This was consistent with previous studies reporting glycosylation of OA 92 by endogenous enzymes in this plant. Interestingly, we also observed glycosylation products of naringenin chalcone 97 with HuTKS4, suggesting that this enzyme can accept aromatic substrates in addition to aliphatic forms (Figures 5A and 15A-15B). However, we did not observe CBGA1 or its glycosylated forms with HuCBGAS4. This is likely due to the low availability of OA92 and its rapid glycosylation in planta. When leaves expressing HuCBGAS4 were infiltrated with OA92 and GPP, CBGA1 and Glc-CBGA109 were observed (Figures 5B, 15A, and 15C).
[0550] We also reconstituted the cannabinoid pathway by expressing the HuCoAT6, HuTKS4, CsOAC, and HuCBGAS4 genes in S. cerevisiae. We observed the production of OA92, CBGA1, and PCP95 without the need for precursors (Figures 5C, 15D, and 15E). Similar to the in vitro assay, we also observed peaks of HTAL and PDAL, which are not present in planta (Figure 15F). When cells were supplemented with OA92 and GPP, significantly higher amounts of CBGA1 were produced (Figure 5D).
[0551] While the present invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, the present invention is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and broad scope of the appended claims.
Claims
1. 1. An isolated DNA molecule comprising at least a first nucleic acid sequence encoding a first protein and at least a second nucleic acid sequence encoding a second protein, wherein the first protein and the second protein are derived from Helichrysum umbraculigerum and belong to an enzyme family selected from the group consisting of acyl-activating enzymes (AAEs), polyketide synthases (PKSs), polyketide cyclases (PKCs), prenyltransferases (PTs), and cannabichromenic acid synthases (CBCASs), and wherein the first protein and the second protein belong to different enzyme families.
2. 2. The isolated DNA molecule of claim 1, further comprising at least a third nucleic acid sequence derived from H. umbraculigerum and encoding a third protein belonging to an enzyme family selected from the group consisting of AAE, PKS, PKC, PT, and CBCAS, wherein the first protein, the second protein, and the third protein belong to different enzyme families.
3. 3. The isolated DNA molecule of claim 2, further comprising at least a fourth nucleic acid sequence encoding a fourth protein derived from H. umbraculigerum, the fourth protein belonging to an enzyme family selected from the group consisting of AAE, PKS, PKC, PT, and CBCAS, and wherein the first protein, the second protein, the third protein, and the fourth protein belong to different enzyme families.
4. 4. The isolated DNA molecule of claim 3, further comprising at least a fifth nucleic acid sequence derived from H. umbraculigerum and encoding a fifth protein belonging to an enzyme family selected from the group consisting of AAE, PKS, PKC, PT, and CBCAS, wherein the first protein, the second protein, the third protein, the fourth protein, and the fifth protein belong to different enzyme families.
5. 5. The isolated DNA molecule of any one of claims 1 to 4, further comprising a nucleic acid sequence encoding a protein belonging to an enzyme family selected from the group consisting of uridine diphosphate (UDP)-glycosyltransferases (UGTs), alcohol acyltransferases (AATs), and both, derived from H. umbraculigerum.
6. 6. The isolated DNA molecule of any one of claims 1 to 5, a. The AAE is encoded by a nucleic acid sequence having at least 89% identity to any one of SEQ ID NOs: 1-11, and any combination thereof; b. the PKS is encoded by a nucleic acid sequence having at least 83% identity to any one of SEQ ID NOs:23-26, and any combination thereof; c. The PKC is encoded by a nucleic acid sequence having at least 88% identity to any one of SEQ ID NOs: 31-38, and any combination thereof; d. The PT is encoded by a nucleic acid sequence having at least 91% identity to any one of SEQ ID NOs: 47-58, and any combination thereof; e. the CBCAS is encoded by a nucleic acid sequence having at least 82% identity to any one of SEQ ID NOs: 71-79, and any combination thereof; or f. An isolated DNA molecule that is any combination of (a) through (e).
7. 7. The isolated DNA molecule of claim 5 or 6, a. the UGT is encoded by a nucleic acid sequence having at least 87% identity to any one of SEQ ID NOs: 89-101, and any combination thereof; b. The AAT is encoded by a nucleic acid sequence having at least 87% identity to any one of SEQ ID NOs: 115-129, and any combination thereof; or c. An isolated DNA molecule that is both (a) and (b).
8. 8. The isolated DNA molecule of any one of claims 1 to 7, a. the AAE comprises an amino acid sequence having at least 93% homology to any one of SEQ ID NOs: 12-22; b. the PKS comprises an amino acid sequence having at least 93% identity to any one of SEQ ID NOs:27-30; c. the PKC comprises an amino acid sequence having at least 87% identity to any of SEQ ID NOs: 39-46; d. the PT comprises an amino acid sequence having at least 92% homology to any one of SEQ ID NOs: 59-70; e. the CBCAS comprises an amino acid sequence having at least 86% identity to any one of SEQ ID NOs: 80-88; or f. An isolated DNA molecule that is any combination of (a) through (e).
9. 9. The isolated DNA molecule of any one of claims 5 to 8, a. the UGT comprises an amino acid sequence having at least 90% identity to any one of SEQ ID NOs: 102-114; b. The AAT comprises an amino acid sequence having at least 91% identity to any one of SEQ ID NOs: 130-144; or c. An isolated DNA molecule that is both (a) and (b).
10. 10. The isolated DNA molecule of any one of claims 1 to 9, a. The AAE consists of the amino acid sequence of any one of SEQ ID NOs: 12 to 22; b. the PKS consists of the amino acid sequence of any one of SEQ ID NOs: 27-30; c. The PKC consists of the amino acid sequence of any one of SEQ ID NOs: 39 to 46; d. The PT consists of the amino acid sequence of any one of SEQ ID NOs: 59-70; e. the CBCAS consists of the amino acid sequence of any one of SEQ ID NOs: 80-88; or f. An isolated DNA molecule that is any combination of (a) through (e).
11. 11. The isolated DNA molecule of any one of claims 5 to 10, a. the UGT consists of the amino acid sequence of any one of SEQ ID NOs: 102-114; b. The AAT consists of the amino acid sequence of any one of SEQ ID NOs: 130-144; or c. An isolated DNA molecule that is both (a) and (b).
12. 12. The isolated DNA molecule of any one of claims 1 to 11, comprising a plurality of isolated DNA molecule forms.
13. 13. The isolated DNA molecule of claim 12, wherein each type of the plurality of isolated DNA molecule types encodes one or more proteins belonging to different enzyme families.
14. An artificial nucleic acid molecule comprising the isolated DNA molecule according to any one of claims 1 to 13.
15. A plasmid or agrobacterium comprising the artificial nucleic acid molecule of claim 14.
16. A transgenic cell comprising: a. the isolated DNA molecule of any one of claims 1 to 13; b. the artificial nucleic acid molecule of claim 14; c) the plasmid or Agrobacterium of claim 15; or d. Any combination of (a) to (c) A transgenic cell comprising:
17. 17. The transgenic cell of claim 16, which is one of a unicellular organism, a cell of a multicellular organism, and a cell in culture.
18. 18. The transgenic cell of claim 17, wherein the unicellular organism comprises a fungus or a bacterium.
19. 19. The transgenic cell of claim 18, wherein the fungus is a yeast cell.
20. 20. The transgenic cell of claim 19, which is a transgenic Cannabis sativa cell.
21. An extract derived from the transgenic cells of any one of claims 19 to 20, or any fraction thereof.
22. 22. The extract of claim 21, comprising a cannabinoid, a precursor thereof, or a combination thereof.
23. A transgenic plant, transgenic plant tissue or plant part, comprising: a. the isolated DNA molecule of any one of claims 1 to 13; b. the artificial nucleic acid molecule of claim 14; c) the plasmid or Agrobacterium of claim 15; d. The transgenic cell of any one of claims 16 to 20; or e. Any combination of (a) to (d) 10. A transgenic plant, transgenic plant tissue or plant part comprising:
24. 24. The transgenic plant, transgenic plant tissue or plant part of claim 23, wherein the plant is a transgenic C. sativa plant.
25. 1. A composition comprising: a. the isolated DNA molecule of any one of claims 1 to 13; b. the artificial nucleic acid molecule of claim 14; c) the plasmid or Agrobacterium of claim 15; d. The transgenic cell of any one of claims 16 to 20; e) the extract of claim 21 or 22; f. A transgenic plant, transgenic plant tissue or plant part according to claim 23 or 24; or g. Any combination of (a) to (f); and an acceptable carrier.
26. 1. A method for synthesizing a cannabinoid, a precursor thereof, or a combination thereof, comprising: a) Providing a transgenic cell or a cell transfected with an isolated DNA molecule according to any one of claims 1 to 13 or an artificial nucleic acid molecule according to claim 14; b. Culturing the transgenic or transfected cells of step (a) so that at least the first protein and the second protein encoded by the artificial nucleic acid molecule are expressed, thereby synthesizing the cannabinoid, its precursor, or any combination thereof.
27. 27. The method of claim 26, wherein the precursor is selected from the group consisting of acyl-coenzyme A (CoA), polyketides, resorcinoid precursors, and any combination thereof.
28. 28. The method of claim 27, wherein the acyl is a C1-C8 alkyl.
29. 29. The method of claim 27 or 28, wherein the acyl-CoA is hexanoyl-CoA.
30. 30. The method of any one of claims 27 to 29, wherein the polyketide is a tetraketide.
31. 31. The method of claim 30, wherein the tetraketide is a linear tetraketide.
32. 32. The method of any one of claims 27 to 31, wherein the resorcinoid precursor is olivetolic acid.
33. 33. The method of any one of claims 26 to 32, wherein the cannabinoid is cannabigerolic acid (CBGA), CBCA, or both.
34. The method according to any one of claims 26 to 33, wherein the artificial nucleic acid molecule is an expression vector.
35. The method of any one of claims 26 to 34, wherein the transgenic or transfected cell is a prokaryotic or eukaryotic cell.
36. The method of any one of claims 26 to 35, wherein the transgenic or transfected cell is a C. sativa cell.
37. The method of any one of claims 26 to 36, further comprising, prior to step (a), introducing or transfecting the artificial nucleic acid molecule into a cell, thereby obtaining the transgenic cell or the transfected cell.
38. 38. The method of any one of claims 27 to 37, further comprising the step of extracting said transgenic or transfected cells, thereby obtaining an extract from said transgenic or transfected cells.
39. 39. An extract of transgenic or transfected cells obtained according to the method of claim 38.
40. 40. The extract of claim 39, comprising a cannabinoid, a precursor thereof, or any combination thereof.
41. 41. The extract of claim 39 or 40, comprising CBGA, CBCA, or both.
42. A composition comprising the extract of any one of claims 39 to 41 and an acceptable carrier.