Uridine diphosphate-glycosyltransferase and transgenic cells, tissues, and organisms containing the same
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- YEDA RES & DEV CO LTD
- Filing Date
- 2023-04-13
- Publication Date
- 2026-05-19
AI Technical Summary
There is a need to develop methodologies for large-scale glycosylation of cannabinoids or their precursors to improve their processability and applicability, as existing enzymes from non-cannabinoid producing plants are not well characterized for these substrates.
Identification of a UGT gene/enzyme from Helichrysum that catalyzes the glycosylation of olivetolic acid (OA) and cannabinoids, which shows higher efficiency compared to non-specific enzymes from rice and Stevia.
The use of the Helichrysum-derived UGT enzyme enables efficient glycosylation of cannabinoids, potentially enhancing their water solubility and oral bioavailability, thereby improving their therapeutic potential.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Reference to Electronic Sequence Listing The contents of the electronic sequence listing (YEDA-P-013-PCTST26.xml; size: 45,465 bytes; creation date: April 13, 2023) are incorporated herein by reference in their entirety.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 330,490, entitled "URIDINE DIPHOSPHATE-GLYCOSYL TRANSFERASE AND A TRANSGENIC CELL, TISSUE, AND ORGANISM COMPRISING SAME," filed April 13, 2022, the contents of which are incorporated herein by reference in their entirety.
[0003] The present invention relates to uridine diphosphate (UDP)-glycosyltransferases (UGTs), and transgenic cells, tissues and organisms comprising uridine diphosphate (UDP)-glycosyltransferases (UGTs), such as polynucleotides encoding uridine diphosphate (UDP)-glycosyltransferases (UGTs), and methods of using uridine diphosphate (UDP)-glycosyltransferases (UGTs), such as to produce glycosylated cannabinoids or precursors thereof. [Background technology]
[0004] Cannabinoids are unique to Cannabis sativa L. (Cannabis), but some specific compounds have also been identified in other flowering plants, mosses, and fungi. One of these plants is Helichrysum umbraculigerum Less (Helichrysum). This perennial South African plant is the only known plant outside of Cannabis that produces cannabigerolic acid (CBGA), the five-carbon alkyl precursor of all major cannabinoids.
[0005] In the past few years, the therapeutic use of cannabinoids has leapt significantly as new reports highlight their potential in various medical purposes. However, one of the main challenges in the research and development of cannabinoid medicines is to improve their water solubility and their oral bioavailability and absorption into the bloodstream. A possible strategy to overcome this challenge could be the glycosylation of cannabinoids. In plants, uridine diphosphate (UDP)-glycosyltransferases (UGTs) catalyze the covalent addition of sugars to a wide range of lipophilic molecules, improving their solubility. Notably, endogenous cannabinoid glycosylation activity has been shown in several plant species that do not naturally synthesize cannabinoids. Moreover, UGT genes from rice (Oryza sativa) and stevia (Stevia rebaudiana) have been identified in recent years to glycosylate cannabinoids and reported in research papers and patents.
[0006] However, these enzymes were identified in plants that do not produce cannabinoids, and therefore their ability to glycosylate other substrates, such as cannabinoids or their precursors, that are not natural substrates, is unknown.
[0007] There is therefore a need to develop methodologies that allow for the large-scale glycosylation of cannabinoids or their precursors, which could improve their processability and / or applicability. Summary of the Invention
[0008] The present invention, in some embodiments, is based in part on the identification of glycosylated forms of cannabinoids and their precursor, olivetolic acid (OA), in plants. Due to sequence similarity with UGT enzymes from Arabidopsis thaliana, previously identified to glycosylate 2,4-dihydroxybenzoic acid (2,4-DHBA, a compound structurally similar to OA), the inventors have discovered UGT genes / enzymes from Helichrysum that catalyze the glycosylation of OA and cannabinoids. These identified UGTs that naturally glycosylate cannabinoids in plants are likely more efficient with these compounds than the non-specific enzymes currently proposed for use, such as those from rice and stevia.
[0009] According to a first aspect, there is provided an isolated DNA molecule comprising a nucleic acid sequence having at least 87% homology to SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, or any combination thereof.
[0010] According to another aspect, there is provided an artificial nucleic acid molecule comprising an isolated DNA molecule disclosed herein.
[0011] According to another aspect, there is provided a plasmid or an agrobacterium comprising an artificial nucleic acid molecule disclosed herein.
[0012] According to another aspect, there is provided an isolated protein encoded by any one of: (a) an isolated DNA molecule disclosed herein; (b) an artificial vector disclosed herein; and (c) a plasmid or agrobacterium disclosed herein.
[0013] According to another aspect, there is provided a transgenic cell comprising (a) an isolated DNA molecule disclosed herein, (b) an artificial nucleic acid molecule disclosed herein, (c) a plasmid or an Agrobacterium disclosed herein, (d) an isolated protein disclosed herein, or (e) any combination of (a)-(d).
[0014] According to another aspect, there is provided an extract derived from the transgenic cells disclosed herein, or any fraction thereof.
[0015] According to another aspect, there is provided a transgenic plant, transgenic plant tissue or plant part comprising (a) an isolated DNA molecule disclosed herein, (b) an artificial vector disclosed herein, (c) a plasmid or agrobacterium disclosed herein, (d) an isolated protein disclosed herein, (e) a transgenic cell disclosed herein, or (f) any combination of (a)-(e).
[0016] According to another aspect, there is provided a composition comprising (a) an isolated DNA molecule as disclosed herein, (b) an artificial vector as disclosed herein, (c) a plasmid or agrobacterium as disclosed herein, (d) an isolated protein as disclosed herein, (e) a transgenic cell as disclosed herein, (f) an extract as disclosed herein, (g) a transgenic plant tissue or plant part as disclosed herein, or (h) any combination of (a)-(g) and an acceptable carrier.
[0017] According to another aspect, a method for glycosylation of a cannabinoid or a precursor thereof is provided, the method comprising: (a) providing a cell comprising an artificial vector comprising a nucleic acid sequence having at least 87% homology to SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13; and (b) culturing the cell of step (a) such that a protein encoded by the artificial vector is expressed, thereby glycosylating the cannabinoid or a precursor thereof.
[0018] According to another aspect, there is provided an extract of a cell obtained according to the methods disclosed herein.
[0019] According to another aspect, there is provided a medium, or a portion thereof, separated from cultured cells obtained according to the methods disclosed herein.
[0020] According to another aspect, there is provided a composition comprising: (a) an extract as disclosed herein; (b) a medium as disclosed herein or a portion thereof; or (c) a combination of (a) and (b) and an acceptable carrier.
[0021] According to another aspect, there is provided a method for glycosylation of a cannabinoid or precursor thereof, the method comprising contacting a cannabinoid or precursor thereof with an effective amount of a protein comprising an amino acid sequence having at least 90% homology to SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26, thereby glycosylating the cannabinoid or precursor thereof.
[0022] In some embodiments, the nucleic acid sequence has at least 87% homology to any one of SEQ ID NOs: 1-13 and is 700-1,800 nucleotides in length.
[0023] In some embodiments, the nucleic acid sequence encodes a protein that is a uridine 5'-diphospho (UDP)-glucuronosyltransferase (UGT).
[0024] In some embodiments, the isolated protein comprises an amino acid sequence having at least 90% homology to SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26.
[0025] In some embodiments, the isolated protein consists of the amino acid sequence of SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26.
[0026] In some embodiments, the isolated protein is characterized in that it is capable of glycosylating a cannabinoid or a precursor thereof.
[0027] In some embodiments, the transgenic cell is a cell of a unicellular organism, a cell of a multicellular organism, or a cell in culture.
[0028] In some embodiments, the unicellular organism comprises a fungus or a bacterium.
[0029] In some embodiments, the fungus is a yeast cell.
[0030] In some embodiments, the extract comprises isolated DNA molecules, isolated proteins, or both.
[0031] In some embodiments, the transgenic plant is a Cannabis sativa plant.
[0032] In some embodiments, the cell is a transgenic cell, or a cell transfected with an isolated DNA molecule disclosed herein or an artificial vector disclosed herein.
[0033] In some embodiments, the protein is characterized in that it is capable of transferring a glucuronic acid moiety of UDP-glucuronic acid to a cannabinoid or a precursor thereof.
[0034] In some embodiments, the culturing comprises supplementing the cells with an effective amount of UDP.
[0035] In some embodiments, the artificial vector is an expression vector.
[0036] In some embodiments, the cell is a prokaryotic or eukaryotic cell.
[0037] In some embodiments, the method further comprises a step (c) comprising extracting the cells, thereby obtaining an extract of the cells.
[0038] In some embodiments, the method further comprises a step preceding step (c) comprising separating the cultured cells from the medium in which the cells are cultured.
[0039] In some embodiments, the method further comprises a step preceding step (a) comprising introducing or transfecting an artificial vector into the cell.
[0040] In some embodiments, the contacting is in a cell-free system.
[0041] In some embodiments, the cannabinoid is CBDA, CBGA, heliCBGA, or any combination thereof.
[0042] In some embodiments, the cannabinoid precursor is olivetolic acid (OA).
[0043] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention pertains.Although methods and materials similar or equivalent to those described herein can be used to carry out or test embodiments of the present invention, exemplary methods and / or materials are described below.In case of conflict, the present specification, including definitions, will take precedence.In addition, the materials, methods, and examples are only illustrative and are not necessarily intended to be limiting.
[0044] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description set forth below. It should be understood, however, that while the detailed description and specific examples indicate preferred embodiments of the invention, they are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description. [Brief description of the drawings]
[0045] [Figure 1A-1C] Graphs showing identification of compounds in ethanol extracts of Helichrysum: (1A) cannabigerolic acid (CBGA, 359.222 Da), (1B) helicannabigerolic acid (heliCBGA, 393.206 Da), and (1C) olivetolic acid (OA, 223.097 Da) extracted ion current (XIC) chromatograms and MS / MS spectra matching of standards or authentic compounds versus Helichrysum samples. [Figures 2A-2E] Table and chemical structure of CBGA and heliCBGA by 1D and 2D NMR. (Figure 2A) 1H and 13C chemical shift assignments, (Figure 2B and Figure 2C) atomic numbers and COSY correlations, (Figure 2D and Figure 2E) HMBC correlations of CBGA and heliCBGA, respectively. The carbon on the carboxyl of heliCBGA was not observed in the NMR spectrum, but the LC / HRMSMS spectrum and chemical formula confirm the presence of this group. [Figure 3A-3E]Table and chemical structure of Glc-OA and Glc-DHSA by 1D and 2D NMR. (Figure 3A) 1H and 13C chemical shift assignments, (Figure 3B-3C) atomic numbers and COSY correlations, (Figure 3D-3E) HMBC correlations of Glc-DHSA and Glc-DHSA, respectively. The carbon on the carboxyl was not observed in the NMR spectrum. However, the LC / HRMSMS spectrum and chemical formula confirm the presence of this group. [Figure 4A-4G]Graphs and chemical structure illustrations showing the identification of glucosylated intermediates, cannabinoids, and amorfurthins in ethanol extracts of Helichrysum. (Figure 4A) Comparison of MS / MS spectra in negative polarity for Glc-OA and Glc-DHSA versus OA and DHSA, respectively. As shown, the glucosylated compound showed a neutral loss of 162.053 Da corresponding to the loss of a hexose, and similar fragments as the non-glucosylated compound. The difference in relative abundance of fragment ions is likely due to the larger mass difference between the glucosylated and non-glucosylated compounds. Extracted ion current (XIC) chromatograms of m / z 357.119, 371.136, 385.150, and 399.166 (Figure 4B) as well as (Figure 4C) MS / MS spectra of unlabeled and isotopically labeled (Figure 4D) glucosylated alkyl intermediates in negative polarity. The marked peaks in each chromatogram correspond to the glucosylated intermediates detected. As shown, the alkyl congeners elute from the reversed-phase column in order of chain length as a result of increasing lipophilicity. For all alkyl congeners, appropriate m / z shifts were observed in the MS / MS spectra of all product ions containing the alkyl chain. Isomers with similar mass and MS / MS fragmentation were observed for Glc-OA and Glc-HA, which we assigned as branched short-chain FAs according to feeding experiments, consistent with the identified cannabinoids. (Figure 4E) MS / MS spectrum and proposed fragmentation structure of Glc-OA with labeling. The fragments colored in blue correspond to the m / z of specific fragments in the compound labeled with hexanoic acid D11. The fragments are similar to those observed in the non-glucosylated compound. (Figure 4F) Negative polarity MS / MS spectra of Glc-CBPA, Glc-CBGA, Glc-heliCBPA, and Glc-heliCBGA. Identification was achieved by MS / MS fragmentation and relative RT similar to the non-glycosylated compounds. (FIG. 4G) Structures of isotopically labeled precursors of observed and identified compounds. IP, isoprenyl; MP, monoprenyl. [Figure 5A-5G] Graphs and micrographs showing CBGA and Glc-OA content in plant tissues and CBGA localization to glandular trichomes of Helichrysum leaves and flowers. (Fig. 5A) Absolute concentration of CBGA (%w / w, n=4) in different Helichrysum plant tissues, and peak area of Glc-OA (by UPLC-qTOF peak area at m / z 385.15). Optical images of leaf and receptacle cross sections (Fig. 5B and Fig. 5D) and MALDI-MSI (Fig. 5C and Fig. 5E) of m / z 361.24 ± 0.01 Da (corresponding to the protonated mass of CBGA). Glandular trichomes in Fig. 5B and Fig. 5D are marked for better interpretation. CBGA localizes to stalked glandular trichomes. Optical images of (Fig. 5F) flower head and (Fig. 5G) excised receptacle for reference. Sessile glandular hairs (yellow arrows) and mechanical glandular trichomes (white arrows) can be observed at the surface border of the tissue in (5G). The white dashed lines in (Figure 5C) and (Figure 5E) indicate the analyzed area. Scale bars: 100 μm (5B); 500 μm (5C); 200 μm (5D); 1,000 μm (5E); 1,000 μm (5F); and 500 μm (5G). [Figure 6A-6C] Graphs showing expression profiling (UMI-aware 3'Trans-seq) of Helichrysum genes. (Figure 6A) PCA plot of normalized gene expression distribution of 21 sequenced samples. (Figure 6B) Over-representation analysis (ORA) of genes belonging to each co-expression module obtained with the CEMI tool. Module number M4 contains genes enriched in glandular trichomes and leaves (6C). Normalized expression of genes belonging to module number 4. [Figure 7] Graph showing expression profile (UMI-aware 3'Trans-seq) of selected UGT Helichrysum genes. CPM normalized expression of selected UGT genes with expression patterns correlated with CBGA accumulation. The secondary axis containing quantification of CBGA is to the right of the plot. [Figure 8A-8C]Chemical structures and graphs showing activity of lysates containing (Fig. 8A) OA, (Fig. 8B) CBGA, (Fig. 8C) heliCBGA as substrates and UDP-glc and HuUGT as glycosyl donors. Reactions show different substrate specificities and types of products produced. Peaks are from chromatogram annotation for HuUGT1. The most abundant product is indicated with an asterisk. EV, empty vector. [Figure 9] Spectra showing the functional properties of UGTs. Extracted ion chromatograms of monoglucosides observed according to the theoretical m / z values after enzymatic assays with purified enzymes (HuUGT1, HuUGT6, HuUGT11, HuUGT13, OsUGT and SrUGT) in the presence of UDP and either OA, DHSA, CBGA, heliCBGA, CBDA, Δ9-THCA, CBCA, olivetol, CBG, CBD or Δ9-THC. One to three glucosylated compounds were observed for each substrate according to the possible sites of glucosylation marked on each structure. Monoglucosides were identified according to MS / MS fragmentation and assigned by the fragmentation patterns (Figure 10). [Figure 10] MS / MS spectra of glucosylated compounds observed after in vitro assays with UGTs from Helichrysum, Stevia, and Rice. Assignment of peaks (1-3) was according to MS / MS fragmentation patterns and m / z difference between parent and fragment 1. Retention times of peaks 1 and 2 were constant, whereas peak 3 eluted at a different relative RT. XIC chromatograms of each substrate after reaction with SrUGT or HuUGT6 are shown as reference. [Figure 11] LC / MS chromatograms of observed diglucosides after enzymatic assay with purified enzyme in the presence of UDP-glc and cannabinoid acceptor. All LC / MS chromatograms were selected for the theoretical m / z of each compound of interest. [Figure 12A-12B]Curves and table showing a comparison of steady-state kinetic analysis of HuUGT11 and HuUGT13 with OsUGT and SrUGT using olivetolic acid and UDP-glc. (12A) Michaelis-Menten Km values for each enzyme were calculated using varying concentrations (0.5 μM-3 mM) and a constant concentration (1 mM) of olivetolic acid and GPP (n=3 technically independent samples, measurements plotted separately). V0 was calculated using the calibration curve for OA since there was no analytical standard available for Glc-OA. (Figure 12B) Summary of the results presented in Figure 12A. V0 and Vmax were calculated using the calibration curve for OA since there was no analytical standard available for Glc-OA. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0046] The present invention, in some embodiments, relates to a polynucleotide sequence encoding a protein or a plurality of proteins derived from Helichrysum ambraculigerum and belonging to the uridine diphosphate (UDP) glycosyltransferase (UGT) family.
[0047] According to some embodiments, there is provided a polynucleotide comprising a nucleic acid sequence comprising any one of SEQ ID NOs: 1-13, and any combination thereof.
[0048] In some embodiments, the polynucleotide is an isolated polynucleotide. In some embodiments, the polynucleotide is a DNA molecule. In some embodiments, the polynucleotide is an isolated DNA molecule. In some embodiments, the DNA molecule is an isolated DNA molecule. In some embodiments, the DNA molecule is a complementary DNA (cDNA) molecule.
[0049] As used herein, the terms "isolated polynucleotide" and "isolated DNA molecule" refer to a nucleic acid molecule that is essentially free of contaminating cellular components, such as carbohydrate, lipid, or other protein impurities associated with naturally occurring nucleic acids. Typically, an isolated DNA or RNA preparation contains nucleic acid in a highly purified form, e.g., at least about 80% pure, at least about 90% pure, at least about 95% pure, more than 95% pure, or more than 99% pure. In some embodiments, the isolated polynucleotide is any one of DNA, RNA, and cDNA. In some embodiments, the isolated polynucleotide is a synthetic polynucleotide. Polynucleotide synthesis is well known in the art and may be performed, for example, by ligating or covalently linking multiple nucleic acid molecules together with a primer linker.
[0050] The term "nucleic acid" is well known in the art. As used herein, "nucleic acid" generally refers to any molecule (e.g., a chain) of DNA, RNA or derivatives or analogs thereof that contains nucleotides. Nucleotides are composed of a nucleoside and a phosphate group. The nitrogenous bases of the nucleoside include the natural purine or pyrimidine nucleosides found, for example, in DNA (e.g., adenine "A", guanine "G", thymine "T", or cytosine "C") or RNA (e.g., A, G, uracil "U", or C).
[0051] The term "nucleic acid molecule" includes, but is not limited to, single-stranded RNA (ssRNA), double-stranded RNA (dsRNA), single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), small RNA, circular nucleic acids, fragments of genomic DNA or RNA, degraded nucleic acids, amplification products, modified nucleic acids, plasmids or organelle nucleic acids, and artificial nucleic acids such as oligonucleotides.
[0052]
[0053] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 77%, at least 79%, at least 85%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:1. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 77%-95%, 78%-100%, 79%-99%, or 77%-100% homology or identity to SEQ ID NO:1. Each possibility represents a separate embodiment of the present invention.
[0054]
[0055] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 76%, at least 77%, at least 85%, at least 93%, at least 97%, or at least 99%, or any value and range therebetween, homology or identity to SEQ ID NO:2. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 76%-95%, 77%-98%, 80%-99%, or 76%-100% homology or identity to SEQ ID NO:2. Each possibility represents a separate embodiment of the present invention.
[0056] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequence: (SEQ ID NO:3).
[0057] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 78%, at least 80%, at least 85%, at least 95%, or at least 99%, or any value and range therebetween, homology or identity to SEQ ID NO:3. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 79%-95%, 78%-100%, 80%-99%, or 79%-100% homology or identity to SEQ ID NO:3. Each possibility represents a separate embodiment of the present invention.
[0058]
[0059] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 87%, at least 92%, at least 97%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:4. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 87%-100%, 88%-99%, 89%-99%, or 87%-100% homology or identity to SEQ ID NO:4. Each possibility represents a separate embodiment of the present invention.
[0060]
[0061] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 87%, at least 92%, at least 97%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:5, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 87%-100%, 88%-99%, 89%-99%, or 87%-100% homology or identity to SEQ ID NO:5. Each possibility represents a separate embodiment of the present invention.
[0062]
[0063] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 80%, at least 87%, at least 93%, at least 97%, or at least 99%, or any value and range therebetween, homology or identity to SEQ ID NO:6. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 80%-98%, 81%-99%, 85%-99%, or 80%-100% homology or identity to SEQ ID NO:6. Each possibility represents a separate embodiment of the present invention.
[0064] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequence: (SEQ ID NO:7).
[0065] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 77, at least 85%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:7, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 77%-95%, 82%-97%, 81%-98%, or 77%-100% homology or identity to SEQ ID NO:7. Each possibility represents a separate embodiment of the present invention.
[0066]
[0067] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 82%, at least 85%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99%, or any value and range therebetween, homology or identity to SEQ ID NO:8. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 82%-95%, 83%-98%, 82%-99%, or 82%-100% homology or identity to SEQ ID NO:8. Each possibility represents a separate embodiment of the present invention.
[0068]
[0069] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 79, at least 85%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:9, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 79%-95%, 82%-97%, 81%-98%, or 79%-100% homology or identity to SEQ ID NO:9. Each possibility represents a separate embodiment of the present invention.
[0070]
[0071] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 78, at least 85%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO: 10, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 78%-95%, 82%-97%, 81%-98%, or 78%-100% homology or identity to SEQ ID NO: 10. Each possibility represents a separate embodiment of the present invention.
[0072]
[0073] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 82, at least 85%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:11, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 82%-95%, 82%-97%, 83%-98%, or 82%-100% homology or identity to SEQ ID NO:11. Each possibility represents a separate embodiment of the present invention.
[0074]
[0075] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 74, at least 80%, at least 85%, at least 87%, at least 93%, or at least 99%, or any value and range therebetween, homology or identity to SEQ ID NO: 12. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 74%-95%, 75%-97%, 76%-98%, or 74%-100% homology or identity to SEQ ID NO: 12. Each possibility represents a separate embodiment of the present invention.
[0076]
[0077] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 80, at least 85%, at least 95%, or at least 99%, or any value and range therebetween, homology or identity to SEQ ID NO: 13. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 80%-95%, 82%-97%, 81%-98%, or 80%-100% homology or identity to SEQ ID NO: 13. Each possibility represents a separate embodiment of the present invention.
[0078] In some embodiments, a polynucleotide of the invention comprises between 700 and 1,800 nucleotides. In some embodiments, a polynucleotide of the invention is between 730 and 1,730 nucleotides in length.
[0079] In some embodiments, 700-1,800 nucleotides includes at least 705 nucleotides, at least 750 nucleotides, at least 800 nucleotides, at least 900 nucleotides, at least 1,000 nucleotides, at least 1,150 nucleotides, at least 1,400 nucleotides, at least 1,600 nucleotides, at least 1,700 nucleotides, or at least 1,750 nucleotides, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, 700-1,800 nucleotides includes 710-1,750 nucleotides, 720-1,760 nucleotides, 730-1,780 nucleotides, or 740-1,700 nucleotides. Each possibility represents a separate embodiment of the present invention.
[0080] In some embodiments, the polynucleotide comprises a plurality of polynucleotides. In some embodiments, the polynucleotide comprises a plurality of types of polynucleotides. As used herein, the term "plurality" includes any integer greater than or equal to 2. In some embodiments, the polynucleotide comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, or 13 different nucleic acid sequences, or any value and range therebetween, where each of the different nucleic acid sequences is selected from SEQ ID NOs: 1-13. Each possibility represents a separate embodiment of the invention. In some embodiments, the polynucleotide comprises 2-13, 2-10, 2-8, 2-5, 3-7, 3-9, 3-12, 5-10, 5-12, or 3-13 different nucleic acid sequences, where each of the different nucleic acid sequences is selected from SEQ ID NOs: 1-13.
[0081] In some embodiments, the polynucleotide is or comprises a plurality of polynucleotide molecules, each of the plurality of polynucleotide molecules comprising a different nucleic acid sequence, wherein each of the different nucleic acid sequences is selected from SEQ ID NOs: 1-13.
[0082] In some embodiments, the polynucleotide encodes a protein characterized by a catalytic activity that transfers the glucuronic acid moiety of UDP-glucuronic acid to a small hydrophobic molecule (e.g., a UGT). In some embodiments, the polynucleotide encodes a protein characterized by a glycosyltransferase catalytic activity. In some embodiments, the polynucleotide encodes a protein characterized by being capable of transferring the glucuronic acid moiety of UDP-glucuronic acid to a cannabinoid or a precursor thereof. In some embodiments, the polynucleotide encodes a protein characterized by having a catalytic activity that glycosylates a cannabinoid or a precursor thereof. In some embodiments, the polynucleotide encodes a UGT enzyme.
[0083] In some embodiments, the UGT is a UGT from Helichrysum ambulatory. As used herein, the term "UGT" encompasses any enzyme derived from H. ambulatory and having or characterized as having an activity as described herein.
[0084] According to some embodiments, there is provided an artificial nucleic acid molecule comprising a polynucleotide disclosed herein.
[0085] In some embodiments, the artificial vector comprises a plasmid. In some embodiments, the artificial vector comprises an Agrobacterium comprising the artificial nucleic acid molecule or is an Agrobacterium comprising the artificial nucleic acid molecule. In some embodiments, the artificial vector is an expression vector. In some embodiments, the artificial vector is a plant expression vector. In some embodiments, the artificial vector is for use in expressing a nucleic acid sequence encoding a UGT disclosed herein. In some embodiments, the artificial vector is for use in heterologous expression of a nucleic acid sequence encoding a UGT disclosed herein in a cell, tissue, or organism.
[0086] It is well known to those skilled in the art to express polynucleotides in cells. It can be carried out by transfection, viral infection, or direct modification of the genome of cells, among many other methods. In some embodiments, the polynucleotide is in an expression vector, such as a plasmid or a viral vector. The vector nucleic acid sequence generally comprises at least a replication origin for propagation in cells, and optionally additional elements, such as heterologous polynucleotide sequences, expression control elements (e.g., promoters, enhancers), selection markers (e.g., antibiotic resistance), polyadenine sequences, etc.
[0087] The vector can be a DNA plasmid delivered via non-viral or viral methods. The viral vector can be a retroviral vector, a herpes virus vector, an adenovirus vector, an adeno-associated virus vector, a virgaviridae virus vector or a pox virus vector. Barley stripe mosaic virus (BSMV), tobacco rattle virus, cabbage leaf curl geminivirus (CbLCV) can also be used. The promoter can be active in plant cells. The promoter can be a viral promoter.
[0088] In some embodiments, the polynucleotide disclosed herein is operably linked to a promoter. The term "operably linked" is intended to mean that the nucleotide sequence of interest is linked to a regulatory element(s) in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). In some embodiments, a promoter is operably linked to the polynucleotide of the present invention. In some embodiments, the promoter is a heterologous promoter. In some embodiments, the promoter is an endogenous promoter.
[0089] In some embodiments, the vector is introduced into the cell by standard methods including electroporation (e.g., as described in From et al., Proc. Natl. Acad. Sci. USA 82, 5824 (1985)), heat shock, infection with viral vectors, high velocity ballistic penetration by small particles carrying nucleic acid either within the matrix or on the surface of small beads or particles (Klein et al., Nature 327.70-73 (1987)), biolistec use of coated particles, and needle-shaped particles, Agrobacterium Ti plasmids, and / or the like.
[0096] As used herein, the term "promoter" refers to a group of transcriptional control modules that are integrated around the initiation site of RNA polymerase, i.e., RNA polymerase II. Promoters are composed of separate functional modules, each consisting of about 7-20 bp of DNA and containing one or more recognition sites for transcriptional activator or repressor proteins. Promoters can extend upstream or downstream from the transcription start site and can be any size, ranging from a few base pairs to several kilobases.
[0090] In some embodiments, the polynucleotides are transcribed by RNA polymerase II (RNAP II and Pol II), an enzyme found in eukaryotic cells that is known to catalyze the transcription of DNA to synthesize precursors of mRNA and most snRNAs and microRNAs.
[0091] In some embodiments, a plant expression vector is used. In one embodiment, the expression of the polypeptide coding sequence is driven by multiple promoters. In some embodiments, viral promoters are used, such as the 35S RNA promoter and the 19S RNA promoter of CaMV [Brisson et al., Nature 310:511-514 (1984)], or the coat protein promoter for TMV [Takamatsu et al., EMBOJ.6:307-311 (1987)]. In another embodiment, a plant promoter is used, such as the small subunit of RUBISCO [Coruzzi et al., EMBOJ.3:1671-1680 (1984); and Brogli et al., Science 224:838-843 (1984)] or a heat shock promoter, such as soybean hspl7.5-E or hspl7.3-B [Gurley et al., Mol. Cell. Biol.6:559-565 (1986)]. In one embodiment, constructs are introduced into plant cells using Ti plasmid, Ri plasmid, plant virus vector, direct DNA transformation, microinjection, electroporation and other techniques well known to those skilled in the art.See, for example, Weissbach&Weissbach [Methods for Plant Molecular Biology, Academic Press, NY, Section VIII, pp421-463(1988)].Other expression systems, such as insect and mammalian host cell systems, well known in the art, can also be used in the present invention.
[0092] In some embodiments, expression vectors containing regulatory elements from eukaryotic viruses, such as retroviruses, are used according to the invention. SV40 vectors include pSVT7 and pMT2. In some embodiments, vectors derived from bovine papilloma virus include pBV-1MTHA, and vectors derived from Epstein-Barr virus include pHEBO and p205. Other exemplary vectors include pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo-5, baculovirus pDSVE, and any other vector that allows expression of proteins under the direction of the SV-40 early promoter, SV-40 late promoter, metallothionein promoter, mouse mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown to be effective for expression in eukaryotic cells.
[0093] In some embodiments, recombinant viral vectors that provide advantages such as systemic infection and target specificity are used for in vivo expression. In one embodiment, systemic infection is a process that is inherent in, for example, the life cycle of retroviruses, whereby a single infected cell produces many progeny virions that infect neighboring cells. In one embodiment, this results in a large area becoming rapidly infected, most of which were not initially infected by the original viral particle. In one embodiment, a viral vector that cannot spread systemically is produced. In one embodiment, this feature can be useful when the desired purpose is to introduce a specific gene into only a local number of target cells.
[0094] In some embodiments, a plant viral vector is used. In some embodiments, a wild type virus is used. In some embodiments, a deconstructed virus, such as those known in the art, is used. In some embodiments, Agrobacterium is used to introduce the vectors of the invention into the virus.
[0095] A variety of methods can be used to introduce the expression vector of the present invention into cells. Such methods are generally described in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Springs Harbor Laboratory, New York (1989, 1992); Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989); Chang et al., Somatic Gene Therapy, CRC Press, Ann Arbor, Mich. (1995); Vega et al., Gene Targeting, CRC Press, Ann Arbor Mich. (1995); Vectors: A Survey of Molecular Cloning Vectors and Their Uses, Butterworths, Boston Mass. (1988) and Gilboa et al. [Biotechniques 4(6):504-512, 1986], including, for example, stable or transient transfection, lipofection, electroporation, infection with Agrobacterium Ti plasmids and recombinant viral vectors. Additionally, see U.S. Patent Nos. 5,464,764 and 5,487,992 for positive-negative selection methods.
[0096] It will be understood that, in addition to containing the necessary elements for the transcription and translation of the inserted coding sequence (encoding a polypeptide), the expression constructs of the present invention may also contain sequences engineered to optimize the stability, production, purification, yield or activity of the expressed polypeptide.
[0097] In some embodiments, the artificial vector comprises a polynucleotide encoding a protein comprising an amino acid sequence described herein.
[0098] According to some embodiments, there is provided a protein encoded by (a) a polynucleotide disclosed herein; (b) an artificial vector disclosed herein; or a plasmid or agrobacterium disclosed herein.
[0099] In some embodiments, the protein is encoded by a polynucleotide comprising or consisting of SEQ ID NOs: 1-13.
[0100] In some embodiments, the protein comprises an amino acid sequence having at least 90%, at least 92%, at least 93%, at least 95%, at least 97%, or at least 99%, or any value and range therebetween, homology or identity to any one of SEQ ID NOs: 14-26. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 90-100%, 93-100%, 95-100%, or 97-100% homology or identity to any one of SEQ ID NOs: 14-26. Each possibility represents a separate embodiment of the present invention.
[0101] In some embodiments, the protein is an isolated protein.
[0102] The terms "peptide", "polypeptide" and "protein" as used herein are interchangeable and refer to a polymer of amino acid residues. In another embodiment, the terms "peptide", "polypeptide" and "protein" as used herein encompass natural peptides, peptidomimetics (typically containing non-peptide bonds or other synthetic modifications), and peptide analogs peptoids and semipeptoids, or any combination thereof. In another embodiment, the described peptides, polypeptides, and proteins have modifications that make them more stable while in an organism or more likely to enter a cell. In one embodiment, the terms "peptide", "polypeptide" and "protein" apply to naturally occurring amino acid polymers. In another embodiment, the terms "peptide", "polypeptide" and "protein" apply to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acid.
[0103] As used herein, the term "isolated protein" refers to a protein that is essentially free of contaminating cellular components, such as carbohydrates, lipids, or other proteinaceous impurities associated with natural nucleic acids. Typically, an isolated protein preparation comprises a highly purified form of the protein, e.g., at least about 80% pure, at least about 90% pure, at least about 95% pure, more than 95% pure, or more than 99% pure. In some embodiments, the isolated protein is a synthesized protein. Protein synthesis is well known in the art and can be carried out, for example, by heterologous expression in transformed cells, such as those exemplified herein.
[0104] In some embodiments, the protein comprises or consists of the following amino acid sequence: MTNSELVFIPSPGAGHLPPTVELAKLLLHREPQLSVTIIIMNLPHETKPTTETRMSTPRLRFIDIPKDESTKDLISRHTFISAFLEHQKPHVRNIVRSITESDSVRLVGFVVDMFCIAMMDVANELGAPTYLYFTSSAASLGLMFCLQAKRDDEEFDVTELKDKDSELSIPCYTNPLPAKLLPSVLFDKRGGSKTFIDLARKYRESRGIVVN TFQELESYAIEYLASSNANVPPVFPVGAILNQEKKVNDDKTEEIMTWLNEQPESSVVFLCFGSMGSFGEDQIKEIALAIEESGQRFLWSLRRPPSNENKYPKEYENFGEVLPEGFLERTSVGKVIGWAPQMAVLSHSSVGGFVSHCGWNSTLESIWCGVPVAAWPLYAEQQLNAFKLVVELGLAVEIKIDYRSENEIILTSKEIESGIRRLMNDEELRMKVKEMKGNSRFAVSEGGSSYVSIRRFIDLVMTKE (SEQ ID NO: 14).
[0105] In some embodiments, the protein comprises an amino acid sequence having at least 75%, at least 85%, at least 92%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO: 14. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 75%-99%, 76%-98%, or 75%-100% homology or identity to SEQ ID NO: 14. Each possibility represents a separate embodiment of the present invention.
[0106] In some embodiments, the protein comprises or consists of the following amino acid sequence: MPTSELVFIPSPGVGHLSPTIELVNQLLHRDQRLSVTIIVMKFSLESKHDTETPTSTPRLRFIDIPYDESAMALINPNTFLSAFVEHNKPHVRNIVRDISESNSVRLAGFVVDMFCVAMTDVVNEFEIPTYIYFTSTANLLGLMFYLQAKRDDEGFDVTVLKDSESEFLSVPSYVNPVPAKVLPDAVLDKNGGSQMCLDLAKGFRESKGIIVNTFQE LERRGIEHLLSSNMNLPPVFPVGPILNLRNAPNDGKTADIMTWLNDHPENSVVFLCFGSMGSFEKEQVKEIAIAIEQSGQRFLWSLRRPTSLEKFEFPKDYENPEEVLPKGFLERTKGVGKVIGWAPQMAVLSHPSVGGFVSHCGWNSTLESIWCGVPIAAWPLYAEQKINAFQLVVEMGMAAEIRIDYRTNTRPGGGKEMMVMAEEIESGIRKLMSDDEMRKKVKGMKDKSRAAVLEGGSSHTSIGILIENLVSITI (SEQ ID NO: 15).
[0107] In some embodiments, the protein comprises an amino acid sequence having at least 76%, at least 85%, at least 92%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO: 15, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 76%-99%, 80%-98%, or 76%-100% homology or identity to SEQ ID NO: 15. Each possibility represents a separate embodiment of the present invention.
[0108] In some embodiments, the protein comprises or consists of the following amino acid sequence:MVGLKCFWILQKGFRESKGIIVNTFQELERRGIEHLLSSNMDLPPVFPVGPILNLRNARNDGKMADIMTWLNDQPENSVVFLCFGSRGSFKEEQVKEIAIAIEQSGQRFLWSLRRPTSIETFFEFPKYYENPEEVLPKGFLERTKSVGKVIGWAPQMAVLSHPSVGGFVSHCGWNSTLESIWCGVPIAAWPLYAEQQTNAFQLVVEMGMAAEIRIDYRTNTPLVGGKDMMVTAEEIERGIRKLMSDDEMRKKVKDMKDKSRGAVLEGGSSHTSIGNLIDVLVSITI (SEQ ID NO: 16).
[0109] In some embodiments, the protein comprises an amino acid sequence having at least 77%, at least 85%, at least 92%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO: 16, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 77%-99%, 79%-98%, or 77%-100% homology or identity to SEQ ID NO: 16. Each possibility represents a separate embodiment of the present invention.
[0110] In some embodiments, the protein comprises or consists of the following amino acid sequence: MATNNLHFLLIPHIGPGHTIPMIDMAKLLAKQPNVMVTIATTPLNITRYGHTLADAINSFRFFEVPFPAVEAGLPEGCESTDKIPSMDLVPNFLTAIGMLEQKLEEHFHLLEPRPNCIISDKYMSWTGDFADKYRIPRIMFDGMSCFNELCYNNLYENKVFEGMHETEPFVVPGLPDKIELTRKQLPPEFNPSSIDTSEFRQRARDAEVRAYGVVINSFEELE QEYVNEYKKLRKGKVWCIGPLSLCNSDNSDKAQRGNIASVDEEKCLKWLDSHEADSVVYACFGSLVRVNTPQLIELGLGLEASNRPFIWVVRSVHREKEVEEWLVESGFEERIKDRGLIIRGWAPQVLILSHPSIGGFLTHCGWNSTLESVCAGVPMITWPQFAEQFINEKLIVQVLGIGVGVGVDSVVHVGEEDRSGVKVKRESVTKAIEKVMDDEIDGNERRRRSKEFGKIANNAIKEGGSSYLNLTLLIQDIMRYANADASS (SEQ ID NO: 17).
[0111] In some embodiments, the protein comprises an amino acid sequence having at least 88%, at least 92%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO: 17, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 88%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO: 17. Each possibility represents a separate embodiment of the present invention.
[0112] In some embodiments, the protein comprises or consists of the following amino acid sequence: MEKTPHIAIVPSPGMGHLIPLVEFAKKLKNHHNIHATFIIPNDGPLSISQKVFLDSLPNGLNYLILPPVNFDDLPQDTQIETRISLMVTRSLDSLREVFKSLVVEKNMVALFIDLFGTDAFDVAIEFGVSPYVFFPSTAMALSLFLYLPKLDQMVSCEYRELPEPVQIPGCIPVRGQDLVDPVQDRKNDAYKWVLHNAKKYSMAKGIAVNSFK ELEGGALNALLEDEPGKPKVYPVGPLVQTGFSCDVDSIECLKWLDGQPCGSVLYISFGSGGTLSSSQLNELAMGLELSEQRFIWVVRSPNDQPNATYFDSHGHKDPLGFLPKGFLERTKGIGFVIPSWAPQAQILSHSATGGFLTHCGWNSILETVVHGVPVIAWPLYAEQKMNAVSLTEGIKMALRPTVGENGIVGRLEVARVVKSLLEGEEGKAIRSRVRDLKDAAANVLSKDGSSTKTLDQLAVQLKKQELS (SEQ ID NO: 18).
[0113] In some embodiments, the protein comprises an amino acid sequence having at least 90%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO: 18. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 90%-100%, 93%-100%, 95%-100%, or 97%-100% homology or identity to SEQ ID NO: 18. Each possibility represents a separate embodiment of the present invention.
[0114] In some embodiments, the protein comprises or consists of the following amino acid sequence: MTQKQMQMQPHFLLVTYPAQGHINPSLQFAERLIRLGVKVTFTTTVSAYRRMSKAGNISEFLNFAAFSDGFDDGFNFETDDHGLFLTQLRSRGKDSLKETILSNAKNGTPISCLVYTLLLPWAPEVARGLNVPSAFLWIQPASVLRLYYYYFNGYNELIGDDCNEPSWSIQLPGLPLLKS (SEQ ID NO: 19).
[0115] In some embodiments, the protein comprises an amino acid sequence having at least 77%, at least 85%, at least 92%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO: 19. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 77%-100%, 79%-100%, 80%-100%, or 90%-100% homology or identity to SEQ ID NO: 19. Each possibility represents a separate embodiment of the present invention.
[0116] In some embodiments, the protein comprises or consists of the following amino acid sequence: MTKIQQQPHFLLVTYPAQGHINPSLRFAERLIRLGVKVTFTITVSAYRRMSKAGHISEFLNFAVFSDGFDDGFNSKTDDYGLFLTQFRSRGKDSLKETILSNAKNGTPVSCLVYTLLLPWAPEVARGLNVPSAFLWIQPASVLRLYYYYFNGYNELIGDDCNEPSWSIQLPGLPLLKSRDLPSFCLPSNPYADVLTLVKEHLDVLDLEEKPKILVNSFDELEREALNEIDGKLKMVAVGPLIPSAFFGWTGCI (SEQ ID NO: 20).
[0117] In some embodiments, the protein comprises an amino acid sequence having at least 73%, at least 85%, at least 92%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:20. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 73%-100%, 77%-100%, 85%-100%, or 90%-100% homology or identity to SEQ ID NO:20. Each possibility represents a separate embodiment of the present invention.
[0118] In some embodiments, the protein comprises or consists of the following amino acid sequence: MGSWRNSRTTSTKFLWLILPLMVVTVIIGVKKSNYGSKYNYPWVWSSVINSYSSSAVKEDVTVVAEGPVESFGLRSTVVNGGGVVAEGPSEDFGFNSSYPPLAMEDEMDVELPAIAKEDDLNATLSGPDLFVSANQTGGLHVDIGINSKYTSLDKLEARLGQVRAAIKEAESGNRTYDPDYVPEGPMYWHAASFHRSYLEMEKQFKVFVYEEGEPPIFHNGPCKNIYAMEGNFIYHMETTKFRTKNPEKAHTFFLPMSAAMMVRFIFERDPNVDHWRPMKQTIKDYVDLVGGKYPFWNRSLGADHFTVACHDWVSKVFYPIIFMLLLVFIFRMSTGC ((SEQ ID NO: 21).
[0119] In some embodiments, the protein comprises an amino acid sequence having at least 81%, at least 90%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:21, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 81%-100%, 85%-100%, 87%-100%, or 91%-100% homology or identity to SEQ ID NO:21. Each possibility represents a separate embodiment of the present invention.
[0120] In some embodiments, the protein comprises or consists of the following amino acid sequence: MSTVEVAKLLVNRDHRLFITFLIIQPPSSGSGSAITTYIESLAEKAMDRISFIELPQDKIPPPRYPKSLPTAESKAHPLIFMIEFIKCHCKYVRNIVSDMISQPSSGRVAGLVIDMLCFSMMDVANEFNIPTYVFVTSNAAFLGFYLYVQILSNDQNQDVVELSKSDTEISVPGFVKPVPTKVFWTVVRTKEGLDFVLSSAQKLRQAKAIMVNT FLELETHAIKSLSDDTSIPPVYPVGPILNLEGGAGKTFDNDISRWLDSQPPSSVVFLCFGSHGCFDEIQVKEIAHALEQSGHRFLWSLRRPPSDQTLKVPGDYEDPGVVLPEGFLERTAGRGKVIGWAPQVMVLAHRAVGGFVSHCGWNSLLESLWFGVPTATWPIYAEQQMNAFEMVVELGLAVEITLDYRNDMDMFIVTAQEIESGIRKVMEDNEVRTKVKERSEKSRAAVAEGGSSYASVGHLIKEFTGNIS (sequence number 22).
[0121] In some embodiments, the protein comprises an amino acid sequence having at least 74%, at least 85%, at least 92%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:22. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 74%-100%, 79%-100%, 85%-100%, or 90%-100% homology or identity to SEQ ID NO:22. Each possibility represents a separate embodiment of the present invention.
[0122] In some embodiments, the protein comprises or consists of the following amino acid sequence: MSSFINFVESTTQLQPQFEQLIQTLLPITAIISDGFLMWTQDSAEKFNIPRLVFYGTNIFFMTMCNIMAQFKPHAAVNSDDEAFDVPGFTRFKLTANDFEPPFNEVEPKGSMLDFLLEQQKAMVRSHGLVVNSFYEIEHEFNVYWNQNYGPKAWLMGPFCVAKPYASNVMDSEI STKVVKKSAWIQWLDRKLAANEPVLYISFGTQAEASMEHLHEVAIGLERSNVSFIWVVKAKQMQLIGAGFEERVKGRGKVVTEWVDQMEILKHEIVSGFLSHCGWNSLLESMCVGVPVLAMPLMADQLLNARLVVEEIGMGLRLWPRGMVARGIVGAEEVEKMVVELMEGEGGRRVRKRVIEVREMAYGAMKEGGSSSRTLDSLIDHVCEAFHKTV (sequence number 23).
[0123] In some embodiments, the protein comprises an amino acid sequence having at least 76%, at least 85%, at least 92%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:23. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 76%-100%, 80%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO:23. Each possibility represents a separate embodiment of the present invention.
[0124] In some embodiments, the protein comprises or consists of the following amino acid sequence:MGSLKKGAHILIFPFPAQGHMLPLLDLTHHLATNGLTITILVTPKNLPILNPLLSSSPNIQPLVFPFPPHPRLPPHVENVKDIGNHANVPITNSLAKLQDQIIQWFNSHHNPPVAIISDFFLGWTQHLANKLGIPRVGFFSSGAYLTAVLDYVCHNIKTVRSQEETVFHDLPNSPCFKFEHLPGLAQIYKESDPEWELVLDGHIANGLS WGWIVNTFDGLESRYMEYLTKKMGVGRVFGVGPVNLLNGSDPMTRGKSESGSDSGVLNWLDGKPDGSVLYVCFGSQKFLTNDQMEGLSIGLEQSGVHYVWVVKDEQGDAIRSGSGRGLVVTGWAPQVSILGHGAVGGFLSHCGWNSVLEAIVNGVMILAWPMEADQFVNAKLLVDDHGIGVWVCEGPNTVPDSTELARKIGESMSTDKSEKVKAKEMKNKANEAVKEGGSSSMELSRLVKELSNFETNGP (SEQ ID NO: 24).
[0125] In some embodiments, the protein comprises an amino acid sequence having at least 81%, at least 92%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:24. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 81%-100%, 85%-100%, 90%-100%, or 93%-100% homology or identity to SEQ ID NO:24. Each possibility represents a separate embodiment of the present invention.
[0126] In some embodiments, the protein comprises or consists of the following amino acid sequence: (SEQ ID NO:25).
[0127] In some embodiments, the protein comprises an amino acid sequence having at least 71%, at least 85%, at least 92%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:25, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 71%-100%, 77%-100%, 85%-100%, or 90%-100% homology or identity to SEQ ID NO:25. Each possibility represents a separate embodiment of the present invention.
[0128] In some embodiments, the protein comprises or consists of the following amino acid sequence: MSLVTNNPHLLVYPLPTSGHIIPLLDLTDLLLRRGLTITVVISTTDLTLLDTLLSSHPTSLHKLYFPDPEIGPSSHPVIARIIATQKLFDPIVKWFESHPSPPVAIISDFFLGWTNELASRLGIRRVVFSPSGALGHSILQSLWRDVAEINAKNVDGNGNYSISFTDIPNSPEFHWWQLSQLLRVHREGDPDFEFFRNGMLANTKSWG IVYNTFERIEKVYIDHVKKQIGHDRVWAIGPLLPEEHGPVGSTARGGSSVVPPHDLLTWLDKKPHDSVVYICFGSRLTLSEKQMSALASALELSNVDFILCVKASGSSFIPSGFEDRVVGRGFVIKGWA PQLAILRHRAVGSFVTHCGWNSTLEGVSSGVMMLTWPMGADQYANAKLLVDQLGVGKRVCEGGPESVPDSTELARLLEESLSGDTSERVKVKELSREANTAVKEGTSIRDLNMFVNLLSEL (SEQ ID NO: 26).
[0129] In some embodiments, the protein comprises an amino acid sequence having at least 78%, at least 85%, at least 92%, at least 95%, or at least 99%, or any value and range of homology or identity to SEQ ID NO:26. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 78%-100%, 85%-100%, 90%-100%, or 93%-100% homology or identity to SEQ ID NO:26. Each possibility represents a separate embodiment of the present invention.
[0130] In some embodiments, the protein comprises an amino acid sequence represented by SEQ ID NO:14-15, 17-20, 24, or 26.
[0131] In some embodiments, the protein comprises an amino acid sequence represented by SEQ ID NO:14, 19, 24, or 26.
[0132] The terms "homology" or "identity", used interchangeably herein, refer to the sequence identity between two amino acid sequences or two nucleic acid sequences, with identity being the more strict comparison. The phrases "percent identity or homology" and "% identity or homology" refer to the percentage of sequence identity found in the comparison of two or more amino acid sequences or nucleic acid sequences. The two or more sequences can be 0-100% identical, or any value in between. Identity can be determined by comparing positions in each sequence that can be aligned for purposes of comparison to a reference sequence. If a position in the compared sequences is occupied by the same nucleotide base or amino acid, the molecules are identical at that position. The degree of identity of amino acid sequences is a function of the number of identical amino acids at positions shared by the amino acid sequences. The degree of identity between nucleic acid sequences is a function of the number of identical or matching nucleotides at positions shared by the nucleic acid sequences. The degree of homology of amino acid sequences is a function of the number of amino acids at positions shared by the polypeptide sequences.
[0133] The following is a non-limiting example for calculating the homology or sequence identity (these terms are used interchangeably herein) between two sequences. The sequences are aligned to perform optimal comparison (e.g., gaps may be introduced into one or both of the first and second amino acid or nucleic acid sequences for optimal alignment, and non-homologous sequences may be ignored for comparison purposes). Optimal alignment is determined as the highest score using the GAP program of the GCG software package, using a Blossum62 scoring matrix with a gap penalty of 12, a gap extension penalty of 4, and a frameshift gap penalty of 5. The amino acid residues or nucleotides at corresponding amino acid or nucleotide positions are then compared. If a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, the molecules are identical at that position. The percent identity between two sequences is a function of the number of identical positions shared by the sequences.
[0134] In some embodiments, the % homology or % identity described herein is calculated or determined using BLAST (basic local alignment search tool). In some embodiments, the % homology or % identity described herein is calculated or determined using the Blossum62 scoring matrix.
[0135] According to another embodiment, compounds and / or salts thereof, and / or decarboxylated derivatives thereof are provided, wherein the compounds comprise glycosylated cannabinoids and / or glycosylated cannabinoid precursors (such as glycosylated OA). In some embodiments, the compounds of the invention are isolated compounds. In some embodiments, the compounds of the invention are natural or synthetic compounds. In some embodiments, the compounds of the invention are a single compound or multiple chemically distinct compounds.
[0136] As used herein, the term "isolated compound" refers to a compound that is essentially free of contaminating cellular components, such as carbohydrate, lipid, or other proteinaceous impurities naturally associated with nucleic acids. Typically, preparations of isolated compounds include compounds in highly purified form, e.g., at least about 80% pure, at least about 90% pure, at least about 95% pure, greater than 95% pure, or greater than 99% pure.
[0137] In some embodiments, the compounds of the invention are chemically pure (e.g., substantially devoid of one or more impurities, including any organic compounds). In some embodiments, the compounds of the invention are characterized by a chemical purity of at least 70%, at least 80%, at least 90%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% (including any ranges therebetween). In some embodiments, the compounds of the invention are characterized by a chemical purity of up to 99.99%, up to 99.9%, up to 99%, up to 95%, up to 90% (including any ranges therebetween).
[0138] In some embodiments, the glycosylated cannabinoid of the present invention comprises: [ka] and any salts and / or decarboxylated derivatives thereof, each R is independently H or a sugar moiety; R1 is, [ka] and at least one R is a sugar moiety. In some embodiments, the sugar moiety is or includes a deoxy monosaccharide or deoxy disaccharide. In some embodiments, the term "deoxy" refers to a monosaccharide or disaccharide having a bond in place of one of the hydroxy groups (i.e., a monosaccharide or disaccharide lacking one of its hydroxy groups). In some embodiments, the sugar moiety is or includes a deoxyhexose. In some embodiments, the sugar moiety is or includes a deoxyglucose (e.g., 2-deoxy-D-glucose, and / or enantiomers thereof).
[0139] In some embodiments, the glycosylated cannabinoid of the present invention is or comprises any one of the following, including any salts and / or decarboxylated derivatives thereof, wherein R is as described herein: [ka] In some embodiments, R is [ka] It is.
[0140] According to some embodiments, there is provided a transgenic cell comprising: (a) a polynucleotide disclosed herein; (b) an artificial nucleic acid molecule disclosed herein; (c) a plasmid or an Agrobacterium disclosed herein; (d) an isolated protein disclosed herein, or any combination thereof.
[0141] The term "transgenic cell" as used herein refers to any cell that has undergone human manipulation at the genome or gene level. In some embodiments, a transgenic cell has an exogenous polynucleotide introduced therein, such as an isolated DNA molecule disclosed herein. In some embodiments, a transgenic cell includes a cell into which an artificial vector has been introduced. In some embodiments, a transgenic cell is a cell that has undergone a mutation or modification of its genome. In some embodiments, a transgenic cell is a cell that has undergone CRISPR genome editing. In some embodiments, a transgenic cell is a cell that has undergone a targeted mutation of at least one base pair of its genome. In some embodiments, an exogenous polynucleotide (e.g., an isolated DNA molecule disclosed herein) or vector is stably integrated into the cell. In some embodiments, a transgenic cell expresses a polynucleotide of the invention. In some embodiments, a transgenic cell expresses a vector of the invention. In some embodiments, a transgenic cell expresses a protein of the invention. In some embodiments, a transgenic cell is a cell that lacks a polynucleotide of the invention and has been transformed or genetically modified to include a polynucleotide of the invention. In some embodiments, CRISPR technology is used to modify the genome of a cell, as described herein.
[0142] In some embodiments, the cells include cells of a single-celled organism, a cell of a multicellular organism, or a cell in culture.
[0143] In some embodiments, the unicellular organism comprises a fungus or a bacterium.
[0144] In some embodiments, the fungus is a yeast cell.
[0145] In some embodiments, the cell is an arthropod cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell comprises an insect cell line.
[0146] The types of insect cell lines suitable for transformation and / or heterologous expression are common and will be apparent to one of skill in the art. Non-limiting examples of such insect cell lines include, but are not limited to, Sf-9 cells, SR+Schneider cells, S2 cells, etc.
[0147] According to some embodiments, there is provided an extract derived from the transgenic cells disclosed herein, or any fraction thereof.
[0148] In some embodiments, the extract comprises a polynucleotide of the present invention, an isolated DNA molecule disclosed herein, an isolated protein disclosed herein, or any combination thereof.
[0149] According to some embodiments, there is provided a homogenate, a lysate, an extract, any combination thereof, or any fraction thereof derived from the transgenic cells disclosed herein.
[0150] Methods and / or means for extracting, lysing, homogenizing, fractionating, or any combination thereof of cells or cultures thereof are common and will be apparent to those skilled in the art of cell biology and biochemistry. Non-limiting examples include, but are not limited to, pressure lysis (e.g., using a French press), enzymatic lysis, soluble-insoluble phase separation (e.g., to obtain a supernatant and a pellet), detergent-based lysis, solvents (e.g., polar or non-polar solvents), liquid chromatography mass spectrometry, and the like.
[0151] According to some embodiments, a transgenic plant, a transgenic plant tissue or a plant part is provided. In some embodiments, a transgenic plant, or any part, seed, tissue or organ thereof, is provided, which comprises at least one transgenic plant cell of the present invention. In some embodiments, the transgenic plant, the transgenic plant tissue or a plant part comprises (a) the polynucleotide disclosed herein, (b) the artifact disclosed herein, (c) the plasmid or Agrobacterium disclosed herein, (d) the isolated protein of the present invention, (e) the transgenic cell disclosed herein, or any combination thereof.
[0152] In some embodiments, the transgenic plant, transgenic plant tissue, or plant part comprises transgenic plant cells of the invention. In some embodiments, the transgenic plant, transgenic plant tissue, or plant part comprises at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99%, or any value and range therebetween, transgenic cells of the invention. Each possibility represents a separate embodiment of the invention. In some embodiments, the transgenic plant, transgenic plant tissue, or plant part comprises 20%-50%, 20%-60%, 20%-70%, 20%-80%, 20%-90%, or 20%-100% transgenic cells of the invention. Each possibility represents a separate embodiment of the invention.
[0153] In some embodiments, the transgenic plant, transgenic plant tissue, or plant part is a Cannabis sativa plant or is derived from a Cannabis sativa plant. In some embodiments, the transgenic plant is a C. sativa plant.
[0154] In some embodiments, the transgenic plant, transgenic plant tissue, or plant part is hemp or is derived from hemp. In some embodiments, the C. sativa includes or is hemp.
[0155] According to some embodiments, there is provided a composition comprising any one of (a) the polynucleotides (e.g., isolated DNA molecules) of the invention, (b) artificial vectors, (c) plasmids or Agrobacteria, (d) isolated proteins of the invention, (e) transgenic cells, (f) extracts, (g) transgenic plant tissues or plant parts, and (h) any combination of (a)-(g) disclosed herein, and an acceptable carrier.
[0156] The terms "carrier", "excipient", or "adjuvant" as used herein refer to any component of a composition that is not an active agent, such as a pharmaceutical or nutraceutical composition. As used herein, the term "pharmaceutical acceptable carrier" refers to a non-toxic, inert solid, semi-solid liquid filler, diluent, encapsulating material, formulation auxiliary of any kind, or simply a sterile aqueous medium such as saline. Some examples of materials which may function as pharma- ceutically acceptable carriers include sugars such as lactose, glucose and sucrose, starches such as corn starch and potato starch, cellulose and its derivatives such as sodium carboxymethylcellulose, ethylcellulose and cellulose acetate; powdered tragacanth; malt, gelatin, talc; excipients such as cocoa butter and suppository wax; oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; glycols such as propylene glycol, polyols such as glycerin, sorbitol, mannitol, polyethylene glycol; esters such as ethyl oleate, ethyl laurate; agar; buffers such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline, Ringer's solution; ethyl alcohol and phosphate buffers, as well as other non-toxic compatible substances used in pharmaceutical formulations. Some non-limiting examples of materials that can function as carriers herein include sugar, starch, cellulose and its derivatives, powdered tragacanth, malt, gelatin, talc, stearic acid, magnesium stearate, calcium sulfate, vegetable oils, polyols, alginic acid, pyrogen-free water, isotonic saline, phosphate buffer, cocoa butter (suppository base), emulsifiers (e.g., carbomer, hydroxypropylcellulose, sodium lauryl sulfate) and other non-toxic pharmaceutically compatible materials used in pharmaceutical formulations. Wetting agents and lubricants such as sodium lauryl sulfate, as well as colorants, flavorings, excipients, stabilizers, antioxidants, and preservatives may also be present. Any non-toxic, inert, and effective carrier may be used to formulate the compositions contemplated herein.Suitable pharma- ceutically acceptable carriers, excipients and diluents in this regard are well known to those skilled in the art, and can be found, for example, in The Merck Index, Thirteenth Edition, Budavari et al., Eds., Merck&Co., Inc., Rahway, NJ (2001); the CTFA (Cosmetic, Toiletry, and Fragrance Association) International Cosmetic Ingredient Dictionary and Handbook, Tenth Edition (2004); and "Inactive Ingredient Guide" US Food and Drug Administration (FDA) Center for Drug Evaluation and Research (CDER) Office of Management, the contents of all of which are incorporated herein by reference in their entirety.Examples of the pharma-ceutically acceptable carriers, excipients and diluents useful in the present composition include distilled water, physiological saline, Ringer's solution, dextrose solution, Hank's solution and DMSO. These additional inactive ingredients, as well as effective formulation and administration procedures, are well known in the art and are described in standard texts, such as Goodman and Gillman's: The Pharmacological Bases of Therapeutics, 8th Ed., Gilman et al. Eds. Pergamon Press (1990); Remington's Pharmaceutical Sciences, 18th Ed., Mack Publishing Co., Easton, Pa. (1990); and Remington: The Science and Practice of Pharmacy, 21st Ed., Lippincott Williams & Wilkins, Philadelphia, Pa., (2005), each of which is incorporated herein by reference in its entirety.The presently described compositions may also be included in artificially created structures such as liposomes, ISCOMS, delayed release particles, and other vehicles that extend the half-life of peptides or polypeptides in serum. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and the like. Liposomes for use with the presently described peptides are formed from standard vesicle-forming lipids, which generally include neutral and negatively charged phospholipids and sterols such as cholesterol. The choice of lipid is generally determined by considerations such as liposome size and blood stability. A variety of methods are available for preparing liposomes, as reviewed, for example, by Coligan, JE et al, Current Protocols in Protein Science, 1999, John Wiley & Sons, Inc., New York, and see also U.S. Patent Nos. 4,235,871, 4,501,728, 4,837,028, and 5,019,369.
[0157] The carriers may in total constitute from about 0.1% to about 99.99999% by weight of the pharmaceutical compositions presented herein.
[0158] Method of synthesis According to some embodiments, a method is provided for synthesizing a glycosylated cannabinoid or a precursor thereof. According to some embodiments, a method is provided for synthesizing a glycosylated cannabinoid or a precursor thereof.
[0159] According to some embodiments, a method for glycosylation of a cannabinoid or a precursor thereof is provided.
[0160] In some embodiments, the method comprises synthesizing a monoglycosylated cannabinoid or a precursor thereof. In some embodiments, the method comprises monoglycosylating a cannabinoid or a precursor thereof.
[0161] In some embodiments, glycosylating comprises glucosylating. In some embodiments, glucosylating comprises adding glucose to the cannabinoid or precursor thereof. In some embodiments, the cannabinoid or precursor thereof according to the disclosed methods comprises glucose. In some embodiments, glycosylating comprises monoglycosylating. In some embodiments, glycosylating comprises diglycosylating. In some embodiments, the cannabinoid or precursor thereof is monoglycosylated according to the disclosed methods. In some embodiments, the cannabinoid or precursor thereof is diglycosylated according to the disclosed methods.
[0162] According to some embodiments, the method comprises: (a) providing a cell comprising an artificial vector comprising a nucleic acid sequence having at least 87%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99%, or any value and range therebetween, homology or identity to any one of SEQ ID NOs: 1-13, or any combination thereof; and (b) culturing the cell of step (a) such that the protein encoded by the artificial vector is expressed, thereby synthesizing a glycosylated cannabinoid or a precursor thereof. Each possibility represents a separate embodiment of the present invention.
[0163] According to some embodiments, the method comprises: (a) providing a cell comprising an artificial vector comprising a nucleic acid sequence having at least 87%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99%, or any value and range therebetween, homology or identity to any one of SEQ ID NOs: 1-13, or any combination thereof; and (b) culturing the cell of step (a) such that the protein encoded by the artificial vector is expressed, thereby glycosylating the cannabinoid or its precursor. Each possibility represents a separate embodiment of the present invention.
[0164] According to some embodiments, the method comprises contacting a cannabinoid or precursor thereof with an effective amount of a protein comprising an amino acid sequence having at least 90%, at least 93%, at least 95%, at least 99%, or 100%, or any value and range therebetween, homology or identity to any one of SEQ ID NOs: 14-26, thereby glycosylating the cannabinoid or precursor thereof. Each possibility represents a separate embodiment of the present invention.
[0165] According to some embodiments, the method comprises contacting a cannabinoid or precursor thereof with an effective amount of a protein comprising an amino acid sequence having at least 90%, at least 93%, at least 95%, at least 99%, or 100% homology or identity to any one of SEQ ID NOs: 14-26, or any value and range therebetween, thereby synthesizing a glycosylated cannabinoid or precursor thereof. Each possibility represents a separate embodiment of the present invention.
[0166] In some embodiments, the cannabinoid is CBDA, CBGA, HeliCBGA, delta-9-tetrahydrocannabinolic acid (Δ 9 -THCA), Δ 9 -Being or containing THC, CBD, CBG, CBCA, or any combination thereof.
[0167] In some embodiments, the cannabinoid precursor is or comprises olivetolic acid (OA). In some embodiments, the cannabinoid precursor is or comprises olivetol, DHSA, HA, iValA, BA, and VA, including any salts thereof and any combinations thereof.
[0168] According to some embodiments, methods are provided for synthesizing glycosylated or glucosylated phloroglucinoids, flavonoids, or any precursors thereof.
[0169] According to some embodiments, a method for glycosylation of a phloroglucinoid, a flavonoid, or any precursor thereof is provided.
[0170] According to some embodiments, the method comprises the steps of: (a) providing a cell comprising an artificial vector comprising a nucleic acid sequence having at least 87%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99%, or any value and range therebetween, homology or identity to any one of SEQ ID NOs: 1-13, or any combination thereof; and (b) culturing the cell of step (a) such that the protein encoded by the artificial vector is expressed, thereby synthesizing glycosylated phloroglucinoids, flavonoids, or precursors thereof. Each possibility represents a separate embodiment of the present invention.
[0171] According to some embodiments, the method comprises contacting a phloroglucinoid, flavonoid, or any precursor thereof with an effective amount of a protein comprising an amino acid sequence having at least 90%, at least 93%, at least 95%, at least 99%, or 100%, or any value and range therebetween, homology or identity to any one of SEQ ID NOs: 14-26, thereby synthesizing a glycosylated phloroglucinoid, flavonoid, or any precursor thereof. Each possibility represents a separate embodiment of the present invention.
[0172] In some embodiments, the phloroglucinoid, flavonoid, or precursor thereof is selected from 1-(2,4,6-trihydroxyphenylhexane)-1-one, naringenin chalcone, pinocembrin chalcone, or a combination thereof.
[0173] According to some embodiments, a method for obtaining an extract from a transgenic or transfected cell is provided.
[0174] In some embodiments, the method comprises culturing the transgenic or transfected cells in a medium and extracting the transgenic or transfected cells.
[0175] In some embodiments, the method comprises: (a) culturing the transgenic or transfected cells in a medium; and (b) extracting the transgenic or transfected cells, thereby obtaining an extract from the transgenic or transfected cells.
[0176] In some embodiments, the transgenic or transfected cells comprise an artificial vector comprising a nucleic acid sequence having at least 87%, at least 89%, at least 92%, at least 95%, at least 97%, at least 99%, or 100% homology or identity to any one of SEQ ID NOs: 1-13, or any combination thereof, or any value and range therebetween, with each possibility representing a separate embodiment of the present invention.
[0177] In some embodiments, the transgenic or transfected cell comprises a polynucleotide of the invention, or a plurality thereof, as disclosed herein.
[0178] In some embodiments, the transgenic or transfected cells comprise an artificial nucleic acid molecule or vector disclosed herein.
[0179] In some embodiments, the cell is a transgenic cell or a cell transfected with an isolated DNA molecule disclosed herein.
[0180] In some embodiments, culturing comprises supplementing the cells with an effective amount of a cannabinoid or a precursor thereof, in some embodiments, via a growth or culture medium in which the cells are cultured.
[0181] In some embodiments, the method further comprises a step preceding step (a) comprising introducing or transfecting a cell with an artificial nucleic acid molecule or vector disclosed herein.
[0182] Methods for introducing or transfecting an artificial nucleic acid molecule or vector into a cell are common and will be apparent to those skilled in the art.
[0183] In some embodiments, introducing or transfecting comprises introducing an artificial nucleic acid molecule or vector comprising a polynucleotide disclosed herein into the cell, or modifying the genome of the cell to include a polynucleotide disclosed herein. In some embodiments, introducing comprises transfection. In some embodiments, introducing comprises transformation. In some embodiments, introducing comprises lipofection. In some embodiments, introducing comprises nucleofection. In some embodiments, introducing comprises viral infection.
[0184] As used herein, the terms "transfecting" and "introducing" are interchangeable.
[0185] In some embodiments, the contacting is in a cell-free system.
[0186] The type of cell-free system suitable for utilizing any one of the polynucleotides or multiple of the invention and isolated protein or multiple of the invention disclosed herein will be apparent to one of skill in the art.
[0187] In some embodiments, the method further comprises a step preceding step (b) comprising separating the cultured transgenic or transfected cells from the culture medium.
[0188] Methods for separating cells from the medium are common and may include, but are not limited to, centrifugation, ultracentrifugation, or other methods, as will be apparent to one of skill in the art.
[0189] According to some embodiments, extracts of transgenic or transfected cells obtained according to the methods disclosed herein are provided.
[0190] According to some embodiments, culture medium or a portion thereof separated from cultured transgenic or transfected cells obtained according to the methods disclosed herein is provided.
[0191] According to some embodiments, there is provided a composition comprising: (a) an extract as disclosed herein; (b) a medium as disclosed herein or a portion thereof; or (c) any combination of (a) and (b) and an acceptable carrier as described herein.
[0192] In some embodiments, a portion includes a fraction or fractions.
[0193] General Where a range of values is stated, it is understood that each intervening value between the upper and lower limits of that range, to the nearest tenth of the lower limit, and any other stated or intervening value in that stated range, is included in the invention, unless the context clearly dictates otherwise. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the invention, subject to any limit specifically excluded in the stated range. Where one or both of the upper and lower limits are included in the stated range, ranges excluding either or both of those included limits are also included in the invention.
[0194] As used herein, the term "about" when combined with any value refers to ±10% of the reference value. For example, a length of about 1000 nanometers (nm) refers to a length of 1000 nm ±100 nm.
[0195] It should be noted that, as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to "a polynucleotide" includes a plurality of such polynucleotides, a reference to "the polypeptide" includes a reference to one or more polypeptides known to those of skill in the art and equivalents thereof, and so forth. It should be further noted that claims may be drafted to exclude any optional element. As such, this statement is intended to serve as a prerequisite for using exclusive language, such as "solely," "only," and the like, in connection with the description of claim elements or the use of "negative" limitations.
[0196] Where a convention similar to "at least one of A, B, and C, etc." is applied, such a configuration is generally intended in the sense that one of ordinary skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having A only, B only, C only, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by one of ordinary skill in the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to consider the possibility of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."
[0197] It is understood that certain features of the invention that are described in the context of separate embodiments for clarity may be provided in combination in a single embodiment. Conversely, various features of the invention that are described in the context of a single embodiment for brevity may be provided separately or in any suitable subcombination. All combinations of the embodiments of the invention are specifically embraced by the invention and disclosed herein as if every combination were individually and explicitly disclosed. Moreover, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the invention and disclosed herein as if every such subcombination were individually and explicitly disclosed herein.
[0198] Additional objects, advantages, and novel features of the present invention will become apparent to those skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as described hereinabove and as claimed in the claims section below finds experimental support in the following examples.
[0199] Various embodiments and aspects of the present invention as delineated hereinabove and in the claims section below find experimental support in the following examples.
[0200] Working Example In general, the nomenclature used herein and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological and recombinant DNA techniques. Such techniques are explained in detail in the literature. See, for example, "Molecular Cloning: A Laboratory Manual" Sambrook et al., (1989); "Current Protocols in Molecular Biology" Volumes I-III Ausubel, RM, ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology", John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning", John Wiley&Sons, New York (1988); Watson et al., "Recombinant DNA", Scientific American Books, New York; Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series" 1-4, Cold Spring Harbor Laboratory Press, New York. New York (1998); the methods described in U.S. Pat. Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659 and 5,272,057; "Cell Biology: A Laboratory Handbook", Volumes I-III Cellis, JE, ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique" Freshney, Wiley-Liss, NY (1994), Third Edition; "Current Protocols in Immunology" Volumes I-III Coligan JE, ed. (1994); Stites et al.(eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization - A Laboratory Course Manual", CSHL Press (1996); all of which are incorporated by reference. Other general references are provided throughout this specification.
[0201] material and method chemicals CBGA, Hexanoic Acid D 11 (D>98%), and uridine 5'-diphosphoglucose (UDP) disodium salt were purchased from Sigma-Aldrich (Rehovot, Israel). HeliCBGA (NP009525, 90%) was purchased from Analyticon Discovery GmbH (Potsdam, Germany). OA was purchased from Cayman Chemical (Ann Arbor, MI, USA). Phenylalanine-D5 (D>98%) and Phenylalanine- 13 C9, 15 N1( 13 C, 15 Pentanoic acid-D9 (D>98%), heptanoic acid-D5 (D>99%), and isocaproic acid-D 11 (D>98%) was purchased from C / D / N isotopes (Quebec, Canada). Naringenin chalcone, pinocembrin chalcone, and hexanoylphloroglucinol (95%) were purchased from Wuhan ChemFaces Biochemical Co Ltd. (Hubei, China).
[0202] Plant sources and growing conditions Helichrysum plants were grown in normal soil and fertilized with 18-18-18 NPK-Mg fertilizer. Plants were grown in a greenhouse at the Weizmann Institute (Rehovot) under natural light supplemented with HPS artificial lighting for 16 h of light per day.
[0203] UPLC-qTOF analysis of cannabinoids from Helichrysum and Cannabis tissues Fresh samples of six different tissues: young leaves, old leaves, florets and thalamus, stems, and roots were collected from plants at the flowering stage. The florets and thalamus were separated using a scalpel and extracted separately. All tissues were flash frozen in liquid N2 and ground to a fine powder in a mortar. Next, 100 mg of frozen powder from all plant tissues was extracted with 1 ml of ethanol, vortexed, and sonicated for 20 min at room temperature. Finally, the extracts were centrifuged at 14,000 x g for 15 min and the solvent was filtered through a 0.22 μm filter.
[0204] Samples were analyzed using a high-resolution ultra-performance liquid chromatography-tandem quadrupole time-of-flight (UPLC-qTOF) system consisting of a UPLC (Waters Acquity) equipped with a diode array detector coupled to either a XEVO G2-S QTof (Waters) or a Synapt HDMS (Waters). Chromatographic separation of compounds was performed on a 100 mm × 2.1 mm id, 1.7 μm UPLC BEH C18 column (Waters Acquity). The mobile phase consisted of acetonitrile:water with 0.1% formic acid (5:95, v / v, phase A) and acetonitrile with 0.1% formic acid (phase B). The flow rate was 0.3 ml min -1and the column temperature was maintained at 35°C. Cannabinoids were analyzed using a 29-minute multi-step gradient method. Initial conditions were 40% B for 1 minute, ramped to 100% B by 23 minutes, held at 100% B for 3.8 minutes, and ramped down to 40% B by 27 minutes. 40% B was held until 29 minutes for system re-equilibration. Intermediates and glycosylated compounds were analyzed using a 40-minute multi-step gradient method: ramped from 0% to 28% B over 22 minutes, ramped to 100% B by 36 minutes, held at 100% B for 2 minutes, ramped down to 0% B by 38.5 minutes, and held at 40% B until 40 minutes for system re-equilibration. Electrospray ionization (ESI) was used in negative ionization with an m / z range of 50-1,000 Da. The mass of eluted compounds was analyzed using the following settings: capillary 1 kV, source temperature 140 °C, desolvation temperature 450 °C, and desolvation gas flow rate 800 lh. -1 Detection was performed at 100 Hz. Argon was used as the collision gas. MS / MS experiments were performed in negative ionization mode according to the deprotonated mass observed. The following settings were used: 1 kV capillary spray, 30 eV cone voltage, 15-50 eV collision energy ramp.
[0205] Purification of compounds for NMR analysis A total of 86 g of fresh leaves were flash frozen in liquid N2, ground to a fine powder using a motorized grinder, extracted with 600 ml of ethanol, sonicated in an ultrasonic bath for 20 min, and agitated on an orbital shaker for 30 min at 25° C. The supernatant was then pressure filtered and the ethanol evaporated using a rotary evaporator at 40° C., followed by lyophilization to remove residual water. The final extract was reconstituted in 25 ml of acetonitrile and used either for direct purification (after 10-fold dilution) or for prefractionation by medium pressure liquid chromatography (MPLC).
[0206] MPLC was performed on a Buchi Sepacore system equipped with two C-605 pump modules, a C-620 control unit, a C-660 fraction collector, a C-640 UV photometer (Buchi Labortechnik AG, Switzerland), and a C18 manually packed column. The mobile phase consisted of acetonitrile:water (5:95, v / v, phase A) and acetonitrile (phase B) using the following multi-step gradient method: initial conditions were 0% B for 10 min, increasing to 99% B by 530 min, and slowly increasing to 100% B by 660 min. The flow rate was 15 ml min -1 The injection volume was 15 ml and the wavelengths used to monitor the acquisition were: 210, 224, 270, and 350 nm. Fractions of 100 ml were collected during the run, resulting in 99 tubes. The fractions were analyzed by UPLC-qTOF to select specific compounds for purification. The selected fractions were evaporated using a rotary evaporator at 40° C., lyophilized to remove residual water, reconstituted in methanol, and filtered through a 0.22 μm syringe filter.
[0207] Purification of compounds was performed either on an Agilent 1290 Infinity II UPLC system equipped with a quaternary pump, autosampler, and diode array detector, a Bruker / Spark Prospekt II LC-SPE system (Spark), and an Impact HD UHR-QqTOF MS (Bruker) connected via a Bruker NMR MS interface (BNMI-HP) (general instrument setup from Jozwiak et al.), or on a UPLC system (Waters Acquity) equipped with a binary pump, autosampler, fraction manager, and diode array detector. Triggering on both instruments was performed using specific UV wavelengths depending on the compound. The mobile phase consisted of acetonitrile:water with 0.1% formic acid (5:95, v / v, phase A) and acetonitrile with 0.1% formic acid (phase B).
[0208] Bruker system method development was performed by acquiring both MS and UV signals. MS spectra were acquired from m / z 50 to 1,700 in negative full scan mode. Chromatographic separation was performed using XBridge (BEHC18, 250 × 4.6 mm id, 5 μm; Waters) or Luna (C18, 250 × 4.6 mm id, 5 μm; Phenomenex) HPLC columns, and conditions were tuned and optimized for each compound. With this system, the eluent containing the compounds of interest was eluted at 1.8 ml min -1 The eluates were mixed with 1000 μl of make-up water and then captured on solid phase extraction (SPE) cartridges (10×2 mm Hysphere resin GP cartridges). Each cartridge was loaded four times with the same compound, and approximately 60 cartridges were used to capture one compound, depending on the concentration of the injected sample. Prior to NMR measurements, the SPE cartridges were dried with a stream of N2 and fractions from each cartridge were eluted with a total of 150 μl of MeOH into a 96-well plate. The eluates containing the same compound were pooled, dried under a stream of N2, and stored at −20° C. until NMR analysis.
[0209] Chromatographic separation on the Waters system was performed on a Luna phenylhexyl column (150 mm × 2 mm id, 3 μm, Phenomenex). The flow rate was 0.3 ml min -1 and the column temperature was maintained at 35°C. All other conditions were adjusted and optimized according to the sample. The eluates containing the target compounds were collected in 2 ml HPLC vials. The eluates containing the same compounds were pooled, dried under N2 flow, lyophilized, and stored at -20°C until NMR analysis.
[0210] NMR method The purified compounds were resuspended in 300 μl of MeOD-D4, dried under a stream of N2 to remove traces of 1H from the previous solvent, and diluted with 0.01% 3-propionic acid-2,2,3,3-D4 sodium salt (which is 1 H and 13The compounds were reconstituted in 70 μl of MeOD-D4 containing 1.2 mM NaCl (used as internal chemical shift reference for C spectra) and transferred to a 1.7 mm NMR test tube for structure elucidation. NMR spectra were recorded on a Bruker AVANCE NEO-600 NMR spectrometer equipped with a 5 mm TCI-xyz CryoProbe. All spectra were obtained at 25° C. The structures of the different compounds were determined by one-dimensional (1D) NMR spectra and various two-dimensional (2D) NMR spectra: 1 H- 1 H COSY (Correlation Spectroscopy), 1 H- 1 H TOCSY (Total Correlation Spectroscopy), 1 H- 1 H ROESY(Rotating Frame Nuclear Overhauser Spectroscopy) 1 H- 13 C HSQC (Heteronuclear Single Quantum Coherence), and 1 H- 13 C HMBC (Heteronuclear Multiple Bond Correlation) spectrum.
[0211] one dimensional 1H NMR spectra were acquired using 16,384 data points and a recycle delay of 2.5 seconds. 2D COSY, TOCSY, and ROESY spectra were acquired using data points of 16,384–8,192 (t2) × 400–512 (t1). 2D TOCSY spectra were obtained using isotropic mixing times of 100–300 ms. T-ROESY spectra were recorded using spin-locked pulses of 100–400 ms. 2D HSQC and 2D HMBC spectra were recorded using data points of 4,096 (t2) × 400–512 (t1). Multiplicity-edited HSQC allows one to distinguish between methyl and methine groups, which give rise to positive correlations, and methylene groups, which appear as negative peaks. The HMBC delay for the occurrence of long-range couplings was adjusted to 100–200 ms using a 100-ms recycle delay of 2.5 seconds. H,C = 8Hz long-range coupling was observed.
[0212] 1 H and 13 C chemical shift assignments were based on information obtained from all NMR spectra. Proton assignments as axial or equatorial were based on observed vicinal J couplings. Large values (>10 Hz) indicate axial protons, which is further supported by correlations observed in the ROESY spectrum. 1 H- 13 The C correlations are indicated by arrows and are observed in the COSY spectrum. 1 H- 1 The H correlation is shown as a dashed line.
[0213] Absolute quantification of CBGA Samples were extracted as described above. The final volume was diluted 10,000-fold to match the linear range of the calibration curve. Injections were performed on a UPLC (Waters) connected to a triple quadrupole detector (TQ-S, Waters) in multiple reaction monitoring (MRM) mode. Chromatographic separation was achieved using similar columns and mobile phases as described above. A short 7-min method was established using the following multi-step gradient program: initial conditions were 57% B ramped to 85% B until 4 min, ramped to 100% B until 4.2 min, held at 100% B until 6 min, ramped down to 67% B until 6.2 min, and held at 67% B until 7 min for system re-equilibration. 0.6 mL min -1 A flow rate of 1000 µL was used, the column temperature was 40 °C, and the injection volume was 1 µL. The instrument was operated in negative mode with a capillary voltage of 1.5 kV and a cone voltage of 40 V. Absolute quantification of CBGA was performed by external calibration using two different transitions (359.3 > 191.2, 32 V for quantification and 359.3 > 315.4, 21 V for qualification).
[0214] MALDI Imaging For localization of terpenes to individual glandular trichomes, fresh leaves and flowers were embedded in Peel-A-Way disposable embedding molds (Peel-A-Way Scientific) with M1 embedding matrix (Thermo Scientific) and frozen on dry ice. Embedded tissues were transferred to a cryostat (Leica CM3050) and allowed to thermally equilibrate at −17°C for at least 2 h. Frozen tissues were sliced into 40 μm thick sections. Sections were thaw-mounted onto Superfrost Plus slides (Fisher Scientific), vacuum dried in a desiccator, and photographed with a Nikon DS-Ri2 microscope. 2,5-dihydroxybenzoic acid (DHB, 40 mg ml ) was added to the 2,5-dihydroxybenzoic acid solution using a TM nebulizer (HTX Technologies). -1 The plant tissue was coated with DHB matrix solution (dissolved in 70% MeOH containing 0.2% trifluoroacetic acid). The nozzle temperature was set at 70°C, and the DHB matrix solution was pumped at a flow rate of 120 cm min-1 Linear velocity, 50 μl min -1 The solution was sprayed 16 times onto the tissue sections at a flow rate of 100 Hz. MALDI imaging was performed using a 7T Solarix FT-ICR (Fourier transform ion cyclotron resonance) mass spectrometer (Bruker Daltonics). Data sets were collected in positive ion mode with lock mass calibration (DHB matrix peak: [3DHB+H-3H2O]+, m / z 409.055408) at a frequency of 1 kHz, 40% laser power, and 200 laser shots per pixel at pixel sizes of 15 μm or 25 μm for leaf and flower cross sections, respectively. Each mass spectrum was recorded in the range of m / z 150–3,000 in broadband mode with an acquisition time domain of 1 M, giving an estimated resolution of 115,000 at m / z 400. The acquired spectra were processed using Flex-Imaging software 4.0 (Bruker Daltonics). Spectra were normalized to root-mean-square intensity and MALDI images were plotted with pixel interpolation on at theoretical m / z ±0.005%.
[0215] Isolation of glandular trichomes Young leaves were harvested, immersed in ice-cold distilled water, and then polished using a BeadBeater machine (Biospec Products, Bartlesville, OK). Polycarbonate chambers were filled with 15 g of plant material and half the volume was filled with glass beads (0.5 mm diameter), XAD-4 resin (1 g / g plant material), and 80% ethanol to the full volume. Leaves were beaten with 2 to 4 pulses of 1 min each. The procedure was performed at 4°C, and the chamber was cooled on ice after each pulse. After abrasion, the contents of the chamber were filtered first through a kitchen mesh strainer and then through a 100 μm nylon mesh to remove the plant material, glass beads, and XAD-4 resin. The remaining plant material and beads were scraped off the mesh and rinsed twice with additional 80% ethanol, which was also passed through the 100 μm mesh. The presence of concentrated glandular trichome secretory cells was confirmed by visualization with an inverted light microscope.
[0216] Helichrysum genome sequencing and assembly The genome size of Helichrysum was estimated by flow cytometry. Briefly, nuclei were isolated by chopping young leaf tissues of Helichrysum and tomato (used as a known reference) in separation buffer. Samples were stained with propidium iodide, at least 10,000 nuclei were analyzed by flow cytometer, and the ratio of G1 peak average between both samples was calculated. High molecular weight DNA was extracted from frozen young leaves and sent to the Genome Center of UC Davis for sequencing. DNA quality was confirmed by TapeStation trace and Qubit fluorometer (Thermo Fisher). Sequencing was performed on a Pacbio Sequel II platform, and a DNA SMRT bell library of approximately 12 kilobases was prepared according to the manufacturer's protocol. Three different SMRT 8M cells were used to generate a total of 57.8 Gb of HiFi data (approximately 44-fold haploid coverage). In addition to the Pacbio HiFi data, 200M reads of PE 2x150 Illumina Hi-C data were acquired by Phase Genomics. Both the Pacbio HiFi and HiC data were integrated to generate a haplotype-resolved assembly at the chromosome scale using Hifiasm software.
[0217] Further scaffolding of the primary assembly was performed using Hi-C data and SALSA software. Ragtags were used for a final ordering round using the primary assembly as a reference to arrive at a syntenic scaffold for each haplotype. Visualization of Hi-C data was performed using Juicer, and whole genome alignments were performed using the pafr package (dwinter.github.io / pafr / ). Finally, the assembly was soft-masked for repetitive elements using EDTA.
[0218] RNA sequencing and genome annotation of Helichrysum RNA was extracted from seven tissues: young leaves, old leaves, florets and receptacles, stems, roots, and glandular trichomes. RNA integrity was confirmed using a TapeStation instrument. Paired-end Illumina libraries were prepared for five tissues and sequenced on an Illumina HiSeq 3000 instrument (PE2x150, ~40M reads per sample). Random sequence errors were corrected using Rcorrector and uncorrectable reads were removed. Adapter and quality trimming was performed using TrimGalore! with the following parameters: --length 36-q5--stringency 1-e0.1 (github.com / FelixKrueger / TrimGalore). Ribosomal RNA was filtered by discarding reads mapping to the SILVA_132_LSURef and SILVA_138_SSURef non-redundant databases using bowtie2--very-sensitive-local mode. Fastq quality checks for each step were performed using MultiQC. The remaining reads were pooled and used for genome-guided de novo transcriptome assembly using Trinity. Iso-Seq data were obtained from four tissues and processed using isoseq3 and the cDNA Cupcake ToFU pipeline (github.com / Magdoll / cDNA_Cupcake). Fused and unspliced transcripts were removed and only polyA positive transcripts were kept as a unique set of high-quality isoforms. Iso-Seq and Trinity transcripts were aligned to the assembly using minimap2 and the BAM files were used in the PASA pipeline to generate RNA-based gene model structures. Further, novo gene structures were obtained using the software breaker2 and the aforementioned BAM files as external training evidence. Finally, the ab initio and RNA-based gene models were combined using EvidenceModeler and the final round of the PASA pipeline.Gene function annotation was performed for predicted mature transcripts using TransDecoder (github.com / TransDecoder / TransDecoder), considering HMMER hits against PFAM and BLASTP hits against the UniProt database as similarity retention criteria. Further annotation of protein-coding transcripts was performed by BLASTP searches against a curated plant protein database, and GO and KEGG terms were obtained using Trianotate.
[0219] UMI-based 3'RNAseq of three replicates out of seven tissues were obtained similarly to that described. Adapter and quality trimming was performed using TrimGalore! in two steps, including PolyA trimming mode. Reads were mapped to the genome using STAR, UMI duplicates removed using umitools, and counts were obtained with featureCounts. Normalization was performed using the varianceStabilizingTransformation algorithm in DESeq2, and the CEMItool package was used for co-expression analysis (dissimilarity threshold of 0.6, p-value of 0.1). Genes in modules with expression profiles consistent with metabolites of interest were analyzed. Candidate genes were selected based on functional annotation and blast hits with known UGT enzymes.
[0220] UGT expression in E. coli BL21(DE3) cells and protein purification Selected UGT genes from E. helichrysum were individually cloned into the pET28b vector and expressed in E. coli BL21(DE3) cells. Bacterial starters were grown overnight at 37 °C in LB medium, diluted 1:100 in fresh LB, and incubated again at 37 °C. When the cultures reached an A600 = 0.6, protein expression was induced with 400 μM isopropyl-1-thio-β-d-galactopyranoside (IPTG) overnight at 15 °C. Bacterial cells were incubated in 50 mM Tris-HCl pH 8, 0.5 mM phenylmethylsulfonyl fluoride (PMSF, Sigma Aldrich) solution in isopropanol, 10% glycerol and protease inhibitor cocktail (Sigma Aldrich), and 1 mg ml -1 Cells were lysed by sonication in lysozyme (Sigma Aldrich). Whole cell extracts were either saved for functional activity or used for protein purification. Protein purification was performed with Ni-NTA agarose beads (Adar Biotech). Proteins were eluted with 200 mM imidazole (Fluka) in a buffer containing 50 mM NaH2PO4, pH 8 and 0.5 M NaCl. Protein concentrations of the eluted fractions were measured using Pierce™ 660 nm protein assay reagent (Thermo Scientific).
[0221] β-Glucosidase assay for preparation of DHSA The two MPLC fractions (50 ml each) containing Glc-OA and Glc-DHSA were evaporated as described above and each reconstituted in 15 ml of McIlvaine buffer (pH 5.0). The reactions were carried out in separate 20 ml vials and incubated at 45 °C for 24 h. Each reaction contained 6 ml of McIlvaine buffer (pH 5.0), 0.1 mg ml -1 of almond β-glucosidase solution (≧6U mg -1The reaction consisted of 100 μl of OA and 1.5 ml of Glc-DHSA-containing fractions (Sigma Aldrich) and 1.5 ml of Glc-DHSA-containing fractions. The compounds were extracted using 3 volumes of ethyl acetate:diethyl ether 1:1, evaporated using a rotary evaporator, and reconstituted in 5 ml of methanol. The products from the reaction included a mixture of both glucosylated and deglucosylated OA and DHSA. DHSA was therefore purified using a Waters apparatus as previously described and reconstituted in 100 μl of methanol for the enzyme assay. The purified DHSA was analyzed by UPLC-qTOF to confirm that the purified fractions were free of Glc-DHSA.
[0222] UGT enzyme assay Recombinant UGT assays with different aromatic substrates were performed by mixing 1.5 μl UDP solution (80 mM, final concentration: 2.5 mM), 27.5 μl Tris buffer (100 mM), 1 μl of each substrate (50 mM, final concentration: 1 mM), and 20 μl of lysate enzyme solution. The reactions were incubated at 30° C. for 1 h. To stop the reactions, 50 μl of methanol was added to each tube, vortexed for 10 s, and centrifuged at maximum speed for 10 min, after which the supernatant was collected and used for UPLC-qTOF analysis. Assays with purified UGTs were performed by mixing 2 μl of cannabinoid acceptor (OA, DHSA, CBGA, heliCBGA, CBDA, ΔC, ΔH, ΔH) in the presence of 1.5 μl UDP 80 mM, 46.5 μl Tris buffer (100 mM, pH 8.0), and 1 μl of each enzyme. 9 -THCA, CBCA, Olivetol, CBG, CBD or Δ 9-THC, hexanoylphloroglucinol, naringenin chalcone or pinocembrin chalcone). To stop the reaction, 100 μl of methanol was added to each tube and the compounds were extracted and analyzed as described above. Kinetic assays were performed with purified enzyme (1.5 μg / μl) dissolved in 45 μl of Tris buffer (100 mM, pH 8.0), and substrates were added with varying concentrations (0.5 μM to 3 mM) and a constant concentration (1 mM) of OA and UDP, with a total reaction volume of 50 μl. To stop the reaction, 100 μl of methanol was added to each tube and the compounds were extracted and analyzed as described above.
[0223] Example 1 UPLC-qTOF profiling and RNA-Seq transcriptome of Helichrysum tissues First, we profiled six tissues of Helichrysum (young leaves, old leaves, florets, and thalamus, stems, and roots) using UPLC-qTOF. CBGA and its phenethyl analog, heliCBGA, were observed in all tissues except roots. These compounds were identified by comparison with analytical standards or authentic compounds (Figure 1A and Figure 1B). According to the metabolic pathways of cannabinoids and amorfurthins in other plants, CBGA and heliCBGA are biosynthesized from the intermediates OA and dihydrostilbenic acid (DHSA), respectively. These compounds and their glycosylated forms (Glc-OA and Glc-DHSA, respectively) were also identified in Helichrysum extracts according to UPLC-qTOF. OA was identified by comparison with its analytical standard (Figure 1C), and DHSA was identified by its MS / MS spectrum and relative retention time (RT) to OA. To confirm the assignments, CBGA, heliCBGA, Glc-OA, and Glc-DHSA were purified and analyzed by one- and two-dimensional nuclear magnetic resonance (NMR) (Figures 2-3). Several additional glycosylated compounds were identified in Helichrysum extracts by UPLC-qTOF, including C3-C6 alkyl chain intermediates, glycosylated CBGA, and heliCBGA (Glc-CBGA and Glc-heliCBGA, respectively), as well as two similar compounds with isoprenyl instead of monoprenyl (Glc-CBPA and Glc-heliCBPA, respectively, Figure 4). Glycosylated compounds were also observed in Cannabis flowers and leaves, including C3-C6 alkyl chain intermediates, Glc-CBGA, and glycosylated cannabidiolic acid (Glc-CBDA).
[0224] Among the tissues examined in Helichrysum, leaves and flowers showed the highest accumulation of CBGA, while roots contained no CBGA (Figure 5A). A similar trend was also observed for Glc-OA (Figure 5A). CBGA was further localized to glandular trichomes of transversely cut leaves and flowers by MALDI-MSI (Figure 5B-G). Using this information, RNA-seq transcriptome analysis was performed on these tissues, and candidate genes were selected based on their high expression profiles in leaves and isolated glandular trichomes compared to roots.
[0225] Example 2 Functional characteristics of UGTs Functional annotation of Helichrysum and Cannabis genes was inferred by sequence similarity with known UGT enzymes. Over 100 genes encoding UGTs have been identified in Arabidopsis, some of which have established substrate specificity. In particular, genes encoding the enzymes AtUGT89B1 and AtUGT89A2 have previously been found to catalyze the glycosylation of hydroxybenzoic acid (HBA) and some dihydroxybenzoic acids (DHBA). These molecules are structurally similar to OA and were therefore used to select candidate UGT enzymes in Helichrysum. Additional candidates were identified following a positive correlation between gene expression and metabolite accumulation (Figures 6 and 7). In total, we selected 13 (HuUGT1-13) candidate genes encoding putative UGTs.
[0226] Next, we recombinantly expressed 11 of the 13 UGTs from Helichrysum in E. coli and used crude lysates to examine the activity of the proteins using OA, CBGA, and heliCBGA in reactions involving UDP-Glc as the sugar donor. Several enzymes showed activity towards different substrates, including HuUGT1-2 (SEQ ID NO:14-15), HuUGT4-7 (SEQ ID NO:17-20), HuUGT11 (SEQ ID NO:24), and HuUGT13 (SEQ ID NO:26, FIG. 8). We purified the four most active enzymes (HuUGT1 (SEQ ID NO:14), HuUGT6 (SEQ ID NO:19), HuUGT11 (SEQ ID NO:24), and HuUGT13 (SEQ ID NO:26), together with previously characterized enzymes from Stevia and rice (SrUGT and OsUGT, respectively), and tested them in situ with a range of cannabinoid substrates, both natural and unnatural, for Helichrysum. In vitro assays were performed (Figure 9). We cleaned the LC / HRMS chromatograms for glucosylated and diglucosylated products according to the m / z theoretical values. Two major monoglucosides were observed for each substrate, assigned according to UPLC-qTOF as glucosides, with one of the hydroxyl groups in each molecule glucosylated (Figures 9 and 10). All enzymes were active and had various substrate specificities and products produced. For example, HuUGT11 and HuUGT13 were highly active on cannabinoid intermediates, but were almost inactive on prenylated compounds. Disaccharides of acidic compounds were observed only in the case of HuUGT6, while olivetol, cannabidiol (CBD), and cannabigerol (CBG) were disaccharided by different HuUGTs depending on the compound (Figure 11).
[0227] Interestingly, UGTs from Helichrysum also glucosylated precursors of phloroglucinoids and flavonoids naturally occurring in plants, even though the glucosylated forms were not observed in the extracts (Figure 9). Kinetic assays of HuUGT11, HuUGT13, OsUGT, and SrUGT showed the superior catalytic advantage of HuUGT11 with OA and UDP over all other enzymes (Figure 12). HuUGT11 is therefore a very promising enzyme that can produce large amounts of Glc-OA and Glc-DHSA, as shown in plants.
[0228] While the present invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and broad scope of the appended claims.
Claims
1. An isolated DNA molecule containing a nucleic acid sequence having at least 87% homology to SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, or any combination thereof.
2. (i) The nucleic acid sequence having at least 87% homology to any one of Sequence IDs 1 to 13 is 700 to 1,800 nucleotides long, (ii) The nucleic acid sequence encodes a protein which is uridine 5'-diphospho(UDP)-glucuronosyltransferase (UGT), or (iii) both (i) and (iii).
3. An artificial nucleic acid molecule comprising an isolated DNA molecule as described in claim 1.
4. A plasmid or Agrobacterium comprising the artificial nucleic acid molecule described in claim 3.
5. a. The isolated DNA molecule according to claim 1; b. The artificial vector according to claim 3; and c. The plasmid or Agrobacterium described in claim 4. An isolated protein encoded by one of the following.
6. The isolated protein according to claim 5, comprising, or consisting of, an amino acid sequence having at least 90% homology to SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, or SEQ ID NO:
26.
7. The isolated protein according to claim 6, characterized by the ability to glycosylate cannabinoids or their precursors.
8. a. The isolated DNA molecule according to claim 1; b. The artificial nucleic acid molecule according to claim 3; c. The plasmid or Agrobacterium described in claim 4; or d. Any combination of (a) to (c) Transgenic cells containing, Optionally, the transgenic cell is one of a unicellular organism, a cell of a multicellular organism, or a cell in a culture, and optionally, the unicellular organism includes a fungus or a bacterium, and optionally, the fungus is a yeast cell.
9. An extract obtained from the transgenic cells described in claim 8, or any fraction thereof, wherein the extract optionally contains the isolated DNA molecule.
10. a. The isolated DNA molecule according to Claim 1; b. The artificial vector according to claim 3; c. The plasmid or Agrobacterium described in claim 4; or d. Any combination of (a) to (c) Transgenic plants, transgenic plant tissues, or plant parts, including Optionally, the transgenic plant is a Cannabis sativa plant, transgenic plant tissue, or plant part.
11. a. The isolated DNA molecule according to Claim 1; b. The artificial vector according to claim 3; c. The plasmid or Agrobacterium described in claim 4; d. The extract according to claim 9; or e. Any combination of (a) to (d), and acceptable carriers A composition containing the following:
12. A method for glycosylation of a cannabinoid or its precursor, a. To provide cells containing an artificial vector having a nucleic acid sequence having at least 87% homology to SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, or SEQ ID NO: 13, b. Culturing the cells from step (a) so that the protein encoded by the artificial vector is expressed. A method comprising, thereby glycosylation of cannabinoids or their precursors.
13. (i) The cells are transgenic cells, or cells that have been transfected with an isolated DNA molecule or an artificial vector containing a nucleic acid sequence having at least 87% homology to SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, or any combination thereof. (ii) The protein is characterized by being able to transfer the glucuronic acid component of UDP-glucuronic acid into the cannabinoid or its precursor, (iii) The culturing includes supplying the cells with an effective amount of UDP, (iv) The artificial vector is an expression vector, (v) The cell is a prokaryotic cell or a eukaryotic cell, (vi) The cell is a prokaryotic cell or a eukaryotic cell. (vii) The method further comprises step (c) comprising extracting the cells, thereby obtaining an extract of the cells, and optionally the method further comprises a step preceding step (c) comprising separating the cultured cells from the culture medium in which the cells are cultured. (viiii) The method further includes a step preceding step (a) which includes introducing or transfecting the cells with the artificial vector, (ix) The cannabinoid is CBGA, heliCBGA, CBDA, or any combination thereof. (x) The cannabinoid precursor is olivetolic acid (OA), and Any combination of (xi)(i) to (x), The method according to claim 12, wherein the method is one of the following.
14. An extract of cells obtained according to the method of Claim 12.
15. A method for glycosylation of a cannabinoid or its precursor, comprising contacting the cannabinoid or its precursor with an effective amount of a protein having an amino acid sequence having at least 90% homology to SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, or SEQ ID NO: 26, thereby glycosylation of the cannabinoid or its precursor, wherein the contact is optionally performed in a cell-free system.