Prenyltransferase, and transgenic cells, tissues, and organisms containing it.

JP2026143769APending Publication Date: 2026-09-08YEDA RES & DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026099772
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-03-10
Filing Date
2026-06-16
Publication Date
2026-09-08

Smart Images

  • Figure 2026143769000001_ABST
    Figure 2026143769000001_ABST
Patent Text Reader

Abstract

To provide a novel enzyme capable of promoting the prenylation of aromatic compounds. [Solution] An isolated DNA molecule comprising a nucleic acid sequence derived from Helichrysum ambracrigerum that has at least 85% identity with SEQ ID NO: 7 and encodes a protein that is a prenyltransferase, and an isolated protein comprising an amino acid sequence that has at least 80% identity with SEQ ID NO: 18.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference to Related Applications This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 159,028, filed on March 10, 2021, entitled "PRENYLTRANSFERASE AND A TRANSGENIC CELL, TISSUE, AND ORGANISM COMPRISING SAME", the entire content of which is incorporated herein by reference.

[0002] Field of the Invention The present invention relates to PT comprising a polynucleotide encoding prenyltransferase (PT), and methods of using the same. Background Art

[0003] Background The development of novel methodologies related to the chemistry and biosynthesis of natural products has been attracting increasing attention. Prenylated aromatic natural products are considered to be a very promising class of therapeutically active compounds. Prenylation of aromatic compounds often results in significant alterations to the biological activity profile of the compound, both by creating new C-C bonds in the framework of the final product and introducing one or more double bonds. Such compounds can affect a wide variety of biological systems in mammals, and can act as antioxidants, anti-inflammatory agents, antiviral drugs, antiproliferative agents, and anticancer agents.

[0004] Prenyltransferases are ubiquitous enzymes that catalyze the alkylation of electron-rich prenyl acceptors by the alkyl moiety of allylic isoprene diphosphates. Prenyltransferases use isoprenoid diphosphates as substrates to catalyze the addition of acyclic prenyl moieties to isopentenyl diphosphate (IPP), higher-order prenyl diphosphates, aromatic-rich molecules and proteins.

[0005] Prenyltransferase (PT) may also be useful in the synthesis of cannabinoid analogs and analogs of cannabinoid precursors. Cannabinoid analogs have been synthesized and may be useful as pharmaceutical products. Identifying enzymes involved in the prenylation of aromatic receptor molecules, such as aromatic polyketides, and the nucleotide sequences encoding such enzymes, such as PT, remains a need in the art. [Overview of the project]

[0006] In addition to novel enzymes capable of promoting the prenylation of aromatic compounds, there is a need in the art for the identification of compounds that can modulate the prenylation of aromatic compounds. The present invention addresses these and other needs, as described in more detail herein and in the subsequent claims.

[0007] overview According to the first embodiment, an isolated DNA molecule is provided that contains nucleic acid sequences having at least 91% homology to SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, and 11, or any combination thereof.

[0008] In another embodiment, an artificial nucleic acid molecule comprising the isolated DNA molecule of the present invention is provided.

[0009] In another embodiment, a plasmid or Agrobacterium comprising an artificial nucleic acid molecule disclosed herein is provided.

[0010] In another embodiment, an isolated protein encoded by any one of the following is provided: (a) an isolated DNA molecule of the present invention; (b) an artificial vector disclosed herein; and (c) a plasmid or Agrobacterium disclosed herein.

[0011] In another embodiment, a transgenic cell is provided comprising (a) an isolated DNA molecule of the present invention; (b) an artificial nucleic acid molecule disclosed herein; (c) a plasmid or Agrobacterium disclosed herein; (d) an isolated protein disclosed herein; or (e) any combination of (a) to (d).

[0012] In another embodiment, extracts derived from transgenic cells disclosed herein, or any fraction thereof, are provided.

[0013] In another embodiment, a transgenic plant, transgenic plant tissue, or plant part is provided, comprising (a) an isolated DNA molecule of the present invention; (b) an artificial vector disclosed herein; (c) a plasmid or Agrobacterium disclosed herein; (d) an isolated protein disclosed herein; (e) a transgenic cell disclosed herein; or (f) any combination of (a) to (e).

[0014] In another embodiment, a composition is provided comprising (a) an isolated DNA molecule of the present invention; (b) an artificial vector disclosed herein; (c) a plasmid or Agrobacterium disclosed herein; (d) an isolated protein disclosed herein; (e) a transgenic cell disclosed herein; (f) an extract disclosed herein; (g) a transgenic plant tissue or plant part disclosed herein; or (h) any combination of (a) to (g) and an acceptable carrier.

[0015] In another embodiment, Equation II: [ka] A method is provided for synthesizing a compound represented by formula II (wherein (i) R1 is selected from the group consisting of C1-C8 alkyl and alpha-unsaturated phenylalkyl carboxylic acids and R2 is OH, or (ii) R1 is OH, R2 is selected from the group consisting of C1-C8 alkyl and alpha-unsaturated phenylalkyl carboxylic acids, R3 is a prenyl group, and R4 is hydrogen or a prenyl group), comprising: (a) providing cells containing an artificial vector having at least 91% homology to SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, or SEQ ID NO: 11; and (b) culturing the cells from step (a) to express a protein encoded by the artificial vector, thereby synthesizing a compound represented by formula II.

[0016] In another embodiment, a method for synthesizing a compound represented by formula II, comprising contacting a substrate molecule with an effective amount of a protein containing an amino acid sequence having at least 92% homology to SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, or SEQ ID NO: 22, thereby synthesizing a compound represented by formula II, wherein the substrate molecule is defined by formula I: [ka] A method is provided in which the formula is represented by (i) R1 is selected from the group consisting of C1-C8 alkyl and alpha-unsaturated phenylalkyl carboxylic acids and R2 is OH, or (ii) R1 is OH and R2 is selected from the group consisting of C1-C8 alkyl and alpha-unsaturated phenylalkyl carboxylic acids.

[0017] According to another aspect, there is provided a method for obtaining an extract from a transgenic cell or a transfected cell, the method comprising: (a) culturing the transgenic cell or transfected cell in a medium, wherein the transgenic cell or transfected cell comprises an artificial vector comprising a nucleic acid sequence having at least 91% homology to SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, or SEQ ID NO: 11; and (b) extracting the transgenic cell or transfected cell, thereby obtaining an extract from the transgenic cell or transfected cell.

[0018] According to another aspect, an extract of a transgenic cell or a transfected cell obtained according to the method disclosed herein is provided.

[0019] According to another aspect, there is provided a medium or a portion thereof isolated from cultured transgenic cells or cultured transfected cells obtained according to the method disclosed herein.

[0020] According to another aspect, there is provided a composition comprising: (a) the extract disclosed herein; (b) the medium or a portion thereof disclosed herein; or a combination of (a) and (b); and an acceptable carrier.

[0021] In some embodiments, the nucleic acid sequence has at least 80% homology to any one of SEQ ID NOs: 1 to 11, and is 950 to 1,750 nucleotides in length.

[0022] In some embodiments, the nucleic acid sequence encodes a protein that is prenyltransferase.

[0023] In some embodiments, the transgenic cell is any one of a cell of a unicellular organism, a cell of a multicellular organism, and a cultured cell.

[0024] In some embodiments, the unicellular organism comprises a fungus or a bacterium.

[0025] In some embodiments, the fungus is a yeast cell.

[0026] In some embodiments, the extract comprises an isolated DNA molecule, an isolated protein, or both.

[0027] In some embodiments, the isolated protein comprises an amino acid sequence having at least 92% homology to SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, or SEQ ID NO: 22.

[0028] In some embodiments, the isolated protein consists of the amino acid sequence of SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, or SEQ ID NO: 22.

[0029] In some embodiments, the isolated protein is characterized in that it is capable of transferring a prenyl group to a substrate molecule.

[0030] In some embodiments, the prenyl group is selected from the group consisting of dimethylallyl diphosphate, geranyl diphosphate, farnesyl diphosphate, and geranylgeranyl diphosphate.

[0031] In some embodiments, the substrate molecule is represented by Formula I.

[0032] In some embodiments, the α,β-unsaturated phenylalkyl carboxylic acid comprises cinnamic acid or a derivative thereof.

[0033] In some embodiments, the cinnamic acid derivative is a hydroxylated derivative of cinnamic acid.

[0034] In some embodiments, the hydroxylated derivative of cinnamic acid is coumaric acid.

[0035] In some embodiments, the transgenic plant is a Cannabis sativa plant.

[0036] In some embodiments, the protein is characterized by its ability to transfer a prenyl group to a substrate molecule.

[0037] In some embodiments, culturing involves supplying cells with an effective amount of substrate molecules.

[0038] In some embodiments, the artificial vector is an expression vector.

[0039] In some embodiments, the cells are prokaryotic or eukaryotic.

[0040] In some embodiments, the cells are transgenic cells having the isolated DNA molecule of the present invention or the artificial vector disclosed herein, or cells transfected with the isolated DNA of the present invention or the artificial vector disclosed herein.

[0041] In some embodiments, the method further includes a step preceding step (a), which includes introducing or transfecting cells with an artificial vector.

[0042] In some embodiments, contact occurs in a cell-free system.

[0043] In some embodiments, the substrate molecule is selected from the group consisting of resorcinoid precursors, stilbene acid precursors, acylphloroglucinoid precursors, and chalcone precursors.

[0044] In some embodiments, the substrate molecule is [ka] It is represented by a formula selected from the group consisting of (wherein R1 is a C1-C8 alkyl group, and R2 is an alpha-unsaturated phenylalkyl carboxylic acid or an alpha-saturated phenylalkyl carboxylic acid).

[0045] In some embodiments, the substrate molecule is [ka] It is selected from the group consisting of the following.

[0046] In some embodiments, the compound is selected from the group consisting of cannabinoids, amorphurtins, acylphlorogluconoids, and prenyl chalcones.

[0047] In some embodiments, the compound is [ka] The formula is selected from the group consisting of (wherein R1 is a C1-C8 alkyl group, R2 is an alpha-unsaturated phenylalkyl carboxylic acid or an alpha-saturated phenylalkyl carboxylic acid, R3 is a prenyl group, and R4 is hydrogen or a prenyl group).

[0048] In some embodiments, the compound is [ka] It is selected from the group consisting of the following.

[0049] In some embodiments, the compound is [ka] That is the case.

[0050] In some embodiments, the method further includes a step preceding step (b), which includes separating cultured transgenic cells or cultured transfected cells from the culture medium.

[0051] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as those commonly understood by those skilled in the art in which the invention pertains. Similar or identical methods and materials to those described herein may be used in the practice or testing of embodiments of the invention, but exemplary methods and / or materials are described below. In case of any inconsistency, the specification of the invention, including definitions, shall prevail. In addition, materials, methods, and examples are illustrative and not necessarily intended to be limiting.

[0052] Further embodiments and the full scope of the present invention will be obvious from the detailed description provided hereafter herein. However, it should be understood that the detailed description and specific examples illustrate preferred embodiments of the present invention, as various changes and improvements within the spirit and scope of the invention are obvious to those skilled in the art from this detailed description. [Brief explanation of the drawing]

[0053] [Figure 1] Figures 1A-1B include graphs showing the identification of CBGA in the Herchrysum ambracrigerum ethanol extract. (1A) Ion current (XIC) chromatogram of the extracted material (359.222 Da), and (1B) MS / MS spectral matching of the CBGA standard compared to the plant extract. [Figure 2] Figure 2 includes a graph showing the in vitro production of cannabigerol acid (CBGA). Purified microsomal fractions from yeast cells expressing prenyltransferase (PT) were used in enzyme assays containing olivetolic acid (OA) and geranyl pyrophosphate (GPP). The Cannabis sativa geranyl pyrophosphate:olivetolic acid geranyltransferase 4 (CsGOT4) enzyme assay was used as a positive control in the experiment. Extracted ion chromatograms (EICs) are shown. LC-MS was used for assay product analysis. Standard substance - STD; negative control - empty vector. [Figure 3] Figure 3 includes phylogenetic trees of functionally characterized aromatic PTs from Helichrysum ambracrigerum and other plants discussed in de Brujin et al. (2020). Sequences were aligned using MUSCLE, and a maximum likelihood phylogenetic tree using the JTT distance matrix method was constructed using MEGA11 software. Bootstrap values ​​are shown at the nodes of each branch (100 replicates). For comparison, GOT4 from cannabis is represented (*). [Figure 4] Figure 4 includes graphs showing the steady-state reaction kinetics of Helichrysum umbracrigerum PT1 (HuPT1; SEQ ID NO: 1), HuPT3 (SEQ ID NO: 3), and HuPTx (SEQ ID NO: 7) using olivetolic acid and GPP. The Michaelis-Menten Km values ​​for each enzyme were calculated using various olivetolic acid concentrations (0.5 μM to 1.5 mM) and a constant GPP concentration (1 mM) (n=3 technically independent samples; measurements were plotted individually). [Modes for carrying out the invention]

[0054] Detailed description In some embodiments, the present invention relates to polynucleotide sequences encoding proteins or multiple proteins belonging to the prenyltransferase (PT) family, derived from Helichrysum ambracrigerum.

[0055] According to some embodiments, a polynucleotide is provided which includes a nucleic acid sequence containing any one of sequence numbers 1 to 11, or any combination thereof.

[0056] In some embodiments, the polynucleotide is an isolated polynucleotide. In some embodiments, the polynucleotide is a DNA molecule. In some embodiments, the polynucleotide is an isolated DNA molecule. In some embodiments, the DNA molecule is an isolated DNA molecule. In some embodiments, the DNA molecule is a complementary DNA (cDNA) molecule.

[0057] As used herein, the terms “isolated polynucleotide” and “isolated DNA molecule” refer to nucleic acid molecules that are essentially free from contaminating cellular components such as carbohydrates, lipids, or other proteinaceous impurities that naturally associate with nucleic acids. Typically, the preparation of isolated DNA or RNA contains nucleic acids in a highly purified form, e.g., at least about 80% pure, at least about 90% pure, at least about 95% pure, more than 95% pure, or more than 99% pure. In some embodiments, the isolated polynucleotide is one of DNA, RNA, and cDNA. In some embodiments, the isolated polynucleotide is a synthetic polynucleotide. The synthesis of polynucleotides is well known in the art and can be carried out, for example, by ligating multiple nucleic acid molecules together with a primer linker or by covalent linking.

[0058] The term “nucleic acid” is well known in the art. As used herein, “nucleic acid” generally refers to any molecule (e.g., a chain) of DNA, RNA, or their derivatives or analogs containing nucleotides. A nucleotide consists of a nucleoside and a phosphate group. Examples of nitrogenous bases of nucleosides include naturally occurring purines or pyrimidine nucleosides found in DNA (e.g., adenine “A”, guanine “G”, thymine “T”, or cytosine “C”) or RNA (A, G, uracil “U”, or C).

[0059] The term "nucleic acid molecule" includes, but is not limited to, single-stranded RNA (ssRNA), double-stranded RNA (dsRNA), single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), small RNA, circular nucleic acids, fragments of genomic DNA or RNA, degraded nucleic acids, amplification products, modified nucleic acids, plasmids or organelle nucleic acids, and artificial nucleic acids such as oligonucleotides.

[0060] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequences:

[0061] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 75%, at least 79%, at least 85%, at least 95%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 1. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 75% to 100%, 80% to 100%, 90% to 100%, or 95% to 100% homology or identity with respect to SEQ ID NO: 1. Each possibility represents another embodiment of the present invention.

[0062] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequences:

[0063] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 2. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 80% to 100%, 85% to 100%, 90% to 100%, or 95% to 100% homology or identity with respect to SEQ ID NO: 2. Each possibility represents another embodiment of the present invention.

[0064] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequences:

[0065] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 75%, at least 79%, at least 85%, at least 95%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 3. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 75% to 100%, 80% to 100%, 90% to 100%, or 95% to 100% homology or identity with respect to SEQ ID NO: 3. Each possibility represents another embodiment of the present invention.

[0066] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequences:

[0067] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 91%, at least 92%, at least 95%, at least 97%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 4. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 91% to 100%, 93% to 100%, 95% to 100%, or 97% to 100% homology or identity with respect to SEQ ID NO: 4. Each possibility represents another embodiment of the present invention.

[0068] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequences:

[0069] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 91%, at least 92%, at least 95%, at least 97%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 5. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 91% to 100%, 93% to 100%, 95% to 100%, or 97% to 100% homology or identity with respect to SEQ ID NO: 5. Each possibility represents another embodiment of the present invention.

[0070] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequences:

[0071] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 90%, at least 92%, at least 95%, at least 97%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 6. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 90% to 100%, 92% to 100%, 95% to 100%, or 97% to 100% homology or identity with respect to SEQ ID NO: 6. Each possibility represents another embodiment of the present invention.

[0072] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequences:

[0073] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 77%, at least 79%, at least 85%, at least 95%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 7. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 77% to 100%, 85% to 100%, 90% to 100%, or 95% to 100% homology or identity with respect to SEQ ID NO: 7. Each possibility represents another embodiment of the present invention.

[0074] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequences:

[0075] In some embodiments, the polynucleotide comprises a nucleic acid sequence having homology or identity with respect to SEQ ID NO: 8 of at least 89%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or any value and range in between. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having homology or identity with respect to SEQ ID NO: 8 of 89% to 100%, 92% to 100%, 94% to 100%, or 97% to 100%. Each possibility represents another embodiment of the present invention.

[0076] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequences:

[0077] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 76%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% homology or identity with respect to SEQ ID NO: 9, or any value and range between these. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 76% to 100%, 83% to 100%, 90% to 100%, or 95% to 100% homology or identity with respect to SEQ ID NO: 9. Each possibility represents another embodiment of the present invention.

[0078] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequences:

[0079] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 10. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 75% to 100%, 80% to 100%, 90% to 100%, or 95% to 100% homology or identity with respect to SEQ ID NO: 10. Each possibility represents another embodiment of the present invention.

[0080] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequences:

[0081] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 76%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% homology or identity with respect to SEQ ID NO: 11, or any value and range between these. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 76% to 100%, 85% to 100%, 90% to 100%, or 96% to 100% homology or identity with respect to SEQ ID NO: 11. Each possibility represents another embodiment of the present invention.

[0082] In some embodiments, the polynucleotide comprises or consists of the following nucleic acid sequences:

[0083] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 77%, at least 79%, at least 85%, at least 95%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 23. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 77% to 100%, 85% to 100%, 90% to 100%, or 95% to 100% homology or identity with respect to SEQ ID NO: 23. Each possibility represents another embodiment of the present invention.

[0084] In some embodiments, the polynucleotide of the present invention contains 950 to 1750 nucleotides. In some embodiments, the polynucleotide of the present invention is 1,100 to 1,500 nucleotides long.

[0085] In some embodiments, 950–1,750 nucleotides include at least 970 nucleotides, at least 1,000 nucleotides, at least 1,100 nucleotides, at least 1,150 nucleotides, at least 1,250 nucleotides, at least 1,400 nucleotides, at least 1,500 nucleotides, at least 1,600 nucleotides, or at least 1,730 nucleotides, or any value and range in between. Each possibility represents another embodiment of the invention. In some embodiments, 950–1,750 nucleotides include 950–1,250 nucleotides, 1,100–1,350 nucleotides, 970–1,325 nucleotides, 1,150–1,400 nucleotides, or 1,170–1,490 nucleotides. Each possibility represents another embodiment of the invention.

[0086] In some embodiments, the polynucleotide comprises multiple polynucleotides. In some embodiments, the polynucleotide comprises multiple types of polynucleotides. As used herein, the term “multiple” includes any integer of two or more. In some embodiments, the polynucleotide comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or 8 different nucleic acid sequences, or any value and range between them, each of which is selected from SEQ ID NOs: 1-11 and 23. Each possibility represents another embodiment of the present invention. In some embodiments, the polynucleotide comprises 2-3, 2-4, 2-5, 2-8, 2-11, 3-7, 3-9, 3-11, 4-10, 4-11, 6-8, 6-11, 7-10, 7-11, 8-10, 9-11, or 10-11 different nucleic acid sequences, each of which is selected from SEQ ID NOs: 1-11 and 23.

[0087] In some embodiments, the polynucleotide is a plurality of polynucleotide molecules, or comprises a plurality of polynucleotide molecules, each of which comprises a different nucleic acid sequence, the different nucleic acid sequences being selected from SEQ ID NOs: 1-11 and 23.

[0088] In some embodiments, the polynucleotide encodes a protein characterized by prenyl transfer activity. In some embodiments, the polynucleotide encodes a protein that is a prenyltransferase (PT). In some embodiments, the PT is a PT derived from Helichrysum ambracrigerum. As used herein, the terms “prenyltransferase” and “PT” encompass any enzyme derived from H. ambracrigerum that has, or is, a functional analogue of “geranyl pyrophosphate:olibetolate geranyltransferase” or “GOT” of Cannabis sativa. In some embodiments, GOT is GOT4 or CsGOT4.

[0089] As used herein, the terms “prenyltransferase” and “PT” are interchangeable and refer to any peptide, polypeptide, or protein capable of transferring an allyl prenyl group to a receptor molecule. In some embodiments, PT activity includes cyclization. In some embodiments, PT activity includes transferring an allyl prenyl group to a receptor molecule.

[0090] According to several embodiments, artificial nucleic acid molecules comprising the polynucleotides disclosed herein are provided.

[0091] In some embodiments, the artificial vector comprises a plasmid. In some embodiments, the artificial vector comprises or is an Agrobacterium containing an artificial nucleic acid molecule. In some embodiments, the artificial vector is an expression vector. In some embodiments, the artificial vector is a plant expression vector. In some embodiments, the artificial vector is for use in the expression of PT-coding nucleic acid sequences as disclosed herein. In some embodiments, the artificial vector is for use in heterologous expression of PT-coding nucleic acid sequences as disclosed herein in cells, tissues, or organisms.

[0092] The expression of polynucleotides within cells is well known to those skilled in the art. It can be carried out by transfection, viral infection, or direct modification of the cell genome, among many other methods. In some embodiments, the polynucleotides are contained within an expression vector, such as a plasmid or viral vector. The nucleic acid sequence of the vector generally contains at least one origin of replication for cell proliferation, and optionally additional elements, such as heterologous polynucleotide sequences, expression regulatory elements (e.g., promoters, enhancers), selection markers (e.g., antibiotic resistance), and polyadenine sequences.

[0093] The vector may be a DNA plasmid delivered via a non-viral or viral method. Viral vectors may be retroviral vectors, herpesvirus vectors, adenovirus vectors, adeno-associated virus vectors, Virugaviridae virus vectors, or poxvirus vectors. Barley stripe mosaic virus (BSMV), tobacco rattle virus, and cabbage leaf curl geminivirus (CbLCV) may also be used. The promoter may be active in plant cells. The promoter may be a viral promoter.

[0094] In some embodiments, the polynucleotides disclosed herein are operably ligated to a promoter. The term “operably ligated” means that the nucleotide sequence in question is ligated to a regulatory element or a set of regulatory elements in a manner that enables the expression of the nucleotide sequence (for example, in an in vitro transcription / translation system when the vector is introduced into a host cell, or in the host cell). In some embodiments, the promoter is operably ligated to the polynucleotide of the present invention. In some embodiments, the promoter is a heterologous promoter. In some embodiments, the promoter is an endogenous promoter.

[0095] In some embodiments, vectors are introduced into cells by electroporation (e.g., as described in From et al., Proc. Natl. Acad. Sci. USA 82, 5824 (1985)), heat shock, infection with viral vectors, high-velocity ballistic penetration of nucleic acid-containing particles into or onto the matrix or surface of small beads or particles (Klein et al., Nature 327. 70-73 (1987)), bioristic use of coated particles and needle-shaped particles, etc., and by standard methods including Agrobacterium Ti plasmids and / or similar.

[0096] As used herein, the term "promoter" refers to a group of transcriptional regulatory modules that assemble around the start site of RNA polymerase, i.e., RNA polymerase II. A promoter consists of separate functional modules, each consisting of approximately 7–20 bp of DNA and containing one or more recognition sites for transcriptional activation or repressor proteins. A promoter may extend upstream or downstream of the transcription start site and may be of any size ranging from 2-3 base pairs to several kilobases.

[0097] In some embodiments, polynucleotides are transcribed by RNA polymerase II (RNAP II and Pol II). RNAP II is an enzyme found in eukaryotic cells that is known to catalyze DNA transcription to synthesize mRNA and precursors of most snRNAs and microRNAs.

[0098] In some embodiments, plant expression vectors are used. In one embodiment, polypeptide coding sequence expression is driven by multiple promoters. In some embodiments, viral promoters such as the 35S RNA and 19S RNA promoters of CaMV [Brisson et al., Nature 310:511-514 (1984)], or the coat protein promoter of TMV [Takamatsu et al., EMBO J. 6:307-311 (1987)] are used. In another embodiment, for example, a small subunit of RUBISCO [Coruzzi et al., EMBO J. 3: 1671-1680 (1984); and Brogli et al., Science 224:838-843 (1984)], or a plant promoter such as a heat shock promoter, e.g., soybean hspl7.5-E or hspl7.3-B [Gurley et al., Mol. Cell. Biol. 6:559-565 (1986)], may be used. In one embodiment, the construct is introduced into plant cells using Ti plasmids, Ri plasmids, plant virus vectors, direct DNA transformation, microinjection, electroporation, and other techniques well known to those skilled in the art. See, for example, Weissbach & Weissbach [Methods for Plant Molecular Biology, Academic Press, NY, Section VIII, pp 421-463 (1988)]. Other expression systems, such as insect and mammalian host cell systems, which are well known in the art, can also be used in accordance with the present invention.

[0099] In some embodiments, expression vectors containing regulatory elements from eukaryotic viruses such as retroviruses are used in the present invention. Examples of SV40 vectors include pSVT7 and pMT2. In some embodiments, examples of vectors derived from bovine papillomavirus include pBV-lMTHA, and examples of vectors derived from Epstein-Barr virus include pHEBO and p205. Other exemplary vectors include pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo-5, baculovirus pDSVE, and other vectors that enable protein expression under the direction of the SV-40 early promoter, SV-40 late promoter, metallothionein promoter, murine mammary tumor virus promoter, lotus sarcoma virus promoter, polyhedrin promoter, or other promoters that have been shown to be effective for expression in eukaryotic cells.

[0100] In some embodiments, recombinant viral vectors that offer advantages such as systemic infection and target specificity are used for in vivo expression. In one embodiment, systemic infection is inherent in the life cycle of, for example, retroviruses, where a single infected cell generates many virions, which then infect adjacent cells. In one embodiment, this results in rapid infection of large areas, most of which were not initially infected by the original viral particles. In one embodiment, a viral vector that cannot be transmitted systemically is produced. In one embodiment, this feature may be useful when the desired objective is to introduce a specified gene into only a very limited number of target cells.

[0101] In some embodiments, a plant virus vector is used. In some embodiments, a wild-type virus is used. In some embodiments, a degraded virus known in the art is used. In some embodiments, Agrobacterium is used to introduce the vector of the present invention into a plant.

[0102] Various methods can be used to introduce the expression vector of the present invention into cells. Such methods are generally described in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Springs Harbor Laboratory, New York (1989, 1992), Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989), Chang et al., Somatic Gene Therapy, CRC Press, Ann, Vega et al., Gene Targeting, CRC Press, Ann Arbor Mich. (1995), Vectors: A Survey of Molecular Cloning Vectors and Their Uses, Butterworths, Boston Mass. (1988), and Gilboa et at. [Biotechniques 4 (6): 504-512, 1986], and include, for example, stable or transient transfection, lipofection, electroporation, and infection with Agrobacterium Ti plasmids and recombinant viral vectors. In addition, for information on the selection of positive / negative, please refer to U.S. Patents No. 5,464,764 and No. 5,487,992.

[0103] It will be recognized that, in addition to containing elements necessary for the transcription and translation of the inserted coding sequence (encoding the polypeptide), the expression construct of the present invention may also include sequences manipulated to optimize the stability, generation, purification, yield, or activity of the expressed polypeptide.

[0104] In some embodiments, the artificial vector comprises a polynucleotide encoding a protein containing the amino acid sequence described herein.

[0105] According to several embodiments, proteins encoded by (a) polynucleotides disclosed herein; (b) artificial vectors disclosed herein; or plasmids or Agrobacterium disclosed herein are provided.

[0106] In some embodiments, the protein is encoded by polynucleotides including, or consisting of, sequence numbers 1-11 and 23.

[0107] In some embodiments, the protein comprises an amino acid sequence having at least 92%, at least 93%, at least 95%, at least 97%, or at least 99% homology or identity with any one of SEQ ID NOs. 12-22 and 24.

[0108] In some embodiments, the protein is an isolated protein.

[0109] As used herein, the terms “peptide,” “polypeptide,” and “protein” are interchangeable and refer to polymers of amino acid residues. In other embodiments, the terms “peptide,” “polypeptide,” and “protein” as used herein encompass natural peptides, peptide mimetic (typically including non-peptide bonds or other synthetic modifications), as well as peptide analogs peptoids and semipeptoids, or any combination thereof. In other embodiments, the described peptides, polypeptides, and proteins have modifications that make them more stable in living organisms or that allow for better penetration into cells. In one embodiment, the terms “peptide,” “polypeptide,” and “protein” apply to naturally derived amino acid polymers. In another embodiment, the terms “peptide,” “polypeptide,” and “protein” apply to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of corresponding naturally derived amino acids.

[0110] As used herein, the term “isolated protein” refers to a protein that is essentially free from contaminating cellular components such as carbohydrates, lipids, or other proteinaceous impurities that naturally associate with nucleic acids. Typically, the preparation of an isolated protein involves containing the protein in a highly purified form, for example, at least about 80% pure, at least about 90% pure, at least about 95% pure, more than 95% pure, or more than 99% pure. In some embodiments, the isolated protein is a synthetic protein. Protein synthesis is well known in the art and may be carried out, for example, by heterologous expression in transformed cells as illustrated herein.

[0111] In some embodiments, the protein comprises or consists of the following amino acid sequence: MELSLSSSSSSSLPQLHTHPSSSSSSSHYIKKSPFFINKFNNHTKCKFHNSSALRTNFFYTTITKTSSSRFVLNKNPNQFSVKACSQVGSAGSDPALNKVADFKDAFWRFLRPHTIRGTALGSVSLVTRALLENPNLIRWSLLLKAFSGLVALICGNGYIVGINQIYDIGIDKVNKPYLPIAAGDLSVQSAWFLVLAFAMVGVIIVGMNFGPFITSLYSLGLFLGTIYSVPPLRMKRFPVVAFLIIATVRGFLLNFGVYYAVRAALGLTFQWSSAVAFITTFVTLFALVIAITKDLPDVEGDRKFQISTFATKLGVRNIALLGSGLLLINYIGSIVAALYMPQAFRSSLMIPLHTILASCLIYQAWILERANYTQEAIAGYYRFVWNLFYSEYIIFPFI (Sequence ID 12).

[0112] In some embodiments, the protein comprises an amino acid sequence having at least 82%, at least 85%, at least 90%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 12. Each possibility represents another embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 82% to 100%, 85% to 100%, 90% to 100%, or 95% to 100% homology or identity with respect to SEQ ID NO: 12. Each possibility represents another embodiment of the present invention.

[0113] In some embodiments, the protein comprises or consists of the following amino acid sequence: MATMASSLLNPLSCSIKPNSNRLPLPTPISLSRSCRRLTIKATETDANEVKPKAPEKAPAASGSGFNQILGIKGAKQETNKWKIRVQLTKPVTWPPLIWGVVCGAAASGNFQWTVEDVAKSIVCMLMSGPFLTGYTQTINDWYDRDIDAINEPYRPIPSGAISENEVITQIWVLLLGGIGLAGILDVWAGHKSPTIFYLALGGSLLSYIYSAPPLKLKQNGWIGNFALGASYISLPWWAGQALFGTLTPDIVVLTLLYSIAGLGIAIVNDFKSVEGDRKMGLQSLPVAFGEETAKWICVGAIDITQLSIAGYLLGSGKPYYALALVGLIVPQIFFQFKYFLKDPVKYDVKYQASAQPFLILGLLVTALATSH (Sequence ID 13).

[0114] In some embodiments, the protein comprises an amino acid sequence having at least 92%, at least 93%, at least 95%, at least 97%, at least 98%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 13. Each possibility represents another embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 92% to 100%, 93% to 100%, 95% to 100%, or 97% to 100% homology or identity with respect to SEQ ID NO: 13. Each possibility represents another embodiment of the present invention.

[0115] In some embodiments, the protein comprises or consists of the following amino acid sequence: MKSLIIGSFSNKVSCYSPSLPDSSSSLIPTGCYHVSLRTFQRNRAIQAQSSLVRCNIGKFNETLLLSRKRSTKHVACAVSEQPIEPDATNPQSSLPNALDAFYRFSRPHTVIGTALSIVSVSLLAVQKLSDFSPLFFIGVFEAIVAAFFMNIYIVGLNQLSDIEIDKVNKPYLPLASGEYSVQTGIIIVSSFAVMSFWLGWIVGSWPLFWALFISFLLGTAYSINIPMLRWKRFALVAAMCILAVRAIIVQVAFYLHIQTFVYGRLAVFPKPVIFATGFMSFFSVVIALFKDIPDIVGDKIFGIQSFTVRMGQKRVFWICILLLEIAYGVAILVGASSPFLWSRYITVLGHAILGLILWGRAKSTDLESKSAITSFYMFIWQLFYAEYLLIPLVR (Sequence ID 14).

[0116] In some embodiments, the protein comprises an amino acid sequence having at least 89%, at least 90%, at least 95%, or at least 99% homology or identity with respect to SEQ ID NO: 14, or any value and range in between. Each possibility represents another embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 89% to 100%, 92% to 100%, 94% to 100%, or 96% to 100% homology or identity with respect to SEQ ID NO: 14. Each possibility represents another embodiment of the present invention.

[0117] In some embodiments, the protein comprises or consists of the following amino acid sequence: MELSLSSSSSSSLPQLHTHPSSSSSSSHYIKKSPFFINKFNNHTKCKFHNSSALRTNFFYTTITKTSSSRFVLNKNPNQFSVKACSQVGSAGSDPALNKVADFKDAFWRFLRPHTIRGTALGSVSLVTRALLENPNLIRWSLLLKAFSGLVALICGNGYIVGINQIYDIGIDKVNKPYLPIAAGDLSVQSAWFLVLAFAMVGVIIVGMNFGPFITSLYSLGLFLGTIYSVPPLRMKRFPVVAFLIIATVRGFLLNFGVYYAVRAALGLTFQWSSAVAFITTFVTLFALVIAITKDLPDVEGDRKFQISTFATKLGVRNIALLGSGLLLINYIGSIVAALYMPQAFRSSLMIPLHTILASCLIYQAWILERANYTQRSQYFDMSSCRRR (Sequence ID 15).

[0118] In some embodiments, the protein comprises an amino acid sequence having at least 81%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 15. Each possibility represents another embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 81% to 100%, 85% to 100%, 88% to 100%, or 93% to 100% homology or identity with respect to SEQ ID NO: 15. Each possibility represents another embodiment of the present invention.

[0119] In some embodiments, the protein comprises or consists of the following amino acid sequence: MELSLSSSSSSSLPQLHTHPSSSSSSSHYIKKSPFFINKFNNHTKCKFHNSSALRTNFFYTTITKTSSSRFVLNKNPNQFSVKACSQVGSAGSDPALNKVADFKDAFWRFLRPHTIRGTALGSVSLVTRALLENPNLIRWSLLLKAFSGLVALICGNGYIVGINQIYDIGIDKVNKPYLPIAAGDLSVQSAWFLVLAFAMVGVIIVGMNFGPFITSLYSLGLFLGTIYSVPPLRMKRFPVVAFLIIATVRGFLLNFGVYYAVRAALGLTFQWSSAVAFITTFVTLFALVIAITKDLPDVEGDRKFQISTFATKLGVRNIALLGSGLLLINYIGSIVAALYMPQVKTTSIDHYRPYSFLVDLPGQNGITLAA (Sequence ID 16).

[0120] In some embodiments, the protein comprises an amino acid sequence having at least 81%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 16. Each possibility represents another embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 81% to 100%, 85% to 100%, 90% to 100%, or 95% to 100% homology or identity with respect to SEQ ID NO: 16. Each possibility represents another embodiment of the present invention.

[0121] In some embodiments, the protein comprises or consists of the following amino acid sequence: MATMASSLLNPLSCSIKPNSNRLPLPLPIPISLSRSCRRLTIKATETDANEVKPKAPEKAPAASGSGFNQILGIKGAKQETNKWKIRVQLTKPVTWPPLIWGVVCGAAASGNFQWTVEDVAKSIVCMLMSGPFLTGYTQTINDWYDRDIDAINEPYRPIPSGAISENEVITQIWVLLLGGIGLAGILDVWAGHKSPTIFYLALGGSLLSYIYSAPPLKLKQNGWIGNFALGASYISLPWWAGQALFGTLTPDIVVLTLLYSIAGLGIAIVNDFKSVEGDRKMGLQSLPVAFGEETAKWICVGAIDITQLSIAGYLLGSGKPYYALALVGLIVPQIFFQFKYFLKDPVKYDVKYQASAQPFLILGLLVTALATSH (Sequence ID 17).

[0122] In some embodiments, the protein comprises an amino acid sequence having at least 92%, at least 93%, at least 95%, at least 97%, at least 98%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 17. Each possibility represents another embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 92% to 100%, 93% to 100%, 96% to 100%, or 98% to 100% homology or identity with respect to SEQ ID NO: 17. Each possibility represents another embodiment of the present invention.

[0123] In some embodiments, the protein comprises or consists of the following amino acid sequence: MASLAIGSLGSPSSRQCSSPVASSSSFAIGSQIASKFLRISKFDKTKNSPLTLQQKHINKSIDQSFFEPLPLHKINKDKFKLYATSTNNPQFDATHDLKTPEVSIINFVDALYRLIRPYTAVVTIVSVVAMSLLTVNSLSDFSPLFFIKVVQALIGGIFMQMYVSGFNQICDIELDKVNKQSLPLAAGELSMKTAIVIASLSAIMSLSIGWFVGSPPLLWCLVWWFIVGTAYSANVLPYLRWKRFPFTAAFCAMTSRALVLPIGYYLHMQNSIPGVSALLSRPILFAVAMLSAFSLSAMFFKDIPDIKGDRMHGIKSLAIKLGEKRVYWISISIIEIAYIAAAFIGATSPISWSKYVTIIGHLGMGLLLWVRARSVDPTNTVAVQSMYMFLIKLVYAEYGLISLVR (Sequence ID 18).

[0124] In some embodiments, the protein comprises an amino acid sequence having at least 71%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 18. Each possibility represents another embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 71% to 100%, 75% to 100%, 80% to 100%, or 90% to 100% homology or identity with respect to SEQ ID NO: 18. Each possibility represents another embodiment of the present invention.

[0125] In some embodiments, the protein comprises or consists of the following amino acid sequence: MKSLIIGSFSNKVSCYSPSLPDSSSSLIPTGCYHVSLRTFQRNRAIQAQSSLVRCNIGKFNETLLLSRKRSTKHVACAVSEQPIEPDATNPQSSLPNALDAFYRFSRPHTVIGTALSIVSVSLLAVQKLSDFSPLFFIGVFEAIVAAFFMNIYIVGLNQLSDIEIDKVNKPYLPLASGEYSVQTGIIIVSSFAVMSFWLGWIVGSWPLFWALFISFLLGTAYSINIPMLRWKRFALVAAMCILAVRAIIVQVAFYLHIQTFVYGRLAVFPKPVIFATGFMSFFSVVIALFKDIPDIVGDKIFGIQSFTVRMGQKRVFWICILLLEIAYGVAILVGASSPFLWSRYITVLGHAILGLILWGRAKSTDLESKSAITSFYMFIWQLFYAEYLLIPLVR (Sequence ID 19).

[0126] In some embodiments, the protein comprises an amino acid sequence having at least 89%, at least 90%, at least 95%, or at least 99% homology or identity with respect to SEQ ID NO: 19, or any value and range between those. Each possibility represents another embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 89% to 100%, 92% to 100%, 95% to 100%, or 97% to 100% homology or identity with respect to SEQ ID NO: 19. Each possibility represents another embodiment of the present invention.

[0127] In some embodiments, the protein comprises or consists of the following amino acid sequence: MLIHHEHFLTTGFESSNDRAAYSINFSKQHHLHMASIATGSLCRPTSHQFSIPVASSSSFATGSQFASKFLHISISAKKSSLTLQQRHIHKNIDQSFLKPLALQKLNKDKF KLNGTSPDNPQFDATHDLKTQIESTINFVDVLYRLLRPYALLQMGLCVVTMSLLTVESLSDFSPLFFVKVAQALIGGIFMQMYVNGFNQICDIELDKVNKPSLPLASGELS KTTTIVVSSLSAITSLSIGWFVGSPPLLWSLVVWFIAGTTYSANLPYLRWKRFPFTNMFCNLTMALVVPIGTYLHMENSIHGVSTLLSRPLLFTVAMCTVFPVSIILFKDIPDIKGDRMHGMKSLAIILGEKRTYWICIWILEITYIAAAFFGATSPISWSKYVTIISHLGMGFLLWLRSKSVDVKNTVAVQSMYMFLWKLLYAEYGLILLVR (Sequence ID 20).

[0128] In some embodiments, the protein comprises an amino acid sequence having at least 68%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% homology or identity with respect to SEQ ID NO: 20, or any value and range between these. Each possibility represents another embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 68% to 100%, 75% to 100%, 80% to 100%, or 90% to 100% homology or identity with respect to SEQ ID NO: 20. Each possibility represents another embodiment of the present invention.

[0129] In some embodiments, the protein comprises or consists of the following amino acid sequence: MFIHHEQFLTTGFESSNDRAAYSINFLKQHHLHMVSIATGSLCRPTSHRFSIPVASSSSFATGSQFASISAKKSSLTLKQRHTHKNIDQSFFKPLALQKMNKGKFKLNA TSPDNSQLDATHDLKTQIESIINFVDVLYRLIRPYVVLGMGVTIVTMCLLTVDSLSDFSPLFFVKVAQALIGSIFMAMYVNSFNEICDIELDKVNKPSLPLASGELSMTT AIVVSSLSAIMSLSIGWFVGSPPLLWSLVVWFILGTAYSANLPYLRWKRFPLTTLSSALTMGALVIPIGNYMHMENSIRGVTTLLSRPLLFAVAMCAAFHVSTILFKDIPDIKGDRMHGMKSLAIKLGEKRMYWICIWILEIAYIAAAFFGATSPISWSKYVTIISHLGMGFLLWLRSKSVDVKNTVAVQSMYMFLWKLFYVEHGLILLVR (Sequence ID 21).

[0130] In some embodiments, the protein comprises an amino acid sequence having at least 66%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% homology or identity with respect to SEQ ID NO: 21, or any value and range in between. Each possibility represents another embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 66% to 100%, 75% to 100%, 85% to 100%, or 90% to 100% homology or identity with respect to SEQ ID NO: 21. Each possibility represents another embodiment of the present invention.

[0131] In some embodiments, the protein comprises or consists of the following amino acid sequence: MASIATGSLCRPTSHRFSIHVASSSSFATGSQFASKILQISISAKKSSLTLQQRHIHKNIDQSFFKPLALQKMNKDKFKLNATSPDNPQFDATRDLKTQIESIIKFVDVLYRLLRPYAILEMGLSVVTMSLLTVESLSDFSPLFFVKVAQALIGGIFMQMYVNGFNQICDIELDKVNKPSLPLASGELSTTTTIVVSSLSAIMSLSIGWFVGSPPLLWSLVVWFIVGTTYSTNLPYLRWKRFPFTAMFCNLTRALVVPIGTYLHMKNSIHEVSTLLSRPLLFAVAMCTVFPISIILFKDIPDIKGDRMHGMKSLAIILGEERTYWICIWILEIAYIAAAFFGATSPISWSKYVMIISHLGMGFLLWLRSKSVDVKNTVAVQSMYMFLWKLLYAEYGLILLVR (Sequence ID 22).

[0132] In some embodiments, the protein comprises an amino acid sequence having at least 68%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% homology or identity with respect to SEQ ID NO: 22, or any value and range in between. Each possibility represents another embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 68% to 100%, 75% to 100%, 85% to 100%, or 95% to 100% homology or identity with respect to SEQ ID NO: 22. Each possibility represents another embodiment of the present invention.

[0133] In some embodiments, the protein comprises or consists of the following amino acid sequence: MASLAIGSLGSPSSRQCSSPVASSSSFAIGSQIASKFLRISKFDKTKNSPLALQQKHINKSIDQSFFEPLPLHKINKDKFKLYATSTNNPQFDATHDLKTPEVSIINFVDALYRLIRPYTAVVTIVSVVAMSLLTVNSLSDFSPLFFIKVVQALIGGIFMQMYVSGFNQICDIELDKVNKQSLPLAAGELSMKTAIVIASLSAIMSLSIGWFVGSPPLLWCLVWWFIVGTAYSANVLPYLRWKRFPFTAAFCAMTSRALVLPIGYYLHMQNSIPGVSALLSRPILFAVAMLSAFSLSAMFFKDIPDIKGDRMHGIKSLAIKLGEKRVYWISISIIEIAYIAAAFIGATSPISWSKYVTIIGHLGMGLLLWVRARSVDPTNTVAVQSMYMFLIKLVYAEYGLISLVR (Sequence ID 24).

[0134] In some embodiments, the protein comprises an amino acid sequence having at least 71%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99%, or any value and range in between, homology or identity with respect to SEQ ID NO: 24. Each possibility represents another embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 71% to 100%, 80% to 100%, 90% to 100%, or 95% to 100% homology or identity with respect to SEQ ID NO: 24. Each possibility represents another embodiment of the present invention.

[0135] The terms “homology” and “identity,” used interchangeably herein, refer to sequence identity between two amino acid sequences or between two nucleic acid sequences, while identity is a more rigorous comparison. The phrases “percent identity or homology” and “% identity or homology” refer to the percentage of sequence identity found in a comparison of two or more amino acid sequences or nucleic acid sequences. Two or more sequences may be % identical to any value between 0% and 100%, or any value in between. Identity can be determined by comparing the positions of each sequence that can be aligned for comparison with a reference sequence. If the positions of the compared sequences are occupied by the same nucleotide base or amino acid, then the molecules are identical at those positions. The degree of identity of amino acid sequences is a function of the number of identical amino acids at positions shared by the amino acid sequences. The degree of identity between nucleic acid sequences is a function of the number of identical or matching nucleotides at positions shared by the nucleic acid sequences. The degree of homology of amino acid sequences is a function of the number of amino acids at positions shared by polypeptide sequences.

[0136] The following are non-restrictive examples for calculating homology or sequence identity between two sequences (these terms are used interchangeably herein): The sequences are aligned for optimal comparison (for example, gaps may be introduced in one or both of the first and second amino acid or nucleic acid sequences for optimal alignment, and non-homologous sequences can be ignored for comparison purposes). Optimal alignment is determined as the best score using the GAP program in the GCG software package with a Blossum 62 score matrix with a gap penalty 12, a gap extension penalty 4, and a frameshift gap penalty 5. The amino acid residues or nucleotides at corresponding amino acid or nucleotide positions are then compared. If a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, the molecules are identical at that position. Percent identity between two sequences is a function of the number of identical positions shared by the sequences.

[0137] In some embodiments, the % homology or identity described herein is calculated or determined using the Basic Local Alignment Search Tool (BLAST). In some embodiments, the % homology or identity described herein is calculated or determined using the Blossum62 score matrix.

[0138] In some embodiments, the protein contains or is characterized by prenyl transfer activity as described herein. In some embodiments, the protein is characterized by its ability to transfer a prenyl group to a substrate molecule. In some embodiments, the protein is characterized by its ability to transfer an allyl prenyl group to an acceptor molecule. In some embodiments, the protein is a prenyl diphosphate synthase. In some embodiments, the protein is a trans-prenyltransferase. In some embodiments, the protein is a cis-prenyltransferase.

[0139] In some embodiments, the prenyl group is selected from dimethylallyl diphosphate, geranyl diphosphate, farnesyl diphosphate, or geranylgeranyl diphosphate.

[0140] In some embodiments, the substrate molecule is of formula I: [ka] (The formula is represented by (i) R1 is selected from C1-C8 alkyl, alpha-unsaturated phenylalkyl carboxylic acid, or alpha-saturated phenylalkyl carboxylic acid, and R2 is OH, or (ii) R1 is OH, and R2 is selected from C1-C8 alkyl, alpha-unsaturated phenylalkyl carboxylic acid, or alpha-saturated phenylalkyl carboxylic acid).

[0141] In some embodiments, the alpha-unsaturated phenylalkyl carboxylic acid includes cinnamic acid or a derivative thereof.

[0142] In some embodiments, the cinnamic acid derivative is a hydroxylated derivative of cinnamic acid, or comprises such hydroxylated derivative.

[0143] In some embodiments, the hydroxylated derivative of cinnamic acid is coumaric acid, or comprises coumaric acid.

[0144] According to several embodiments, transgenic cells are provided comprising (a) polynucleotides disclosed herein; (b) artificial nucleic acid molecules disclosed herein; (c) plasmids or Agrobacterium disclosed herein; (d) isolated proteins disclosed herein; or any combination thereof.

[0145] As used herein, the term “transgenic cell” refers to any human cell that has been manipulated at the genomic or genetic level. In some embodiments, transgenic cells are cells into which an exogenous polynucleotide, such as an isolated DNA molecule disclosed herein, has been introduced. In some embodiments, transgenic cells include cells having an artificial vector introduced into the cell. In some embodiments, transgenic cells are cells that have undergone genomic mutation or modification. In some embodiments, transgenic cells are cells that have undergone CRISPR genome editing. In some embodiments, transgenic cells are cells that have undergone targeted mutation of at least one base pair in their genome. In some embodiments, the exogenous polynucleotide (e.g., an isolated DNA molecule disclosed herein) or vector is stably incorporated into the cell. In some embodiments, transgenic cells express the polynucleotide of the present invention. In some embodiments, transgenic cells express the vector of the present invention. In some embodiments, transgenic cells express the protein of the present invention. In some embodiments, transgenic cells are cells lacking the polynucleotides of the present invention that have been transformed to contain or genetically modified to contain the polynucleotides of the present invention. In some embodiments, CRISPR technology is used to modify the genome of the cells as described herein.

[0146] In some embodiments, cells are single-celled organisms, cells of multicellular organisms, and cells in cultures.

[0147] In some embodiments, the single-celled organism includes fungi or bacteria.

[0148] In some embodiments, the fungus is a yeast cell.

[0149] In some embodiments, the cells are insect cells. In some embodiments, the cells include insect cell lines.

[0150] Types of insect cell lines suitable for transformation and / or heterologous expression are common and will be obvious to those skilled in the art. Non-limiting examples of such insect cell lines include, but are not limited to, Sf-9 cells, SR+ Schneider cells, S2 cells, and others.

[0151] According to several embodiments, extracts derived from transgenic cells or any fraction thereof as disclosed herein are provided.

[0152] In some embodiments, the extract comprises the polynucleotides of the present invention, isolated DNA molecules disclosed herein, isolated proteins disclosed herein, or any combination thereof.

[0153] According to some embodiments, homogenates, lysates, extracts, or any combination thereof derived from transgenic cells disclosed herein, or any fraction thereof, are provided.

[0154] Methods and / or means for extracting, lysing, homogenizing, fractionating, or any combination thereof of cells or their cultures are common and will be obvious to those skilled in the art of cell biology and biochemistry. Non-limiting examples include, but are not limited to, pressure digestion (e.g., using a French press), enzymatic digestion, soluble-insoluble phase separation (e.g., to obtain supernatant and pellet), digestion based on washing agents, solvents (e.g., polar or nonpolar solvents), liquid chromatography-mass spectrometry, or others.

[0155] According to several embodiments, transgenic plants, transgenic plant tissues, or plant parts are provided. In several embodiments, transgenic plants, or any part thereof, seeds, tissues, or organs comprising at least one transgenic plant cell of the present invention are provided. In several embodiments, the transgenic plants, transgenic plant tissues, or plant parts comprise (a) polynucleotides disclosed herein; (b) artifacts disclosed herein; (c) plasmids or Agrobacterium disclosed herein; (d) isolated proteins of the present invention; (e) transgenic cells disclosed herein; or any combination thereof.

[0156] In some embodiments, a transgenic plant, transgenic plant tissue, or plant portion comprises the transgenic plant cells of the present invention. In some embodiments, a transgenic plant, transgenic plant tissue, or plant portion contains at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99%, or any value and range in between, of the transgenic cells of the present invention. Each possibility represents another embodiment of the present invention. In some embodiments, a transgenic plant, transgenic plant tissue, or plant portion contains 20%-50%, 20%-60%, 20%-70%, 20%-80%, 20%-90%, or 20%-100% of the transgenic cells of the present invention. Each possibility represents another embodiment of the present invention.

[0157] In some embodiments, the transgenic plant, transgenic plant tissue, or plant part is a Cannabis sativa plant or derived from a Cannabis sativa plant. In some embodiments, the transgenic plant is a C. sativa plant.

[0158] In some embodiments, the transgenic plant, transgenic plant tissue, or plant part is hemp or derived from hemp. In some embodiments, C. sativa contains hemp or is hemp.

[0159] According to several embodiments, a composition is provided comprising (a) polynucleotides of the present invention (e.g., isolated DNA molecules); (b) artificial vectors; (c) plasmids or Agrobacterium; (d) isolated proteins of the present invention; (e) transgenic cells; (f) extracts; (g) transgenic plant tissue or plant parts; and (h) any combination of (a) to (g) and an acceptable carrier.

[0160] As used herein, the terms “carrier,” “excipient,” or “adjuvant” refer to any component of a composition, such as a pharmaceutical or dietary supplement, that is not an activator. As used herein, the term “pharmaceutically acceptable carrier” refers to a non-toxic and inert solid, semi-solid, or liquid filler, diluent, encapsulating material, any type of formulation aid, or simply a sterile aqueous medium such as saline solution. Some examples of materials that can serve as pharmaceutically acceptable carriers include sugars such as lactose, glucose, and sucrose; starches such as corn starch and potato starch; celluloses and their derivatives such as sodium carboxymethylcellulose, ethylcellulose, and cellulose acetate; tragacanth powder; malt; gelatin; talc; excipients such as cocoa butter and suppository waxes; oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; glycols such as propylene glycol; polyols such as glycerin, sorbitol, mannitol, and polyethylene glycol; esters such as ethyl oleate and ethyl laurate; agar; buffers such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline, Ringer's solution; ethyl alcohol and phosphate buffer solution; and other non-toxic, suitable substances used in pharmaceutical formulations. Some non-limiting examples of substances that may serve as carriers as described herein include sugars, starches, cellulose and their derivatives, tragacanth powder, malt, gelatin, talc, stearic acid, magnesium stearate, calcium sulfate, vegetable oils, polyols, alginic acid, pyrogen-free water, isotonic saline, phosphate buffer solution, cocoa butter (suppository base), emulsifiers (e.g., carbomer, hydroxypropylcellulose, sodium lauryl sulfate), and other non-toxic pharmaceutically acceptable substances used in other pharmaceutical formulations. In addition to wetting and lubricating agents such as sodium lauryl sulfate, colorants, flavoring agents, excipients, stabilizers, antioxidants, and preservatives may also be present. Any non-toxic, inert, and effective carrier may be used to formulate the compositions intended herein.In this regard, suitable pharmaceutically acceptable carriers, excipients, and diluents are well known to those skilled in the art, for example, those described in The Merck Index, Thirteenth Edition, Budavari et al., Eds., Merck & Co., Inc., Rahway, NJ (2001); the CTFA (Cosmetic, Toiletry, and Fragrance Association) International Cosmetic Ingredient Dictionary and Handbook, Tenth Edition (2004); and the “Inactive Ingredient Guide,” US Food and Drug Administration (FDA) Center for Drug Evaluation and Research (CDER) Office of Management, all of which are incorporated herein by reference as a whole. Examples of pharmaceutically acceptable excipients, carriers, and diluents useful in the compositions of the present invention include distilled water, physiological saline, Ringer's solution, dextrose solution, Hanks' solution, and DMSO. In addition to these additional inactive ingredients, effective formulations and administration procedures are well known in the art and are incorporated herein by reference in standard textbooks such as Goodman and Gillman's: The Pharmacological Bases of Therapeutics, 8th Ed., Gilman et al. Eds. Pergamon Press (1990); Remington's Pharmaceutical Sciences, 18th Ed., Mack Publishing Co., Easton, Pa. (1990); and Remington: The Science and Practice of Pharmacy, 21st Ed., Lippincott Williams & Wilkins, Philadelphia, Pa., (2005).The compositions described herein may also be contained within artificially constructed structures such as liposomes, ISCOMS, sustained-release particles, and other vehicles that increase the half-life of peptides or polypeptides in serum. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and similar structures. Liposomes for use with the peptides described herein are generally formed from standard vesicle-forming lipids, including neutral and negatively charged phospholipids, as well as sterols such as cholesterol. The choice of lipids is generally determined considering factors such as liposome size and stability in blood. Various methods are available for preparing liposomes, as discussed, for example, by Coligan, JE et al, Current Protocols in Protein Science, 1999, John Wiley & Sons, Inc., New York; see also U.S. Patents 4,235,871, 4,501,728, 4,837,028 and 5,019,369.

[0161] The carrier may constitute, in total, about 0.1% to about 99.99999% by weight of the pharmaceutical composition presented herein.

[0162] Synthesis method According to several embodiments, Formula II: [ka] A method is provided for synthesizing compounds represented by the formula (i) R1 is selected from C1-C8 alkyl, alpha-unsaturated phenylalkyl carboxylic acid, or alpha-saturated phenylalkyl carboxylic acid, and R2 is OH, or (ii) R1 is OH, R2 is selected from C1-C8 alkyl, alpha-unsaturated phenylalkyl carboxylic acid, or alpha-saturated phenylalkyl carboxylic acid, R3 is a prenyl group, and R4 is hydrogen or a prenyl group).

[0163] According to several embodiments, the method comprises (a) providing cells comprising an artificial vector comprising a nucleic acid sequence having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 97%, or at least 99%, or any value and range between thereof, or any combination thereof, to any one of SEQ ID NOs: 1-11 and 23; and (b) culturing the cells from step (a) to thereby express the protein encoded by the artificial vector, thereby synthesizing a compound represented by formula II. Each possibility represents another embodiment of the present invention.

[0164] According to some embodiments, the method is based on formula I: [ka] The present invention comprises contacting a substrate molecule represented by formula (i) R1 is selected from C1-C8 alkyl, alpha-unsaturated phenylalkyl carboxylic acid, or alpha-saturated phenylalkyl carboxylic acid and R2 is OH, or (ii) R1 is OH and R2 is selected from C1-C8 alkyl, alpha-unsaturated phenylalkyl carboxylic acid, or alpha-saturated phenylalkyl carboxylic acid) with an effective amount of a protein containing an amino acid sequence having homology or identity of at least 92%, at least 93%, at least 95%, at least 99%, or 100%, or any value and range between thereof, with any one of SEQ ID NOs: 12-22 and 24, thereby synthesizing a compound represented by formula II. Each possibility represents another embodiment of the present invention.

[0165] According to several embodiments, a method is provided for obtaining an extract from transgenic cells or transfected cells.

[0166] In some embodiments, the method includes culturing transgenic cells or transfected cells in a culture medium and extracting the transgenic cells or transfected cells.

[0167] In some embodiments, the method comprises (a) culturing transgenic cells or transfected cells in a culture medium; and (b) extracting the transgenic cells or transfected cells to obtain an extract from the transgenic cells or transfected cells.

[0168] In some embodiments, the transgenic or transfected cells include an artificial vector comprising a nucleic acid sequence having at least 91%, at least 93%, at least 95%, at least 97%, at least 99%, or 100%, or any value and range between them, homology or identity to any one of SEQ ID NOs. 1-11 and 23, or any combination thereof. Each possibility represents another embodiment of the present invention.

[0169] In some embodiments, the transgenic cells or transfected cells comprise the polynucleotides or a plurality thereof of the present invention, as disclosed herein.

[0170] In some embodiments, the transgenic cells or transfected cells include artificial nucleic acid molecules or vectors as disclosed herein.

[0171] In some embodiments, the cells are transgenic cells having the isolated DNA disclosed herein, or cells transfected with the isolated DNA.

[0172] In some embodiments, culturing involves supplying cells with an effective amount of a substrate molecule represented by formula I. In some embodiments, supplying is This occurs through the growth or culture medium in which cells are cultured.

[0173] In some embodiments, the substrate molecule is selected from resorcinoid precursors, stilbenic acid precursors, acylphloroglucinoid precursors, or chalcone precursors.

[0174] In some embodiments, the substrate molecule is [ka] It is represented by a formula selected from the following: (wherein R3 is a C1-C8 alkyl group, and R4 is an alpha-unsaturated phenylalkyl carboxylic acid or an alpha-saturated phenylalkyl carboxylic acid).

[0175] In some embodiments, the substrate molecule is [ka] Selected from.

[0176] In some embodiments, the substrate molecule is [ka] That is the case.

[0177] In some embodiments, the compound represented by formula II is selected from cannabinoids, amorphurtins, acylphlorogluconoids, or prenyl chalcones.

[0178] In some embodiments, the compound represented by formula II is [ka] (wherein R1 is a C1-C8 alkyl group, R2 is an alpha-unsaturated phenylalkyl carboxylic acid or an alpha-saturated phenylalkyl carboxylic acid, R3 is a prenyl group, and R4 is hydrogen or a prenyl group) is selected from these.

[0179] In some embodiments, the prenyl group is selected from dimethylallyl diphosphate, geranyl diphosphate, farnesyl diphosphate, or geranylgeranyl diphosphate.

[0180] In some embodiments, the compound represented by formula II is [ka] Selected from.

[0181] In some embodiments, the compound is [ka] That is the case.

[0182] In some embodiments, the method further includes a step preceding step (a) which comprises introducing or transfecting cells with an artificial nucleic acid molecule or vector disclosed herein.

[0183] Methods for introducing or transfecting cells with artificial nucleic acid molecules or vectors are common and would be obvious to those skilled in the art.

[0184] In some embodiments, introduction or transfection includes transferring an artificial nucleic acid molecule or vector containing the polynucleotides disclosed herein into a cell; or modifying the genome of a cell to contain the polynucleotides disclosed herein. In some embodiments, transfer includes transfection. In some embodiments, transfer includes transformation. In some embodiments, transfer includes lipofection. In some embodiments, transfer includes nucleofection. In some embodiments, transfer includes viral infection.

[0185] The terms "transfect" and "introduce" as used herein are interchangeable.

[0186] In some embodiments, contact occurs in a cell-free system.

[0187] Suitable cell-free systems for utilizing any one of the polynucleotides or a plurality of the present invention as disclosed herein, and any one of the isolated proteins or a plurality of the present invention, will be obvious to those skilled in the art.

[0188] In some embodiments, the method further includes a step preceding step (b), which includes separating cultured transgenic cells or cultured transfected cells from the culture medium.

[0189] Methods for separating cells from the culture medium are common and obvious to those skilled in the art, and may include, but are not limited to, centrifugation, ultracentrifugation, or other methods.

[0190] According to some embodiments, extracts of transgenic or transfected cells obtained according to the methods disclosed herein are provided.

[0191] According to some embodiments, a culture medium or a portion thereof is provided that is isolated from cultured transgenic cells or cultured transfected cells obtained according to the methods disclosed herein.

[0192] According to several embodiments, a composition is provided comprising (a) an extract disclosed herein; (b) a culture medium disclosed herein or a portion thereof; or (c) any combination of (a) and (b) and an acceptable carrier, as described herein.

[0193] In some embodiments, some include fractions or a plurality of fractions.

[0194] If a range of values ​​is provided, it is understood that, unless the context otherwise explicitly indicates, each intervention value up to 10 times the lower limit, the range between the upper and lower limits, and any other stated values ​​or intervention values ​​within the stated range are included in the invention. The upper and lower limits of these smaller ranges may independently be included within smaller ranges, which are also included in the invention and subject to any specifically excluded limitations within the stated range. If a stated range includes one or both of the limitations, the range excluding one or both of those included limitations is also included in the invention.

[0195] In this specification, the term "approximately" when used in combination with a value refers to a range of ±10% of the reference value. For example, a length of approximately 1,000 nanometers (nm) refers to a length of 1,000 nm ± 100 nm.

[0196] As used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context otherwise explicitly indicates otherwise. Thus, for example, a reference to “polynucleotide” includes a plurality of such polynucleotides, and a reference to “polypeptide” includes references to one or more polypeptides and their equivalents known to those skilled in the art, and so on. Furthermore, it should be noted that the claims may be drafted to exclude any element by choice. This statement is therefore intended to serve as a prerequisite for the use of exclusive terminology such as “only,” “only,” etc., along with the use of enumeration of elements in the claims or “negative” limitations.

[0197] In examples where a similar convention to "at least one of A, B, and C" is used, such constructions are generally intended to be understood by those skilled in the art (for example, "a system having at least one of A, B, and C" includes, but is not limited to, systems having only A, only B, only C, together with A and B, together with A and C, together with B and C, and / or systems having A, B, and C together). Furthermore, it will be understood by those skilled in the art that virtually any separate word and / or phrase indicating two or more alternative terms, whether in a patent specification, claims, or drawings, must be understood to contemplate the possibility of including one of the terms, either of the terms, or both of the terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B".

[0198] For clarity, it will be recognized that certain features of the Invention described in the context of a different embodiment may also be provided in combination with one embodiment. Conversely, for brevity, various features of the Invention described in the context of one embodiment may also be provided individually or in any suitable partial combination. All combinations of embodiments relating to the Invention are specifically encompassed and disclosed herein as if every single combination were disclosed individually and explicitly. In addition, all partial combinations of various embodiments and their elements are also specifically encompassed and disclosed herein as if every single such partial combination were disclosed individually and explicitly.

[0199] Additional objectives, advantages, and novel features of the present invention will become apparent to those skilled in the art through the examination of the following embodiments, but the embodiments are not intended to be limiting. In addition, each of the various embodiments and aspects of the present invention described earlier in this specification and claimed in the following claims section will find experimental support in the following examples.

[0200] The various embodiments and aspects of the present invention described earlier in this specification and claimed in the following claims section will be experimentally supported in the following examples. [Examples]

[0201] The nomenclature used herein and the laboratory procedures used in this invention generally include molecular, biochemical, microbiological, and recombinant DNA techniques. Such techniques are fully described in the literature. For example, all of the following are incorporated by reference: "Molecular Cloning: A Laboratory Manual" Sambrook et al., (1989); "Current Protocols in Molecular Biology" Volumes I-III Ausubel, RM, ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology", John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning", John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA", Scientific American Books, New York; Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series", Vols. 1-4, Cold Spring Harbor Laboratory Press, New York (1998); Methodologies described in U.S. Patents No. 4,666,828; No. 4,683,202; No. 4,801,531; No. 5,192,659 and No. 5,272,057; “Cell Biology: A Laboratory Handbook”, Volumes I-III Cellis, JE, ed. (1994); “Culture of Animal Cells - A Manual of Basic Technique” by Freshney, Wiley-Liss, NY (1994), Third Edition; “Current Protocols in Immunology” Volumes I-III Coligan, JE, ed.(1994); See Stites et al. (eds), “Basic and Clinical Immunology” (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), “Strategies for Protein Purification and Characterization - A Laboratory Course Manual” CSHL Press (1996). Other general references are provided throughout this document.

[0202] material and method UPLC-qTOF analysis of cannabigerol acid (CBGA) Fresh samples of six different tissues—young leaves, old leaves, tubular florets and receptacles, stems, and roots—were collected from the plants at the flowering stage. The tubular florets and receptacles were dissected using a surgical scalpel and extracted separately. All tissues were rapidly frozen in liquid N2, ground into a fine powder with a mortar and pestle, and extracted with 1 ml of ethanol as previously described.

[0203] Samples were analyzed using a high-resolution ultrafast liquid chromatography-tandem quadrupole time-of-flight (UPLC-qTOF) system consisting of a UPLC (Waters Acquity) with a diode array detector coupled to either a XEVO G2-S QTof (Waters) or Synapt HDMS (Waters). Chromatographic separation of compounds was performed on a 100 mm × 2.1 mm id (inner diameter), 1.7 μm UPLC BEH C18 column (Waters Acquity). The mobile phase consisted of 0.1% formic acid in acetonitrile:water (5:95 v / v) (phase A) and 0.1% formic acid in acetonitrile (phase B). The flow rate was 0.3 ml / min, and the column temperature was maintained at 35°C. Plant extracts were analyzed using a 29-minute multi-step gradient method, with an initial condition of 40% B for 1 minute, increased to 100% B by 23 minutes, held at 100% B for 3.8 hours, decreased to 40% B by 27 minutes, and held at 40% B until 29 minutes for system reequilibrium. Products from the enzyme assay were analyzed in a shorter second step (B increased from 40% to 100% over 13 minutes). Electrospray ionization (ESI) was used for anionization in the m / z range of 50–1,000 Da. The mass of the eluted compound was detected under the following settings: capillary 1 kV, source temperature 140 °C, desolvation temperature 450 °C, and desolvation gas flow rate 800 l / h. Argon was used as the collision gas. MS / MS was performed in anionization mode according to the observed deprotonation mass. The following settings were used: 1kV capillary spray; 30eV cone voltage; collision energy gradient of 15-50eV.

[0204] Triple Quad analysis Injection was performed using a UPLC (Waters) connected to a Triple Quad detector (TQ-S, Waters) in multiple reaction monitoring (MRM) mode. Chromatographic separation was achieved using a column and mobile phase similar to those previously described. A 7-minute short-time method was established using the following multi-step gradient program: the initial condition was 57%B, increased to 85%B up to 4 min, increased to 100%B up to 4.2 min, held at 100%B up to 6 min, decreased to 67%B up to 6.2 min, and held at 67%B up to 7 min for system reequilibrium. A flow rate of 0.6 ml / min was used, the column temperature was 40°C, and the injection volume was 1 μl. The instrument was operated in negative mode with a capillary voltage of 1.5 kV and a cone voltage of 40 V. Two different transitions were used in CBGA analysis (quantitatively: 359.3 > 191.2, 32V; qualitatively: 359.3 > 315.4, 21V).

[0205] Isolation of trichomes Young leaves were collected, immersed in ice-cold distilled water, and then crushed using a BeadBeater apparatus (Biospec Products, Bartlesville, OK). 15 g of plant material was packed into a polycarbonate chamber, half the volume of glass beads (0.5 mm diameter) and XAD-4 resin (1 g / g plant material) were added, and the chamber was filled to full volume with 80% ethanol. The leaves were beaten repeatedly with 2-4 pulses of 1 minute each. This procedure was performed at 4°C, and after each pulse, the chamber was cooled in ice. After crushing, the contents of the chamber were first filtered through a cooking mesh strainer, and then filtered through a 100 μm nylon mesh to remove the plant material, glass beads, and XAD-4 resin. The remaining plant material and beads were scraped off the mesh, rinsed twice with additional 80% ethanol, and passed through the same 100 μm mesh. The presence of high concentrations of glandular ciliate secretory cells was confirmed by visualization using an inverted optical microscope.

[0206] Helichrysum genome sequencing and assembly The genome size of Helichrysum was estimated by flow cytometry. Briefly, nuclei were isolated by cleaving young leaf tissue of Helichrysum and tomato (used as a known reference) in isolation buffer. Samples were stained with propidium iodide, and at least 10,000 nuclei were analyzed by flow cytometry. The ratio of the mean G1 peaks between both samples was calculated. High molecular weight DNA was extracted from young frozen leaves and sent to the Genome Center at UC Davis for sequencing. DNA properties were confirmed by TapeStation trace and Qubit fluorometry (Thermo Fisher). Sequencing was performed on the Pacbio Sequel II platform, and a DNA SMARTbell library of approximately 12 kilobases was prepared according to the manufacturer's protocol. Using three different SMART 8M cells, 57.8 Gb of HiFi data (approximately 44 × haploid coverage) was obtained. In addition to Pacbio HiFi data, 200M reads of PE 2×150 Illumina Hi-C data were obtained by Phase Genomics. Both Pacbio HiFi and HiC data were integrated using Hifiasm software to generate chromosome-scale haplotype-resolved assemblies.

[0207] Further scaffolding of the primary assembly was performed using Hi-C data and SALSA software. Ragtag was used for ordering the final round, using the primary assembly as a reference to reach the synteny scaffold for each haplotype. Visualization of the Hi-C data was performed with Juicer, and whole-genome alignment was performed with the pafr package (https: / / dwinter.github.io / pafr / ). Finally, soft masking of repeating elements was performed on the assembly using EDTA.

[0208] RNA sequencing and genome annotation of helichrysum RNA was extracted from seven different tissues: young leaves, old leaves, tubular florets and receptacles, stems, roots, and trichomes. RNA integrity was checked using a TapeStation instrument. Paired-end Illumina libraries were prepared for five of the tissues and sequenced using an Illumina HiSeq 3000 instrument (PE 2×150, approximately 40M reads / sample). Random sequencing errors were corrected using Rcorrector, and uncorrectable reads were removed. Adapter and quality trimming were performed using TrimGalore! with the following parameters: --length 36 -q 5 --stringency 1 -e 0.1 (https: / / github.com / FelixKrueger / TrimGalore). Ribosomal RNA was filtered by discarding reads that mapped to the SILVA_132_LSURef and SILVA_138_SSURef non-redundant databases using bowtie2 --very-sensitive-local mode. Fastq quality checks were performed for each step using MultiQC. The remaining reads were pooled and used for genome-guided de novo transcriptome assembly using Trinity. Iso-Seq data were obtained from four tissues and processed using isoseeq3 and the cDNA Cupcake ToFU pipeline (https: / / github.com / Magdoll / cDNA_Cupcake). Unfuse and unspliced ​​transcripts were removed, and only polyA-positive transcripts were retained for a unique set of high-quality isoforms. Iso-Seq and Trinity transcripts were aligned to the assembly using minimap2, and RNA-based gene model structures were constructed using the PASA pipeline with BAM files. In addition, the novo gene structure was obtained using the software breaker2 and the BAM file mentioned as evidence of exogenous training. Finally, the ab initio and RNA-based gene models were combined using EvidenceModeler and the final round of the PASA pipeline.Functional annotation of genes was performed on predicted mature transcripts using TransDecoder (https: / / github.com / TransDecoder / TransDecoder), which considers HMMER hits against the PFAM database and BLASTP hits against the Uniprot database in terms of similarity retention criteria. Further annotation of protein-coding transcripts was performed by BLASTP searches against selected plant protein databases, and GO and KEGG terms were obtained using Triannotate.

[0209] 3' RNA sequencing based on UMI was obtained from seven tissues through three replicate experiments using a method similar to that described earlier. Adapter and quality trimming were performed using a two-step TrimGalore! including poly-A trimming mode. Reads were mapped to the genome using STAR, UMI duplication was removed using umitools, and counts were obtained using featureCount. Normalization was performed using the DESeq2 varianceStabilizingTransformation algorithm, and the CEMItools package was used for co-expression analysis (difference threshold 0.6, p-value 0.1). Genes in modules with expression profiles consistent with the presence of target metabolites were analyzed. Candidate genes were selected based on functional annotation and blast hits of known enzymes.

[0210] In vitro enzyme assay in PT The HuPT1 (SEQ ID NO: 1), HuPT2 (SEQ ID NO: 2), HuPT3 (SEQ ID NO: 3), and HuPTx (SEQ ID NO: 7) genes from H. ambracrigerum, and the GOT4 gene from Cannabis sativa, were separately cloned into pESC-HIS vectors. Microsome preparation from yeast cells transformed with pESC-HIS vectors was performed as described in Jozwiak et al. (2020). The PT enzyme assay was performed as previously described for CsGOT4 (Luo et al., 2019).

[0211] HuPT kinetic assay Microsomes (2 μl) were dissolved in reaction buffer (50 mM Tris-HCl, 10 mM MgCl2, pH 8.5), and substrates were added to a total volume of 50 μl [0.5 μM~1.5 mM μm olivetolic acid (Cayman Chemicals), 1 mM geranyl pyrophosphate (GPP, Sigma Aldrich)]. The samples were incubated at 30°C for 15 minutes. After extracting the samples with 100 μl ethanol, the mixture was stirred with a vortex mixer and centrifuged. The organic layer was filtered and analyzed via ULC-qTOF and Triple Quad (TQ) instruments.

[0212] Example 1 Functional characterization of prenyltransferase (PT) from H. ambracrigerum The inventors profiled six tissues from H. ambracrigerum (young leaves, old leaves, tubular florets and receptacle, stem, and root) using UPLC-qTOF and identified CBGA in all tissues except the root (Figure 1). CBGA is Δ 9- is a known central precursor of tetrahydrocannabinolic acid (THCA), cannabidiolic acid (CBDA), and several other cannabinoids. It is produced from olivetolic acid (OA) and geranyl pyrophosphate (GPP) by an enzymatic reaction catalyzed by geranyl pyrophosphate:olivetolic acid geranyltransferase 4 (GOT4). To identify genes within H. ambracrigerum associated with GOT-like activity, we searched for candidate prenyltransferase genes in the helichrysum transcriptome. Four prenyltransferase-like genes, namely HuPT1 (SEQ ID NO: 1), HuPT2 (SEQ ID NO: 2), HuPT3 (SEQ ID NO: 3), and HuPTx (SEQ ID NO: 7), were selected for further characterization based on their differential expression profiles in leaves compared to other tissues. All of these PTs shared less than 40% homology with CsGOT4, which is known to be involved in cannabinoid biosynthesis. Regarding functional expression in yeast, the inventors removed the N-terminal plastid-targeting sequence from all four PTs. Subsequently, each candidate PT expression cassette was introduced into yeast. Furthermore, microsomal fractions were purified from yeast cells expressing the candidate PTs, and PT activity was examined using OA and GPP as substrates. Purified yeast microsomal fractions containing GOT4 from Cannabis sativa were used together with OA and GPP in the positive control reaction. Of the four candidates tested, assays using PT1, PT3, and PTx enzymes from H. ambracrigerum showed clear generation of CBGA, similar to that observed in the positive control reaction (Figure 2). Active HuPT clustered with plastid PTs that prenylated various substrates, while HuPT2 clustered with mitochondrial PTs (Figure 3).

[0213] Furthermore, in in vitro assays using purified microsomal fractions containing HuPT1, HuPT3, and HuPTx, the inventors confirmed the activity of these enzymes in CBGA production. Reaction rate assays revealed that HuPTx exhibited very high catalytic activity (Km 2.0±0.8 μM, Figure 4). Other HuPTs, HuPT1 and HuPT3, also showed catalytic activity (Km 19.0±6.3 μM and Km 109.9±44.3 μM, respectively). The activity of HuPTx was similar to that of GOT4 from Cannabis sativa (Km CsGOT4=6.7±0.3 μM, Luo et al., 2019). These results demonstrate that HuPTs, such as HuPTx, can therefore be excellent enzymes for cannabinoid production in heterologous systems.

[0214] While the present invention has been described in conjunction with specific embodiments, it will be apparent that many alternative, modified, and variant forms are obvious to those skilled in the art. Therefore, the present invention is intended to encompass all such alternative, modified, and variant forms that fall within the spirit and broad scope of the appended claims.

Claims

1. An isolated DNA molecule containing a nucleic acid sequence that has at least 85% identity with SEQ ID NO: 7 and encodes a protein that is a prenyltransferase.

2. A plasmid or Agrobacterium comprising the artificial nucleic acid molecule of claim 1.

3. An isolated protein characterized by the ability to transfer a prenyl group to a substrate molecule, a. The isolated DNA molecule of claim 1; and b. Plasmid or Agrobacterium according to claim 2; An isolated protein encoded by one of the following.

4. The isolated protein according to claim 3, comprising an amino acid sequence having at least 80% identity with SEQ ID NO: 18, or consisting of an amino acid sequence having at least 80% identity with SEQ ID NO:

18.

5. The isolated protein according to claim 3, wherein the prenyl group is selected from the group consisting of dimethylallyl diphosphate, geranyl diphosphate, farnesyl diphosphate, and geranylgeranyl diphosphate.

6. The aforementioned substrate molecule is defined by formula I: 【Chemistry 1】 (In the formula, (i) R 1 However, it is selected from the group consisting of C1-C8 alkyl and alpha-unsaturated phenylalkyl carboxylic acids, and R 2 Is it OH or (ii)R 1 is OH and R 2 However, it is represented by (selected from the group consisting of C1-C8 alkyl and alpha-unsaturated phenylalkyl carboxylic acids), The alpha-unsaturated phenylalkyl carboxylic acid may include cinnamic acid or a derivative thereof, and the cinnamic acid derivative may be a hydroxylated derivative of cinnamic acid, and the hydroxylated derivative of cinnamic acid may be coumaric acid, or any combination thereof. The isolated protein according to claim 3.

7. A transgenic cell comprising the isolated DNA molecule of claim 1.

8. A transgenic plant, transgenic plant tissue, or plant part comprising the transgenic cells of claim 7, wherein the transgenic plant may be a Cannabis sativa plant.

9. A composition comprising the transgenic cells of claim 7 and an acceptable carrier.

10. Formula II: 【Chemistry 2】 (wherein (i) R 1 is selected from the group consisting of C1-C8 alkyl and α-unsaturated phenylalkylcarboxylic acids, and R 2 is OH; or (ii) R 1 is OH, and R 2 is selected from the group consisting of C1-C8 alkyl and α-unsaturated phenylalkylcarboxylic acids, R 3 is a prenyl group, and R 4 is hydrogen or a prenyl group), which is a method for synthesizing a compound represented by a. A step of providing cells containing an artificial vector having a nucleic acid sequence having at least 85% identity with SEQ ID NO: 7; and b. A step of culturing the cells from step (a) so that the protein encoded by the artificial vector is expressed; A method comprising, thereby synthesizing a compound represented by formula II.

11. A method according to claim 10, characterized in that the protein is capable of transferring a prenyl group to a substrate molecule, wherein the culture may include supplying the cells with an effective amount of the substrate molecule.

12. The aforementioned substrate molecule is defined by formula I: 【Transformation 3】 (In the formula, (i) R 1 However, it is selected from the group consisting of C1-C8 alkyl and alpha-unsaturated phenylalkyl carboxylic acids, and R 2 Is it OH or (ii)R 1 is OH and R 2 The method according to claim 10 or 11, wherein the representation is selected from the group consisting of C1-C8 alkyl and alpha-unsaturated phenylalkyl carboxylic acids.

13. The method according to claim 10, wherein the cell is a transgenic cell having a nucleic acid sequence having at least 85% identity with SEQ ID NO: 7 and encoding a protein that is a prenyltransferase, or a cell that has been transfected with a nucleic acid sequence having at least 85% identity with SEQ ID NO: 7 and encoding a protein that is a prenyltransferase.

14. The method according to claim 10, further comprising a step preceding step (a), which includes introducing or transfecting the cells with the artificial vector.

15. Formula II: 【Chemistry 4】 (In the formula, (i) R 1 However, it is selected from the group consisting of C1-C8 alkyl and alpha-unsaturated phenylalkyl carboxylic acids, and R 2 Is it OH or (ii)R 1 is OH and R 2 However, R is selected from the group consisting of C1-C8 alkyl and alpha-unsaturated phenylalkyl carboxylic acids. 3 This is a prenyl group, and R 4 A method for synthesizing a compound represented by (where is hydrogen or a prenyl group), the method comprising contacting a substrate molecule with an effective amount of a protein having an amino acid sequence having at least 80% identity with SEQ ID NO: 18, thereby synthesizing a compound represented by formula II, wherein the substrate molecule is represented by formula I: 【Transformation 5】 (In the formula, (i) R 1 However, it is selected from the group consisting of C1-C8 alkyl and alpha-unsaturated phenylalkyl carboxylic acids, and R 2 Is it OH or (ii)R 1 is OH and R 2 A method represented by (selected from the group consisting of C1-C8 alkyl and alpha-unsaturated phenylalkyl carboxylic acids).

16. The method according to claim 15, wherein the contact is performed in a cell-free system.

17. (i) The substrate molecule is selected from the group consisting of resorcinoid precursors, stilbenic acid precursors, acylphloroglucinoid precursors, and chalcone precursors; (ii) The substrate molecule 【Transformation 6】 It is expressed by an expression selected from the group consisting of R, where R 1 is a C1-C8 alkyl group, and R 2 It is an alpha-unsaturated phenylalkyl carboxylic acid; (iii) The substrate molecule is 【Transformation 7】 Selected from the group consisting of; or (iv) The substrate molecule is any combination of (i) to (iii); The method according to claim 10.

18. The method according to claim 10, wherein the compound is selected from the group consisting of cannabinoids, amorphurtins, acylphlorogluconoids, and prenyl chalcones, the compound is 【Transformation 8】 (In the formula, R 1 is a C1-C8 alkyl group, and R 2 R is an alpha-unsaturated phenylalkyl carboxylic acid, 3 This is a prenyl group, and R 4 A method in which (where is a hydrogen atom or a prenyl group) may be selected from the group.

19. The method according to claim 10, wherein the prenyl group is selected from the group consisting of dimethylallyl diphosphate, geranyl diphosphate, farnesyl diphosphate, and geranylgeranyl diphosphate.

20. The method according to claim 10, wherein the alpha-unsaturated phenylalkyl carboxylic acid comprises cinnamic acid or a derivative thereof, wherein the cinnamic acid derivative may be a hydroxylated derivative of cinnamic acid, and the hydroxylated derivative of cinnamic acid may be coumaric acid.

21. The aforementioned compound, 【Chemistry 9】 A method according to claim 10, selected from the group consisting of, the compound is 【Chemistry 10】 A good method.