Polyketide synthases, and transgenic cells, tissues, and organisms containing the same
The invention addresses the need for aromatic polyketide synthesis by providing nucleic acid sequences encoding polyketide synthases, enabling the efficient production of polyketides and cannabinoid analogs in transgenic cells and plants.
Patent Information
- Application Number
- JP2025514560
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-08
- Filing Date
- 2023-09-07
- Publication Date
- 2025-10-01
AI Technical Summary
There is a need in the art to identify enzymes and nucleotide sequences involved in the synthesis of aromatic polyketides, particularly for the production of cannabinoid analogs and precursors, which have potential pharmaceutical applications.
The invention provides isolated DNA molecules, artificial nucleic acid molecules, and plasmids or Agrobacterium containing nucleic acid sequences with at least 83% homology to specific sequences (SEQ ID NOs: 1 to 4), which encode proteins with polyketide synthase activity, enabling the synthesis of polyketides in transgenic cells and plants.
These sequences enable the efficient synthesis of polyketides, including tetraketides, in transgenic cells and plants, providing a basis for producing cannabinoid analogs and other valuable pharmaceutical compounds.
Smart Images

Figure 2025532528000001_ABST
Abstract
Description
[Technical Field]
[0001] Reference to electronic sequence listing The contents of the electronic sequence listing (YEDA-P-005-PCT ST26.xml; size: 14,356 bytes; and creation date: August 31, 2023) are incorporated herein by reference in their entirety.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 404,604, filed September 8, 2022, entitled "POLYKETIDE SYNTHASE AND A TRANSGENIC CELL, TISSUE, AND ORGANISM COMPRISING SAME."
[0003] FIELD OF THE INVENTION The present invention relates to polyketide synthesizing enzymes (PKS), polynucleotides encoding same, and methods of using same, for example, to produce polyketides. [Background technology]
[0004] Polyketides are a large, structurally diverse family of natural products that possess a wide range of biological activities, including antibacterial and pharmacological properties.
[0005] For example, polyketides are exemplified by antibiotics such as tetracycline and erythromycin, anti-cancer drugs including daunomycin, immunosuppressants such as FK506 and rapamycin, and veterinary products such as monensin and avermectins.
[0006] Polyketides occur in most biological groups and are particularly abundant in the class Actinomycetes, a class of mycelial bacteria that produce a wide variety of polyketides. Polyketide synthases (PKSs) are multifunctional enzymes related to fatty acid-1-synthase (FAS). PKSs catalyze the biosynthesis of polyketides through repeated (decarboxylative) Claisen condensations between acyl thioesters, usually acetyl, propionyl, malonyl, or methylmalonyl thioesters. After each condensation, they introduce structural diversity into the products by catalyzing all, some, or none of the reductive cycles, including ketoreduction, dehydration, and enoylreduction, on the β-keto group of the growing polyketide chain. In addition to varying the various condensation cycles, PKSs incorporate structural diversity into their products by controlling the overall chain length, the selection of primer and extender units, and, particularly in the case of aromatic polyketides, the regiospecific cyclization of the nascent polyketide chain. After the carbon chain has grown to a length characteristic of each specific product, it is released from the synthase by thiolysis or acyltransfer. Thus, PKSs consist of a family of enzymes that work together to produce a given polyketide. Controlled variations in chain length, choice of chain building units, and the reduction cycle, genetically programmed into each PKS, contribute to the variation observed among naturally occurring polyketides. Two general classes of PKSs exist. These classifications are well known.
[0007] Polyketide synthases are known to be involved in the biosynthesis of cannabinoids.
[0008] Genes encoding the enzymes of cannabinoid biosynthesis may also be useful for the synthesis of cannabinoid analogs and analogs of cannabinoid precursors. Cannabinoid analogs have previously been synthesized and may be useful as pharmaceutical products. There remains a need in the art to identify enzymes, and the nucleotide sequences encoding such enzymes, involved in the synthesis of aromatic polyketides. Summary of the Invention
[0009] In a first aspect, there is provided an isolated DNA molecule comprising a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1 to 4, or any combination thereof.
[0010] In another aspect, there is provided an artificial nucleic acid molecule comprising a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1 to 4, or any combination thereof.
[0011] In another aspect, there is provided a plasmid or Agrobacterium comprising a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1 to 4, or any combination thereof.
[0012] In another aspect, there is provided an isolated protein encoded by any one of: (a) an isolated DNA molecule of the invention; (b) an artificial vector disclosed herein, and a plasmid or Agrobacterium disclosed herein.
[0013] In another aspect, a transgenic cell is provided comprising: (a) a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1-4, or any combination thereof; (b) an artificial nucleic acid molecule disclosed herein; (b) a plasmid or Agrobacterium disclosed herein; (c) an isolated protein disclosed herein; or (d) any combination of (a)-(d).
[0014] In another aspect, there is provided an extract derived from the transgenic cells disclosed herein, or any fraction thereof.
[0015] In another aspect, there is provided a transgenic plant, transgenic plant tissue, or plant part comprising: (a) a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1-4, or any combination thereof; (b) an artificial vector disclosed herein; (b) a plasmid or Agrobacterium disclosed herein; (c) an isolated protein disclosed herein; (d) a transgenic cell disclosed herein; or (e) any combination of (a)-(e).
[0016] In another aspect, there is provided a composition comprising: (a) an isolated DNA molecule of the present invention; (b) an artificial vector disclosed herein; (c) a plasmid or Agrobacterium disclosed herein; (d) an isolated protein disclosed herein; (e) a transgenic cell disclosed herein; (f) an extract disclosed herein; (g) a transgenic plant tissue or plant part disclosed herein; or (h) any combination of (a)-(g) and an acceptable carrier.
[0017] In another aspect, a method for synthesizing a polyketide is provided, comprising: (a) providing a cell containing an artificial vector comprising a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1-4; and (b) culturing the cell from step (a) so that a protein encoded by the artificial vector is expressed, thereby synthesizing a polyketide.
[0018] In another aspect, a method for synthesizing a polyketide is provided, comprising contacting a diketide substrate with an effective amount of a protein comprising an amino acid sequence having at least 93% homology to any one of SEQ ID NOs: 5-8, thereby synthesizing the polyketide.
[0019] In another aspect, a method for obtaining an extract from a transgenic or transfected cell is provided, comprising: (a) culturing a transgenic or transfected cell in a culture medium, wherein the transgenic or transfected cell comprises a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1-4; and (b) extracting the transgenic or transfected cell, thereby obtaining an extract from the transgenic or transfected cell.
[0020] In another aspect, there is provided an extract of a transgenic or transfected cell obtained by the methods disclosed herein.
[0021] In another aspect, there is provided a culture medium or a portion thereof separated from cultured transgenic or transfected cells obtained by the methods disclosed herein.
[0022] In another aspect, there is provided a composition comprising: (a) an extract disclosed herein; (b) a medium disclosed herein or a portion thereof; or (c) a combination of (a) and (b), and an acceptable carrier.
[0023] In some embodiments, the nucleic acid sequence has at least 83% homology to any one of SEQ ID NOs: 1-4 and is 1,000-1,400 nucleotides in length.
[0024] In some embodiments, the nucleic acid sequence encodes a protein that is a polyketide synthase.
[0025] In some embodiments, the isolated protein comprises an amino acid sequence having at least 93% homology to any one of SEQ ID NOs:5-8.
[0026] In some embodiments, the isolated protein consists of the amino acid sequence of any one of SEQ ID NOs:5-8.
[0027] In some embodiments, the isolated protein is characterized by having the activity to polymerize a diketide substrate into a polyketide.
[0028] In some embodiments, the diketide substrate is obtained by coupling an acyl-CoA start site.
[0029] In some embodiments, the acyl-CoA start site is selected from the group consisting of: acetyl-CoA, butyryl-CoA, hexanoyl-CoA, octanoyl-CoA, cinnamoyl-CoA, coumaroyl-CoA, and any combination thereof.
[0030] In some embodiments, the acyl-CoA is hexanoyl-CoA, cinnamoyl-CoA, or both.
[0031] In some embodiments, the polyketide comprises a tetraketide.
[0032] In some embodiments, the transgenic cell is one of: a unicellular organism, a cell of a multicellular organism, and a cell in culture.
[0033] In some embodiments, the unicellular organism comprises a fungus or a bacterium.
[0034] In some embodiments, the fungus is a yeast cell.
[0035] In some embodiments, the extract comprises isolated DNA molecules, isolated proteins, or both.
[0036] In some embodiments, the transgenic plant is a Cannabis sativa plant.
[0037] In some embodiments, the protein is characterized by having the activity to polymerize a diketide substrate into a polyketide.
[0038] In some embodiments, the culturing comprises supplementing the cells with a diketide substrate, an acyl-CoA start site, or both.
[0039] In some embodiments, the artificial vector is an expression vector.
[0040] In some embodiments, the cell is a prokaryotic or eukaryotic cell.
[0041] In some embodiments, the cell is a transgenic cell, or a cell transfected with an isolated DNA molecule of the invention or an artificial vector disclosed herein.
[0042] In some embodiments, the method further comprises a step preceding step (a) comprising introducing an artificial vector into the cell or transfecting the cell with the artificial vector.
[0043] In some embodiments, the contacting is in a cell-free system.
[0044] In some embodiments, the method further comprises a step preceding step (b) comprising separating the cultured transgenic or transfected cells from the medium.
[0045] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the present invention, exemplary methods and / or materials are described below. In case of conflict, the present specification, including definitions, will control. Furthermore, the materials, methods, and examples are merely illustrative and are not intended to be necessarily limiting.
[0046] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description set forth below. It should be understood, however, that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description. [Brief explanation of the drawings]
[0047] [Figure 1] Figures 1A-1B include schemes and graphs of ion abundance. (1A) Scheme showing the steps in an in vitro recombinant enzyme assay using Helichrysum umbraculigerum polyketide synthase (HuPKS) with or without Cannabis sativa olivetolic acid cyclase (CsOAC) and the types of by-products produced. (1B) Peak areas of ion abundances of products from in vitro recombinant enzyme assays with or without CsOAC. PKS, polyketide synthase; PDAL, pentyl diacetic acid lactone; HTAL, triacetic acid lactone; OA, olivetolic acid; THPH, hexanoylphloroglucinol; EV, crude protein from E. coli cells transformed with an empty vector. [Figure 2]Figures 2A-2B contain extracted ion chromatograms and tandem mass spectrometry (MS / MS) spectra showing the products of a recombinant enzyme assay of HuPKS4 using either an empty vector (EV, crude protein obtained from E. coli cells transformed with the empty vector) or CsOAC in the presence of hexanoyl-CoA and malonyl-CoA. (2A) The extracted ion chromatogram shows the formation of olivetol ([M+H] = 181.123 Da), pentyl diacetic acid lactone (PDAL; [M-H] = 181.087 Da), hexanoyl triacetic acid lactone (HTAL), and olivetolic acid (OA) and hexanoyl phloroglucinol (THPH) ([M-H] = 223.097 Da). Standards of olivetol, OA, and THPH are shown for reference. A magnified view of the extracted ion chromatogram, normalized to the maximum value in each plot, is shown for improved interpretation. (2B) MS / MS spectrum of the produced OA compared with the standard (Std). DETAILED DESCRIPTION OF THE INVENTION
[0048] Detailed Description The present invention, in some embodiments, is directed to polynucleotide sequences encoding a protein or proteins from Helichrysum umbraculigerum that belong to the polyketide synthase (PKS) family.
[0049] In some embodiments, a polynucleotide ("polynucleotide of the invention") is provided that comprises a nucleic acid sequence comprising any one of SEQ ID NOs: 1-4, or any combination thereof.
[0050] In some embodiments, the polynucleotide is an isolated polynucleotide. In some embodiments, the polynucleotide is a DNA molecule. In some embodiments, the polynucleotide is an isolated DNA molecule. In some embodiments, the DNA molecule is an isolated DNA molecule. In some embodiments, the DNA molecule is a complementary DNA (cDNA) molecule.
[0051] As used herein, the terms "isolated polynucleotide" and "isolated DNA molecule" refer to a nucleic acid molecule that is essentially free from contaminating cellular components, such as carbohydrates, lipids, or other proteinaceous impurities naturally associated with the nucleic acid. Typically, an isolated DNA or RNA preparation contains a highly purified form of nucleic acid, e.g., at least about 80% pure, at least about 90% pure, at least about 95% pure, more than 95% pure, or more than 99% pure. In some embodiments, the isolated polynucleotide is any one of DNA, RNA, and cDNA. In some embodiments, the isolated polynucleotide is a synthetic polynucleotide. Polynucleotide synthesis is well known in the art and can be performed, for example, by ligating multiple nucleic acid molecules together or covalently linking them with a primer linker.
[0052] The term "nucleic acid" is well known in the art. As used herein, "nucleic acid" generally refers to a molecule (e.g., a chain) of DNA, RNA, or any of their derivatives or analogs containing nucleotides. A nucleotide is composed of a nucleoside and a phosphate group. The nitrogenous base of a nucleoside includes, for example, naturally occurring purine or pyrimidine nucleosides found in DNA (e.g., adenine "A," guanine "G," thymine "T," or cytosine "C") or purine or pyrimidine nucleosides found in RNA (e.g., A, G, uracil "U," or C).
[0053] The term "nucleic acid molecule" includes, but is not limited to, single-stranded RNA (ssRNA), double-stranded RNA (dsRNA), single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), small RNA, circular nucleic acids, fragments of genomic DNA or RNA, degraded nucleic acids, amplification products, modified nucleic acids, plasmid or organelle nucleic acids, and artificial nucleic acids such as oligonucleic acids.
[0054] In some embodiments, the polynucleotide comprises the nucleic acid sequence: comprising or consisting of:
[0055] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 83%, at least 85%, at least 87%, at least 95%, or at least 99% homology or identity to SEQ ID NO:1, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 83%-100%, 88%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO:1. Each possibility represents a separate embodiment of the present invention.
[0056] In some embodiments, the polynucleotide comprises the nucleic acid sequence: comprising or consisting of:
[0057] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 83%, at least 85%, at least 87%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:2, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 83%-100%, 87%-100%, 90%-100%, or 93%-100% homology or identity to SEQ ID NO:2. Each possibility represents a separate embodiment of the present invention.
[0058] In some embodiments, the polynucleotide comprises the nucleic acid sequence: comprising or consisting of:
[0059] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 83%, at least 87%, at least 89%, at least 90%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:3, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 83%-100%, 88%-100%, 93%-100%, or 95%-100% homology or identity to SEQ ID NO:3. Each possibility represents a separate embodiment of the present invention.
[0060] In some embodiments, the polynucleotide comprises the nucleic acid sequence: comprising or consisting of:
[0061] In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 82%, at least 85%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to SEQ ID NO:4, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence having 82%-100%, 86%-100%, 90%-100%, or 95%-100% homology or identity to SEQ ID NO:4. Each possibility represents a separate embodiment of the present invention.
[0062] In some embodiments, a polynucleotide of the invention comprises 1,000 to 1,400 nucleotides. In some embodiments, a polynucleotide of the invention is 1,000 to 1,400 nucleotides in length.
[0063] In some embodiments, 1,000 to 1,400 nucleotides includes at least 1,000 nucleotides, at least 1,200 nucleotides, at least 1,250 nucleotides, at least 1,300 nucleotides, at least 1,350 nucleotides, at least 1,370 nucleotides, or at least 1,390 nucleotides, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, 1,000 to 1,400 nucleotides includes 1,000 to 1,050 nucleotides, 1,000 to 1,175 nucleotides, 1,000 to 1,250 nucleotides, 1,000 to 1,300 nucleotides, 1,000 to 1,350 nucleotides, 1,000 to 1,370 nucleotides, or 1,000 to 1,390 nucleotides. Each possibility represents a separate embodiment of the present invention.
[0064] In some embodiments, a polynucleotide comprises a plurality of polynucleotides. In some embodiments, a polynucleotide comprises a plurality of types of polynucleotides. As used herein, the term "plurality" includes any integer greater than or equal to two. In some embodiments, a polynucleotide comprises at least two, or at least three different nucleic acid sequences, or any value and range therebetween, where each different nucleic acid sequence is selected from SEQ ID NOs: 1-4. Each possibility represents a separate embodiment of the invention. In some embodiments, a polynucleotide comprises two to three, two to four, or three to four different nucleic acid sequences, where each different nucleic acid sequence is selected from SEQ ID NOs: 1-4.
[0065] In some embodiments, the polynucleotide comprises a plurality of polynucleotide molecules, each of the plurality of polynucleotide molecules comprising a different nucleic acid sequence, wherein the different nucleic acid sequences are selected from SEQ ID NOs: 1-4.
[0066] In some embodiments, the polynucleotide encodes a protein characterized by polyketide synthesis activity. In some embodiments, the polynucleotide encodes a protein that is a polyketide synthase (PKS). In some embodiments, the PKS is a PKS from Helichrysum umbraculigerum. As used herein, the terms "polyketide synthase" and "PKS" encompass any enzyme from H. umbraculigerum that has or is characterized by a functional analog of the "olivetol synthase" or "OLS" of Cannabis sativa.
[0067] As used herein, the terms "polyketide synthase" and "PKS" are used interchangeably and refer to any peptide, polypeptide, or protein capable of catalyzing the elongation of a ketide or polyketide chain. In some embodiments, the PKS activity is transacylation. In some embodiments, the PKS activity includes Claisen condensation. In some embodiments, the PKS activity includes the reduction of a β-keto group to a β-hydroxy group. In some embodiments, the PKS activity includes HO splitting, thereby obtaining, providing, or resulting in an α-β-unsaturated alkene. In some embodiments, the PKS activity includes reducing an α-β-double bond to a single bond. In some embodiments, the PKS activity includes hydrolyzing a polyketide chain or completed polyketide chain from the acyl carrier protein domain of the PKS. In some embodiments, the PKS activity includes polymerizing and / or ligating a diketide substrate into a polyketide chain. In some embodiments, the PKS activity includes elongating a diketide into a polyketide chain. In some embodiments, the PKS activity comprises elongating a polyketide chain.
[0068] In some embodiments, an artificial nucleic acid molecule is provided that comprises a polynucleotide disclosed herein.
[0069] In some embodiments, the artificial vector comprises a plasmid. In some embodiments, the artificial vector comprises or is an Agrobacterium comprising the artificial nucleic acid molecule. In some embodiments, the artificial vector is an expression vector. In some embodiments, the artificial vector is a plant expression vector. In some embodiments, the artificial vector is for use in expressing nucleic acid sequences encoding a PKS disclosed herein. In some embodiments, the artificial vector is for use in heterologous expression of nucleic acid sequences encoding a PKS disclosed herein in a cell, tissue, or organism. In some embodiments, the artificial vector is for use in producing or producing a polyketide in a cell, tissue, or organism.
[0070] The expression of polynucleotides in cells is well known to those skilled in the art.It can be achieved by transfection, viral infection, or direct modification of the cell genome, among many other methods.In some embodiments, the polynucleotide is in an expression vector, such as a plasmid vector or a viral vector.The nucleic acid sequence of the vector generally includes at least a replication origin for propagation in cells, and optionally further elements such as heterologous polynucleotide sequences, expression control elements (e.g., promoters, enhancers), selectable markers (e.g., antibiotic resistance), poly-adenine sequences, etc.
[0071] The vector can be a DNA plasmid delivered via non-viral or viral methods. The viral vector can be a retroviral vector, a herpesvirus vector, an adenovirus vector, an adeno-associated virus vector, a virgaviridae virus vector, or a poxvirus vector. Barley stripe mosaic virus (BSMV), tobacco rattle virus (TRV), and cabbage leaf curl geminivirus (CaLCuV) can also be used. The promoter can be active in plant cells. The promoter can be a viral promoter.
[0072] In some embodiments, the polynucleotides disclosed herein are operably linked to a promoter. The term "operably linked" is intended to mean that the nucleotide sequence of interest is linked to control elements in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). In some embodiments, a promoter is operably linked to the polynucleotide of the present invention. In some embodiments, the promoter is a heterologous promoter. In some embodiments, the promoter is an endogenous promoter.
[0073] In some embodiments, vectors are introduced into cells by standard methods, including electroporation (e.g., as described in From et al., Proc. Natl. Acad. Sci. USA 82, 5824 (1985)), heat shock, infection with viral vectors, high-velocity ballistic penetration by small particles carrying nucleic acid either within the matrix or on the surface of small beads or particles (Klein et al., Nature 327, 70-73 (1987)), use of biolistic bombardment of coated particles, e.g., needle-shaped particles, Agrobacterium Ti plasmids, etc.
[0096] As used herein, the term "promoter" refers to a group of transcriptional control modules clustered around the initiation site of RNA polymerase, i.e., RNA polymerase II. Promoters are composed of discrete functional modules, each consisting of approximately 7-20 bp of DNA and containing one or more recognition sites for transcriptional activator or transcriptional repressor proteins. Promoters can extend upstream or downstream from the transcription start site and can be any size, from a few base pairs to several kilobases.
[0074] In some embodiments, the polynucleotide is transcribed by RNA polymerase II (RNAP II and Pol II), an enzyme found in eukaryotic cells that is known to catalyze the transcription of DNA to synthesize precursors of mRNA, as well as most snRNAs and microRNAs.
[0075] In some embodiments, plant expression vectors are used. In one embodiment, expression of the polypeptide coding sequence is driven by a number of promoters. In some embodiments, viral promoters are used, such as the 35S RNA and 19S RNA promoters of CaMV [Brisson et al., Nature 310:511-514 (1984)] or the coat protein promoter for TMV [Takamatsu et al., EMBO J. 6:307-311 (1987)]. In other embodiments, plant promoters are used, such as the small subunit promoter of RUBISCO [Coruzzi et al., EMBO J. 3:1671-1680 (1984); and Brogli et al., Science 224:838-843 (1984)] or heat shock promoters, such as soybean hspl7.5-E or hspl7.3-B [Gurley et al., Mol. Cell. Biol. 6:559-565 (1986)]. In one embodiment, the construct is introduced into plant cells using Ti plasmids, Ri plasmids, plant viral vectors, direct DNA transformation, microinjection, electroporation, and other techniques well known to those skilled in the art. See, for example, Weissbach & Weissbach [Methods for Plant Molecular Biology, Academic Press, NY, Section VIII, pp 421-463 (1988)]. Other expression systems well known in the art, such as insect and mammalian host cell systems, can also be used in accordance with the present invention.
[0076] In some embodiments, expression vectors containing regulatory elements derived from eukaryotic viruses, such as retroviruses, are used in accordance with the present invention. SV40 vectors include pSVT7 and pMT2. In some embodiments, vectors derived from bovine papillomavirus include pBV-lMTHA, and vectors derived from Epstein-Barr virus include pHEBO and p205. Other exemplary vectors include pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo-5, the baculovirus pDSVE, and any other vector that allows protein expression under the direction of the SV-40 early promoter, SV-40 late promoter, metallothionein promoter, mouse mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown to be effective for expression in eukaryotic cells.
[0077] In some embodiments, recombinant viral vectors that offer advantages such as systemic infection and target specificity are used for in vivo expression. In one embodiment, systemic infection is inherent in, for example, the life cycle of retroviruses, a process by which a single infected cell produces many progeny virions that infect neighboring cells. In one embodiment, the result is the rapid infection of a large area, most of which were not originally infected by the original viral particle. In one embodiment, a viral vector that cannot propagate systemically is produced. In one embodiment, this feature can be useful when the desired goal is to introduce a specific gene into only a limited number of target cells.
[0078] In some embodiments, plant viral vectors are used. In some embodiments, wild-type viruses are used. In some embodiments, degraded viruses are used, such as viruses known in the art. In some embodiments, Agrobacterium is used to introduce the vectors of the invention into the viruses.
[0079] Various methods can be used to introduce the expression vector of the present invention into cells. These methods are generally described in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Springs Harbor Laboratory, New York (1989, 1992), in Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989), Chang et al., Somatic Gene Therapy, CRC Press, Ann Arbor, Mich. (1995), Vega et al., Gene Targeting, CRC Press, Ann Arbor, Mich. (1995), Vectors: A Survey of Molecular Cloning Vectors and Their Uses, Butterworths, Boston Mass. (1988), and Gilboa et al. [Biotechniques 4 (6): 504-512,
[0019] and include, for example, stable or transient transfection, lipofection, electroporation, Agrobacterium Ti plasmids, and infection with recombinant viral vectors. Additionally, see U.S. Patent Nos. 5,464,764 and 5,487,992 for positive-negative selection methods.
[0080] It is understood that, in addition to containing the elements necessary for the transcription and translation of the inserted coding sequence (encoding a polypeptide), the expression constructs of the present invention may also contain sequences engineered to optimize the stability, production, purification, yield, or activity of the expressed polypeptide.
[0081] In some embodiments, the artificial vector comprises a polynucleotide that encodes a protein comprising an amino acid sequence described herein.
[0082] In some embodiments, provided are proteins encoded by: (a) a polynucleotide disclosed herein; (b) an artificial vector disclosed herein; or a plasmid or Agrobacterium disclosed herein.
[0083] In some embodiments, the protein is encoded by a polynucleotide comprising or consisting of SEQ ID NOs: 1-4.
[0084] In some embodiments, the protein comprises an amino acid sequence having at least 88%, at least 90%, at least 93%, at least 95%, at least 97%, or at least homology or identity to any one of SEQ ID NOs: 5-8.
[0085] In some embodiments, the protein is an isolated protein.
[0086] As used herein, the terms "peptide," "polypeptide," and "protein" are used interchangeably to refer to a polymer of amino acid residues. In another embodiment, as used herein, the terms "peptide," "polypeptide," and "protein" encompass native peptides, peptidomimetics (which typically contain non-peptide bonds or other synthetic modifications), and peptide analogs peptoids and semipeptoids, or any combination thereof. In another embodiment, the described peptides, polypeptides, and proteins possess modifications that render them more stable in organisms or more permeable to cells. In one embodiment, the terms "peptide," "polypeptide," and "protein" refer to naturally occurring amino acid polymers. In another embodiment, the terms "peptide," "polypeptide," and "protein" refer to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids.
[0087] As used herein, the term "isolated protein" refers to a protein that is essentially free from contaminating cellular components, such as carbohydrates, lipids, or other proteinaceous impurities naturally associated with nucleic acids. Typically, a preparation of isolated protein comprises the protein in a highly purified form, e.g., at least about 80% pure, at least about 90% pure, at least about 95% pure, greater than 95% pure, or greater than 99% pure. In some embodiments, the isolated protein is a synthesized protein. Protein synthesis is well known in the art and can be performed, for example, by heterologous expression in transformed cells, as exemplified herein.
[0088] In some embodiments, the protein has the amino acid sequence: MASSINISKIREAQRAQGPASILAVGTANPSNCVYQADYPDYYFRITKSEHMVDLKRKFKRMCDQSMIRKRYMQITEEYLKENPNICEYMAPSLDARQDVVVVEVPKLGKEAATKAIKEWGQPKSKITHLIFCTTSGVDMPGADYQLTKLLGLCPSVKRFMMYQQGCFAGGTVLRLAKDIAENNKGARVLVVCSEITAVIFRGPNDTHLDSLIGQALFGDGASSVIVGSDPDLTTERPLFEIISAAQTILPDSEGAIDGHLREAGLTFHLLKDVPRLISKNIEKALTQAFSPLGISDWNSIFWVTHPGGPAILDQVELKLGLKEEKMRTTRHVLSEYGNMSSACVFFVLDEMRKRSAKGGARTTGEGLDWGVLFGFGPGLTVETVVLHSLPTTMSIAT (SEQ ID NO: 5) It comprises or consists of:
[0089] In some embodiments, the protein comprises an amino acid sequence having at least 92%, at least 96%, at least 98%, or at least 99% homology or identity to SEQ ID NO:5, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 92%-100%, 95%-100%, 96%-100%, or 98%-100% homology or identity to SEQ ID NO:5. Each possibility represents a separate embodiment of the present invention.
[0090] In some embodiments, the protein is MASSINISKIREAQRAQGPASILAVGTANPSNCVYQADYPDYYFRITKSEHMVDLKEKFQRMCDKSMIRKRHIHITEEFLKENPNLCEYMAPSLDTRQDVVVVEVPKLGKEAATKAIKEWGQPKSKITHLIFCTTSGVDMPGADYQLTKLLGLHPSVKRFMMYQQGCFAGGTVLRLAKDLAENNKGARVLAVCSEITAVTFRGPNDTHIDSLVGQALFGDGAAAVIVGSDPDLTTERPLFEIISAAQTILPNSEGAIDGHVREVGVTIHILKDVPVLISKNIEKALTQAFSPLGISDWNSIFWVVHPGGPAILDQVELKLGLKEEKMRTTRHVLSEYGNMSSACVFFVLDEMRKRSAKGGARTTGEGLDWGVLFGFGPGLTVETVVLHSLPTTMSIAT (SEQ ID NO: 6) It comprises or consists of:
[0091] In some embodiments, the protein comprises an amino acid sequence having at least 91%, at least 94%, at least 95%, or at least 97% homology or identity to SEQ ID NO:6, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 91%-100%, 94%-100%, 97%-100%, or 97%-100% homology or identity to SEQ ID NO:6.
[0092] In some embodiments, the protein has the amino acid sequence: MASSINISKIREAQRAQGPASILAVGTANPSNCVYQADYPNYYFRITKSEHMVDLKRKFKRMCDQSMIRKRYMQITEEYLKENPNICEYMAPSLDARQDVVVVEVPKLGKEAATKAIKEWGQPKSKITHLIFCTTSGVDMPGADYQLTKLLGLCPSVKRFMMYQQGCFAGGTVLRLAKDIAENNKGARVLVVCSEITAVIFRGPNDTHLDSLIGQALFGDGASSVIVGSDPDLTTERPLFEIISAAQTILPDSEGAIDGHLREAGLTFHLLKDVPGLISKNIEKALTQAFSPLGISDWNSIFWVTHPGGPAILDQVELKLGLKEEKMRASRHVLSEYGNMSSACVFFILDEMRKKSDEDGAPTTGEGLDWGVLFGFGPGLTVETVVLHSLPTTMSIAT (SEQ ID NO: 7) It comprises or consists of:
[0093] In some embodiments, the protein comprises an amino acid sequence having at least 93%, at least 95%, or at least 97% homology or identity to SEQ ID NO:7, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 93%-100%, 94%-100%, 96%-100%, or 98%-100% homology or identity to SEQ ID NO:7. Each possibility represents a separate embodiment of the present invention.
[0094] In some embodiments, the protein has the amino acid sequence: MASSINISKIREAQRAQGPASILAVGTANPSNYEIQADFPDYYFRVTKSEHMADMKGTFQRMCDKSMIRKRHMLITEEFLKENPNLCEYMAPSLDTRQDVVVVEVPKLGKEAATKAIKEWGQPKSKITHLIFCTTTGVDMPGADYQLTKLLGLAPSVKRFMIYQQGCFAGGTVLRLAKDIAENNKGARVLAVCSEITAMSFRGPNDTHVDSLVGQALFGDGAAAVIVGSDPDLTTERPLFEIISAAQTILPNSEGAIDGHVREVGLTIHILKDVPVLISKNIEKALTQAFSPLGISDWNSIFWIVHPGGPAILDQVELKVGLKKEKMATSRHVLSEYGNMSSACVFFIMDEMRKRSAKGGARTTGEGLDWGVLFGFGPGLTVETVVLHSLPTTM (SEQ ID NO: 8) It comprises or consists of:
[0095] In some embodiments, the protein comprises an amino acid sequence having at least 88%, at least 92%, at least 95%, or at least 97% homology or identity to SEQ ID NO:8, or any value and range therebetween. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein comprises an amino acid sequence having 88%-100%, 91%-100%, 93%-100%, or 95%-100% homology or identity to SEQ ID NO:8. Each possibility represents a separate embodiment of the present invention.
[0096] The terms "homology" and "identity," when used interchangeably herein, refer to the sequence identity between two amino acid sequences or two nucleic acid sequences, with identity being a more strict comparison. The phrases "percent identity or homology" and "% identity or homology" refer to the percentage of sequence identity found in a comparison of two or more amino acid sequences or nucleic acid sequences. Two or more sequences can be 0 to 100% identical, or any value therebetween. Identity can be determined by comparing positions in each sequence that can be aligned for comparison purposes to a reference sequence. If a position in the compared sequences is occupied by the same nucleotide base or amino acid, the molecules are identical at that position. The degree of identity of amino acid sequences is a function of the number of identical amino acids at positions shared by the amino acid sequences. The degree of identity between nucleic acid sequences is a function of the number of identical or matching nucleotides at positions shared by the nucleic acid sequences. The degree of homology of amino acid sequences is a function of the number of amino acids at positions shared by the polypeptide sequences.
[0097] The following are non-limiting examples for calculating homology or sequence identity between two sequences (these terms are used interchangeably herein). The sequences are aligned for optimal comparison purposes (e.g., gaps may be introduced into one or both of the first and second amino acid or nucleic acid sequences for optimal alignment, and non-homologous sequences may be ignored for comparison purposes). Optimal alignment is determined as the best score using the GAP program in the GCG software package, with a Blossum62 scoring matrix, a gap penalty of 12, a gap length penalty of 4, and a frameshift gap penalty of 5. The amino acid residues or nucleotides at corresponding amino acid or nucleotide positions are then compared. If a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, the molecules are identical at that position. The percent identity between two sequences is a function of the number of identical positions shared by the sequences.
[0098] In some embodiments, the % homology or identity described herein is calculated or determined using the Basic Local Alignment Search Tool (BLAST). In some embodiments, the % homology or identity described herein is calculated or determined using the Blossum62 scoring matrix.
[0099] In some embodiments, the protein comprises or is characterized by polyketide synthesis activity, as described herein, hi some embodiments, the protein is characterized by having the activity to polymerize diketide substrates into polyketides.
[0100] In some embodiments, the diketide substrate is obtained by coupling an acyl-CoA start site.
[0101] In some embodiments, the acyl-CoA start site is selected from acetyl-CoA, butyryl-CoA, hexanoyl-CoA, octanoyl-CoA, cinnamoyl-CoA, coumaroyl-CoA, or any combination thereof.
[0102] In some embodiments, the acyl-CoA is or comprises hexanoyl-CoA, cinnamoyl-CoA, or both.
[0103] In some embodiments, the acyl-CoA is hexanoyl-CoA.
[0104] In some embodiments, the polyketide comprises a tetraketide. In some embodiments, the polyketide comprises a linear polyketide. In some embodiments, the polyketide comprises a linear tetraketide.
[0105] In some embodiments, a transgenic cell is provided that comprises: (a) a polynucleotide disclosed herein; (b) an artificial nucleic acid molecule disclosed herein; (c) a plasmid or agrobacterium disclosed herein; (d) a protein disclosed herein; or any combination thereof.
[0106] As used herein, the term "transgenic cell" refers to any cell that has been artificially manipulated at the genomic or genetic level. In some embodiments, a transgenic cell is a cell into which an exogenous polynucleotide, such as an isolated DNA molecule disclosed herein, has been introduced. In some embodiments, a transgenic cell includes a cell into which an artificial vector has been introduced. In some embodiments, a transgenic cell is a cell that has undergone genomic mutation or modification. In some embodiments, a transgenic cell is a cell that has undergone CRISPR genome editing. In some embodiments, a transgenic cell is a cell that has undergone targeted mutation of at least one base pair of its genome. In some embodiments, an exogenous polynucleotide (e.g., an isolated DNA molecule disclosed herein) or vector is stably integrated into the cell. In some embodiments, a transgenic cell expresses a polynucleotide of the invention. In some embodiments, a transgenic cell expresses a vector of the invention. In some embodiments, a transgenic cell expresses a protein of the invention. In some embodiments, a transgenic cell is a cell that lacks a polynucleotide of the invention and has been transformed or genetically modified to include a polynucleotide of the invention. In some embodiments, CRISPR technology is used to modify the genome of a cell, as described herein.
[0107] In some embodiments, the cells are cells of unicellular organisms, cells of multicellular organisms, and cells in culture.
[0108] In some embodiments, the unicellular organism comprises a fungus or a bacterium.
[0109] In some embodiments, the fungus is a yeast cell.
[0110] In some embodiments, the cell is an insect cell. In some embodiments, the cell comprises an insect cell line.
[0111] The variety of insect cell lines suitable for transformation and / or heterologous expression is common and will be apparent to those skilled in the art. Non-limiting examples of such insect cell lines include, but are not limited to, Sf-9 cells, SR+ Schneider cells, S2 cells, etc.
[0112] In some embodiments, an extract derived from the transgenic cells disclosed herein or any fraction thereof is provided.
[0113] In some embodiments, the extract comprises a polynucleotide of the present invention, an isolated DNA molecule disclosed herein, an isolated protein disclosed herein, or any combination thereof.
[0114] In some embodiments, a homogenate, a lysate, an extract, any combination thereof, or any fraction thereof from the transgenic cells disclosed herein is provided.
[0115] Methods and / or means for extracting, lysing, homogenizing, fractionating, or any combination thereof, cells or cell cultures are common and will be apparent to those skilled in the art of cell biology and biochemistry. Non-limiting examples include, but are not limited to, pressure lysis (e.g., using a French press), enzymatic lysis, soluble-insoluble phase separation (e.g., to obtain a supernatant and a pellet), detergent-based lysis, solvents (e.g., polar or non-polar solvents), liquid chromatography mass spectrometry, etc.
[0116] In some embodiments, a transgenic plant, a transgenic plant tissue, or a plant part is provided. In some embodiments, a transgenic plant, or any part, seed, tissue, or organ thereof, is provided, comprising at least one transgenic plant cell of the present invention. In some embodiments, the transgenic plant, the transgenic plant tissue, or a plant part comprises: (a) a polynucleotide disclosed herein; (b) an artificial gene disclosed herein; (c) a plasmid or Agrobacterium disclosed herein; (d) an isolated protein of the present invention; (e) a transgenic cell disclosed herein; or any combination thereof.
[0117] In some embodiments, the transgenic plant, transgenic plant tissue, or plant part comprises transgenic plant cells of the present invention. In some embodiments, the transgenic plant, transgenic plant tissue, or plant part comprises at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99%, or any value and range therebetween, transgenic cells of the present invention. Each possibility represents a separate embodiment of the present invention. In some embodiments, the transgenic plant, transgenic plant tissue, or plant part comprises 20%-50%, 20%-60%, 20%-70%, 20%-80%, 20%-90%, or 20%-100% transgenic cells of the present invention. Each possibility represents a separate embodiment of the present invention.
[0118] In some embodiments, the transgenic plant, transgenic plant tissue, or plant part is a Cannabis sativa plant or is derived from a Cannabis sativa plant. In some embodiments, the transgenic plant is a Cannabis sativa plant.
[0119] In some embodiments, the transgenic plant, transgenic plant tissue, or plant part is hemp or is derived from hemp. In some embodiments, C. sativa includes or is hemp.
[0120] In some embodiments, compositions are provided comprising any one of: (a) a polynucleotide (e.g., an isolated DNA molecule) of the invention; (b) an artificial vector; (c) a plasmid or Agrobacterium; (d) an isolated protein of the invention; (e) a transgenic cell; (f) an extract; (g) a transgenic plant tissue or plant part; and (h) any combination of (a)-(g) as disclosed herein, and an acceptable carrier.
[0121] As used herein, the term "carrier," "excipient," or "adjuvant" refers to any component of a composition, such as a pharmaceutical composition or a dietary supplement, that is not an active ingredient. As used herein, the term "pharmaceutically acceptable carrier" refers to a non-toxic, inert solid, semi-solid, liquid filler, diluent, encapsulating material, any type of formulation auxiliary, or simply a sterile aqueous medium, such as physiological saline. Some examples of materials which can serve as pharmaceutically acceptable carriers are sugars such as lactose, glucose, sucrose, and the like; starches such as corn starch and potato starch; cellulose and its derivatives such as sodium carboxymethylcellulose, ethylcellulose, and cellulose acetate; powdered tragacanth; malt, gelatin, talc; excipients such as cocoa butter and suppository waxes; oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; glycols such as propylene glycol, polyols such as glycerin, sorbitol, mannitol, and polyethylene glycol; esters such as ethyl oleate and ethyl laurate, agar; buffers such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline, Ringer's solution; ethyl alcohol, and phosphate buffers, as well as other non-toxic, compatible substances used in pharmaceutical formulations. Some non-limiting examples of substances that can function as carriers herein include sugars, starches, cellulose and its derivatives, powdered tragacanth, malt, gelatin, talc, stearic acid, magnesium stearate, calcium stearate, vegetable oils, polyols, alginic acid, pyrogen-free water, isotonic saline, phosphate buffer, cocoa butter (suppository base), emulsifiers (e.g., carbomer, hydroxypropyl cellulose, sodium lauryl sulfate), and other non-toxic, pharmaceutically compatible substances used in other pharmaceutical formulations. Wetting agents and lubricants such as sodium lauryl sulfate, as well as colorants, flavorings, excipients, stabilizers, antioxidants, and preservatives may also be present. Any non-toxic, inert, and effective carrier may be used to formulate the compositions contemplated herein.In this respect, suitable pharmaceutically acceptable carriers, excipients and diluents are well known to those skilled in the art, such as those listed in The Merck Index, Thirteenth Edition, Budavari et al., Eds., Merck & Co., Inc., Rahway, NJ (2001); the CTFA (Cosmetic, Toiletry, and Fragrance Association) International Cosmetic Ingredient Dictionary and Handbook, Tenth Edition (2004); and the "Inactive Ingredient Guide," US Food and Drug Administration (FDA) Center for Drug Evaluation and Research (CDER) Office of Management (all of which are incorporated herein by reference in their entirety).Examples of pharmaceutically acceptable carriers, carriers and diluents useful in the present composition include distilled water, physiological saline, Ringer's solution, dextrose solution, Hank's solution and DMSO. These additional inactive ingredients, as well as techniques for effective formulation and administration, are well known in the art and are described in standard textbooks, such as Goodman and Gillman et al.: The Pharmacological Bases of Therapeutics, 8th Ed., Gilman et al. Eds. Pergamon Press (1990); Remington's Pharmaceutical Sciences, 18th Ed., Mack Publishing Co., Easton, Pa. (1990); and Remington: The Science and Practice of Pharmacy, 21st Ed., Lippincott Williams & Wilkins, Philadelphia, Pa., (2005), each of which is incorporated herein by reference in its entirety.The presently described compositions may also be contained in artificially engineered structures such as liposomes, ISCOMS, sustained-release particles, and other vehicles that increase the half-life of peptides or polypeptides in serum. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and the like. Liposomes for use with the presently described peptides are formed from standard vesicle-forming lipids, which generally include neutral and negatively charged phospholipids and a sterol, such as cholesterol. The choice of lipid is generally determined by considerations such as liposome size and stability in the blood. A variety of methods are available for preparing liposomes, as reviewed, for example, by Coligan, JE et al., Current Protocols in Protein Science, 1999, John Wiley & Sons, Inc., New York; see also U.S. Pat. Nos. 4,235,871, 4,501,728, 4,837,028, and 5,019,369.
[0122] The carriers may comprise, in total, from about 0.1% to about 99.99999% by weight of the pharmaceutical compositions presented herein.
[0123] Synthesis method In some embodiments, methods for synthesizing polyketides are provided.
[0124] In some embodiments, the method includes the steps of: (a) providing a cell containing an artificial vector comprising a nucleic acid sequence having at least 82%, at least 85%, at least 89%, at least 92%, at least 95%, at least 97%, or at least 99% homology or identity to any one of SEQ ID NOS: 1-4, or any combination thereof, and (b) culturing the cell from step (a), thereby expressing the protein encoded by the artificial vector, thereby synthesizing a polyketide. Each possibility represents a separate embodiment of the present invention.
[0125] In some embodiments, the method comprises contacting a diketide substrate with an effective amount of a protein comprising an amino acid sequence having at least 88%, at least 91%, at least 95%, or at least 97% homology or identity to any one of SEQ ID NOs: 5-8, or any value and range therebetween, thereby synthesizing a polyketide. Each possibility represents a separate embodiment of the present invention.
[0126] In some embodiments, methods are provided for obtaining extracts from transgenic or transfected cells.
[0127] In some embodiments, the method comprises culturing the transgenic or transfected cells in a medium and extracting the transgenic or transfected cells.
[0128] In some embodiments, the method comprises: (a) culturing the transgenic or transfected cells in a culture medium; and (b) extracting the transgenic or transfected cells, thereby obtaining an extract from the transgenic or transfected cells.
[0129] In some embodiments, the transgenic or transfected cells comprise an artificial vector comprising a nucleic acid sequence having at least 82%, at least 87%, at least 91%, at least 93%, at least 95%, at least 97%, at least 99%, or 100% homology or identity to any one of SEQ ID NOS: 1-4, or any value and range therebetween, or any combination thereof. Each possibility represents a separate embodiment of the present invention.
[0130] In some embodiments, the transgenic or transfected cell comprises a polynucleotide or multiple polynucleotides of the invention as disclosed herein.
[0131] In some embodiments, the transgenic or transfected cells comprise an artificial nucleic acid molecule or vector disclosed herein.
[0132] In some embodiments, the cell is a transgenic cell or a cell transfected with a polynucleotide disclosed herein.
[0133] In some embodiments, the culturing comprises supplementing the cells with an effective amount of a diketide substrate, an acyl-CoA initiation site, or both, hi some embodiments, the supplementing is via the growth or culture medium in which the cells are cultured.
[0134] In some embodiments, the diketide substrate is obtained by coupling of an acyl-CoA start site. In some embodiments, the diketide substrate is a substrate for a protein encoded by a polynucleotide disclosed herein. In some embodiments, the diketide substrate is a substrate for a protein disclosed herein. In some embodiments, the diketide is a substrate for a PKS enzyme disclosed herein (e.g., a protein encoded by a polynucleotide of the invention and / or a protein of the invention).
[0135] In some embodiments, the acyl-CoA is selected from acetyl-CoA, butyryl-CoA, hexanoyl-CoA, octanoyl-CoA, cinnamoyl-CoA, coumaroyl-CoA, or any combination thereof.
[0136] In some embodiments, the acyl-CoA is or comprises hexanoyl-CoA.
[0137] In some embodiments, the method further comprises a step preceding step (a) comprising introducing into or transfecting a cell an artificial nucleic acid molecule or vector disclosed herein.
[0138] Methods for introducing artificial nucleic acid molecules or vectors into cells or transfecting cells with them are common and will be apparent to those skilled in the art.
[0139] In some embodiments, introducing or transfecting comprises delivering an artificial nucleic acid molecule or vector comprising a polynucleotide disclosed herein to the cell; or modifying the genome of the cell to comprise a polynucleotide disclosed herein. In some embodiments, delivering comprises transfection. In some embodiments, delivering comprises transformation. In some embodiments, delivering comprises lipofection. In some embodiments, delivering comprises nucleofection. In some embodiments, delivering comprises viral infection.
[0140] As used herein, the terms "transfecting" and "introducing" are interchangeable.
[0141] In some embodiments, the contacting is in a cell-free system.
[0142] The type of cell-free system suitable for synthesizing a polyketide using any one of the polynucleotide(s) of the invention and the protein(s) of the invention disclosed herein will be apparent to one of skill in the art.
[0143] In some embodiments, the method further comprises a step preceding step (b) comprising separating the cultured transgenic or transfected cells from the culture medium.
[0144] Methods for separating cells from the medium are common and may include, but are not limited to, centrifugation, ultracentrifugation, or other methods apparent to one of skill in the art.
[0145] In some embodiments, extracts of transgenic or transfected cells obtained by the methods disclosed herein are provided.
[0146] In some embodiments, culture medium or a portion thereof separated from cultured transgenic or transfected cells obtained by the methods disclosed herein is provided.
[0147] In some embodiments, (a) an extract disclosed herein; (b) an extract disclosed herein Compositions are provided that include the disclosed media or portions thereof; or (c) any combination of (a) and (b) and an acceptable carrier as described herein.
[0148] In some embodiments, a portion includes a fraction or a plurality thereof.
[0149] Where a range of values is provided, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limit of that range, and any other stated or intervening value within that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0150] As used herein, the term "about" when used in conjunction with a value refers to ±10% of the reference value. For example, a length of about 1,000 nanometers (nm) refers to a length of 1,000 nm ±100 nm.
[0151] It should be noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to "a polynucleotide" includes a plurality of such polynucleotides, a reference to "the polynucleotide" includes a reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth. It should be further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as a predicate for using exclusive terminology, such as "solely," "only," and the like, in connection with the recitation of claim elements or the use of a "negative" limitation.
[0152] When a convention similar to "at least one of A, B, and C, etc." is used, such a configuration is generally intended in the sense that one of ordinary skill in the art would understand that convention (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). One of ordinary skill in the art will further understand that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" is understood to include the possibilities of "A" or "B," or "A and B."
[0153] It will be understood that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination. All combinations of the embodiments belonging to the invention are specifically embraced by the invention and disclosed herein, just as if all combinations were individually and explicitly disclosed. Furthermore, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the invention and disclosed herein, just as if all such subcombinations were individually and explicitly disclosed herein.
[0154] Additional objects, advantages, and novel features of the present invention will become apparent to those skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as described hereinabove and as claimed in the claims section below finds experimental support in the following examples.
[0155] Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples.
[0156] Example Generally, the nomenclature used herein and the research procedures utilized in the present invention include molecular, biochemical, microbiological, and recombinant DNA techniques. Such techniques are fully explained in the literature. See, e.g., "Molecular Cloning: A Laboratory Manual" by Sambrook et al. (1989); "Current Protocols in Molecular Biology" Volumes I-III, Ausubel, R.M., ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology," John Wiley & Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning," John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA," Scientific American Books, New York; and Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series," Vols. 1-4, Cold Spring Harbor Laboratory Press, New York. (1998); methods described in U.S. Patent Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659, and 5,272,057; "Cell Biology: A Laboratory Handbook," Volumes I-III, Cellis, JE, ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique" by Freshney, Wiley-Liss, NY (1994), Third Edition; "Current Protocols in Immunology," Volumes I-III, Coligan, JE, ed. (1994); Stites et al.(eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization - A Laboratory Course Manual" CSHL Press (1996), all of which are incorporated by reference. Other general references are provided throughout this document.
[0157] material and method Chemicals and Reagents Unless otherwise stated, all analyzed compounds were ≥95% pure. Malonyl-CoA (≥90%), hexanoyl-CoA (≥85%), and olivetol were purchased from Sigma-Aldrich (Rehovot, Israel). OA (>90%) was purchased from Cayman Chemical (Ann Arbor, MI, USA). Hexanoylphloroglucinol (THPH) was purchased from Wuhan ChemFaces Biochemical Co., Ltd. (Hubei, China).
[0158] Trichome isolation Young leaves were harvested, immersed in ice-cold distilled water, and then ground using a BeadBeater machine (Biospec Products, Bartlesville, OK). A polycarbonate chamber was filled with 15 g of plant material, half the volume of glass beads (0.5 mm diameter), XAD-4 resin (1 g / g plant material), and 80% ethanol to capacity. The leaves were ground using 2–4 pulses of 1 min each. This procedure was performed at 4°C, and the chamber was cooled on ice after each pulse. After grinding, the contents of the chamber were first filtered through a kitchen mesh strainer and then through a 100 μm nylon mesh to remove the plant material, glass beads, and XAD-4 resin. The remaining plant material and beads were scraped from the mesh, rinsed twice with additional 80% ethanol, and then passed through a 100 μm mesh. The presence of concentrated glandular trichome secretory cells was confirmed by visualization under an inverted light microscope.
[0159] Genome sequencing and assembly of H. umbraculigerum The genome size of H. umbraculigerum was estimated by flow cytometry. Briefly, nuclei were isolated by mincing young leaf tissues of Helichrysum and tomato (used as a known reference material) in isolation buffer. Samples were stained with propidium iodide, and at least 10,000 nuclei were analyzed by flow cytometry. The ratio of the mean G1 peaks between the two samples was calculated. High-molecular-weight DNA was extracted from frozen young leaves and sent for sequencing at the UC Davis Genome Center. DNA quality was confirmed using TapeStation traces and a Qubit fluorometer (Thermo Fisher). Sequencing was performed on a Pacbio Sequel II platform, and a DNA SMRT bell library of approximately 12 kilobase pairs was prepared according to the manufacturer's protocol. Three different SMRT 8M cells were used to obtain 57.8 Gb of HiFi data (approximately 44× haploid coverage). In addition to Pacbio HiFi data, 200M reads of PE 2x150 Illumina Hi-C data were obtained from Phase Genomics. Both Pacbio HiFi and HiC data were integrated to generate a chromosome-scale and haplotype-resolved assembly using Hifiasm software.
[0160] Further scaffolding of the primary assembly was performed using Hi-C data and SALSA software. Ragtag was used for a final round of ordering, using the primary assembly as a reference to arrive at a syntenic scaffold for each haplotype. Hi-C data visualization was performed with Juicer, and whole-genome alignment was performed with the pafr package (https: / / dwinter.github.io / pafr / ). Finally, the assembly was soft-masked for repetitive elements using EDTA.
[0161] RNA sequencing and genome annotation of H. umbraculigerum RNA was extracted from seven different tissues: young leaves, old leaves, flower florets and receptacles, stems, roots, and trichomes. RNA integrity was confirmed using a TapeStation instrument. Paired-end Illumina libraries were prepared for five of the tissues and sequenced on an Illumina HiSeq 3000 instrument (PE 2x150, approximately 40M reads per sample). Random sequencing errors were corrected using Rcorrector18, and uncorrectable reads were filtered out. Adapter and quality trimming was performed using TrimGalore! with the following parameters: length 36, q 5, stringency 1, e 0.1 (github.com / FelixKrueger / TrimGalore). Ribosomal RNA was filtered using bowtie2 --very-sensitive-local mode by discarding reads that mapped to the nonredundant databases SILVA_132_LSURef and SILVA_138_SSURef. Fastq quality checks at each step were performed using MultiQC. The remaining reads were pooled and used for genome-guided de novo transcriptome assembly using Trinity. Iso-Seq data were obtained from four of the tissues and processed using isoseq3 and the cDNA Cupcake ToFU pipeline (github.com / Magdoll / cDNA_Cupcake). Fused and unspliced transcripts were filtered out, and only polyA-positive transcripts were retained for a unique set of high-quality isoforms. Iso-Seq and Trinity transcripts were aligned to the assembly using minimap2, and the BAM files were used with the PASA pipeline to generate RNA-based gene model structures. Additionally, de novo gene structures were obtained using the software braker2 and the aforementioned BAM files as external training evidence. Finally, the ab initio and RNA-based gene models were combined using EvidenceModeler and a final round of the PASA pipeline.Functional annotation of genes was performed on predicted mature transcripts using TransDecoder (github.com / TransDecoder / TransDecoder), which considers HMMER hits against PFAM and BLASTP hits against the UniProt database as similarity retention criteria. Further annotation of protein-coding transcripts was performed by BLASTP searches against curated plant protein databases, and GO and KEGG terms were obtained with Triannotate.
[0162] UMI-based 3' RNAseq of three replicates of seven tissues was performed similarly as described. Adapter and quality trimming was performed using TrimGalore! in two steps, including PolyA trimming mode. Reads were mapped to the genome using STAR, de-duplicated by UMI using umitools, and measured with featureCounts. Normalization was performed with the varianceStabilizingTransformation algorithm in DESeq2, and the CEMItools package was used for co-expression analysis (dissimilarity threshold 0.6, p-value 0.1). Genes in modules with expression profiles consistent with the presence of metabolites of interest were analyzed. Candidate genes were selected based on functional annotation and blast hits with known enzymes.
[0163] Cloning of PKS candidate genes The coding sequences of candidate HuPKSs from H. umbraculigerum were amplified for heterologous expression in SoluBL21 E. coli. Due to the high sequence similarity of the coding sequences, HuPKS2-4 were synthesized by Twist Biosciences. Additionally, a PKC known to be involved in the biosynthesis of olivetolic acid in Cannabis sativa (i.e., CsOAC) was also amplified. The PCR product was purified from agarose gel and ligated into the pOPINF vector (digested with HindIII and KpnI) using the ClonExpress II one-step cloning kit (Vazyme, Germany). The fusion reaction was transformed into competent E. coli Stellar cells (Clontech Takara). Recombinant colonies were selected on LB agar plates supplemented with ampicillin (100 μg / mL). Positive clones were confirmed by Sanger sequencing.
[0164] Protein expression and purification The gene cloned into the pOPINF vector was transformed into the SoluBL21 E. coli strain for protein expression. A single colony grown on LB agar was picked and grown overnight in 10 mL of LB medium at 37°C. The next day, 1 mL of the overnight culture was used to inoculate 100 mL of LB medium. The fresh culture was grown at an OD of 0.6-0.8. 600 The cells were grown at 37°C until reaching 100 μM, then induced with 600 μM IPTG and incubated overnight at 16°C. For purification, cells were harvested by centrifugation (3200 × g for 10 min) and resuspended in 50 mM Tris-HCl pH 8, 0.5 mM phenylmethylsulfonyl fluoride (PMSF, Sigma Aldrich) solution in isopropanol, 10% glycerol and protease inhibitor cocktail (Sigma Aldrich), and 0.1 mg ml -1The recombinant enzyme was dissolved in 1000 kDa lysozyme (Sigma-Aldrich) by sonication. Protein purification was performed using Ni-NTA agarose beads (Adar Biotech). Protein was eluted with 200 mM imidazole (Fluka) in a buffer containing 50 mM NaH2PO4, pH 8.0, and 0.5 M NaCl. Dialysis and buffer exchange were performed using 20 mM HEPES pH 7.2 in a centrifugal concentrator with a size exclusion of 3 KDa or 10 KDa, depending on the protein size. The recombinant enzyme was confirmed by SDS-PAGE analysis, and protein concentration was measured using Pierce™ 660 nm Protein Assay Reagent (Thermo Scientific).
[0165] HuPKS enzyme assay for olivetolic acid production Individual and coupled HuPKS and CsOAC assays were performed as described by Gagne et al. (2012) with minor modifications. Enzyme assays were performed in 50 μL using 20 mM HEPES, 5 mM DTT, 1.8 mM malonyl-CoA, and 0.6 mM hexanoyl-CoA at pH 7.2. HuPKS (5 μg) and CsOAC (10 μg) were added individually or in combination. The reaction mixture was incubated at 30°C for 3 hours. The reaction was stopped by extraction with 100 μL of MeOH, vortexing, and centrifugation at 15,000 g for 10 minutes. The supernatant was filtered and analyzed by ultra-performance liquid chromatography coupled to a quadrupole time-of-flight system (UPLC-qTOF) or a Triple Quad system. Chromatographic separation was performed on a 100 mm × 2.1 mm id, 1.7 μm UPLC BEH C18 column (Waters Acquity). The mobile phase consisted of 0.1% formic acid in acetonitrile:water (5:95, v / v; phase A) and 0.1% formic acid in acetonitrile (phase B). The flow rate was 0.3 ml min -1The column temperature was maintained at 35°C. Compounds were analyzed using an 11-minute multistep gradient method: initial conditions were 10% B, increased to 70% by 6 minutes, increased to 100% B by 6.2 minutes, held at 100% B until 8 minutes, decreased to 10% B by 8.5 minutes, and held at 10% B until 11 minutes for system re-equilibration. Electrospray ionization (ESI) was used with negative or positive ionization in the m / z range of 50-1,000 Da. Mass of eluted compounds was detected with the following settings: source temperature 140°C, desolvation temperature 450°C, and desolvation gas flow 800 l h -1 Capillary voltage was 1.0 kV in negative mode and 1.5 kV in positive mode. Argon was used as the collision gas. MS / MS experiments were performed in negative or positive ionization mode depending on the specific mass of the deprotonated or protonated compound. The following settings were used: a cone voltage of 30 eV, a collision energy ramp of 15-50 eV (negative mode) and 10-45 eV (positive mode).
[0166] Triple-Quad analyses were performed on a TQ-S system in MRM mode using the same columns and gradients as previously described. The instrument was operated in both positive and negative modes at capillary voltages of 3.5 kV or 1.5 kV, and cone voltages of 40 V or 20 V, respectively. Two different transitions were used for the analysis: olivetolic acid (223.1 > 179.1, 15.0 V; 223.1 > 137.1, 20.0 V); PDAL (181.2 > 137.1, 10.0 V; 181.2 > 97.1, 20.0 V); HTAL (223.1 > 179.1, 10.0 V; 223.1 > 125.1, 10.0 V); THPH (223.1 > 179.1, 20.0 V; 223.1 > 81.0, 25.0 V); negative mode; and olivetol (181.1 > 111.0, 10.0 V; 181.1 > 71.2, 10.0 V) (positive mode).
[0167] Example 1 Polyketide synthase (PKS) of H. umbraculigerum OA is the first key intermediate in the cannabinoid biosynthetic pathway in cannabis (Cannabis sativa). The biosynthesis of OA from hexanoyl-CoA involves two enzymatic reactions catalyzed by the type III PKS, olivetol synthase (CsOLS). This enzyme first converts hexanoyl-CoA to a tetraketide intermediate, which is further modified into OA by the polyketide cyclase (PKC)-type olivetolic acid cyclase (CsOAC) (Taura et al., 2009; Gagne et al., 2012). To identify genes in H. umbraculigerum associated with CsOLS-like activity, we searched for candidate type III PKS genes in the H. umbraculigerum genome. Four PKS-like genes, namely HuPKS1 (SEQ ID NO: 1), HuPKS2 (SEQ ID NO: 2), HuPKS3 (SEQ ID NO: 3), and HuPKS4 (SEQ ID NO: 4) enzymes, were selected for further characterization based on their differential expression profiles in leaves and trichomes compared to other tissues.
[0168] Next, we expressed the HuPKS1-4 and CsOAC enzymes and tested their ability to form OA using hexanoyl-CoA and malonyl-CoA, either alone or together, in in vitro assays. In these assays, derailment of unstable intermediates occurred, producing additional by-products not naturally identified in plant extracts: olivetol, pentyl acyl diacetic acid lactone (PDAL), and hexanoyl acyl triacetic acid lactone (HTAL). PDAL and HTAL are produced by spontaneous lactonization of unstable triketide and tetraketide intermediates, whereas CsOLS produces olivetol in the absence of CsOAC in an aldol decarboxylation cyclization reaction similar to the production of resveratrol by stilbene synthase (Figure 1A).
[0169] In this study, we demonstrate that, like CsOLS, all HuPKSs produce PDAL and HTAL byproducts in the absence of CsOAC, whereas HuPKS1, HuPKS2, and HuPKS4 also produce olivetol (Fig. 1B). Verification of OA and olivetol production using LC-HRMS is shown in Fig. 2. When the reaction was performed in the presence of CsOAC, olivetol decreased and OA increased, particularly for HuPKS2 and HuPKS4 (Fig. 1B). Notably, all HuPKSs also produced the phloroglucinoid precursor hexanoylphloroglucinol (THPH), regardless of CsOAC (Fig. 1B), indicating that the same HuPKS enzyme can mediate both aldol and clausen cyclization reactions (Fig. 1A).
[0170] While the present invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, the present invention is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and broad scope of the appended claims.
Claims
1. An isolated DNA molecule comprising a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1 to 4, or any combination thereof.
2. 2. The isolated DNA molecule of claim 1, wherein the nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1 to 4 is 1,000 to 1,400 nucleotides in length.
3. 3. The isolated DNA molecule of claim 1 or 2, wherein the nucleic acid sequence encodes a protein that is a polyketide synthase.
4. An artificial nucleic acid molecule comprising a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1 to 4, or any combination thereof.
5. A plasmid or Agrobacterium comprising a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1 to 4, or any combination thereof.
6. a. The isolated DNA molecule of any one of claims 1 to 3; b. an artificial vector according to claim 4, and c) The plasmid or Agrobacterium according to claim 5. An isolated protein encoded by any one of
7. 7. The isolated protein of claim 6, comprising an amino acid sequence having at least 93% homology to any one of SEQ ID NOs: 5 to 8.
8. 8. The isolated protein according to claim 6 or 7, which consists of the amino acid sequence of any one of SEQ ID NOs: 5 to 8.
9. 9. The isolated protein of claim 8, characterized by having the activity of polymerizing a diketide substrate into a polyketide.
10. 10. The isolated protein of claim 9, wherein the diketide substrate is obtained by coupling an acyl-CoA initiation site.
11. 11. The isolated protein of claim 10, wherein the acyl-CoA start site is selected from the group consisting of acetyl-CoA, butyryl-CoA, hexanoyl-CoA, octanoyl-CoA, cinnamoyl-CoA, coumaroyl-CoA, and any combination thereof.
12. 12. The isolated protein of claim 10 or 11, wherein the acyl CoA is hexanoyl CoA, cinnamoyl CoA, or both.
13. The isolated protein of any one of claims 9 to 12, wherein the polyketide comprises a tetraketide.
14. a. a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1-4, or any combination thereof; b. The artificial nucleic acid molecule of claim 4. c. The plasmid or Agrobacterium of claim 5, d. An isolated protein according to any one of claims 6 to 13, or e. Any combination of (a) to (d) A transgenic cell comprising:
15. 15. The transgenic cell of claim 14, which is any one of a unicellular organism, a cell of a multicellular organism, and a cell in culture.
16. The transgenic cell of claim 15 , wherein the unicellular organism comprises a fungus or a bacterium.
17. The transgenic cell of claim 16 , wherein the fungus is a yeast cell.
18. An extract derived from the transgenic cell according to any one of claims 14 to 17, or any fraction thereof.
19. 20. The extract of claim 18, comprising the isolated DNA molecule, the isolated protein, or both.
20. a. a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1-4, or any combination thereof; b. The artificial vector of claim 4. c. The plasmid or Agrobacterium of claim 5, d. The isolated protein of any one of claims 6 to 13. e. The transgenic cell of any one of claims 14 to 17, or f. Any combination of (a) to (e) A transgenic plant, transgenic plant tissue, or plant part comprising:
21. 21. The transgenic plant of claim 20, which is a Cannabis sativa plant.
22. a. The isolated DNA molecule of any one of claims 1 to 3; b. The artificial vector of claim 4. c. The plasmid or Agrobacterium of claim 5, d. The isolated protein of any one of claims 6 to 13. e. The transgenic cell according to any one of claims 14 to 17. f. The extract according to claim 18 or 19, g. A transgenic plant tissue or plant part according to claim 20 or 21, or h. Any combination of (a) to (g); Acceptable carriers and A composition comprising:
23. 1. A method for synthesizing a polyketide, comprising: a. Providing a cell containing an artificial vector comprising a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1-4; b. Culturing the cells from step (a) so that the protein encoded by the artificial vector is expressed. thereby synthesizing a polyketide, method.
24. 24. The method of claim 23, wherein the protein is characterized by having the activity of polymerizing a diketide substrate into a polyketide.
25. 25. The method of claim 23 or 24, wherein the polyketide comprises a tetraketide.
26. 26. The method of claim 24 or 25, wherein the diketide substrate is obtained by coupling of an acyl-CoA start site.
27. 27. The method of any one of claims 23-26, wherein the culturing step comprises supplementing the cells with an effective amount of a diketide substrate, an acyl-CoA initiation site, or both.
28. 28. The method of claim 26 or 27, wherein the acyl-CoA start site is selected from the group consisting of acetyl-CoA, butyryl-CoA, hexanoyl-CoA, octanoyl-CoA, cinnamoyl-CoA, coumaroyl-CoA, and any combination thereof.
29. 29. The method of any one of claims 26-28, wherein the acyl-CoA start site is hexanoyl-CoA, cinnamoyl-CoA, or both.
30. The method according to any one of claims 23 to 29, wherein the artificial vector is an expression vector.
31. The method of any one of claims 23 to 30, wherein the cell is a prokaryotic or eukaryotic cell.
32. 32. The method of any one of claims 23 to 31, wherein the cell is a transgenic cell or a cell transfected with the isolated DNA molecule of any one of claims 1 to 3 or the artificial vector of claim 4.
33. 33. The method of any one of claims 23 to 32, further comprising a step preceding step (a), comprising introducing the artificial vector into the cell or transfecting the cell with the artificial vector.
34. A method for synthesizing a polyketide, comprising contacting a diketide substrate with an effective amount of a protein comprising an amino acid sequence having at least 93% homology to any one of SEQ ID NOs: 5-8, thereby synthesizing the polyketide.
35. 35. The method of claim 34, wherein the polyketide comprises a tetraketide.
36. 36. The method of claim 34 or 35, wherein the diketide substrate is obtained by coupling of an acyl-CoA start site.
37. 37. The method of claim 36, wherein the acyl-CoA start site is selected from the group consisting of acetyl-CoA, butyryl-CoA, hexanoyl-CoA, octanoyl-CoA, cinnamoyl-CoA, coumaroyl-CoA, and any combination thereof.
38. 38. The method of claim 36 or 37, wherein the acyl-CoA start site is hexanoyl-CoA, cinnamoyl-CoA, or both.
39. The method of any one of claims 34 to 38, wherein the contacting is in a cell-free system.
40. 1. A method for obtaining an extract from a transgenic or transfected cell, comprising: a. culturing a transgenic or transfected cell in a culture medium, wherein the transgenic or transfected cell comprises a nucleic acid sequence having at least 83% homology to any one of SEQ ID NOs: 1-4; b. Extracting the transgenic or transfected cells thereby obtaining an extract from said transgenic or transfected cells. method.
41. 41. The method of claim 40, further comprising a step preceding step (b) comprising separating the cultured transgenic or transfected cells from the medium.
42. An extract of transgenic or transfected cells obtained by the method of claim 40 or 41.
43. 42. A culture medium or a portion thereof separated from cultured transgenic or transfected cells obtained by the method of claim 41.
44. a. the extract of claim 42, b. The medium of claim 43 or a part thereof, or c. A combination of (a) and (b); Acceptable carriers and A composition comprising: