A bioenzyme, nucleic acid molecule, biomaterial and its application that catalyzes the C-2 acetyl group of taxane compounds
Patent Information
- Application Number
- CN202610912087.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-18
AI Technical Summary
但是由于紫杉醇生物合成网络繁琐复杂,其涉及多种代谢中间体
1.发明人筛选到针对C-2位的乙酰基转移酶候选基因,经饲喂紫杉素确定催化C-2位乙酰基化的T2AT酶,例如表3中不同来源的T2AT酶;
Smart Images

Figure CN122772831A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of synthetic biology, specifically to a bioenzyme, nucleic acid molecule, biomaterial, and its application that catalyzes the C-2 acetyl group of taxane compounds. Background Technology
[0002] Taxanes are a series of derivatives synthesized from antitumor active ingredients isolated from plants through structural modification. Taxane drugs mainly include paclitaxel, docetaxel, cabazitaxel, and derivatives with a taxane skeleton. Paclitaxel was first isolated from the bark of the yew tree (Taxus brevifolia) and is now widely used in the clinical treatment of various cancers.
[0003] However, the yield of paclitaxel in yew cells is extremely low, and the paclitaxel synthesis pathway is complex with numerous byproducts. Improving the paclitaxel yield has become a current challenge and focus of attention in this field. The low yield of paclitaxel biosynthesis is mainly due to its long biosynthetic pathway, involving more than 20 steps, the heterogeneity of enzymes involved in biosynthesis causing metabolic flux to be diverted, and poor enzyme activity. In particular, enzyme heterogeneity not only leads to a network-like structure in the paclitaxel biosynthetic pathway but also results in the generation of numerous byproducts. Therefore, screening for specific enzymes or specifically modifying enzymes has become a focus of paclitaxel biosynthesis. Using synthetic biology techniques and metabolic engineering strategies to reconstruct the paclitaxel synthesis pathway in plant or microbial chassis, and to mass-produce paclitaxel through cell culture, holds promise for completely solving this problem. The biosynthesis of paclitaxel involves multiple steps, including hydroxylation, acetylation, oxidation, and epoxidation, as well as C-2 modification. C-2 modification plays a crucial role in the anticancer activity of paclitaxel. In paclitaxel, C-2 exists in the form of a benzoyl group. However, due to the complex and intricate biosynthetic network of paclitaxel, which involves multiple metabolic intermediates, baccatin IV, with its C-2 acetyl group, can be converted from an acetyl group to a benzoyl group after being catalyzed by the acyltransferase TBT, thereby generating baccatin VI, a paclitaxel biosynthetic intermediate that can be used in subsequent paclitaxel biosynthesis. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, this application provides a biological enzyme, nucleic acid molecule, biomaterial, and its applications. The biological enzyme of this invention can catalyze the conversion of taxane compounds at the C-2 position to acetyl taxane compounds. Successful resolution of this biological enzyme facilitates the biosynthesis of taxane compounds, such as 2α-acetoxytaxusin and / or baccatin IV, and further synthesizes baccatin VI, baccatin III, and paclitaxel, thereby increasing the yield of taxane compounds.
[0005] A first aspect of the present invention provides a biological enzyme comprising at least one of the following sequences: a1) Any of the amino acid sequences shown in SEQ ID NO: 1-7; a2) Proteins with the same function obtained by substituting and / or deleting and / or adding one or more amino acid residues of any of the amino acid sequences shown in SEQ ID NO: 1-7. a3) Proteins that have at least 50% homology with and have the same function as any of the amino acid sequences shown in SEQ ID NO: 1-7; A fusion protein is obtained by attaching a tag to the N-terminus and / or C-terminus of at least one of the amino acid sequences shown in a4)(a1)(-a3).
[0006] Preferably, substitution and / or deletion and / or addition of one or more amino acid residues includes substitution and / or deletion and / or addition of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30 or more amino acid residues.
[0007] Preferably, having at least 50% homology includes having at least 50%, at least 50.913%, at least 56%, at least 56.027%, at least 57%, at least 57.955%, at least 60%, at least 60.403%, at least 65%, at least 70%, at least 75%, at least 77%, at least 77.263%, at least 80%, at least 85%, at least 90%, at least 94%, at least 94.42%, at least 98.5%, at least 98.6%, at least 98.7%, at least 98.8%, at least 98.9%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% homology.
[0008] Preferably, proteins having at least 50% homology with the amino acid sequence shown in SEQ ID NO: 1 and having the same function include SEQ ID NO: 2-7.
[0009] Preferably, proteins that have at least 60% homology with the amino acid sequence shown in SEQ ID NO: 1 and have the same function include SEQ ID NO: 2, 5, and 7.
[0010] Preferably, proteins that have at least 70% homology with the amino acid sequence shown in SEQ ID NO: 1 and have the same function include SEQ ID NO: 2 and 5.
[0011] Preferably, proteins having at least 90% homology to the amino acid sequence shown in SEQ ID NO: 1 and having the same function include SEQ ID NO: 2.
[0012] Preferably, the homologous sequences are derived from gymnosperms, such as plants from the Pinaceae, Cupressaceae, Taxaceae, Ginkgoaceae, and Taxaceae families. More preferably, they are derived from the Taxaceae family and the Taxus genus (including but not limited to *Taxus baccata* L., *Taxus brevifolia* Nutt., *Taxus calcicola* L. M. Gao & Mich. Möller, *Taxus canadensis* Marshall, *Taxus chinensis* (Pilg.) Rehd., *Taxus contorta* Griff., *Taxus cuspidata* Siebold & Zucc., *Taxus floridana* Nutt. ex Chapm., *Taxus florinii* Spjut., *Taxus globosa* Schltdl., and *Taxus mairei*). Plants of the genus *Taxus phytonii* (Lemée & Lév.) SY Hu, *Taxus phytonii* Spjut, *Taxus sumatrana* (Miq.) de Laub., *Taxus wallichiana* Zucc., *Taxus Huangshan type*, *Taxus Emei type*, *Taxus Qinling type*, *Taxus yunnanensis* WC Cheng & L. K. Fu, *Taxus ×media* Rehder, *Taxus × hunnewelliana* Rehder, or plants of the family Taxaceae and the genus *Taxus* (including but not limited to *Taxus yunnanensis*). Pseudotaxus chienii (WC Cheng) WC Cheng).
[0013] Preferably, the bioenzyme is capable of catalyzing the C-2 acetylation of compounds with a taxane skeleton structure (e.g., taxane).
[0014] Preferably, the protein having the same function refers to a protein that retains the acetyltransferase activity of the biological enzyme, and whose activity is higher, equivalent to or lower than that of SEQ ID NO: 1-7, for example, an activity of 10%-1000% of that of SEQ ID NO: 1-7 (e.g., 10%, 50%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900% or 1000%).
[0015] Preferably, the label includes an MBP label or a GST label.
[0016] Preferably, the substitution includes conservative amino acid substitution.
[0017] Preferably, the bioenzyme is an acetyltransferase (an enzyme that catalyzes the production of acetyl groups), which is also named T2AT in this application.
[0018] Preferably, the bioenzyme uses a compound of formula I or a pharmaceutically acceptable salt, isomer or solvate thereof; or a compound of formula II or a pharmaceutically acceptable salt, isomer or solvate thereof; or a compound of formula III or a pharmaceutically acceptable salt, isomer or solvate thereof as a substrate to catalyze the acetylation at the C-2 position.
[0019] The bioenzymes described in this application are isolated bioenzymes or artificially synthesized bioenzymes, such as prokaryotic or eukaryotic expression.
[0020] In a second aspect, the present invention provides a biological enzyme composition comprising the above-described biological enzyme.
[0021] Preferably, the bioenzyme composition further includes taxane 2α-hydroxylase (T2αH).
[0022] Preferably, the bioenzyme composition further includes one or more of the following: taxane 7β-hydroxylase (T7βH), taxane 7β-oxo-acetyltransferase (T7AT), taxane C4-C20 oxobutane synthase (TOT), and taxane 1β-hydroxylase (T1βH).
[0023] In one specific embodiment of the present invention, the bio-enzyme composition includes the above-mentioned bio-enzyme and T2αH.
[0024] In one specific embodiment of the present invention, the bio-enzyme composition includes the above-mentioned bio-enzyme, T2αH, T7βH, T7AT, TOT and T1βH.
[0025] In some embodiments, the natural amino acid sequences of the aforementioned various biological enzymes may be variant sequences derived from different species and / or different strains of the genus *Taxus*. The artificial amino acid sequences of the aforementioned various biological enzymes may be variant sequences obtained by appropriately modifying the natural amino acid sequences. These modifications include, but are not limited to, substitution and / or deletion and / or addition of appropriate amino acids that do not affect the biological activity of the target protein, codon optimization suitable for host cell preferences, tag addition, fusion, etc.
[0026] In a third aspect, the present invention provides a nucleic acid molecule comprising a nucleotide sequence encoding the above-mentioned biological enzyme or a nucleotide sequence encoding the above-mentioned biological enzyme composition.
[0027] Preferably, the nucleic acid molecule comprises at least one of the following sequences: b1) Any of the nucleotide sequences shown in SEQ ID NO: 8-14; b2) A complementary, degenerate, or homologous sequence of any of the nucleotide sequences shown in SEQ ID NO: 8-14, wherein the homologous sequence is a nucleotide sequence that has at least 77% homology with any of the nucleotide sequences shown in SEQ ID NO: 8-14; b3) Nucleotide sequences that encode proteins with the same function through substitution and / or deletion and / or addition of one or more amino acid residues caused by artificial modification or natural nucleic acid polymorphism; b4) A nucleotide sequence that hybridizes under strict conditions with the nucleotide sequence shown in b1), b2), or b3) and is capable of encoding a protein with the same function.
[0028] The complementary sequence is a complementary sequence formed by the principle of base complementarity pairing. Preferably, the complementary sequence can be a nucleotide sequence that is completely or partially complementary to any of the nucleotide sequences shown in SEQ ID NO: 8-14 and encodes an amino acid sequence having the biological enzyme activity described in the first aspect.
[0029] The degenerate sequence refers to a sequence in which one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more) nucleotides in any of the nucleotide sequences shown in SEQ ID NO: 8-14 are changed, and the types of amino acids encoded by the changed nucleotide sequences remain unchanged, without affecting the function and expression level of the nucleic acid.
[0030] Preferably, the nucleotide sequence having at least 77% homology includes nucleotide sequences having at least 77%, at least 77.622%, at least 77.927%, at least 78%, at least 78.588%, at least 79%, at least 79.061%, at least 80%, at least 85%, at least 87%, at least 87.446%, at least 88%, at least 89%, at least 90%, at least 95%, at least 96%, at least 96.437%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% homology. Examples include mutant genes, alleles, or derivatives.
[0031] Preferably, the sequences homologous to any of the nucleotide sequences shown in SEQ ID NO: 8-14 are derived from gymnosperms, such as plants from the Pinaceae, Cupressaceae, Taxaceae, Ginkgoaceae, and Taxaceae families. More preferably, they are derived from Taxaceae and the Taxus genus (including but not limited to *Taxus baccata* L., *Taxus brevifolia* Nutt., *Taxus calcicola* LM Gao & Mich. Möller, *Taxus scanadensis* Marshall, *Taxus chinensis* (Pilg.) Rehd., *Taxus contorta* Griff., *Taxus cuspidata* Siebold & Zucc., *Taxus floridana* Nutt. ex Chapm., *Taxus florinii* Spjut., and *Taxus globosa*). Plants of the following genera and species: *Taxus mairei* (Lemée & Lév.) SY Hu, *Taxus phytonii* Spjut, *Taxus sumatrana* (Miq.) de Laub., *Taxus wallichiana* Zucc., *Taxus Huangshan type*, *Taxus Emei type*, *Taxus Qinling type*, *Taxus yunnanensis* WCCheng & L. K. Fu, *Taxus × media Rehder*, *Taxus ×hunnewelliana Rehder*, or plants of the Taxaceae family and the *Taxus* genus (including but not limited to *Taxus × hunnewelliana*). Pseudotaxus chienii (WC Cheng) WC Cheng).
[0032] Preferably, the nucleic acid molecule further comprises a nucleotide sequence encoding taxane 2α-hydroxylase (T2αH); Preferably, the nucleic acid molecule further comprises a nucleotide sequence encoding at least one of the following biological enzyme compositions: taxane 7β-hydroxylase (T7βH), taxane 7β-oxo-acetyltransferase (T7AT), taxane C4-C20 oxobutane synthase (TOT), or taxane 1β-hydroxylase (T1βH).
[0033] In some embodiments, the natural nucleotide sequences of the relevant nucleic acid molecules in this invention can be variant sequences from different genera of Taxaceae, or different species and / or different strains of the same genus. The artificial nucleotide sequences can be variant sequences obtained by appropriately modifying the natural nucleotide sequences. The modifications include, but are not limited to, appropriate nucleotide substitutions and / or additions and / or deletions that do not affect the biological activity of the target protein, and codon optimization suitable for host cell preferences.
[0034] Preferably, the nucleic acid molecule can be synthetic and / or isolated.
[0035] A fourth aspect of the present invention provides a biomaterial comprising any one of the following c1)-c6): c1) The above-mentioned nucleic acid molecules; c2) Expression cassettes containing the aforementioned nucleic acid molecules; c3) A recombinant vector containing the above-mentioned nucleic acid molecules or the expression cassette described in c2); c4) A host cell containing at least one of the above-described nucleic acid molecules, the expression cassette described in c2), or the recombinant vector described in c3); c5) Recombinant plant cells and / or recombinant animal cells and / or recombinant algal cells, wherein the recombinant plant cells and / or recombinant animal cells and / or recombinant algal cells contain at least one of the above-mentioned nucleic acid molecules or the expression cassette of c2) or the recombinant vector of c3); c6) A cell-free system containing the above-mentioned biological enzymes, or a composition of the above-mentioned biological enzymes.
[0036] Preferably, the host cell includes at least one of microbial cells, plant cells, animal cells, or algal cells.
[0037] Preferably, the microbial cells include, but are not limited to, bacteria or fungi.
[0038] Preferably, the microbial cells include at least one selected from Escherichia coli, yeast cells, Streptomyces, Pseudomonas, and Bacillus. Preferably, the microbial cells can be different combinations of microorganisms, such as combinations of different bacteria, combinations of different fungi, or combinations of bacteria and fungi, etc. For example, a combination of Escherichia coli and yeast. Combinations of different microorganisms can achieve the purpose of this invention through co-culture, multi-segment culture, etc.
[0039] Preferably, the plant cells include, but are not limited to, at least one of the following: tobacco cells, yew cells, white yew cells, artemisia cells, Arabidopsis cells, *Sclerotium spp.* cells, *Hemerocallis fulva* cells, *Ligustrum lucidum* cells, tomato cells, ginseng cells, American ginseng cells, cotton cells, sugarcane cells, potato cells, corn cells, wheat cells, rice cells, radish cells, duckweed cells, lettuce cells, rapeseed cells, or lettuce cells. Preferably, the plant cells can be protoplasts or suspension cells, etc.
[0040] Preferably, the animal cells include, but are not limited to, at least one of insect cells, mammalian cells, nematode cells, or fish cells.
[0041] Preferably, insect cells include, but are not limited to, S2 Drosophila cells or Sf9 cells. Preferably, mammalian cells include, but are not limited to, fibroblasts, lymphocytes, epithelial cells, and myelocytes. Examples include HEK293 cells, CHO cells, COS cells, BHK cells, HeLa cells, Chinese hamster ovary cells, Vero cells, SP2 / 0 cells, NS / 0 myeloma cells, hamster kidney cells, human B cells, human T cells (Jurkat), neurons, CV-I / EBNA cells, L cells, 3T3 cells, HEPG2 cells, and MDCK cells. Preferably, fish cells include, but are not limited to, zebrafish cells. It should be understood that recombinant plant cells and / or recombinant animal cells include recombinant plant cells and / or recombinant animal cells that have undergone non-biological transformation. Optionally, recombinant plant cells or recombinant animal cells can express biological enzymes and produce target products through cell culture.
[0042] Preferably, the algal cells include, but are not limited to, at least one of cyanobacteria, green algae, Synechococcus, Synechococcus, Anabaena or Chlamydomonas reinhardtii.
[0043] In an embodiment of the present invention, the expression cassette further includes regulatory elements, which include at least one of a promoter, enhancer, leader sequence, transposon, terminator, and marker gene.
[0044] In embodiments of the present invention, the expression cassette includes a natural promoter and / or a heterologous promoter.
[0045] In some implementations, the expression of nucleic acids in the expression cassette can be regulated by manipulating the copy number of the gene or operon in the cell. In some implementations, the expression of nucleic acids in the expression cassette can be regulated by manipulating the order of nucleic acids within the module. In some implementations, the expression of nucleic acids in the expression cassette is regulated by integrating one or more nucleic acids or operons into the chromosome.
[0046] In some embodiments, one or more nucleic acid molecules related to this invention are expressed in an expression vector. As used herein, a "vector" can be any of a wide range of nucleic acids, in which one or more desired sequences can be inserted via restriction enzyme digestion and ligation for delivery in different genetic environments or expression in host cells. Vectors are typically composed of DNA, but can also be composed of RNA.
[0047] Preferably, the carrier can be a eukaryotic carrier or a prokaryotic carrier.
[0048] Preferably, the vector can be a cloning vector or an expression vector.
[0049] Preferably, the vector includes plasmids, chloroplasts, viral vectors, bacteriophages, phage particles, sclerotia, F sclerotia, bacterial phages, or artificial chromosomes. Optionally, the viral vector includes adenovirus vectors, retrovirus vectors, or adeno-associated virus vectors. Optionally, the vector includes bacterial artificial chromosomes (BAC), plasmids, bacterial phage P1-derived vectors (PAC), yeast artificial chromosomes (YAC), or mammalian artificial chromosomes (MAC).
[0050] In one specific embodiment of the present invention, the vectors include, for example, pFastBac1, pYES2, pYES2.1, pESC-Ura, pESC-Trp, pESC-Leu, pESC-His, pPIC3.5K, pGAPZ, pPIC9k, pcDNA3.1, pMAL, pGEX2T, pGEX-4T-1, pBAD-DEST49, pPICZα, PGAPZα, and P. VL1393, PCDNA3.1, PEF1α, PTT5, pTAex3, pUSA, pYMB0, pHT43, pET28a, pET30a, pET32a, pET22b, pET28b, pMAL-c2, pCold-TF, pIJ702, pUCP19, pEAQ-HT, pYMB03, pHT43, etc., preferably pEAQ-HT vector or pESC-Leu vector.
[0051] In one specific embodiment of the present invention, the vector stably integrates the nucleic acid of the present invention into the genome of a cell. In another specific embodiment of the present invention, the nucleic acid of the present application is expressed in a vector.
[0052] In one specific embodiment of the present invention, the vector described in the present invention is applied to a plant or part thereof (e.g., the underside of a leaf) by means of microinjection, microparticle bombardment, viral vector infection, or by spraying, irrigation, or dusting, so that the nucleic acid or expression cassette is stably or unstablely integrated into the genome of the plant or part thereof, or the nucleic acid or expression cassette is expressed in the vector, so that the plant or part thereof transiently expresses the target gene.
[0053] In one specific embodiment of the present invention, the vector is transfected into cells, for example, into bacteria (e.g., Agrobacterium), and then the cells are injected into a plant or a part thereof (e.g., the underside of a leaf), so that the plant or the part thereof transiently expresses the target gene.
[0054] A fifth aspect of the present invention provides a method for producing the above-described biological enzyme, characterized in that at least one of the above-described nucleic acid molecule, the above-described expression cassette in the biological material, or the above-described recombinant vector in the biological material is transformed into a host cell, thereby causing the host cell to produce the above-described biological enzyme.
[0055] The methods described above for transforming nucleic acid molecules, expression cassettes, and recombinant vectors into host cells include, but are not limited to, Agrobacterium-mediated transformation, gene gun transformation, electroporation, polyethylene glycol (PEG) transformation, lipid transfection, heat shock, calcium phosphate precipitation, virus-mediated transformation, microinjection, and gene editing techniques. Host cells include microbial cells, plant cells, animal cells, and / or algal cells.
[0056] Preferably, the nucleic acid is recombinantly expressed in any of the aforementioned cells. Preferably, any of the cells described in this invention can be cultured in any type (rich or basic) and any composition of culture medium. Those skilled in the art will understand that routine optimization will allow the use of various types of culture media (e.g., LB medium (containing kanamycin and rifampin)). Selected culture media can be supplemented with various additional components. Some non-limiting examples of supplementary components include glucose, antibiotics, IPTG for gene induction, ATCC micromineral supplements, and glycolic acid. Preferably, a resuspension of MMA (containing MES, MgCl2, and acetylsuccinone) may also be added during culture.
[0057] Similarly, other aspects of the culture medium and the growth conditions for any of the aforementioned cells can be optimized through routine experiments. For example, pH and temperature are non-limiting examples of optimizable factors. In some embodiments, factors such as culture medium selection, culture medium replenishment, and temperature can influence the production level of taxane-based compounds. In some embodiments, the concentration and amount of supplementary components can be optimized.
[0058] Liquid cultures used for cell growth related to this invention can be stored in any culture vessel known and used in the art. In some embodiments, large-scale production in an aerated reaction vessel (such as a stirred reactor) can be used to generate large quantities of taxane skeleton compounds that can be recovered from the cell culture. In some embodiments, the taxane skeleton compounds are recovered from the gas phase of the cell culture, for example by adding an organic layer such as dodecane to the cell culture and recovering the taxane skeleton compounds from the organic layer.
[0059] A sixth aspect of the present invention provides a method for generating recombinant cells, the method comprising converting the aforementioned nucleic acid molecule, an expression cassette containing the aforementioned nucleic acid molecule, or a recombinant vector containing the aforementioned nucleic acid molecule or the expression cassette into cells. Preferably, the cells may be microbial cells, plant cells, animal cells, or algal cells.
[0060] A seventh aspect of the present invention provides a method for producing plants or plant cells, the method comprising converting at least one of the above-described nucleic acid molecules, or expression cassettes or recombinant vectors in the above-described biological materials, into plants or plant cells.
[0061] An eighth aspect of the present invention provides a plant or a portion thereof, said plant being any one of the following: (1) A plant formed by the growth of plant cells expressing the above-mentioned biological enzymes or biological enzyme compositions; (2) The offspring formed by self-pollination of the plant in (1), and the plant formed by the growth of the offspring; (3) The offspring formed by hybridization of the plant in (1) with other varieties, and the plant formed by the growth of the offspring.
[0062] Preferably, the plant part includes roots, stems, leaves, flowers, fruits, pollen, or seeds.
[0063] In a ninth aspect, the present invention provides a plant or a portion thereof that expresses the above-described biological enzyme or the above-described biological enzyme composition, or the plant or a portion thereof that comprises the above-described nucleic acid molecule, an expression cassette containing the above-described nucleic acid molecule, a recombinant vector containing the above-described nucleic acid molecule or the expression cassette, and / or a transgenic plant of the above-described biological material.
[0064] In a tenth aspect, the present invention provides a method for preparing a plant or a part thereof, the method comprising converting the above-mentioned nucleic acid molecule, an expression cassette containing the above-mentioned nucleic acid molecule, a recombinant vector containing the above-mentioned nucleic acid molecule or the expression cassette, and / or the above-mentioned biological material into a plant.
[0065] Preferably, the method includes steps of tissue culture and / or induction culture.
[0066] Eleventh aspect of the present invention provides a method for manufacturing a commercial product, comprising obtaining a plant or plant part thereof as described in any one of the eighth and ninth aspects, and manufacturing the commercial product from said plant or plant part thereof, wherein said commercial product comprises crude extracts, active pharmaceutical ingredients and / or pharmaceutical preparations of 2α-acetoxytaxusin and its derivatives or baccatin IV and its derivatives or baccatin VI and its derivatives, baccatin III and its derivatives and paclitaxel and its derivatives.
[0067] In a twelfth aspect, the present invention provides the use of at least one of the above-described bioenzyme, the above-described bioenzyme composition, the above-described nucleic acid molecule, and the above-described biomaterial in catalyzing the C-2 acetylation of taxane compounds.
[0068] Preferably, the taxane compound includes at least one of paclitaxel, taxane, docetaxel, or cabazitaxel.
[0069] Preferably, the taxane compounds include compounds of formula I or a pharmaceutically acceptable salt, isomer, or solvation thereof; or compounds of formula II or a pharmaceutically acceptable salt, isomer, or solvation thereof; or compounds of formula III or a pharmaceutically acceptable salt, isomer, or solvation thereof.
[0070] ; ; ; in, R1 is selected from -OAc, H, -OH, -OBz or -OMe; preferably, R1 is selected from -OAc, -OH or -OMe; more preferably, R1 is selected from -OAc or -OMe, and even more preferably, R1 is -OAc.
[0071] R2 is selected from -OAc, -OBz, H, -OH or =O; preferably, R2 is selected from -OH, -OAc or =O; more preferably, R2 is -OAc.
[0072] R3 is selected from =O, -OH, H, -OAc, , , or Preferably, R3 is selected from -OH, -OAc, , , or Further preferred, R3 is -OAc.
[0073] R4 is selected from -OAc, H or -OH; preferably, R4 is H.
[0074] R5 is selected from -OAc, H or -OH, preferably -OAc.
[0075] R6 is selected from -OH, H, -Oxylose, -OAc or -OMe; preferably, R6 is H.
[0076] R7 is selected from -OAc, H or -OH; preferably, R7 is H.
[0077] R8 is selected from -OH or H; R9 is selected from -OH or H.
[0078] Where Ac represents acetyl; Bz represents benzoyl; Me represents methyl; and Oxylose represents xylose.
[0079] In one specific embodiment of the present invention, the stereoisomers of the compound of formula II include: or .
[0080] A twelfth aspect of the present invention provides a method for catalytically acetylating the C-2 position of a compound of formula I or a pharmaceutically acceptable salt, isomer, or solvation thereof; or a compound of formula II or a pharmaceutically acceptable salt, isomer, or solvation thereof; or a compound of formula III or a pharmaceutically acceptable salt, isomer, or solvation thereof, the method comprising catalytically acetylating the C-2 position of a compound of formula I or a pharmaceutically acceptable salt, isomer, or solvation thereof; or a compound of formula II or a pharmaceutically acceptable salt, isomer, or solvation thereof; or a compound of formula III or a pharmaceutically acceptable salt, isomer, or solvation thereof, under the action of the aforementioned bioenzyme or the aforementioned bioenzyme composition, and / or... The method includes introducing at least one of the above-mentioned nucleic acid molecules, expression cassettes in the above-mentioned biological materials, or recombinant vectors in the above-mentioned biological materials into a host cell, and catalyzing the C-2 position acetylation of the compound of formula I or its pharmaceutically acceptable salt, isomer, or solvation; or the compound of formula II or its pharmaceutically acceptable salt, isomer, or solvation; or the compound of formula III or its pharmaceutically acceptable salt, isomer, or solvation under the action of the host cell or its produced biological enzymes.
[0081] In a thirteenth aspect of the present invention, a method for synthesizing taxane compounds with an acetyl group at the C-2 position is provided, the method comprising synthesizing taxane compounds with an acetyl group at the C-2 position using the aforementioned biological enzyme or the aforementioned biological enzyme composition, and / or, The method includes introducing at least one of the above-mentioned nucleic acid molecules, expression cassettes in the above-mentioned biological materials, or recombinant vectors in the above-mentioned biological materials into a host cell, and synthesizing a taxane compound with an acetyl group at the C-2 position under the action of the host cell or its produced biological enzymes.
[0082] Preferably, the taxane compounds with an acetyl group at the C-2 position include 2α-acetoxytaxusin and its derivatives or baccatin IV and its derivatives.
[0083] Preferably, the host cell includes microbial cells, plant cells, animal cells, or algal cells.
[0084] In one specific embodiment of the present invention, the host cell is a microbial cell, and the method for synthesizing taxane compounds with an acetyl group at the C-2 position further includes a fermentation culture step.
[0085] In one specific embodiment of the present invention, the host cell is a plant cell or an animal cell, and the method for synthesizing taxane compounds with an acetyl group at the C-2 position further includes a cell culture step.
[0086] Optionally, the culture method includes immobilized cell culture, two-stage culture, two-phase culture, addition of inducers (e.g., salicylic acid, silver nitrate, methyl jasmonate, arachidonic acid, ammonium citrate, etc.), addition of bypass inhibitors (e.g., chlormequat chloride), addition of substrates (e.g., taxanes), etc.
[0087] In this embodiment of the invention, the host cell is a plant cell, and the method for synthesizing Baccatin III and / or its intermediates further includes the steps of planting the plant, harvesting the plant, and extracting the product.
[0088] In one specific embodiment of the present invention, the host cell is a plant cell, and the method for synthesizing taxane compounds with an acetyl group at the C-2 position further includes the steps of planting plants, harvesting plants, and extracting products.
[0089] In a fourteenth aspect, the present invention provides a method for synthesizing taxane compounds, the method comprising: Taxane compounds with an acetyl group at the C-2 position were synthesized using the method described above; Taxanes with an acetyl group at the C-2 position were used as substrates for further catalytic synthesis of taxanes.
[0090] Preferably, the taxane compound includes at least one of paclitaxel, taxane, docetaxel, or cabazitaxel.
[0091] In a fifteenth aspect, the present invention provides a method for preparing the nucleic acid described in the third aspect above, the method comprising amplifying the genome or cDNA of the gymnosperm using primers, wherein the primer sequences comprise: upstream primer: 5'-ATGGACAAGTTACATGTAAATATCATTGAG-3' (SEQ ID NO: 23), downstream primer: 5'-TCATAACTCAGAGAAACACGCTC-3' (SEQ ID NO: 24). Alternatively, the primers comprise: upstream primer: 5'-ATGGAGAAGGGAAATGCGAG-3' (SEQ ID NO: 25), downstream primer: 5'-TTAGGCTTTCGGCATATATTTAATTATCATGG-3' (SEQ ID NO: 26). Alternatively, the primers comprise: upstream primer: 5'-ATGGAGAAGAGAAATGCGAGT-3' (SEQ ID NO: 27), downstream primer: 5'-TTAGGCTTTAGGCATATATTTAATTATCATGGC-3' (SEQ ID NO: 28). Alternatively, the primers may include: upstream primer: 5'-ATGGTGAAGGGAGGCTCA-3' (SEQ ID NO: 29), downstream primer: 5'-TCACGCTTTAGTTGCATATTTGC-3' (SEQ ID NO: 30).
[0092] In one specific embodiment of the present invention, the synthesis or preparation of the above-mentioned taxane compounds further includes an optimization step. The optimization step includes at least one of the following optimization methods: 1) Gene expression optimization: for example, codon optimization, transcription factor optimization, promoter optimization, fusion protein optimization, etc. Codon optimization (including identifying optimal codons for various organisms and methods for achieving codon optimization) is well known to those skilled in the art and can be performed using standard methods. The transcription factor ORCA3 gene is a MeJA-induced transcription factor that regulates the basic and secondary metabolism of plants. Optimization of gene expression can also be achieved by selecting suitable promoters and ribosome binding sites. In some embodiments, this may include selecting plasmids with high copy numbers, or plasmids with low or medium copy numbers. Gene expression can also be regulated by targeting transcription termination steps, which is achieved by introducing or eliminating structures such as stem loops. 2) Metabolic engineering optimization: for example, methods to increase the yield of secondary metabolites through metabolic engineering, such as overcoming rate-limiting steps, reducing metabolic flux to competing pathways, reducing catabolism, and overexpressing regulatory genes. 3) Optimization of regulatory factors: Add MeJA and its analogues, salicylic acid, arachidonic acid, crocin, etc. to regulate biosynthesis; 4) Optimization of culture methods: such as improving the culture medium, adjusting environmental factors such as culture temperature and time, optimizing the culture process, and combining cultures; for example, Zhou K et al. (DOI: 10.1038 / nbt.3095) gave a culture example of co-culturing Escherichia coli and yeast.
[0093] The beneficial effects of the technical solution of this invention are as follows: 1. The inventors screened candidate genes for acetyltransferases targeting the C-2 position, and identified T2AT enzymes that catalyze acetylation at the C-2 position by feeding them with taxanes, such as T2AT enzymes from different sources in Table 3; 2. By co-expressing T2AT with the baccatin IV biosynthesis genes (including T2αH, T7βH, T7AT, TOT, and T1βH), baccatin IV was detected after feeding with taxanes, thus elucidating the biosynthetic pathway of the taxane compound baccatin IV. Important taxane compounds such as baccatin VI, baccatin III, and paclitaxel can be further synthesized. 3. C-2 acetylation can protect the key active site of paclitaxel in advance, thereby reducing the generation of byproducts in non-paclitaxel biosynthetic pathways and driving metabolic flow to the paclitaxel-generating network. Attached Figure Description
[0094] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings, wherein: Figure 1 : Figure 1 Figure A in the middle is a flowchart; Figure 1 Figure B shows the chromatogram of the extracted ion of 2α-acetyltaxane, catalyzed by T2AT1, T2AT2, T2AT3, or T2AT4 acetylation at the C-2 position, with GFP serving as the control group. 100x refers to the mass spectrometry data being multiplied by 100.
[0095] Figure 2 : Figure 2 The chromatogram of the extracted ion, 2α-acetyltaxane, is obtained by acetylation of taxane at C-2 catalyzed by T2AT5, T2AT6, or T2AT7, with GFP serving as the control group. 100x refers to the mass spectrometry data being multiplied by 100.
[0096] Figure 3 : Figure 3 Figure A in the middle is a flowchart; Figure 3 Figure B shows the extracted ion chromatogram of baccatin IV generated by co-expression of T2AT1 with T2αH, T7βH, T7AT, TOT, and T1βH, with the combination lacking T2AT serving as the control group.
[0097] Figure 4 Secondary fragment ion diagrams of tobacco-derived compound 4 (baccatin IV) and standard baccatin IV.
[0098] Figure 5 : Multiple sequence alignment diagram of T2AT homologous genes. Detailed Implementation
[0099] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0100] The inventors screened candidate genes for acetyltransferases using transcriptome and metabolite co-expression analysis. Subsequently, they screened the T2AT gene and identified its function using transient transformation of tobacco benthamiana and substrate feeding experiments, followed by high-performance liquid chromatography detection.
[0101] Terminology Explanation: In this invention, the term "and / or" encompasses all combinations of all items connected by the term, and each combination should be considered as having been individually listed in this application. For example, "A and / or B" includes "A", "B", and "A and B". As another example, "A, B and / or C" includes "A", "B", "C", "A and B", "A and C", "B and C", and "A and B and C".
[0102] In this invention, the terms "comprising" or "including" are open-ended descriptions, containing the specified ingredients or steps described, as well as other specified ingredients or steps that do not substantially affect them.
[0103] In this invention, the term "bioenzyme" or "enzyme" refers to a biologically active protein or polypeptide, including both naturally occurring proteins and their variants and modified forms. A "variant" here means a substantially similar sequence. As will be readily understood by those skilled in the art, natural proteins may exhibit some differences between different species within the same genus or between different samples of the same species. These differences may include the deletion and / or addition of one or more amino acids at one or more internal sites and / or the substitution of one or more amino acids at one or more sites in a natural polypeptide. However, this does not affect their ability to perform the same or similar functions, and these proteins can be referred to as natural variants of the protein. Modifications of proteins include, but are not limited to, appropriate amino acid substitutions / additions / deletions that do not affect the biological activity of the target protein, truncation of N-terminal amino acids, codon optimization to suit host cell preferences, tag addition, fusion, etc. These proteins can be referred to as artificial variants of the protein. For guidance on appropriate amino acid substitutions that do not affect the biological activity of the target protein, see the model in Dayhoff et al.'s (1978) Atlas of Protein Sequence and Structure (Natl. Biomed. Res. Found., Washington, DC), which is incorporated herein by reference. Conservative substitutions, such as replacing one amino acid with another amino acid having similar properties, may be optimal.
[0104] In this invention, the term "amino acid" refers to any amino acid, including but not limited to α-amino acids, β-amino acids, γ-amino acids, and δ-amino acids. Examples of suitable amino acids are shown in Table 2.
[0105] The term "cell-free system" in this invention refers to a transformation system that simulates metabolic or molecular reactions in vivo without relying on the intact structure of living cells, but can achieve specific biochemical reactions. In layman's terms, T2AT enzyme refers to the purification of the protein, the expression of T2AT in cell extracts, or the in vitro enzymatic reaction of T2AT-expressing cell lysate.
[0106] The bioenzymes related to this invention can be isolated from materials containing the bioenzymes of this application from any source. Any method for obtaining the bioenzymes related to this invention is compatible with this invention. In this invention, the term "isolated" means removed from its natural environment or from other compounds present when the compound was first formed. The term "isolated" includes materials isolated from natural sources and materials recovered after preparation through recombinant expression in host cells (such as nucleic acids and proteins), or chemically synthesized compounds such as nucleic acid molecules, proteins, and peptides.
[0107] The term "T2AT" refers to taxane C-2-oxoacetyltransferase or its gene or the nucleotide molecule or sequence encoding the enzyme, which catalyzes the acetylation of the taxane backbone at the C-2 position.
[0108] The term "T2αH" refers to taxane 2α-hydroxylase or its gene, or the nucleotide molecule or sequence encoding the enzyme. This enzyme catalyzes the hydroxylation of the taxane backbone at the C-2 position.
[0109] The term "T7βH" refers to taxane 7β-hydroxylase or its gene, or the nucleotide molecule or sequence encoding the enzyme. This enzyme catalyzes the hydroxylation of the taxane backbone at the C-7 position.
[0110] The term "T7AT" refers to taxane 7β-acetyltransferase or its gene, or the nucleotide molecule or sequence encoding the enzyme. This enzyme catalyzes the acetylation of the taxane C-7 hydroxyl product.
[0111] The term "TOT" refers to taxane C4-C20 oxetane-forming enzyme or its gene, nucleotide molecule, or sequence encoding the enzyme. This enzyme catalyzes the oxidation of carbon-carbon double bonds in taxane molecules to oxetane and oxetane.
[0112] The term "T1βH" refers to taxane C1 hydroxylase or its gene, or the nucleotide molecule or sequence encoding the enzyme. This enzyme catalyzes the hydroxylation of the taxane backbone at the C-1 position.
[0113] The aforementioned biological enzymes in this application can be identified by those skilled in the art through databases or literature reports as other amino acid sequences or nucleotide sequences with the same function, and the specific meanings can be determined in conjunction with the context, as exemplarily shown in Table 1.
[0114] Table 1 In this invention, the term "nucleic acid" (or "nucleic acid molecule" or "polynucleotide") can refer to a polymeric form of nucleotides, which may include the sense and antisense strands of RNA, cDNA, genomic DNA, as well as the synthetic forms and mixed polymers described above. Nucleotides can refer to ribonucleotides, deoxyribonucleotides, or modified forms of any type of nucleotide. As used in this application, "nucleic acid molecule" is synonymous with "nucleic acid" and "polynucleotide." A nucleic acid molecule is generally at least 10 bases in length, unless otherwise stated. The term can refer to RNA or DNA molecules of indeterminate length. The term includes DNA in both single-stranded and double-stranded forms. Nucleic acid molecules may include one or both of naturally occurring and modified nucleotides linked together by naturally occurring and / or non-naturally occurring nucleotides.
[0115] For nucleotide sequences, molecular biology techniques well known in the art, such as polymerase chain reaction (PCR) and hybridization techniques outlined in this application, can be used to identify naturally occurring variants, i.e., substantially similar sequences. Nucleic acid molecules can be chemically or biochemically modified, or can contain non-natural or derived nucleotide bases. Such modifications include, for example, labeling, methylation, substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications (e.g., uncharged bonds: such as methyl phosphonate, triphosphate, phosphoramide, carbamate, etc.; charged bonds: such as thiophosphate, dithiophosphate, etc.; overhanging portions: such as peptides; intercalating agents: such as acridine, psoralen, etc.; chelating agents; alkylating agents; and modified bonds: such as α-anomeric nucleic acids, etc.). The term "nucleic acid molecule" also includes any topological conformation, including single-stranded, double-stranded, partially double-stranded, triple-stranded, hairpin, circular, and padlock conformations.
[0116] In this invention, the term "homology" is sometimes used to refer to the level of similarity (i.e., sequence similarity or identity) between two or more nucleic acid or amino acid sequences, expressed as a percentage of positional identity. Homology also refers to the concept of evolutionary relevance, typically demonstrated by similar functional properties between different nucleic acids or proteins sharing similar sequences. In some embodiments, the homologous sequences relevant to this invention can be obtained by comparing exemplary sequences in genomic or transcriptomic data of evolutionarily closely related species. For example, between different species within the same genus, or between different strains of the same species, homologous sequences obtained by comparing exemplary sequences in sample genomic or transcriptomic data can be expected by those skilled in the art to have the same or similar functions.
[0117] In this invention, the term "identity" refers to sequence similarity to a natural nucleic acid or amino acid sequence. Identity can be evaluated visually or using computer software. Using computer software, the identity between two or more sequences can be expressed as a percentage (%), which can be used to evaluate the identity between related sequences. Unless otherwise stated, the sequence identity values provided in this application refer to values obtained using the full-length sequence of this invention with the default parameters in Jalview version 2.11.2.7 (DOI: 10.1093 / bioinformatics / btp033) and the multiple alignment software MUSCLE version 3.8.31 (DOI: 10.1093 / nar / gkh340); or any equivalent program.
[0118] Other mathematical algorithms are known in the art and can be used to align two sequences. See, for example, the algorithm in Karlin and Altschul (1990) Proc. Natl. Acad. Sci. USA 87:2264, modified as in Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-5877. This algorithm was incorporated into the BLAST program of Altschul et al. (1990) J. Mol. Biol. 215:403. BLAST nucleotide searches can be performed using the BLASTN program (nucleotide query for nucleotide sequence search) to obtain nucleotide sequences homologous to the nucleic acid molecules of the present invention, or using the BLASTX program (translational nucleotide query for protein sequence search) to obtain protein sequences homologous to the nucleic acid molecules of the present invention. BLAST protein searches can be performed using the BLASTP program (protein query for protein sequence search) to obtain amino acid sequences homologous to the protein molecules of the present invention, or using the TBLASTN program (protein query for translated nucleotide sequences) to obtain nucleotide sequences homologous to the protein molecules of the present invention. For gap alignments to be performed for comparative purposes, Gapped BLAST (in BLAST 2.0) can be used as described in Altschul et al. (1997) Nucleic Acids Res. 25:3389. Alternatively, PSI-Blast can be used to perform iterative searches that detect distant relationships between molecules. See Altschul et al. (1997) supra. When using the BLAST, Gapped BLAST, and PSI-Blast programs, the default parameters of the respective programs (e.g., BLASTX and BLASTN) can be used. Alignments can also be performed manually by inspection.
[0119] In some embodiments, the nucleic acid described in this invention can be cloned from DNA containing said nucleic acid from any source, for example by PCR amplification and / or restriction enzyme digestion. Any method for obtaining the nucleic acid described in this invention is compatible with this invention. In some embodiments, the nucleic acid described in this invention is synthetic. “Synthetic nucleic acid” refers to a polynucleotide (i.e., DNA or RNA) molecule produced by chemical synthesis as an in vitro process. For example, synthetic DNA can be produced in a reaction process within an Eppendorf™ tube, such that the synthetic DNA is enzymatically produced from a native DNA or RNA chain. Other laboratory methods can be used to synthesize polynucleotide sequences. Oligonucleotides can be chemically synthesized on an oligonucleotide synthesizer using solid-phase synthesis with phosphorous acid. Synthetic oligonucleotides can be annealed together as complexes to produce “synthetic” polynucleotides. Other methods for the chemical synthesis of polynucleotides are known in the art and can be readily implemented for use in this application.
[0120] In this invention, the term "gene" refers to a nucleic acid fragment that expresses a specific protein. A "gene" includes the DNA region encoding a gene product, as well as all DNA regions that regulate the production of the gene product, regardless of whether such regulatory sequences are adjacent to coding and / or transcriptional sequences. Therefore, genes include, but are not limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, introns, and locus control regions.
[0121] In this invention, the term "gene product" refers to any product produced by a gene. For example, a gene product can be a direct transcription product of a gene (such as mRNA, tRNA, rRNA, antisense RNA, interfering RNA, ribozyme, structural RNA, or any other type of RNA), or it can be a protein produced by translating mRNA.
[0122] In this invention, the term "expression cassette" refers to a DNA fragment into which nucleic acids or polynucleotides can be inserted at specific restriction sites or through homologous recombination. As described in this application, the DNA fragment contains a polynucleotide encoding a biological enzyme, and the expression cassette and restriction sites are designed to ensure that the expression cassette is inserted into the appropriate reading frame for transcription and translation. In one embodiment, the expression cassette may include a polynucleotide encoding a target polypeptide and, in addition to the polynucleotide, elements that promote the transformation of a specific host cell. In one embodiment, the expression cassette may also include elements (regulatory elements) that allow enhanced expression of the polynucleotide encoding the biological enzyme in the host cell. These elements may include, but are not limited to, promoters, enhancers, response elements, leader sequences, transposons, terminators, marker genes, polyadenylated sequences, etc. Preferably, the promoter may be a natural promoter and / or a heterologous promoter. Preferably, the selection of an operable heterologous promoter may depend on many factors, such as the desired timing, location, and expression pattern, and responsiveness to specific biological or abiotic stimuli. Heterologous promoters include, but are not limited to, inducible promoters (promoters of gene expression induced by light, heat, trauma, fungi, symbiotic bacteria, or chemicals), constitutive promoters, tissue-specific promoters, etc. Examples of such promoters include Trc, T5, Tac, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, pL, T3, GAL1, GAL10, MET17, CUP1, AOX1, polyhedrin, CMV, EF1A, EFS, CAG, PGK1, CBh, SFFV, MSCV, SV40, mPGK, hPGK, UBC, human beta actin, Actin, CaMV35S, TEF1, GPD, GDS, Ubi, ADH1, GAP, actin5C, polyyubiquitin, altubulin, TRE, TRE3G, UAS, Ac5, P1, P10, etc. Preferably, the expression cassette can be an expression cassette containing a modified sequence. The modifications can include adding a 5' leader sequence to enhance the translational expression of the sequence. Alternatively, modifications can include adding regulatory elements or modifying the promoter. The expression cassette may also contain selection marker genes for selecting transformed host cells. Selection marker genes are used to select transformed host cells or tissues. Marker genes include genes encoding antibiotic resistance, such as genes encoding neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT), and genes conferring resistance to herbicides such as glufosinate, bromobenzonitrile, imidazolinone, and 2,4-dichlorophenoxyacetate (2,4-d)). Other selection markers include phenotypic markers such as β-galactosidase and fluorescent proteins such as green fluorescent protein (GFP), cyan fluorescent protein (CYP), and yellow fluorescent protein.
[0123] In this invention, the term "vector" refers to a vector that can introduce a DNA or RNA sequence (e.g., a foreign gene) into a host cell to transform the host and promote the expression (e.g., transcription and translation) of the introduced sequence. "Non-viral vector" refers to any vector that does not contain a virus. In some embodiments of this application, a "vector" is a DNA sequence containing at least one DNA replication origin and at least one selective marker gene. This includes, but is not limited to, plasmids, phages, bacterial artificial chromosomes (BACs), or viruses that carry foreign DNA into host cells. In one specific embodiment of this application, the vector is, for example, the pEAQ-HT vector. The vector may also include one or more genes, antisense molecules, and / or selective marker genes and other genetic elements known in the art. The vector can transduce, transform, or infect host cells, thereby causing the host cells to express nucleic acid molecules and / or proteins encoded by the vector. The term "plasmid" refers to a circular strand of nucleic acid capable of autosomal replication in prokaryotic or eukaryotic host cells. This term includes nucleic acids, which can be DNA or RNA, and can be single-stranded or double-stranded. The plasmid as defined may also include sequences corresponding to bacterial replication origins.
[0124] In some implementations, the "cloning vector" can autonomously replicate or integrate into the host cell genome. It is also characterized by one or more restriction endonuclease sites, which can be used to cleave the vector in a defined manner and ligate the desired DNA sequence into the vector, thus allowing the new vector to retain its ability to replicate in the host cell. When the vector is a plasmid, replication of the desired sequence can occur multiple times as the copy number of the plasmid increases in the host cell (e.g., bacteria), or it can occur only once in each host before the host cell reproduces through mitosis. When the vector is a bacteriophage, replication can occur actively during lysis or passively during lysogenic phase.
[0125] In some embodiments, the "expression vector" can be inserted with a desired DNA sequence through restriction enzyme digestion and ligation to effectively link it with a regulatory sequence and express it as an RNA transcript. The vector may also contain one or more marker sequences suitable for identifying whether cells have been transformed or transfected by the vector. Marker sequences may include, for example, genes encoding proteins that increase or decrease resistance or sensitivity to antibiotics or other compounds, genes encoding enzymes whose activity can be detected by standard methods known in the art (e.g., β-galactosidase, luciferase, or alkaline phosphatase), and genes that have a visible effect on the phenotype of transformed or transfected cells, hosts, colonies, or plaques (e.g., green fluorescent protein). Preferably, the vector is a vector capable of autonomously replicating and expressing the product of a structural gene present in the DNA fragment to which it is effectively linked.
[0126] The term "expression" refers to the biosynthesis of a gene product, including the transcription and / or translation of said gene product. "Expressing" or "producing" a protein or polypeptide from a DNA molecule means transcribing and translating the coding sequence to produce the protein or polypeptide, while "expressing" or "producing" a protein or polypeptide from an RNA molecule means translating the RNA coding sequence to produce the protein or polypeptide.
[0127] Gene expression can be influenced by external signals, such as exposure of cells, tissues, or organisms to a substance that can increase or decrease gene expression. Gene expression can also be regulated anywhere along the process from DNA to RNA to protein. Regulation of gene expression can be achieved through control of transcription, translation, RNA transport and processing, degradation of intermediate molecules (such as mRNA), or through activation, inactivation, segmentation, or degradation of specific protein molecules after their formation, or through a combination thereof. The exact nature of the regulatory sequences required for gene expression may vary between species or cell types, but generally should include, as needed, 5′ non-transcriptional and 5′ non-translational sequences relating to the initiation of transcription and translation, such as TATA boxes, capping sequences, CAAT sequences, etc. In particular, such 5′ non-transcriptional regulatory sequences will include promoter regions that include promoter sequences controlling transcriptional control of genes that are effectively linked. Regulatory sequences may also include enhancer sequences or desired upstream activator sequences. The vectors of the present invention may optionally include 5′ leader or signal sequences. The selection and design of suitable vectors is within the competence and judgment of those skilled in the art.
[0128] In some embodiments, when a nucleic acid molecule encoding the bioenzyme described in this invention is expressed in a host cell, a variety of transcriptional control sequences (e.g., promoter / enhancer sequences) can be used to guide its expression. The promoter can be a natural promoter, i.e., the promoter of a gene in its endogenous environment, which provides normal regulation of gene expression. In some embodiments, the promoter can be constitutive, i.e., the promoter continuously transcribes its associated gene without regulation. Various conditional promoters can also be used, such as promoters controlled by the presence or absence of a molecule. Chemically regulated promoters can be used to regulate gene expression in the host by applying exogenous chemical regulators. Depending on the purpose, the promoter can be a chemically induced promoter, where the application of the chemical substance induces gene expression, or a chemically repressive promoter, where the application of the chemical substance represses gene expression. Chemically regulated promoters are known in the art, including but not limited to the maize In2-2 promoter (activated by a benzenesulfonamide herbicide safener), the maize GST promoter (activated by a hydrophobic electrophilic compound used as a pre-germination herbicide), and the tobacco PR-1a promoter (activated by salicylic acid). Other promoters of interest that regulate chemical substances include glucocorticoid-inducible promoters in steroid-responsive promoters, as well as tetracycline-inducible and tetracycline-repressive promoters.
[0129] Expression vectors containing all the necessary expression elements are commercially available and well-known to those skilled in the art. See, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold SpringHarbor Laboratory Press, 1989. Cellular genetic engineering is performed by introducing exogenous DNA (RNA) into cells. This exogenous DNA (RNA) is placed under the efficient control of transcriptional elements to allow the exogenous DNA to be expressed in the cells.
[0130] In this invention, the term "transformation" encompasses all techniques that introduce nucleic acid molecules into cells. These include, but are not limited to: viral vector transfection, plasmid vector transformation, gene gun transformation, electroporation, polyethylene glycol (PEG) transformation, lipid infection, microinjection, Agrobacterium-mediated transformation, direct DNA uptake, Whiskers™-mediated transformation, micro-projectile lipid transfection, heat shock, calcium phosphate precipitation, and gene editing techniques. These techniques can be used for both stable and transient transformation of cells. "Stable transformation" refers to the introduction of a nucleic acid fragment into the genome of a host organism, resulting in genetic stability. Once stable transformation occurs, the nucleic acid fragment is stably integrated into the genome of the host organism and any subsequent generations. Host organisms containing the transformed nucleic acid fragment are referred to as "transgenic" organisms. "Transient transformation" refers to the introduction of a nucleic acid fragment into a host organism or its nucleus or DNA-containing organelles, resulting in gene expression without genetic stability.
[0131] In some embodiments, to transform a host or host cell, the nucleotide sequence of the present invention can be inserted into any vector known in the art suitable for expressing the nucleotide sequence in a host or host cell using standard techniques. The choice of vector depends on the preferred transformation technique and the target host species to be transformed. The transformation method depends on the host cell to be transformed, the stability of the vector used, the expression level of the gene product, and other parameters.
[0132] In this invention, the term "plant" includes seeds, plant cells, plant protoplasts, plant cell tissue cultures from which regenerate plants, plant callus, plant masses, and whole plant cells in plants or plant parts such as embryos, pollen, ovules, seeds, tubers, propagules, leaves, flowers, branches, fruits, roots, root tips, anthers, etc. Progeny, variants, and mutants of regenerated plants are also included within the scope of this invention, provided that these parts contain introduced polynucleotides. As used herein, unless otherwise expressly stated or apparent from the context of use, "progeny" and "progeny plant" include any subsequent generations of a plant, whether produced by sexual and / or asexual reproduction.
[0133] The term "transgenic plant" refers to an equivalent term for "plant" as described above, wherein the plant comprises a heterologous nucleic acid molecule, heterologous polynucleotide, or heterologous polynucleotide construct introduced into the plant by any stable and transient transformation method disclosed elsewhere in this application or otherwise known in the art. Such transgenic plants also refer to, for example, the plant in which the heterologous nucleic acid molecule, heterologous polynucleotide, or heterologous polynucleotide construct was first introduced, and any of its progeny plants comprising said heterologous nucleic acid molecule, heterologous polynucleotide, or heterologous polynucleotide construct.
[0134] In this invention, the use of the term "nucleic acid" is not intended to limit the invention to polynucleotide molecules comprising DNA or RNA. Those skilled in the art will recognize that the nucleic acids of this invention include polynucleotide molecules composed of deoxyribonucleotides (i.e., DNA), ribonucleotides (i.e., RNA), or combinations thereof. Such deoxyribonucleotides and ribonucleotides include naturally occurring molecules and synthetic analogs, including but not limited to nucleotide analogs or modified backbone residues or linkages that are synthetic, naturally occurring, or non-natural, have binding properties similar to a reference nucleic acid, and are metabolized in a manner similar to a reference nucleotide. Examples of such analogs include, but are not limited to, thiophosphates, aminophosphates, methylphosphonates, chiral methylphosphonates, 2-O-methylribonucleotides, and peptide-nucleic acids (PNAs). The polynucleotide molecules of this invention also include all forms of polynucleotide molecules, including but not limited to single-stranded, double-stranded, hairpin, stem-loop, and other similar structures. Furthermore, the nucleotide sequences disclosed in this application also include complementary sequences to the exemplary nucleotide sequences.
[0135] In this invention, the "strict conditions" can be any of low-strict, medium-strict, or high-strict conditions. "Low-strict conditions" are, for example, 5×SSC, 5×Denhardt solution, 0.5% SDS, 50% formamide, and 32°C. "Medium-strict conditions" are, for example, 5×SSC, 5×Denhardt solution, 0.5% SDS, 50% formamide, and 42°C. "High-strict conditions" are, for example, 5×SSC, 5×Denhardt solution, 0.5% SDS, 50% formamide, and 50°C. Under these conditions, the higher the temperature, the more efficiently high-homogeneity DNA can be obtained. Factors affecting hybridization strictness include temperature, probe concentration, probe length, ionic strength, time, salt concentration, etc. Those skilled in the art can achieve the same strict conditions by appropriately selecting these factors.
[0136] In this invention, the term "solvent" refers to the physical combination of the compound of this invention with one or more solvent molecules. This physical combination includes various degrees of ionic and covalent bonding, such as hydrogen bonding. In some cases, the solvate can be isolated, for example when one or more solvent molecules are incorporated into the crystal lattice of a crystalline solid. Solvents include solution phases and separable solvates. Representative solvates include ethanolides or methanolides, etc.
[0137] In this invention, the "stereoisomer" includes those that exist in the form of enantiomers, diastereomers, and geometric isomers.
[0138] In this invention, "pharmaceutically acceptable" means that the biological activity and properties of the active ingredient of the applied product neither significantly stimulate the organism nor inhibit it.
[0139] In this invention, "conservative amino acid substitution" refers to the substitution of one amino acid by another amino acid within the same category, such as the substitution of one acidic amino acid by another, one basic amino acid by another, or one neutral amino acid by another. For example, amino acids can be grouped according to the properties of common side chains: (1) hydrophobic: leucine, Met, Ala, Val, Leu, Ile; (2) neutral hydrophilic: Cys, Ser, Thr, Asn, Gln; (3) acidic: Asp, Glu; (4) basic: His, Lys, Arg; (5) residues affecting chain orientation: Gly, Pro; (6) aromatic: Trp, Tyr, Phe; The conservative amino acid substitution can also refer to the substitution of one amino acid in the above group by another amino acid in the same group, where the conservative amino acid substitution does not change the activity of the amino acid sequence. Common amino acids and their English abbreviations are shown in Table 2.
[0140] Table 2: Amino Acids and Their English Names and Abbreviations Example 1: Cloning of the taxane C-2 acetyltransferase T2AT gene 1. cDNA Acquisition Southern yew was preserved in the laboratory. After collecting young leaves, they were immediately flash-frozen in liquid nitrogen. A portion of the sample was taken, ground in a mortar and pestle, and 100 mg was added to a 1.5 mL centrifuge tube. Lysis buffer was added, and RNA was extracted (Plant Total RNA Isolation Kit Plus, FOREGENE, China). The mRNA was then reverse transcribed into cDNA using a reverse transcription kit (HiScript III 1st StrandcDNA Synthesis Kit, Vazyme, China).
[0141] 2. Cloning of taxane C-2 acetyltransferase T2AT Genome and transcriptome analyses of *Taxus wallichiana* and *Taxus chinensis* yielded several candidate genes for acetyltransferases from both species. Among them, the amino acid and nucleotide sequences of the T2AT1 reading frame are shown in SEQ ID NO: 1 (homological sequences are shown in SEQ ID NO: 2-7 in the sequence listing) and SEQ ID NO: 8 (homological sequences are shown in SEQ ID NO: 9-14 in the sequence listing), respectively. The origin and homology of each sequence are shown in Table 3, and the homology alignment results are shown in […]. Figure 5 .
[0142] Table 3 3. Carrier Construction Based on the nucleotide sequences of each gene in Table 3, primers for amplifying these genes were designed (see Table 4). These primers contained partial sequences of the tobacco expression vector pEAQ-HT. Using cDNA from *Taxus wallichiana* leaves as templates, the T2AT1, T2AT2, T2AT3, and T2AT4 genes were amplified by PCR. Using cDNA from *Taxus chinensis* leaves as templates, the T2AT5, T2AT6, and T2AT7 genes were amplified by PCR. The PCR products were recovered from the gel and ligated into the linear pEAQ-HT vector digested with RruI and XhoI. Cloning and sequencing were performed using the ClonExpress one-step cloning kit (Novizan). The positive recombinant plasmids were named pEAQ-T2AT1, pEAQ-T2AT2, pEAQ-T2AT3, pEAQ-T2AT4, pEAQ-T2AT5, pEAQ-T2AT6, and pEAQ-T2AT7.
[0143] Of the primers used, P1 and P2 were used to amplify T2AT1, P3 and P4 were used to amplify T2AT2, P5 and P6 were used to amplify T2AT3, and P7 and P8 were used to amplify T2AT4. The primer sequences are shown in Table 4.
[0144] Table 4 Primer Sequences Note: The lowercase letter sequence in the sequence is the vector sequence, and the uppercase letter sequence is the T2AT specific primer sequence.
[0145] Example 2: Functional identification of the taxane C-2 acetyltransferase T2AT gene The constructed expression vectors pEAQ-T2AT1, pEAQ-T2AT2, pEAQ-T2AT3, and pEAQ-T2AT4 were transformed into Agrobacterium GV3101, respectively, to obtain GV3101 / pEAQ-T2AT1, GV3101 / pEAQ-T2AT2, GV3101 / pEAQ-T2AT3, and GV3101 / pEAQ-T2AT4 transgenic Agrobacterium. Positive single colonies were picked and inoculated into LB medium (containing 50 μg / mL kanamycin and 25 μg / mL rifampin) and cultured at 28°C for 24 h. The cultured bacterial solution was then diluted 1:100 in 10 mL of LB medium (containing 50 μg / mL kanamycin and 25 μg / mL rifampin) and cultured overnight at 28°C until OD500. 600 The value was 0.8-1.0; MMA resuspension (10 mM MES, 10 mM MgCl2, 150 μM acetylsuccinone) was added to the final OD value. 600 The concentration is 0.8-1.0, and it should be left to stand at room temperature for 1-2 hours.
[0146] Using a 1 mL syringe with the needle removed, inject the resuspended transgenic Agrobacterium combination (combination 1: T2AT1+T2αH, or combination 2: T2AT2+T2αH, or combination 3: T2AT3+T2αH, or combination 4: T2AT4+T2αH, the nucleotide and amino acid sequences of T2αH are shown in Table 1) into the underside of leaves of 4-6-week-old Nicotiana benthamiana. After drying in light for 1-2 hours, transfer to a dark place for 24 hours of culture, and then transfer to normal light for culture. 3-4 days after injecting the above transgenic Agrobacterium combinations, inject paclitaxel (1) aqueous solution into the leaves that have been injected with the above transgenic Agrobacterium combination 1, combination 2, combination 3 or combination 4, and continue to culture for 24 hours.
[0147] Tobacco samples were ground using a grinder, and the powder was transferred to a 5 mL centrifuge tube and freeze-dried. 20 mg of the powder was taken, 1 mL of methanol was added, and the mixture was ultrasonically extracted for 30 min. After centrifugation at 14000 rpm for 15 min at room temperature, the supernatant was transferred to a new 2 mL centrifuge tube and centrifuged again at 14000 rpm for 15 min at room temperature. 200 μL of the supernatant was transferred to a sample vial for detection by LC-MS. The detection revealed that a C-2 acetylated compound was detected in leaves fed with the substrate taxane in the presence of T2AT1, T2AT2, T2AT3, or T2AT4. LC-MS analysis showed that this product was the target product, 2α-acetoxytaxusin (see [link to LC-MS]). Figure 1 ).
[0148] Example 3: Functional identification of T2AT homologous genes derived from Taxus chinensis The vector pEAQ-T2AT5, gene pEAQ-T2AT6, and pEAQ-T2AT7 constructed in Example 1 were transformed into Agrobacterium GV3101 to obtain GV3101 / pEAQ-T2AT5, GV3101 / pEAQ-T2AT6, and GV3101 / pEAQ-T2AT7 transgenic Agrobacterium. Positive single colonies were picked and inoculated into LB medium (containing 50 μg / mL kanamycin and 25 μg / mL rifampin) and cultured at 28°C for 24 h. The cultured bacterial solution was then diluted 1:100 in 10 mL of LB medium (containing 50 μg / mL kanamycin and 25 μg / mL rifampin) and cultured overnight at 28°C until OD (Organic Dysplasia). 600 The value was 0.8-1.0; MMA resuspension (10 mM MES, 10 mM MgCl2, 150 μM acetylsalicylic acid) was added to the final OD value. 600 The concentration is 0.8-1.0, and it should be left to stand at room temperature for 1-2 hours.
[0149] Using a 1 mL syringe with the needle removed, inject the resuspended transgenic Agrobacterium combination (combination 1: T2AT1+T2αH, or combination 5: T2AT5+T2αH, or combination 6: T2AT6+T2αH, or combination 7: T2AT7+T2αH, the nucleotide and amino acid sequences of T2αH are shown in Table 1) into the underside of leaves of 4-6 weeks-old Nicotiana benthamiana. After drying in light for 1-2 h, transfer to a dark place for 24 h of culture, and then transfer to normal light for culture. 3-4 days after injecting the above transgenic Agrobacterium combinations, inject paclitaxel (1) aqueous solution into the leaves that have been injected with the above transgenic Agrobacterium combinations 1, 5, 6 or 7, and continue to culture for 24 h.
[0150] Tobacco samples were ground using a grinder, and the powder was transferred to a 5 mL centrifuge tube and freeze-dried. 20 mg of the powder was taken, 1 mL of methanol was added, and the mixture was ultrasonically extracted for 30 min. After centrifugation at 14000 rpm for 15 min at room temperature, the supernatant was transferred to a new 2 mL centrifuge tube and centrifuged again at 14000 rpm for 15 min at room temperature. 200 μL of the supernatant was transferred to a sample vial for LC-MS detection. The presence of T2AT1, T2AT5, T2AT6, and T2AT7 indicated the detection of a C-2 acetylated compound in leaves fed with the substrate taxane. LC-MS analysis confirmed that this product was the target product, 2α-acetoxytaxusin (see [link to LC-MS]). Figure 2This indicates that T2AT enzymes derived from Taxus chinensis (such as T2AT5, T2AT6, and T2AT7) also have the function of catalyzing C-2 acetylation.
[0151] Example 4: T2AT1-catalyzed baccatin IV formation Following a similar method to Example 1, GV3101 / PEAQ-T2AT1, GV3101 / PEAQ-T2αH, GV3101 / PEAQ-T7βH, GV3101 / PEAQ-T7AT, GV3101 / PEAQ-TOT, and GV3101 / PEAQ-T1βH were constructed. Positive single clones were picked and inoculated into LB medium (containing 50 μg / mL kanamycin and 25 μg / mL rifampin) and cultured at 28°C for 24 h. The cultured bacterial suspension was then diluted 1:100 in 10 mL of LB medium (containing 50 μg / mL kanamycin and 25 μg / mL rifampin) and cultured overnight at 28°C until the OD600 value reached 0.8-1.0. MMA resuspension (10 mM MES, 10 mM MgCl2, 150 μM...) was added. Acetyleugenol) was added until the final OD600 was 0.8-1.0. Then, the above recombinant Agrobacterium was mixed at a ratio of 1:1 to obtain the recombinant Agrobacterium combination (T2AT1+T2αH+T7βH+T7AT+TOT+T1βH combination, or T2αH+T7βH+T7AT+TOT+T1βH combination), and allowed to stand at room temperature for 1-2 h.
[0152] Using a 1 mL syringe with the needle removed, inject the resuspended recombinant Agrobacterium combination (T2AT1+T2αH+T7βH+T7AT+TOT+T1βH combination, or T2αH+T7βH+T7AT+TOT+T1βH combination, i.e., combination without T2AT1, wherein the nucleotide and amino acid sequences of T2αH, T7βH, T7AT, TOT and T1βH are shown in Table 1, the amino acid sequence of T2αH is SEQ ID NO: 37, the nucleotide sequence is SEQ ID NO: 38; the amino acid sequence of T7βH is SEQ ID NO: 39, the nucleotide sequence is SEQ ID NO: 40) into the underside of leaves of 4-6-week-old Nicotiana benthamiana. After drying under light for 1-2 h, transfer to a dark place for 24 h of culture, and then transfer to normal light for culture. 3-4 days after injection of the recombinant Agrobacterium combination, inject paclitaxel aqueous solution into the leaves injected with the recombinant Agrobacterium combination, and continue to culture for 24 h.
[0153] The tobacco sample was ground using a grinder, and the sample powder was transferred to a 5 mL centrifuge tube and freeze-dried. 20 mg of the powder was taken, 1 mL of methanol was added, and the mixture was ultrasonically extracted for 30 min. After centrifuging at 14,000 rpm for 15 min at room temperature, the supernatant was transferred to a new 2 mL centrifuge tube and centrifuged again at 14,000 rpm for 15 min at room temperature. 200 μL of the supernatant was transferred to a sample vial and detected by LC-MS.
[0154] like Figure 3 As shown, in the presence of T2AT1, feeding with the substrate taxane resulted in the detection of baccatin IV production. Figure 4 The study also showed that, in the presence of T2AT1, the product synthesized in tobacco was consistent with the secondary fragment ions of the baccatin IV standard (standard 4), confirming the acetyl catalytic activity of T2AT1 at the C-2 position and proving that the inventor's T2AT1, in combination with other enzymes, catalyzes the formation of baccatin IV.
[0155] In summary, this invention provides a bioenzyme T2AT (including T2AT1-7) that catalyzes the C-2 acetylation of taxanes, laying the foundation for a more comprehensive understanding of the paclitaxel biosynthetic pathway. It can be better applied to the modification and optimization of the synthetic pathways of paclitaxel and other C-2 related taxane compounds, and has significant economic and ecological value.
[0156] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solution of the present invention, and these simple modifications all fall within the protection scope of the present invention.
[0157] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present invention will not describe the various possible combinations separately.
Claims
1. A biological enzyme, characterized in that, It contains at least one of the following sequences: a1) Any of the amino acid sequences shown in SEQ ID NO: 1-7; a2) Proteins with the same function obtained by substituting and / or deleting and / or adding one or more amino acid residues of any of the amino acid sequences shown in SEQ ID NO: 1-7. a3) Proteins that have at least 50% homology with and have the same function as any of the amino acid sequences shown in SEQ ID NO: 1-7; A fusion protein is obtained by attaching a tag to the N-terminus and / or C-terminus of at least one of the amino acid sequences shown in a4)(a1)(-a3).
2. A biological enzyme composition, characterized in that, The bioenzyme composition comprises the bioenzyme of claim 1.
3. The bioenzyme composition according to claim 2, characterized in that, The bioenzyme composition also includes taxane 2α-hydroxylase (T2αH).
4. The bioenzyme composition according to claim 2 or 3, characterized in that, The bioenzyme composition further includes one or more of the following: taxane 7β-hydroxylase (T7βH), taxane 7β-oxo-acetyltransferase (T7AT), taxane C4-C20 oxobutane synthase (TOT), and taxane 1β-hydroxylase (T1βH).
5. A nucleic acid molecule, characterized in that, The nucleic acid molecule comprises a nucleotide sequence encoding the biological enzyme of claim 1 or a nucleotide sequence encoding the biological enzyme composition of any one of claims 2-4.
6. The nucleic acid molecule according to claim 5, characterized in that, The nucleic acid molecule comprises at least one of the following sequences: b1) Any of the nucleotide sequences shown in SEQ ID NO: 8-14; b2) A complementary, degenerate, or homologous sequence of any of the nucleotide sequences shown in SEQ ID NO: 8-14, wherein the homologous sequence is a nucleotide sequence that has at least 77% homology with any of the nucleotide sequences shown in SEQ ID NO: 8-14; b3) Nucleotide sequences that encode proteins with the same function through substitution and / or deletion and / or addition of one or more amino acid residues caused by artificial modification or natural nucleic acid polymorphism; b4) A nucleotide sequence that hybridizes under strict conditions with the nucleotide sequence shown in b1), b2), or b3) and is capable of encoding a protein with the same function.
7. The nucleic acid molecule according to claim 5 or 6, characterized in that, The nucleic acid molecule also contains a nucleotide sequence encoding taxane 2α-hydroxylase (T2αH); Preferably, the nucleic acid molecule further comprises a nucleotide sequence encoding at least one of the following biological enzyme compositions: taxane 7β-hydroxylase (T7βH), taxane 7β-oxo-acetyltransferase (T7AT), taxane C4-C20 oxobutane synthase (TOT), or taxane 1β-hydroxylase (T1βH).
8. A biomaterial, characterized in that, The biomaterials mentioned include any one of c1)-c6) below: c1) The nucleic acid molecule according to any one of claims 5-7; c2) An expression cassette containing any of the nucleic acid molecules described in claims 5-7; c3) A recombinant vector containing any of the nucleic acid molecules according to claims 5-7 or the expression cassette described in c2); c4) A host cell containing at least one of the nucleic acid molecules of any one of claims 5-7, the expression cassette of c2), or the recombinant vector of c3; c5) Recombinant plant cells and / or recombinant animal cells and / or recombinant algal cells, wherein the recombinant plant cells and / or recombinant animal cells and / or recombinant algal cells comprise a recombinant microorganism containing at least one of the nucleic acid molecules of any one of claims 5-7 or the expression cassette of c2) or the recombinant vector of c3); c6) A cell-free system containing the biological enzyme of claim 1, or containing the biological enzyme composition of any one of claims 2-4.
9. The biomaterial according to claim 8, characterized in that, The host cell includes at least one of microbial cells, plant cells, animal cells, or algal cells; The microbial cells include, but are not limited to, bacteria or fungi; preferably, the microbial cells include at least one of Escherichia coli, yeast cells, Streptomyces, Pseudomonas, Bacillus, etc. Preferably, the plant cells include, but are not limited to, at least one of the following: tobacco cells, yew cells, white yew cells, artemisia cells, Arabidopsis cells, sclerotium cells, liverwort cells, tomato cells, ginseng cells, American ginseng cells, cotton cells, sugarcane cells, potato cells, corn cells, wheat cells, rice cells, radish cells, duckweed cells, lettuce cells, rapeseed cells, or lettuce cells; Preferably, the animal cells include, but are not limited to, at least one of insect cells, mammalian cells, nematode cells, or fish cells; Preferably, the algal cells include, but are not limited to, at least one of cyanobacteria, green algae, Synechococcus, Synechococcus, Anabaena, or Chlamydomonas.
10. A method for producing the biological enzyme of claim 1, characterized in that, At least one of the nucleic acid molecules of any one of claims 5-7, the expression cassette of the biological material of claim 8 or 9, or the recombinant vector of the biological material of claim 8 or 9 is transformed into a host cell, thereby causing the host cell to produce the biological enzyme.
11. The use of at least one of the bioenzyme of claim 1, the bioenzyme composition of any one of claims 2-4, the nucleic acid molecule of any one of claims 5-7, or the biomaterial of claim 8 or 9 in catalyzing the C-2 acetylation of taxane compounds.
12. The application according to claim 11, characterized in that, The taxane compounds include compounds of formula I or a pharmaceutically acceptable salt, isomer or solvate thereof; or compounds of formula II or a pharmaceutically acceptable salt, isomer or solvate thereof; or compounds of formula III or a pharmaceutically acceptable salt, isomer or solvate thereof. ; ; ; in, R1 is selected from -OAc, H, -OH, -OBz, or -OMe; R2 is selected from -OAc, -OBz, H, -OH, or =O; R3 is selected from =O, -OH, H, -OAc, , , or ; R4, R5, and R7 are each independently selected from -OAc, H, or -OH; R6 is selected from -OH, H, -Oxylose, -OAc, or -OMe; R8 is selected from -OH or H; R9 is selected from -OH or H.
13. A method for catalytically acetylating a compound of formula I or a pharmaceutically acceptable salt, isomer, or solvation thereof; or a compound of formula II or a pharmaceutically acceptable salt, isomer, or solvation thereof; or a compound of formula III or a pharmaceutically acceptable salt, isomer, or solvation thereof, characterized in that, The method comprises, under the action of the bioenzyme of claim 1 or the bioenzyme composition of any one of claims 2-4, catalyzing the C-2 position acetylation of a compound of formula I or a pharmaceutically acceptable salt, isomer, or solvation thereof; or a compound of formula II or a pharmaceutically acceptable salt, isomer, or solvation thereof; or a compound of formula III or a pharmaceutically acceptable salt, isomer, or solvation thereof; and / or, The method comprises introducing at least one of the nucleic acid molecule of any one of claims 5-7, the expression cassette of the biological material of claim 8 or 9, or the recombinant vector of the biological material of claim 8 or 9 into a host cell, and catalyzing, under the action of the host cell or a biological enzyme produced therefrom, the C-2 position acetylation of the compound of formula I or a pharmaceutically acceptable salt, isomer or solvation thereof; or the compound of formula II or a pharmaceutically acceptable salt, isomer or solvation thereof; or the compound of formula III or a pharmaceutically acceptable salt, isomer or solvation thereof.
14. A plant or a part thereof, characterized in that, The plant or a portion thereof expresses the biological enzyme of claim 1 or the biological enzyme composition of any one of claims 2-4; or, the plant or a portion thereof contains at least one of the nucleic acid molecule of any one of claims 5-7, the expression cassette of the biological material of claim 8 or 9, or the recombinant vector of the biological material of claim 8 or 9.
15. A method for preparing a plant or a part thereof, characterized in that, The method comprises transforming at least one of the nucleic acid molecules described in any one of claims 5-7, the expression cassette in the biological material of claim 8 or 9, or the recombinant vector in the biological material of claim 8 or 9 into a plant.
16. A method for manufacturing a commercial product, characterized in that, The method includes obtaining a plant or a portion thereof as described in claim 14, and manufacturing the commercial product from said plant or plant portion thereof, wherein said commercial product is a crude extract, active pharmaceutical ingredient, and / or pharmaceutical preparation containing at least one of baccatin III and its derivatives, paclitaxel and its derivatives, taxane and its derivatives, docetaxel and its derivatives, cabazitaxel and its derivatives.