Protein related to phloretin enzyme activity as well as coding gene and application thereof

By expressing a protein with the amino acid sequence sequence 2 in Escherichia coli Rosetta 2 (DE3), the synthesis of phlorizin and trifolin was catalyzed, solving the problem of the lack of effective catalysis for phlorizin synthesis in existing technologies and realizing the possibility of industrial production.

CN122012432APending Publication Date: 2026-05-12INSTITUTE OF CHINESE MATERIA MEDICA CHINA ACADEMY OF CHINESE MEDICAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INSTITUTE OF CHINESE MATERIA MEDICA CHINA ACADEMY OF CHINESE MEDICAL SCIENCES
Filing Date
2024-11-11
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

How to provide proteins with phlorizinase activity that catalyze the production of phlorizin and/or trifolin.

Method used

A protein with the amino acid sequence sequence 2 and its encoding gene Ohi.6618 are provided. The protein is fused with the target protein using DNA recombination technology and expressed in Escherichia coli Rosetta 2 (DE3) to prepare a product that catalyzes the production of phlorizin and/or trifolin.

Benefits of technology

This study achieved the in vitro catalytic production of phlorizin and trifolin, providing the possibility for industrial-scale production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122012432A_ABST
    Figure CN122012432A_ABST
Patent Text Reader

Abstract

The invention discloses a protein related to phloretin enzyme activity as well as a coding gene and application thereof. The technical problem to be solved is how to provide a phloretin enzyme activity related protein. The invention specifically discloses any one of the following applications: A1) application of the protein in catalyzing phloretin to generate phlorizin and / or trilobatin and / or preparing a product for catalyzing phloretin to generate phlorizin and / or trilobatin; a2) application in production of phlorizin and / or preparation of a product for producing phlorizin; and A3) producing trilobatin and / or preparing a product for producing the trilobatin. The protein is any one of the following: B1) a protein with an amino acid sequence as shown in a sequence 2; b2) a protein which is obtained by substitution and / or deletion and / or addition of amino acid residues of the protein B1), has 80% or more of identity with the protein B1) and has the same function as the protein B1); and B3) a fusion protein obtained by connecting the N terminal or / and the C terminal of B1) or B2) with a protein tag. The method can be used for industrial production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention specifically discloses proteins related to phloretinase activity, their encoding genes, and their applications. Background Technology

[0002] Himalayan four o'clock flower (Mirabilis himalaica), also known as mountain four o'clock flower (Oxybaphushimalaicus), is an annual herbaceous plant belonging to the genus Mirabilis in the family Nyctaginaceae. It is mainly distributed in high-altitude areas such as Tibet, Qinghai, Gansu, and Sichuan. The root of the Himalayan four o'clock flower is used medicinally, commonly known by names such as Bazhu, Zhixiga, Xiaruoba, and Axiaganaha. It is recorded in classic Tibetan medical texts such as *Yuewang Yaozhen*, *Jingzhu Bencao*, and *Sibu Yitan*, and is also included in *National Compendium of Chinese Herbal Medicine* and *Drug Standards of the Ministry of Health of the People's Republic of China*.

[0003] Previous studies have found that the roots and leaves of Himalayan four o'clock flower contain dihydrochalcone compounds such as phlorizin, phlorizin and trifolin. These compounds have strong antioxidant activity and can also be used as natural sweeteners. Summary of the Invention

[0004] The technical problem solved by this invention is how to provide a protein with phlorizinase activity that can catalyze the production of phlorizin into phlorizin and / or trifolin.

[0005] To address the aforementioned technical problems, the present invention provides the following applications.

[0006] Proteins can be used in any of the following applications:

[0007] A1) Applications in catalyzing the production of phloretin from phloretin and / or trifolin and / or in the preparation of products catalyzing the production of phloretin from phloretin and / or trifolin.

[0008] A2) Applications in the production of phlorizin and / or in the preparation of products that produce phlorizin;

[0009] A3) Applications in the production of trifolin and / or in the preparation of products that produce trifolin;

[0010] The protein is any one of the following:

[0011] B1) The amino acid sequence is that of the protein described in sequence 2;

[0012] B2) A protein having more than 80% identity and the same function as the protein shown in B1) obtained by substitution and / or deletion and / or addition of amino acid residues of the protein described in B1).

[0013] B3) A fusion protein obtained by attaching a protein tag to the N-terminus and / or C-terminus of B1) or B2).

[0014] In the aforementioned proteins, the protein tag refers to a polypeptide or protein fused with the target protein using in vitro DNA recombination technology for expression, detection, tracing, and / or purification of the target protein. The protein tag may be a Flag tag, His tag, MBP tag, HA tag, myc tag, GST tag, and / or SUMO tag, etc.

[0015] In the above-mentioned proteins, identity refers to the identity of the amino acid sequences. The identity of amino acid sequences can be determined using homology search sites on the Internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, using blastp as the program, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, setting Gapexistencecost, Perresiduegapcost, and Lambdaratio to 11, 1, and 0.85 (default values) respectively, and performing an identity search on a pair of amino acid sequences, the identity value (%) can then be obtained.

[0016] In the aforementioned proteins, the 80% or more identity can be at least 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 95%, 96%, 98%, 99%, or 100% identity.

[0017] Of the proteins described above, sequence 2 (SEQ ID No. 2) consists of 491 amino acid residues. It is named the Ohi.6618 protein, and its encoding gene is the Ohi.6618 gene.

[0018] To facilitate the purification and use of proteins in A1), tags as shown in the table below can be attached to the amino or carboxyl terminus of the protein, which consists of the amino acid sequence shown in positions 423-902 of sequence 2 in the sequence listing.

[0019]

[0020]

[0021] In the above applications, the protein is derived from jasmine.

[0022] The jasmine mentioned above may be Mirabilis jalapa.

[0023] Any of the following applications of biomaterials related to the above proteins:

[0024] C1) Applications in catalyzing the production of phloretin from phloretin and / or trifolin and / or in the preparation of products catalyzing the production of phloretin from phloretin and / or trifolin.

[0025] C2) Applications in the production of phlorizin and / or in the preparation of products that produce phlorizin;

[0026] C3) Applications in the production of trifolin and / or the preparation of products that produce trifolin;

[0027] C4) Producing the above-mentioned protein or preparing products that produce the above-mentioned protein;

[0028] The biomaterial is any one of the following D1) to D4):

[0029] D1) The nucleic acid molecule that encodes the above-mentioned protein;

[0030] D2) An expression cassette containing the nucleic acid molecules described in D1);

[0031] D3) A recombinant vector containing the nucleic acid molecule described in D1), or a recombinant vector containing the expression cassette described in D2);

[0032] D4) Recombinant microorganisms containing the nucleic acid molecules described in D1), or recombinant microorganisms containing the expression cassette described in D2), or recombinant microorganisms containing the recombinant vector described in D3).

[0033] In the above text, the substance regulating gene expression can be a substance that performs at least one of the following six types of regulation: 1) regulation at the transcriptional level of the gene; 2) post-transcriptional regulation of the gene (i.e., regulation of splicing or processing of the primary transcript of the gene); 3) regulation of RNA transport of the gene (i.e., regulation of mRNA transport of the gene from the nucleus to the cytoplasm); 4) regulation of translation of the gene; 5) regulation of mRNA degradation of the gene; and 6) post-translational regulation of the gene (i.e., regulation of the activity of the protein translated from the gene).

[0034] The aforementioned 80% or higher identity can be 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%.

[0035] In this article, identity refers to the similarity of amino acid or nucleotide sequences. The identity of amino acid sequences can be determined using homology search sites on the internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, using blastp as the procedure, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, setting the Gap existence cost, Per residue gap cost, and Lambda ratio to 11, 1, and 0.85 (default values) respectively, and performing a search to calculate the identity of amino acid sequences, then the identity value (%) can be obtained.

[0036] In D4) above, the recombinant vector can be a microbial expression vector. The microbial expression vector can be the recombinant vector pMAL-c6T-Ohi.6618.

[0037] As a specific embodiment, the recombinant vector described above is the recombinant vector pMAL-c6T-Ohi.6618. The recombinant vector pMAL-c6T-Ohi.6618 is obtained by replacing the small fragment between the NotI and EcoRI recognition sites of the restriction endonuclease in the pMAL-c6T vector with the nucleotides shown in Sequence 1 of the sequence listing, while keeping the other sequences of the pMAL-c6T vector unchanged. This recombinant vector can express the fusion protein shown in Sequence 2 of the sequence listing.

[0038] The microorganism mentioned in D5 above can be Escherichia coli. The Escherichia coli can be Escherichia coli Rosetta2 (DE3).

[0039] In the above applications, the nucleic acid molecule is the DNA molecule described in Sequence 1.

[0040] To address the aforementioned technical problems, the present invention also provides a preparation method.

[0041] The preparation methods include methods for preparing phlorizin, methods for preparing trifolin, or methods for preparing phlorizin and trifolin.

[0042] The method for preparing phlorizin includes the step of reacting the above-mentioned protein or the product for producing phlorizin with phlorizin to obtain phlorizin.

[0043] The method for preparing trifolin includes the step of reacting the above-mentioned protein or the product for producing trifolin with phloretin to obtain trifolin.

[0044] The method for preparing phlorizin and trifolin includes the step of reacting the above-mentioned protein or the product for producing phlorizin and trifolin with phlorizin to obtain phlorizin and trifolin.

[0045] The reaction system described above also contains uridine diphosphate-β-OD-glucose.

[0046] The reaction system described above was carried out at 30°C.

[0047] In the above text, the chemical formula of the phloretin is C. 15 H 14 O5, CAS number 60-82-2.

[0048] In the above text, the chemical formula of the uridine diphosphate glucose is C0. 15 H 22 N2Na2O 17 P2, CAS number 28053-08-9.

[0049] In the above text, the chemical formula of the phlorizin is C. 21 H 24 O 10 The CAS number is 60-81-1.

[0050] In the above text, the chemical formula of the trifolin is C. 21 H 24 O 10 The CAS number is 4192-90-9.

[0051] In the above method, the reaction system of the reaction also contains uridine diphosphate-β-OD-glucose.

[0052] In the above method, the reaction system is carried out at 30°C.

[0053] To address the aforementioned technical problems, this invention provides reagents or kits.

[0054] The reagent or kit includes phloretin, the aforementioned proteins, or one or more of the aforementioned biological materials.

[0055] As mentioned above, the substances in the kit can be packaged individually.

[0056] To address the aforementioned technical problems, the present invention provides a product.

[0057] The product is any one of the following:

[0058] Y1) The above-mentioned proteins;

[0059] Y2) on biomaterials.

[0060] To address the aforementioned technical problems, the present invention provides a product.

[0061] The product is any one of the following:

[0062] Z1) contains at least one product that catalyzes the formation of phlorizin from phlorisin, as described above;

[0063] Z2) contains at least one product that catalyzes the formation of trifolin from phloretin as described above;

[0064] Z3) contains at least one product that catalyzes the production of phlorizin and / or trifolin as described above.

[0065] Beneficial effects

[0066] This invention discloses a phlorizinase activity-related protein, its encoding gene, and its applications. The technical problem addressed is how to provide a phlorizinase activity-related protein. Specifically, the protein is disclosed for applications in any of the following: A1) catalyzing the production of phlorizin to phloridine and / or trifolin and / or preparing products that catalyze the production of phlorizin to phloridine and / or trifolin; A2) producing phloridine and / or preparing products that produce phlorizin; A3) producing trifolin and / or preparing products that produce trifolin; wherein the protein is any of the following: B1) a protein with the amino acid sequence described in Sequence 2; B2) a protein with more than 80% identity and the same function as the protein shown in B1) obtained by substitution and / or deletion and / or addition of amino acid residues of the protein in B1); B3) a fusion protein obtained by attaching a protein tag to the N-terminus and / or C-terminus of B1) or B2). It can be used for industrial production. Attached Figure Description

[0067] Figure 1 The image shows the UPLC-MS results of the in vitro catalytic reaction of Ohi.6618 crude protein solution. The retention time of 1 mg / mL phlorizin standard solution was 3.81 minutes, and the retention time of 1 mg / mL trifolin standard solution was 4.14 minutes. When using the empty carrier crude protein solution to catalyze the 1 mM substrate phlorizin, neither phlorizin nor trifolin could be detected. However, when using the Ohi.6618 crude protein solution to catalyze the 1 mM substrate phlorizin, both phlorizin and trifolin could be detected simultaneously. Detailed Implementation

[0068] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.

[0069] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.

[0070] Prokaryotic expression vector pMAL-c6T: New England Biolabs, USA, code number: N0378S.

[0071] Escherichia coli Rosetta 2 (DE3): Shanghai Weidi Biotechnology Co., Ltd., catalog number EC1014.

[0072] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the terms and laboratory procedures related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, and immunology used herein are all widely used terms and routine procedures in their respective fields. To better understand this invention, definitions and explanations of relevant terms are provided below.

[0073] Phloretin, chemical formula C 15 H 14 O5, CAS number 60-82-2.

[0074] Uridine diphosphate glucose, chemical formula C 15 H 22 N2Na2O 17 P2, CAS number 28053-08-9.

[0075] Phlorizin, chemical formula C 21 H 24 O 10 The CAS number is 60-81-1.

[0076] Trifolin, chemical formula C 21 H 24 O 10 The CAS number is 4192-90-9.

[0077] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to polymers of amino acids of any length. Polymers may be linear, cyclic, or branched, may contain modified amino acids, particularly conserved modified amino acids, and may be interrupted by non-amino acid components. The term also includes modified amino acid polymers, such as those modified by sulfation, glycosylation, esterification, acetylation, phosphorylation, iodination, methylation, oxidation, proteolytic processing, isopreneation, racemization, selenoylation, transfer-RNA-mediated amino addition such as arginination, ubiquitination, or any other manipulation such as conjugation with a labeled component. As used herein, the term “amino acid” refers to natural and / or non-natural or synthetic amino acids, including glycine and its D or L optical isomers, as well as amino acid analogs and peptide mimics. “Derived from” a specified protein refers to the source of the polypeptide. The term also includes polypeptides expressed by specified nucleic acid sequences.

[0078] The term "amino acid" refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimics that function in a manner similar to naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as subsequently modified amino acids such as hydroxyproline, γ-carboxyglutamic acid, and O-phosphoserine. In this document, amino acids are represented using the commonly used three-letter or single-letter codes recommended by the IUPAC-IUB Committee on Biochemistry Nomenclature. Similarly, nucleotides are represented using their generally accepted single-letter codes.

[0079] The term “gene” refers to a segment of DNA involved in the production of a polypeptide chain; it includes regions before and after the coding region (leader and tail regions) involved in the transcription / translation of the gene product and the regulation of said transcription / translation, as well as insertion sequences (introns) between individual coding regions (exons).

[0080] The term "template" refers to any nucleic acid molecule that can be used for the amplification described in this invention. Non-natural double-stranded RNA or DNA can be prepared into double-stranded DNA for use as double-stranded DNA. Any double-stranded DNA or preparation containing a variety of different double-stranded DNA molecules can be used as template DNA to amplify one or more loci of interest contained within the template DNA.

[0081] The term "primer" refers to an oligonucleotide that can be used in amplification methods such as polymerase chain reaction (PCR) to amplify a nucleotide sequence based on a polynucleotide sequence corresponding to a specific genomic sequence. At least one PCR primer used to amplify the polynucleotide sequence is sequence-specific to that sequence.

[0082] The term "probe" refers to a molecule that binds to a specific sequence or subsequence or other portion of another molecule. Unless otherwise specified, the term "probe" generally refers to a polynucleotide probe that binds to another polynucleotide (often called a "target polynucleotide") through complementary base pairing. A probe can bind to a target polynucleotide that lacks complete sequence complementarity with the probe, depending on the stringency of the hybridization conditions. Probes can be labeled directly or indirectly.

[0083] The term "amplification reaction" refers to a process used for one or more copies of nucleic acids. In embodiments, the amplification methods include, but are not limited to: polymerase chain reaction (PCR), self-sustaining sequencing reaction (SSSR), ligase chain reaction (LCSR), rapid amplification of cDNA ends, PCR and LCSR, Q-β phage amplification, strand displacement amplification, or overlap extension splicing PCR. In some embodiments, single-molecule nucleic acids are amplified, for example, by digital PCR.

[0084] The term "amplification product" refers to nucleic acid products produced through nucleic acid amplification technology.

[0085] The term "kit" refers to any delivery system used to deliver substances. In reaction assays, such delivery systems include systems for storing, transferring, or delivering reaction reagents (e.g., oligonucleotides, enzymes, etc. in appropriate containers) and / or support materials (e.g., buffer solutions, instructions for performing the assay, etc.) from one location to another. For example, a kit may contain one or more housings (e.g., boxes) containing the relevant reaction reagents and / or support materials.

[0086] The terms “include,” “including,” “have,” “contain,” etc., are all open-ended terms, meaning that they include but are not limited to.

[0087] The term "nucleic acid" refers to a polymer consisting of at least two deoxynucleotides or nucleotides, existing in single or double strands. Unless specifically limited, the term encompasses nucleic acids containing known analogs of natural nucleotides, having similar binding properties to reference nucleic acids, and being metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a specific nucleic acid sequence also implicitly includes variants of its conserved modifications (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences, as well as explicitly indicated sequences. Specifically, degenerate codon substitutions can be obtained by generating sequences in which the third position of one or more selected (or all) codons is replaced by a mixture of bases and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19: 5081 (1991); Ohtsukae et al., J. Biol. Chem. 260: 2605-2608 (1985); and Cassole et al. (1992); Rossolinie et al., Mol. Cell. Probes 8: 91-98 (1994)). A “nucleotide” contains a sugar, deoxyribose (DNA) or ribose (RNA), a base, and a phosphate group. Nucleotides are linked together by phosphate groups. "Bases" include purines and pyrimidines, further including natural compounds adenine, thymine, guanine, cytosine, uracil, inosine, and natural analogs, as well as synthetic derivatives of purines and pyrimidines, including, but not limited to, modifications that replace new reactive groups, such as, but not limited to, amines, alcohols, thiols, carboxylates (esters), and alkyl halides. DNA can exist as antisense, plasmid DNA, portions of plasmid DNA, pre-compressed DNA, polymerase chain reaction (PCR) products, vectors (P1, PAC, BAC, YAC, artificial chromosomes), expression cassettes, chimeric sequences, chromosomal DNA, or derivatives of these groups. The terms nucleic acid, gene, cDNA, mRNA encoded by a gene, and interfering RNA molecules may be used interchangeably.

[0088] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. Furthermore, in the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0089] The term "identity" is used to describe the percentage of identical amino acids or nucleotides between two amino acid sequences or nucleic acid sequences relative to a reference sequence, determined by conventional methods, for example, see Ausubel et al., eds. (1995), Current Protocols in Molecule & Lar Biology, Chapter 19 (Greene Publishing and Wiley-Interscience, New York); and the ALIGN program (Dayhoff (1978), Atlas of Protein Sequence and Structure 5: Suppl. 3 (National Biomedical Research Foundation, Washington, DC). There are many algorithms for aligning sequences and determining sequence identity, including the homology alignment algorithm of Needleman et al. (1970) J. Mol. Biol. 48: 443; the local homology algorithm of Smith et al. (1981) Adv. Appl. Math. 2: 482; and the local homology algorithm of Pearson et al. (1988) P… Similarity search methods are described in roc. Natl. Acad. Sci. 85: 2444; the Smith-Waterman algorithm (Meth. Mol. Biol. 70: 173-187 (1997); and the BLASTP, BLASTN, and BLASTX algorithms (see AltschμL et al. (1990) J. Mol. Biol. 215: 403-410). Computer programs utilizing these algorithms are also available, including but not limited to: ALIGN or Megalign (DNASTAR) software, or WU-BLAS. T-2 (AltschμL et al., Meth. Enzym., 266:460-480 (1996)); or GAP, BESTFIT, BLASTAltschμL et al., above, FASTA, and TFASTA, available in Genetics Computing Group (GCG) package, version 8, Madison, Wisconsin, USA; and CLUSTAL in the PC / Gene program provided by Intelligenetics, MountainView, California.

[0090] Example 1: Discovery of Proteins and Genes

[0091] Mountain Four O'Clock: The 2-month-old seedlings of the Mountain Four O'Clock variety (Oxybaphus himalaicus) were obtained by the inventor of this invention through seed cultivation in an artificial climate chamber.

[0092] The inventors of this invention performed extensive targeted metabolome sequencing on the leaves and roots of two-month-old Mirabilis jalapa seedlings, detecting small amounts of phlorizin, phloridine, and trifolin in the leaves. Subsequently, third-generation and second-generation transcriptome sequencing were performed on the leaves and roots of Mirabilis jalapa, revealing a phlorizin glycoside transferase gene that is specifically highly expressed in the leaves. This gene encodes a protein that can be used to prepare phloridine and trifolin. This gene was named Ohi.6618, and the protein was named Ohi.6618 protein.

[0093] The CDS sequence of the Ohi.6618 gene is shown in Sequence 1, as detailed below:

[0094]

[0095] The amino acid sequence of the Ohi.6618 protein is shown in Sequence 2, as follows:

[0096] MERAELVFVPTPGMGHLLSTVELAKLIVSRHPTISVLVLIFKLSVDTTTVDVYVESQSRDCDSTRLTFIVLPPLYNPPGPSTPNFFNTLIALNKPAIKHAVEERVQSGFPKPAGFVLDMFCTS MMDVADEFNVPSYIYFTSGASLLNLFLHFQALADDTGVNIIEFNDSDVEFDIPGFKNRVPCKVIPSVFFDKEWGNKLLLNLARRFRKCNGILVNTFMELESYTIKTLLDQHDEGDIPAVYPVG PILLDSKSRGGSKSKKEEEESIMEWLDEQPDSSVVFLCFGSMGSFDAPQVQEIANGLEHAGHRFLWSIRRPPPEDKKMGIPYGNETYEDALPEGFLERTAGRGKIIGWAPQILILAHRAVGG FVSHCGWNSTLECMWFGVPMATWPMYAEQQLNAFKLVKEMEIAVEIKMDYQRDWKTGKGNMIVTAEEIENGVKKLMSMDKEKKEKWRKMSEEGKKALEENGSSHHWLSCFIDDVLSHSPAKK.

[0097] Example 2: Preparation of recombinant plasmids

[0098] The small fragment between the NotI and EcoRI restriction sites of the prokaryotic expression vector pMAL-c6T was replaced with a foreign DNA molecule to obtain a recombinant plasmid.

[0099] The recombinant vector pMAL-c6T-Ohi.6618 is obtained by replacing the small fragment between the NotI and EcoRI recognition sites of the restriction endonuclease in the pMAL-c6T vector with the nucleotides shown in Sequence 1 of the sequence listing, while keeping the other sequences of the pMAL-c6T vector unchanged. This recombinant vector is named pMAL-c6T-Ohi.6618.

[0100] Example 3: Preparation of crude protein solution

[0101] The recombinant plasmids were the recombinant vector pMAL-c6T-Ohi.6618 prepared in Example 2 and the empty vector plasmid (i.e., the prokaryotic expression vector pMAL-c6T).

[0102] 1. The recombinant plasmid (recombinant vector pMAL-c6T-Ohi.6618) and the empty vector plasmid (prokaryotic expression vector pMAL-c6T) were introduced into *E. coli* Rosetta 2(DE3) to obtain recombinant bacteria Rosetta 2 / recombinant vector pMAL-c6T-Ohi.6618 and recombinant bacteria Rosetta 2 / pMAL-c6T. Recombinant bacteria Rosetta 2 / recombinant vector pMAL-c6T-Ohi.6618 is *E. coli* Rosetta 2(DE3) containing the recombinant vector pMAL-c6T-Ohi.6618; recombinant bacteria Rosetta 2 / pMAL-c6T is *E. coli* Rosetta 2(DE3) containing the prokaryotic expression vector pMAL-c6T.

[0103] 2. Inoculate the recombinant bacteria obtained in step 1 (recombinant bacteria Rosetta 2 / recombinant vector pMAL-c6T-Ohi.6618) into 50 mL of LB liquid medium and culture at 37°C with shaking at 200 rpm until the OD600nm value of the system is 0.6 (0.4-0.8 is acceptable in practical applications). Then add IPTG to a concentration of 500 μM and culture at 16°C with shaking at 100 rpm for 20 hours. Then, centrifuge at 4°C and 3000g for 10 min, discard the supernatant, resuspend the bacterial cells in Tris-EDTA buffer (pH 7.4) and sonicate (90W frequency, 5s interval, 5s disruption, for a total of 5 min). Then centrifuge at 4°C and 10000g for 15 min and collect the supernatant, which is the crude protein solution of Ohi.6618.

[0104] 3. The only difference from step 2 is that the recombinant bacteria are replaced with recombinant Rosetta 2 / pMAL-c6T to obtain an empty crude protein solution.

[0105] The protein concentration of the empty crude protein solution was found to be 5.00 mg / mL. The protein concentration of the Ohi.6618 crude protein solution was 2.73 mg / mL. These were used in the following experiments.

[0106] Example 4: In vitro enzymatic reaction

[0107] 1. Prepare the reaction system and then carry out the reaction.

[0108] Dissolve phlorizin in methanol to make a phlorizin concentration of 100 mM, which is a 100 mM phlorizin solution.

[0109] Dissolve uridine diphosphate glucose in ultrapure water to make the concentration of uridine diphosphate glucose 100mM, which is a 100mM uridine diphosphate glucose solution.

[0110] Dissolve magnesium sulfate in ultrapure water to make the magnesium sulfate concentration 100mM, which is a 100mM magnesium sulfate solution.

[0111] Phlorizin is dissolved in methanol to make a phlorizin concentration of 1 mg / mL, which is a 1 mg / mL phlorizin solution (phlorizin standard).

[0112] Dissolve trifolin in methanol to a concentration of 1 mg / mL, which is a 1 mg / mL trifolin solution (trifolin standard).

[0113] Phloretin, chemical formula C 15 H 14 O5, CAS number 60-82-2.

[0114] Uridine diphosphate glucose, chemical formula C 15 H 22 N2Na2O 17 P2, CAS number 28053-08-9.

[0115] Phlorizin, chemical formula C 21 H 24 O 10 The CAS number is 60-81-1.

[0116] Trifolin, chemical formula C 21 H 24 O 10 The CAS number is 4192-90-9.

[0117] Empty crude protein solution.

[0118] Ohi.6618 crude protein solution.

[0119] Composition of reaction system 1: 500 μL L Hi.6618 crude protein solution, 5 μL 100 mM phloretin solution, 10 μL 100 mM uridine diphosphate glucose solution, and 10 μL 100 mM magnesium sulfate solution. Reaction conditions: 30℃, shaking at 200 rpm for 6 hours. The reaction solution of reaction system 1 was obtained.

[0120] Composition of reaction system 2: 500 μL empty crude protein solution, 5 μL 100 mM phloretin solution, 10 μL 100 mM uridine diphosphate glucose solution, and 10 μL 100 mM magnesium sulfate solution. Reaction conditions: 30℃, 200 rpm shaking for 6 hours. Reaction solution for reaction system 2.

[0121] 2. Add 2 times the volume of methanol to the reaction solution of reaction system 1 and the reaction solution of reaction system 2 respectively, vortex to mix, then centrifuge at 4℃ and 10000g for 10min, collect the supernatant to obtain Ohi.6618 crude protein solution reaction solution and empty crude protein solution reaction solution.

[0122] 3. Perform UPLC-MS detection

[0123] An ACQUITY UPLC HSS T3 column was used. 1.8μm, 2.1mm×50mm).

[0124] The column temperature was 40℃; the injection volume was 2μL; and the mass spectrometry fragment ion peak monitoring conditions were 435.1 / 273.2.

[0125] Mobile phase: a mixture of solution A and solution B; Solution A: ultrapure water containing 0.1% formic acid; Solution B: 100% acetonitrile.

[0126] The mobile phase flow rate was 0.4 mL / min.

[0127] Elution process: 0-0.5 min, the volume fraction of solution A in the mobile phase linearly decreases from 90% to 80%, while the corresponding volume fraction of solution B in the mobile phase linearly increases from 10% to 20%; 0.5-5 min, the volume fraction of solution A in the mobile phase linearly decreases from 80% to 75%, while the corresponding volume fraction of solution B in the mobile phase linearly increases from 20% to 25%; 5-7 min, the volume fraction of solution A in the mobile phase linearly decreases from 75% to 30%, while the corresponding volume fraction of solution B in the mobile phase linearly increases from 25% to 70%; 7-7.2 min, the volume fraction of solution A in the mobile phase linearly decreases from 90% to 80%, while the corresponding volume fraction of solution B in the mobile phase linearly increases from 25% to 70%; 7-7.2 min, the volume fraction of solution A in the mobile phase linearly increases from 90% to 80%, while the corresponding volume fraction of solution B in the mobile phase linearly increases from 10% to 20%; 0.5-5 min, the volume fraction of solution A in the mobile phase linearly decreases from 80% to 75%, while the corresponding volume fraction of solution B in the mobile phase linearly increases from 20% to 25%; 0.5-0. ...75% to 30%, while the corresponding volume fraction of solution B in the mobile phase linearly increases from 25% to 70%; 0.5-0.5 min, the volume fraction of solution B in the mobile phase linearly increases from 75% to 30%, while the corresponding volume fraction of solution B in The volume fraction of the mobile phase decreased linearly from 30% to 5%, while the volume fraction of solution B in the mobile phase increased linearly from 70% to 95%. From 7.2 to 9.2 min, the volume fraction of solution A in the mobile phase was 5%, and the volume fraction of solution B in the mobile phase was 95%. From 9.2 to 10 min, the volume fraction of solution A in the mobile phase increased linearly from 5% to 90%, while the volume fraction of solution B in the mobile phase decreased linearly from 95% to 10%. From 10 to 13 min, the volume fraction of solution A in the mobile phase was 90%, and the volume fraction of solution B in the mobile phase was 10%.

[0128] The Ohi.6618 crude protein solution reaction solution, the empty-load crude protein solution reaction solution, the phlorizin standard, and the trifolin standard were performed in parallel under the chromatographic conditions described above.

[0129] Chromatograms of phlorizin standard (1 mg / mL phlorizin solution), trifolin standard (1 mg / mL trifolin solution), and chromatograms of the empty-load crude protein solution and Ohi.6618 crude protein solution after the above steps are shown below. Figure 1 As shown ( Figure 1In the chromatogram, phlorizin (1 mg / mL phlorizin solution), trifolin (1 mg / mL trifolin solution), and crude protein from the empty carrier (empty carrier crude protein solution), along with Ohi.6618 (Ohi.6618 crude protein solution), were detected by UPLC-MS. Because both target products are located in the 3-5 min range, only this range is shown in the chromatogram. Since this was a combined LC-MS / MS analysis, ion fragment peak screening was performed. Because phlorizin and trifolin are isomers with the same two major ion fragment peaks, 435.1 and 273.2, the analysis was limited to products with the 435.1 / 273.2 ion fragment peaks. Therefore, the peak of the substrate phlorizin could not be displayed in the software, and any other impurities that did not meet the screening criteria were also not displayed. The retention time of phlorizin standard was 3.81 minutes, and the retention time of trifolin standard was 4.14 minutes. The above steps were performed on the empty vector crude protein solution, and no phlorizin or trifolin was found in the reaction product. The results indicate that the crude protein solution obtained from recombinant *E. coli* introduced with the empty vector plasmid cannot catalyze the production of phlorizin and trifolin from phloretin. However, the crude protein solution of Ohi.6618 can catalyze the production of phlorizin and trifolin from phloretin.

[0130] The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims.

Claims

1. The application of proteins in any of the following: A1) Applications in catalyzing the production of phloretin from phloretin and / or trifolin and / or in the preparation of products catalyzing the production of phloretin from phloretin and / or trifolin. A2) Applications in the production of phlorizin and / or in the preparation of products that produce phlorizin; A3) Applications in the production of trifolin and / or in the preparation of products that produce trifolin; The protein is any one of the following: B1) The amino acid sequence is that of the protein described in sequence 2; B2) A protein having more than 80% identity and the same function as the protein shown in B1) obtained by substitution and / or deletion and / or addition of amino acid residues of the protein described in B1). B3) A fusion protein obtained by attaching a protein tag to the N-terminus and / or C-terminus of B1) or B2).

2. The application according to claim 1, characterized in that, The protein is derived from Mirabilis jalapa.

3. Any of the following applications of biomaterials related to the protein described in claim 1 or 2: C1) Applications in catalyzing the production of phloretin from phloretin and / or trifolin and / or in the preparation of products catalyzing the production of phloretin from phloretin and / or trifolin. C2) Applications in the production of phlorizin and / or in the preparation of products that produce phlorizin; C3) Applications in the production of trifolin and / or the preparation of products that produce trifolin; C4) To produce the protein of claim 1 or to prepare a product that produces the protein of claim 1; The biomaterial is any one of the following D1) to D4): D1) A nucleic acid molecule encoding the protein described in claim 1; D2) An expression cassette containing the nucleic acid molecules described in D1); D3) A recombinant vector containing the nucleic acid molecule described in D1), or a recombinant vector containing the expression cassette described in D2); D4) Recombinant microorganisms containing the nucleic acid molecules described in D1), or recombinant microorganisms containing the expression cassette described in D2), or recombinant microorganisms containing the recombinant vector described in D3).

4. The application according to claim 3, characterized in that, The nucleic acid molecule described in claim 3 is the DNA molecule described in sequence 1.

5. The preparation method, characterized in that, The preparation methods include methods for preparing phlorizin, methods for preparing trifolin, or methods for preparing phlorizin and trifolin. The method for preparing phlorizin includes the step of reacting the protein or the product for producing phlorizin as described in claim 1 or 2 with phlorizin to obtain phlorizin. The method for preparing trifolin includes the step of reacting the protein or the product for producing trifolin with phloretin to obtain trifolin. The method for preparing phlorizin and trifolin includes the step of reacting the protein described in claim 1 or 2 or the product for producing phlorizin and trifolin with phlorizin to obtain phlorizin and trifolin.

6. The method as described in claim 5, characterized in that, The reaction system also contains uridine diphosphate-β-OD-glucose.

7. The method as described in claim 5, characterized in that, The reaction system was carried out at 30°C.

8. A reagent or kit, characterized in that, The reagent or kit includes one or more of phloretin, the protein of claim 1 or 2, or the biomaterial of claim 3 or 4.

9. A product characterized in that, The product is any one of the following: Y1) The protein as described in claim 1 or 2; Y2) The biomaterial described in claim 3 or 4.

10. A product characterized in that, The product is any one of the following: Z1) contains a product that catalyzes the production of phlorizin from phlorizin as described in claims 1, 2, 3 and / or 4; Z2) contains the product described in claims 1, 2, 3 and / or 4 that catalyzes the production of trifolin from phloretin; Z3) contains products that catalyze the production of phlorizin and / or trifolin as described in claims 1, 2, 3 and / or 4.