Method and device for tracing producing area of radix tetrastigme

By analyzing the metabolomic data of santocephala in the use of high-throughput real-time plasma mass spectrometer, screening out differential metabolites and training classification models, the problem of insufficient timeliness and accuracy of santocephala in the existing technology is solved, and efficient and accurate tracing of santocephala in the source of santocephala in the process is achieved.

CN120146874APending Publication Date: 2025-06-13INST OF QUALITY STANDARD & TESTING TECH FOR AGRO PROD OF CAAS

Patent Information

Application Number
CN202510615392.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

It is difficult for the prior art to quickly and accurately realize the origin traceability of santophyllum, especially when ensuring timeliness.

Method used

The metabolomic data of Trichophyllum Chrysanthemum was analyzed by high-throughput real-time plasma mass spectrometer (HRT-MS), differential metabolites were screened out, and classification models were used to train the classification model to achieve traceability of Trichophyllum origin.

Benefits of technology

It has achieved rapid and accurate traceability of Sanyeqing's production area, improved the prediction accuracy, and provided an effective means for market security and protection of consumer rights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146874A_ABST
    Figure CN120146874A_ABST
Patent Text Reader

Abstract

The invention provides a tetrastigma hemsleyanum origin tracing method and device, and relates to the technical field of traditional Chinese medicine origin identification. The tetrastigma hemsleyanum origin tracing method provided by the invention comprises the following steps: acquiring metabolome data of tetrastigma hemsleyanum from different origins, and screening out differential metabolites; training a classification model by using the data of the differential metabolites, then classifying the metabolome data of the radix tetrastigme to be tested by using the trained classification model, and predicting the production place of the radix tetrastigme to be tested; different producing areas include Zhejiang, Fujian, Jiangxi, Yunnan, Guangxi and Chongqing. The method is high in prediction accuracy, an efficient and convenient identification method is provided for radix tetrastigme producing area traceability, the market safety of radix tetrastigme is guaranteed, and the legitimate rights and interests of consumers are maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Chinese medicine origin identification, and in particular to a method and device for tracing the origin of Tripterygium wilfordii. Background Art

[0002] Tetrastigma hemsleyanum Diels et Gilg belongs to the genus Tetrastigmaobtectum of the grape family. It is a well-known medicinal and edible plant with the effects of promoting blood circulation, removing blood stasis, detoxifying and resolving phlegm. It is often used in clinical treatment of viral meningitis, Japanese encephalitis, viral pneumonia, icteric hepatitis, etc. In addition, Sanyeqing can also be used as a raw material for functional foods, and is often consumed as functional tea or dietary supplements. Sanyeqing is widely distributed in tropical to subtropical areas, mainly growing in provinces in southern and southwestern China, including Zhejiang, Fujian, Jiangxi, Chongqing, Yunnan and other places. The production area of ​​Chinese medicinal materials is closely related to its quality. The pharmacological activity and price of Sanyeqing in different production areas vary greatly. Many studies have shown that the quality of Sanyeqing from different production areas is different, and the chemical composition and pharmacological activity vary greatly. In addition, a survey found that Sanyeqing in Chongqing, Yunnan and other places in my country has a short growth cycle, large yield and low price; while Sanyeqing in Zhejiang, Fujian and other places has a long growth cycle, low yield, good quality and relatively high price. Driven by economic interests, there are cases in the market where Sanyeqing from lower-quality areas is passed off as Sanyeqing of better quality. The content of chemical components in Sanyeqing from different origins varies greatly. Therefore, tracing the origin of Sanyeqing by detecting differential metabolites can be an effective way to combat adulteration of Sanyeqing's origin.

[0003] Currently, common origin traceability technologies include near-infrared spectroscopy, stable isotope mass spectrometry, mineral element analysis technology, metabolomics technology, chromatographic fingerprint technology, and electronic nose technology, etc. Among them, metabolomics technology is widely used in the research of traditional Chinese medicine origin traceability and has achieved good results. For example: Wang et al. explored the untargeted metabolomics of Cordyceps sinensis from Nagqu, other places in Tibet, Qinghai, Gansu, and Yunnan by UPLC-QTOF-MS and constructed an OPLS-DA model. Markers were successfully identified and verified using blind test samples, and the discrimination accuracy rate reached 100%. Kim et al. used gas chromatography-mass spectrometry to perform targeted metabolite analysis on Perilla frutescens from China and South Korea. The fat-soluble metabolites of Perilla frutescens (including amino acids, organic acids, sugars, sugar alcohols, tocopherols, sterols, shogaols, and fatty acids) were analyzed, and an OPLS-DA discrimination model was established to successfully distinguish Perilla frutescens from China and South Korea, with a discrimination accuracy rate of 100%. When metabolomics is applied to origin traceability research, it targets or non-targets the screening of small molecule metabolites in samples from different regions. Targeted metabolomics focuses on analyzing specific metabolite groups and conducts origin traceability by identifying the levels of specific metabolites, while untargeted metabolomics screens all metabolites of the target and explores the compositional characteristics of metabolites in different origins to distinguish the production areas. Untargeted metabolomics can reflect more comprehensive metabolite information and study the correlations between metabolites. A large number of studies have shown the successful application of untargeted metabolomics technology in the identification and traceability of traditional Chinese medicine.

[0004] In recent years, technologies commonly used in the research of untargeted metabolomics of traditional Chinese medicine include gas chromatography-time of flight mass spectrometry (GC-TOF-MS), ultra-high performance liquid chromatography-quadrupole time of flight mass spectrometry (UPLC-Q-TOF-MS), nuclear magnetic resonance hydrogen spectroscopy (HNMR), etc. However, among these methods, some have too long a measurement time for a single sample and are not suitable for measuring a large number of samples. In addition, there are also problems such as the instrument being prone to breakdown and high maintenance costs. To avoid these problems, this study used a high-throughput real-time inductively coupled plasma mass spectrometer (HRT-MS). HRT-MS uses an inductively coupled plasma ionization source as the ionization device. The sample forms an aerosol state under thermal assistance conditions and is ionized by the plasma, and then enters the time of flight mass spectrometry for qualitative and quantitative screening. HRT-MS does not require complex sample pretreatment and pre-separation and can directly and quickly detect metabolites. The time-consuming for a single sample is only about 1 minute, meeting the timeliness requirements of origin traceability. Using HRT-MS, the information of various differential metabolites can be quickly detected, and it can be brought into the origin traceability model to determine whether there is adulteration. When problems occur, measures can be taken quickly to prevent unqualified products from entering the market and protect the health and safety of consumers. In addition, for the origin traceability of fresh samples, rapid determination of metabolites can reduce the influence of freshness on metabolites, thus enabling more accurate origin traceability.

[0005] Currently, there are more and more studies on origin tracing using metabolomics technology combined with machine learning models. Compared with traditional analytical methods, machine learning has higher data processing capabilities and classification accuracy. With the help of machine learning algorithms, the origin can be more accurately identified. Many studies have also shown that the establishment of an origin tracing model by combining untargeted metabolomics with machine learning algorithms helps to improve the accuracy of origin tracing. However, there is currently no relevant research on using metabolomics technology to establish a machine learning model for origin tracing of Tetrastigma hemsleyanum. In addition, the conventional metabolomics method takes a long time to measure the metabolome and cannot meet the timeliness requirements of origin tracing.

[0006] In view of this, the present invention is specifically proposed. Summary of the Invention

[0007] The first object of the present invention is to provide a method for origin tracing of Tetrastigma hemsleyanum to solve the above technical problems.

[0008] The second object of the present invention is to provide a device for origin tracing of Tetrastigma hemsleyanum.

[0009] The third object of the present invention is to provide an application of metabolites in the origin tracing of Tetrastigma hemsleyanum.

[0010] To achieve the above objects, the following technical solutions are specifically adopted: In the first aspect, the present invention provides a method for origin tracing of Tetrastigma hemsleyanum, including the following steps: a. Obtain the metabolome data of Tetrastigma hemsleyanum from different origins and screen out differential metabolites; b. Use the data of the differential metabolites to train a classification model, and then use the trained classification model to classify the metabolome data of the Tetrastigma hemsleyanum to be tested and predict the origin of the Tetrastigma hemsleyanum to be tested; The different origins include Zhejiang, Fujian, Jiangxi, Yunnan, Guangxi, and Chongqing.

[0011] As a further technical solution, the differential metabolites are histamine, glycolaldehyde dimer, nicotinamide aminopyrine, bifonazole, 5-(6-hydroxy-2,5,7,8-tetramethylchroman-2-yl)-2-methylvaleric acid, sparteine, 1-methyl-2-undecyl-4(1H)-quinolone, 5,6,7-trimethoxyflavone, petunidin, D-erythro-sphingosine, phytosphingosine, guaiol, oleic acid, and icos-19-ene-1,2,4-triol.

[0012] As a further technical solution, the method for obtaining the metabolome data of Tetrastigma hemsleyanum includes the following steps: Extract the metabolites of Tetrastigma hemsleyanum to obtain an extract, and then use a high-throughput real-time inductively coupled plasma mass spectrometer to analyze the metabolites in the extract to obtain the metabolome data of Tetrastigma hemsleyanum.

[0013] As a further technical solution, the classification model includes a support vector machine, a feedforward neural network model, and a random forest model.

[0014] In a second aspect, the present invention provides a device for tracing the origin of Tetrastigma hemsleyanum, including an acquisition module and a classification module; The acquisition module is used to acquire the metabolome data of the Tetrastigma hemsleyanum to be tested; The classification module is used to input the metabolome data of the Tetrastigma hemsleyanum to be tested into a pre-trained classification model, and classify the metabolome data of the Tetrastigma hemsleyanum to be tested through the classification model to predict the origin of the Tetrastigma hemsleyanum to be tested; The classification model is obtained by the following method: a. Obtain the metabolome data of Tetrastigma hemsleyanum from different origins, and screen out differential metabolites; b. Use the data of the differential metabolites to train a classification model to obtain a pre-trained classification model; The different origins include Zhejiang, Fujian, Jiangxi, Yunnan, Guangxi, and Chongqing.

[0015] As a further technical solution, the differential metabolites are histamine, glycolaldehyde dimer, nicotinylaminopyrine, bifonazole, 5-(6-hydroxy-2,5,7,8-tetramethylchroman-2-yl)-2-methylpentanoic acid, sparteine, 1-methyl-2-undecyl-4(1H)-quinolone, 5,6,7-trimethoxyflavone, petunidin, D-erythro-sphingosine, phytosphingosine, guaiol, oleic acid, and icos-19-ene-1,2,4-triol.

[0016] As a further technical solution, the method for obtaining the metabolome data of Tetrastigma hemsleyanum includes the following steps: Extract the metabolites of Tetrastigma hemsleyanum to obtain an extract, and then use a high-throughput real-time inductively coupled plasma mass spectrometer to analyze the metabolites in the extract to obtain the metabolome data of Tetrastigma hemsleyanum.

[0017] As a further technical solution, the classification model includes a support vector machine, a feedforward neural network model, and a random forest model.

[0018] In a third aspect, the present invention provides the application of metabolites in tracing the origin of Tetrastigma hemsleyanum; The metabolites are histamine, glycolaldehyde dimer, nicotinamide aminopyrine, bifonazole, 5-(6-hydroxy-2,5,7,8-tetramethylchroman-2-yl)-2-methylvaleric acid, sparteine, 1-methyl-2-undecyl-4(1H)-quinolone, 5,6,7-trimethoxyflavone, petunidin, D-erythro-sphingosine, phytosphingosine, guaiol, oleic acid, and icos-19-ene-1,2,4-triol.

[0019] As a further technical solution, the producing areas include Zhejiang, Fujian, Jiangxi, Yunnan, Guangxi, and Chongqing.

[0020] Compared with the prior art, the present invention has the following beneficial effects: The method for tracing the origin of Tetrastigma hemsleyanum provided by the present invention analyzes the metabolome data of Tetrastigma hemsleyanum from different producing areas to screen out differential metabolites, and uses the differential metabolites to establish an origin tracing model, so as to realize the origin tracing of Tetrastigma hemsleyanum through the origin tracing model. This method has a high prediction accuracy, provides an efficient and convenient identification method for tracing the origin of Tetrastigma hemsleyanum, helps to ensure the safety of the Tetrastigma hemsleyanum market, and protects the legitimate rights and interests of consumers. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a radar chart of differential metabolites of Tetrastigma hemsleyanum from different producing areas; Figure 2 It is a clustering heat map of differential metabolites from different producing areas; Figure 3 In it, A is the score chart of the PCA model based on the differential metabolites of Tetrastigma hemsleyanum, B is the score chart of the PLS-DA model, C is the score chart of the OPLS-DA model, and D is the cross-validation result calculated by 200 permutation tests of the PLS-DA model; Figure 4 It is the confusion matrix prediction results of the SVM, FNN, and RF models based on differential metabolites (data preprocessing method: Ln); training set (A1, B1, C1); test set (A2, B2, C2); Figure 5 It is the confusion matrix prediction results of the SVM, FNN, and RF models based on differential metabolites (data preprocessing method: none); training set (A1, B1, C1); test set (A2, B2, C2); Figure 6 Confusion matrix prediction results of SVM, FNN, and RF models based on differential metabolites (data preprocessing method: Log); training set (A1, B1, C1); test set (A2, B2, C2); Figure 7 Confusion matrix prediction results of SVM, FNN, and RF models based on differential metabolites (data preprocessing method: square root); training set (A1, B1, C1); test set (A2, B2, C2). Specific implementation manners

[0023] The embodiments of the present invention will be described in detail below in combination with the implementation manners and examples. However, those skilled in the art will understand that the following implementation manners and examples are only used to illustrate the present invention and should not be regarded as limiting the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention. Those not specified in specific conditions are carried out according to conventional conditions or conditions recommended by the manufacturer. Those reagents or instruments not specified in the manufacturer are all conventional products that can be obtained through commercial purchase.

[0024] In a first aspect, the present invention provides a method for tracing the origin of Tetrastigma hemsleyanum, comprising the following steps: a. Obtaining the metabolome data of Tetrastigma hemsleyanum from different origins and screening out differential metabolites; b. Using the data of the differential metabolites to train a classification model, and then using the trained classification model to classify the metabolome data of the Tetrastigma hemsleyanum to be tested, and predicting the origin of the Tetrastigma hemsleyanum to be tested; The different origins include Zhejiang, Fujian, Jiangxi, Yunnan, Guangxi, and Chongqing.

[0025] The method for tracing the origin of Tetrastigma hemsleyanum provided by the present invention analyzes the metabolome data of Tetrastigma hemsleyanum from different origins to screen out differential metabolites, and uses the differential metabolites to establish an origin tracing model, so as to realize the origin tracing of Tetrastigma hemsleyanum through the origin tracing model. This method has a high prediction accuracy, provides an efficient and convenient identification method for tracing the origin of Tetrastigma hemsleyanum, helps to ensure the safety of the Tetrastigma hemsleyanum market, and protects the legitimate rights and interests of consumers.

[0026] In some alternative embodiments, the differential metabolites are histamine, glycolaldehyde dimer, nicotinamide antipyrine, bifonazole, 5-(6-hydroxy-2,5,7,8-tetramethylchroman-2-yl)-2-methylvaleric acid, sparteine, 1-methyl-2-undecyl-4(1H)-quinolone, 5,6,7-trimethoxyflavone, petunidin, D-erythro-sphingosine, phytosphingosine, guaiol, oleic acid, and icos-19-ene-1,2,4-triol.

[0027] In some alternative embodiments, the method for obtaining the metabolome data of Tetrastigma hemsleyanum comprises the following steps: Extract the metabolites of Tetrastigma hemsleyanum to obtain an extract, and then analyze the metabolites in the extract using a high-throughput real-time inductively coupled plasma mass spectrometer to obtain the metabolome data of Tetrastigma hemsleyanum.

[0028] The present invention does not specifically limit the extraction solvent. In some alternative embodiments, an aqueous methanol solution is used to extract Tetrastigma hemsleyanum to obtain an extract.

[0029] In some alternative embodiments, the classification model includes a support vector machine, a feedforward neural network model, and a random forest model, preferably a random forest model. Through research by the inventors, it is found that when using the random forest model for classification, the accuracy rate can reach 100%.

[0030] In a second aspect, the present invention provides a device for tracing the origin of Tetrastigma hemsleyanum, comprising an acquisition module and a classification module; The acquisition module is used to acquire the metabolome data of the Tetrastigma hemsleyanum to be tested; The classification module is used to input the metabolome data of the Tetrastigma hemsleyanum to be tested into a pre-trained classification model, and classify the metabolome data of the Tetrastigma hemsleyanum to be tested through the classification model to predict the origin of the Tetrastigma hemsleyanum to be tested; The classification model is trained through the following method: a. Obtain the metabolome data of Tetrastigma hemsleyanum from different origins, and screen out differential metabolites; b. Use the data of the differential metabolites to train the classification model to obtain a pre-trained classification model; The different origins include Zhejiang, Fujian, Jiangxi, Yunnan, Guangxi, and Chongqing.

[0031] This device can efficiently and accurately trace the origin of the Tetrastigma hemsleyanum to be tested.

[0032] In some optional embodiments, the differential metabolites are histamine, glycolaldehyde dimer, nicotinamide, bifonazole, 5-(6-hydroxy-2,5,7,8-tetramethylchroman-2-yl)-2-methylpentanoic acid, spartoine, 1-methyl-2-undecyl-4(1H)-quinolone, 5,6,7-trimethoxyflavone, petunidin, D-erythrosphingosine, phytosphingosine, guaiac, oleic acid and eicosapentaene-19-ene-1,2,4-triol.

[0033] In some optional embodiments, the method for obtaining the metabolome data of Tripterygium wilfordii comprises the following steps: The metabolites of Tripterygium wilfordii were extracted to obtain an extract, and then the metabolites in the extract were analyzed using a high-throughput real-time plasma mass spectrometer to obtain the metabolome data of Tripterygium wilfordii.

[0034] The present invention does not impose any specific limitation on the extraction solvent. In some optional embodiments, methanol-water solution is used to extract Tripterygium wilfordii to obtain an extract.

[0035] In some optional embodiments, the classification model includes a support vector machine, a feedforward neural network model and a random forest model, preferably a random forest model.

[0036] In a third aspect, the present invention provides the application of metabolites in tracing the origin of Tripterygium wilfordii. The metabolites are histamine, glycolaldehyde dimer, nicotinamide, bifonazole, 5-(6-hydroxy-2,5,7,8-tetramethylchroman-2-yl)-2-methylpentanoic acid, spartocyanine, 1-methyl-2-undecyl-4(1H)-quinolone, 5,6,7-trimethoxyflavone, petunidin, D-erythrosphingosine, phytosphingosine, guaiacylamine, oleic acid and eicos-19-ene-1,2,4-triol.

[0037] According to the research of the inventors, histamine, glycolaldehyde dimer, nicotinamide, bifonazole, 5-(6-hydroxy-2,5,7,8-tetramethylchroman-2-yl)-2-methylpentanoic acid, spartoine, 1-methyl-2-undecyl-4(1H)-quinolone, 5,6,7-trimethoxyflavone, petunidin, D-erythrosphingosine, phytosphingosine, guaiacyl oil, oleic acid and eicosapentaenoic acid-19-ene-1,2,4-triol are used as markers, and the origin of the tested three-leaf green can be accurately traced through a classification model.

[0038] In some optional embodiments, the origin includes Zhejiang, Fujian, Jiangxi, Yunnan, Guangxi and Chongqing.

[0039] The above metabolites can be used as markers for the differentiation and identification of Tetrastigma hemsleyanum from Zhejiang, Fujian, Jiangxi, Yunnan, Guangxi and Chongqing.

[0040] The present invention will be further illustrated by specific examples below. However, it should be understood that these examples are only for more detailed illustration and should not be construed as limiting the present invention in any way.

[0041] Example 1 1. Materials and Methods 1.1 Samples and Reagents A total of 223 samples were collected for the origin traceability study of the metabolomics of Tetrastigma hemsleyanum. These samples came from 6 different origins, including 32 from Zhejiang (ZJ), 50 from Fujian (FJ), 22 from Jiangxi (JX), 51 from Yunnan (YN), 37 from Guangxi (GX), and 31 from Chongqing (CQ). The specific information of the origins is shown in Table 1. Purified water was generated by a Milli-Q water purification system (Millipore Billerica, Massachusetts, USA), methanol was purchased from Fisher Chemical (Pittsburgh, Pennsylvania, USA), and triazophos was purchased from Tanmo Quality Inspection Technology Co., Ltd. (Changzhou, China).

[0042] Table 1 Information Table of the Origins of Tetrastigma hemsleyanum 。

[0043] 1.2 Sample Pretreatment Wash the Tetrastigma hemsleyanum, dry it at 60 °C for 24 h, then peel it, and put it into a high-throughput tissue grinder to grind it into powder. Accurately weigh 100 mg of the sample powder into a centrifuge tube, add 10 mL of methanol-aqueous solution (80:20, v / v), and extract it at room temperature for 10 min using an ultrasonic water bath (500 W, 220 V, 50 Hz). Subsequently, add methanol-aqueous solution (80:20, v / v) to the centrifuge tube to compensate for the weight loss, put it into a centrifuge, and centrifuge it at 10000 rmp at room temperature for 10 min. Take the supernatant as the sample solution for injection. At the same time, take equal amounts of the extracts of the 223 Tetrastigma hemsleyanum samples and mix them to obtain a quality control (QC) sample to monitor the stability and reproducibility of sample analysis.

[0044] 1.3 HRT-MS Data Acquisition The prepared samples were analyzed using a high-throughput real-time inductively coupled plasma mass spectrometer (HRT-MS, HRT-9500P) in positive ion mode. To ensure the stability of the instrument during the experiment, triazophos was used as a reference substance to evaluate the stability of the instrument by comparing its characteristic peaks. One injection of triazophos was made for every 20 samples. When the score of triazophos was lower than 90, the instrument was considered unstable and needed to be retuned. The mass spectrometry parameters during the analysis included a stage temperature of 230 °C, an ion source voltage of 2500 V, and a collision voltage of 55 V. The acquisition time for a single run analysis was 90 s, the relative molecular mass acquisition range was 72 - 1050 Da, and the injection volume for each time was 10 μL.

[0045] 1.4 Data processing and chemometric analysis The raw data was converted into MZML format using the format conversion program provided by the instrument company for subsequent data processing. MS-DIAL 4.60 was used to preprocess the data, including peak alignment, baseline correction, noise removal, normalization, etc. To ensure the reliability of the data used for analysis, peak values with a signal-to-noise ratio (S / N) greater than 30 and a relative standard deviation (RSD) less than 0.03 were screened out using Microsoft Excel 2007. SPSS 22 was used to perform analysis of variance to obtain the P value. SIMICA 14.1 was used to obtain the VIP value of differential metabolites and to perform PCA, PLS-DA, and OPLS-DA analyses. R language was used to visualize the information of differential metabolites and to establish support vector machine (SVM), feedforward neural network (FNN), and random forest (RF) models to distinguish Tetrastigma hemsleyanum from different origins.

[0046] 2. Results and discussion 2.1 General situation of differential metabolites of Tetrastigma hemsleyanum from six origins To distinguish Tetrastigma hemsleyanum from six origins, we performed an analysis using non-targeted metabolomics methods. We analyzed the metabolites in Tetrastigma hemsleyanum from six origins by HRT-MS, and a total of 16950 ion fragments were detected. Through the analysis of the metabolites of Tetrastigma hemsleyanum from six origins, we screened metabolites with a VIP value > 1 and a P value < 0.05 as differential metabolites for subsequent analysis. Fourteen metabolites were preliminarily identified by matching the Massbank, NIST14, and HMBD databases. These metabolites included heterocyclic compounds, alkaloids, flavonoids, amino alcohols, terpenoids, fatty acids, and alcohols (Table 2). The shape of the radar chart reflects the similarity of metabolites in samples from different origins. From Figure 1It can be seen that the radar chart shapes of Yunnan and Chongqing are basically the same, while those of the other four production areas have some differences, indicating that the composition of differential metabolites in Yunnan and Chongqing is relatively similar. Climate change can affect the metabolism of plants. For example, light, temperature, ozone, and ultraviolet radiation can all affect plant metabolism. The similarity of differential metabolites in Yunnan and Chongqing may be due to similar climate conditions. The clustering heatmap ( Figure 2 ) provides an intuitive overview of the data and offers a visual comparison of the distribution of differential metabolites in each production area. Similar colors indicate that the relative contents of differential metabolites in the two production areas are similar. Consistent with the results of the radar chart, the colors of Chongqing and Yunnan in the clustering heatmap are relatively close and they are also clustered into one category, initially showing that the composition of differential metabolites of Tetrastigma hemsleyanum in Yunnan and Chongqing is similar. This may make it difficult to distinguish the production areas of Tetrastigma hemsleyanum using the differential metabolites from these two places.

[0047] Table 2 Information Table of Differential Metabolites .

[0048] The metabolic pathways involved in the differential metabolites of Tetrastigma hemsleyanum include sphingolipid metabolism, cutin, suberine and wax biosynthesis, fatty acid biosynthesis, tropane, piperidine and pyridine alkaloid biosynthesis, and biosynthesis of unsaturated fatty acids. Among them, sphingolipid metabolism involves more differential metabolites, and the sphingolipid metabolism process has a certain connection with environmental changes. Research shows that the regulation of sphingolipid metabolism enables plants to promote cell growth and appropriately respond to biotic and abiotic stresses, manifested as changes in metabolite concentrations to regulate their responses. Among them, sphingolipid metabolism plays a significant role in plants' adaptation to temperature changes. Under cold conditions, the contents of some metabolites in the sphingolipid metabolism process increase, and these changes help the plasma membrane to better carry out hydration, thereby increasing membrane stability during cold stress. Cutin, suberine and wax biosynthesis are closely related to factors such as light, water, and soil nutrients during plant growth. For example, under strong light conditions, plants will enhance the synthesis of cutin and wax to reduce water transpiration loss. Under drought conditions, plants will adapt to water stress by enhancing the synthesis of cutin and wax. Through the analysis of metabolic pathways, it can also be further proved that environmental changes have a certain impact on plant metabolites. Using metabolites to distinguish Tetrastigma hemsleyanum from different production areas is a reliable method.

[0049] 2.2 PCA, PLS-DA, OPLS-DA Classification and Prediction To study the discrimination effect of differential metabolites on Tetrastigma hemsleyanum from different origins, we performed PCA analysis on the differential metabolites. The PCA score plot ( Figure 3 A in) shows that the Tetrastigma hemsleyanum from six origins are roughly divided into three parts. The Tetrastigma hemsleyanum from Jiangxi and Guangxi are in the first quadrant and are completely separated; those from Chongqing and Yunnan are mainly distributed in the second quadrant with severe overlap; a small part of the Tetrastigma hemsleyanum from Zhejiang is in the third quadrant and most of them are in the fourth quadrant, and the Tetrastigma hemsleyanum from Fujian is distributed in the fourth quadrant with a little overlap with that from Zhejiang. The results show that the origins that are not completely separated are all provinces with adjacent geographical locations, which further proves that the differential metabolites of Tetrastigma hemsleyanum are closely related to its geographical origin. The R 2 X of the PCA model in this study is 0.666, and Q 2 is 0.446, indicating that the repeatability and predictability of the model are weak.

[0050] Compared with the PCA model, both PLS-DA and OPLS-DA are supervised models, which can filter out the orthogonal variables irrelevant to the classification variables in the metabolites, thus showing more superior classification ability compared with PCA. Figure 3 B and C in are the score plots of PLS-DA and OPLS-DA of the differential metabolites. The distribution of Tetrastigma hemsleyanum from each origin is consistent with the PCA results, but the distribution of each origin is more concentrated. The Tetrastigma hemsleyanum from Chongqing and Yunnan still overlap severely, and there is a little overlap between those from Zhejiang and Fujian. This echoes the results of Figure 1 and Figure 2 . The radar charts of the differential metabolites of Tetrastigma hemsleyanum from Chongqing and Yunnan are almost the same, and the colors of the clustering heat maps are similar, so they are also clustered together. The result reflected in the score plots of PLS-DA and OPLS-DA is that the two origins cannot be separated at all. The radar charts of Zhejiang and Fujian are slightly different, and the colors of the clustering heat maps are quite different, so they are not clustered together, so there is only a little overlap. To more intuitively observe the prediction effect of the PLS-DA and OPLS-DA models on the origin, we divided the data into a training set and a validation set according to a ratio of 3:1, and then used the untreated data, the data after Ln treatment, the data after Log treatment, and the data after square root treatment to predict the Tetrastigma hemsleyanum from different origins respectively. The prediction results are shown in Table 3. Overall, the PLS-DA model established with the data after Log treatment has the best prediction effect. The prediction accuracy rate of the training set is 96.64%, and the accuracy rate of the validation set is 96.31%. The main parameters of this model are R 2 X = 1, R 2 Y = 0.776, Q 2(cum) = 0.702, and they are all greater than 0.05, indicating that the model is relatively stable and has good explanatory and predictive abilities. 200 permutation tests were conducted on this model, and the results showed that R 2 Y and Q 2 The intercepts of Y are less than 0.3 and 0.05 respectively ( Figure 3 D in

[0051] Table 3 Prediction accuracies of PLS-DA and OPLS-DA models for T. hemsleyanum from different origins 。

[0052] Note: The data unit in the table is %.

[0053] 2.3 Machine learning classification prediction PCA, PLS-DA, and OPLS-DA models established based on differential metabolites could not completely distinguish the origins of T. hemsleyanum. Especially for T. hemsleyanum from Chongqing and Yunnan, they overlapped severely in the score plot and were often misjudged. Nowadays, machine learning models are widely used in the establishment of origin traceability models. Compared with the above classifiers, machine learning algorithms show powerful capabilities in data processing and classification. Machine learning can extract deeper hidden patterns within the data from a large amount of data and use them for prediction or classification. Therefore, in order to better distinguish T. hemsleyanum from different origins, machine learning algorithms need to be used for modeling and discrimination to solve this problem.

[0054] In this study, the data was pre-processed in different ways. The differential metabolite data was divided into a training set and a validation set at a ratio of 3:1. Three machine learning algorithms, support vector machine (SVM), feedforward neural network (FNN), and random forest (RF), were used to classify and predict T. hemsleyanum from different origins. In addition to calculating metrics, the AUC value of the ROC curve was also used as a performance metric for machine learning algorithms. The closer the AUC value is to 1, the better the performance of the model. The specific discrimination results are as follows Figures 4 - 7As shown in the figure. The SVM algorithm has been widely used in classification. The core idea of SVM is to find an optimal hyperplane in the feature space, which can maximize the separation of data points of different categories. In this study, we selected the sigmoid kernel function, and the discriminant model established using the original data had the best effect. The prediction accuracies of the training set and the validation set were 100% and 97.97% respectively, and their AUC values for both the training set and the validation set were 1, indicating that the model had strong prediction ability and was very stable. FNN is a non-linear adaptive model, consisting of an input layer (input of variable values), a middle layer (hidden layer), and an output layer (output of results). It is a method for actively discovering effective mappings and is commonly used for prediction and classification. We used FNN with a sigmoid activation function, and the results showed that the data after square root processing had the best prediction effect. During the construction of the FNN model, the number of hidden layer nodes was 8, and the weight decay size was 0.42. The prediction accuracies of the obtained model for the training set and the validation set were 100% and 96.25% respectively, and their AUC values were both 1, proving that the model had good prediction ability and was very stable. RF is an ensemble learning algorithm that classifies or regresses by constructing multiple decision trees. The number of random trees in the RF model of this study was 500, and the importance order of each variable was: histamine > petunidin > 5-(6-hydroxy-2,5,7,8-tetramethylchroman-2-yl)-2-methylpentanoic acid > D-erythro-sphingosine > guaiol > phytosphingosine > glycolaldehyde dimer > 1-methyl-2-undecyl-4(1H)-quinolone > bifonazole > icosadec-19-ene-1,2,4-triol > oleic acid > sparteine > nicotinylamidopyrine > 5,6,7-trimethoxyflavone. The discriminant effect of the model established using the data after Ln processing was the best, and the accuracies of both the training set and the validation set were 100%, and their corresponding AUC values were both 1, proving that the model could completely distinguish Tetrastigma hemsleyanum from different origins and was very stable.

[0055] Generally speaking, compared with other models, the RF model with data after Ln processing had the best origin prediction effect. Figure 4The confusion matrix shows the specific discrimination of three machine learning methods for the Hedyotis diffusa from each origin after Ln processing. It can be seen that only the RF model correctly discriminated the origins of all Hedyotis diffusa. For other models, like the PLS-DA and OPLS-DA models, there were partial misclassifications between the Hedyotis diffusa from Yunnan and Chongqing, and also between those from Zhejiang and Fujian. The reason for the better origin prediction effect of the RF model may be that, compared with other classification models, the random forest model has the advantages of institutionalization, regularization, connectivity, high preference variation, and prominent selection in the learning model. Compared with the SVM model, the RF model can handle high-dimensional data and is not sensitive to the correlation between features. This may be the reason why the RF model was not affected by the similar metabolic characteristics of the two origins, Yunnan and Chongqing, during prediction classification. In addition, random forest can provide feature importance scores, allowing for a more intuitive observation of the metabolites that have a greater impact on origin prediction.

[0056] 3. Conclusion In this study, the metabolites of Hedyotis diffusa from different origins were rapidly detected by TPAI-TOF technology, and a traceability model for the origin was established using differential metabolites to determine the origin of Hedyotis diffusa. A total of 16,950 ion fragment results were detected in this study, and 14 differential metabolites were screened out. The results showed that the composition of differential metabolites of Hedyotis diffusa was closely related to geographical location, and the composition of differential metabolites in origins with closer distances was also similar. The PCA, PLS-DA, and OPLS-DA models were initially used to distinguish Hedyotis diffusa from different origins. Among them, the PLS-DA model established with the data after Log processing had the best prediction effect. The prediction accuracy of the training set was 96.64%, and the accuracy of the validation set was 96.31%. However, it still could not well distinguish the Hedyotis diffusa from Yunnan and Chongqing. Therefore, we also established a traceability model for the origin using three machine learning algorithms, SVM, FNN, and RF, to explore the prediction effect of the machine learning model on the origin of Hedyotis diffusa. The results showed that the RF traceability model had the best prediction effect, successfully distinguishing the Hedyotis diffusa from six origins completely. The prediction accuracies of the training set and the test set both reached 100%, and the AUC values of the training set and the test set were both 1, indicating that the model was stable. Generally speaking, this study obtained a good classification model, providing an efficient and convenient identification method for the origin traceability of Hedyotis diffusa, aiming to enrich the technical theoretical basis for origin traceability research using the metabolomics of Hedyotis diffusa, realizing the rapid traceability of Hedyotis diffusa, ensuring the safety of the Hedyotis diffusa market in China, and safeguarding the legitimate rights and interests of consumers. However, the HRT-MS technology used in this study has certain limitations and can only detect metabolites in the positive ion mode. Therefore, in future research, while ensuring speed, the instrument should be optimized to add the negative ion mode detection. This can more comprehensively study the metabolome of substances, and the combination of positive and negative ion modes will also make the results more comprehensive and the established traceability model more reliable.

[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for tracing the origin of Tripterygium wilfordii, characterized in that: The following steps are involved: a. Obtain metabolome data of Tripterygium wilfordii from different origins and screen out differential metabolites; b. using the differential metabolite data to train a classification model, and then using the trained classification model to classify the metabolome data of the tested clover green to predict the origin of the tested clover green; The different origins include Zhejiang, Fujian, Jiangxi, Yunnan, Guangxi and Chongqing.

2. The method according to claim 1, characterized in that The differential metabolites were histamine, glycolaldehyde dimer, nicotinamide, bifonazole, 5-(6-hydroxy-2,5,7,8-tetramethylchroman-2-yl)-2-methylpentanoic acid, spartocyanine, 1-methyl-2-undecyl-4(1H)-quinolone, 5,6,7-trimethoxyflavone, petunidin, D-erythrosphingosine, phytosphingosine, guaiacylglycerol, oleic acid and eicos-19-ene-1,2,4-triol.

3. The method according to claim 1, characterized in that The method for obtaining the metabolome data of the three-leaf green vine comprises the following steps: The metabolites of Tripterygium wilfordii were extracted to obtain an extract, and then the metabolites in the extract were analyzed using a high-throughput real-time plasma mass spectrometer to obtain the metabolome data of Tripterygium wilfordii.

4. The method according to claim 1, characterized in that: The classification models include support vector machines, feedforward neural network models and random forest models.

5. A device for tracing the origin of Tripterygium wilfordii, characterized in that: Includes acquisition module and classification module; The acquisition module is used to acquire the metabolomics data of the tested Clerodendrum truncatum; The classification module is used to input the metabolome data of the tested clover to a pre-trained classification model, and classify the metabolome data of the tested clover through the classification model to predict the origin of the tested clover; The classification model is trained by the following method: a. Obtain metabolome data of Tripterygium wilfordii from different origins and screen out differential metabolites; b. training a classification model using the data of the differential metabolites to obtain a pre-trained classification model; The different origins include Zhejiang, Fujian, Jiangxi, Yunnan, Guangxi and Chongqing.

6. The device according to claim 5, characterized in that The differential metabolites were histamine, glycolaldehyde dimer, nicotinamide, bifonazole, 5-(6-hydroxy-2,5,7,8-tetramethylchroman-2-yl)-2-methylpentanoic acid, spartocyanine, 1-methyl-2-undecyl-4(1H)-quinolone, 5,6,7-trimethoxyflavone, petunidin, D-erythrosphingosine, phytosphingosine, guaiacylglycerol, oleic acid and eicos-19-ene-1,2,4-triol.

7. The device according to claim 5, characterized in that The method for obtaining the metabolome data of the three-leaf green vine comprises the following steps: The metabolites of Tripterygium wilfordii were extracted to obtain an extract, and then the metabolites in the extract were analyzed using a high-throughput real-time plasma mass spectrometer to obtain the metabolome data of Tripterygium wilfordii.

8. The device according to claim 5, characterized in that The classification models include support vector machines, feedforward neural network models and random forest models.

9. Application of metabolites in the origin tracing of Tripterygium wilfordii; The metabolites are histamine, glycolaldehyde dimer, nicotinamide, bifonazole, 5-(6-hydroxy-2,5,7,8-tetramethylchroman-2-yl)-2-methylpentanoic acid, spartocyanine, 1-methyl-2-undecyl-4(1H)-quinolone, 5,6,7-trimethoxyflavone, petunidin, D-erythrosphingosine, phytosphingosine, guaiacylamine, oleic acid and eicos-19-ene-1,2,4-triol.

10. The use according to claim 9, characterized in that: The mentioned origins include Zhejiang, Fujian, Jiangxi, Yunnan, Guangxi and Chongqing.

Citation Information

Patent Citations

  • Method for identifying origin of radix tetrastigme

    CN107607485A

  • Dried ginger producing area tracing method based on characteristic metabolites

    CN115586283A

  • Apple producing area tracing method and device

    CN117391724A

  • Raw material production place tracing method for preparing black garlic

    CN118735542A

  • Metabolomics-based method and apparatus for physiological prediction, computer device, and medium

    WO2022121055A1

Cited By

  • Chromatographic fingerprint spectrum-based siraitia grosvenorii producing area traceability classification method

    CN122065142A