Method and system for distinguishing panax quinquefolium by integrating metabolomics and machine learning
By integrating metabolomics and machine learning methods, and using LC-MS detection and machine learning models to distinguish between wild and cultivated American ginseng, the problem of variety identification and origin traceability in existing technologies has been solved, and efficient and accurate identification and quality control of American ginseng varieties has been achieved.
Patent Information
- Application Number
- CN202411934740.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-12-26
AI Technical Summary
Existing technologies are insufficient to efficiently and accurately distinguish between wild and cultivated American ginseng, and American ginseng from different producing areas varies in quality and clinical efficacy. There is a lack of effective methods for variety identification and origin traceability.
By integrating metabolomics and machine learning, we analyzed American ginseng samples using ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry to screen for differential metabolites. We then used a machine learning diagnostic model for classification to establish a method and system for distinguishing between wild and cultivated American ginseng.
It has enabled accurate identification and origin traceability of American ginseng varieties, improved the accuracy and efficiency of market supervision, identified the metabolite differences among American ginseng varieties, and promoted the efficiency and accuracy of variety identification.
Smart Images

Figure CN119861154B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of American ginseng identification, and more particularly to a method and system for distinguishing American ginseng by integrating metabolomics and machine learning. BACKGROUND
[0002] American ginseng is a globally recognized medicinal and dietary plant known for its important clinical uses and health-promoting properties. Wild American ginseng is mainly distributed in North America, particularly in the United States and Canada, between latitudes 40 and 50° N and longitudes 67 and 125° W. However, wild populations are scarce and difficult to maintain, necessitating advanced cultivation techniques to ensure a stable supply.
[0003] Currently, morphological characteristics and microscopic observation remain the primary methods for distinguishing American ginseng varieties. Practical experience has shown that visual observation is unreliable and difficult to distinguish the origin of American ginseng. In recent years, although genomics has been applied to some extent in the identification of American ginseng, it is still time-consuming and labor-intensive due to its complex procedures and technical difficulties. In addition, American ginseng of different varieties and origins varies in quality and clinical efficacy, and the metabolic differences have not been fully explored. Accurate differentiation of wild, cultivated, and different sources of American ginseng is crucial for variety identification, geographical origin tracing, and quality control.
[0004] Non-targeted metabolomics involves using advanced liquid chromatography-mass spectrometry analysis to obtain metabolite profiles in the sample, comparing the differences in metabolite content, and systematically screening specific biomarkers from the differential metabolites, which helps to identify plant phenotypes and analyze metabolic variations.
[0005] Therefore, how to determine the differences in metabolites between different varieties of American ginseng and efficiently and accurately and quickly identify American ginseng varieties is a problem that needs to be solved by those skilled in the art. SUMMARY
[0006] Therefore, the present application provides a method and system for distinguishing American ginseng by integrating metabolomics and machine learning to solve the technical problems mentioned in the background.
[0007] To achieve the above purpose, the present application adopts the following technical solutions:
[0008] A method for distinguishing American ginseng by integrating metabolomics and machine learning, comprising the following steps:
[0009] S1. Prepare wild American ginseng and cultivated American ginseng samples, and prepare a metabolite extract sample solution;
[0010] S2. Detect and analyze the LC-MS data of wild and cultivated American ginseng sample components by ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry analysis on the metabolite extract sample solution;
[0011] S3. Based on metabolomics technology and molecular network strategy, LC-MS data of wild and cultivated Panax quinquefolium samples are subjected to multivariate data analysis, and differential metabolites of wild and cultivated Panax quinquefolium are screened;
[0012] S4. A machine learning diagnostic model of differential metabolites of wild and cultivated Panax quinquefolium is established and trained to distinguish wild and cultivated Panax quinquefolium.
[0013] Preferably, the method for preparing wild and cultivated Panax quinquefolium samples is as follows:
[0014] The wild and cultivated Panax quinquefolium samples are crushed and passed through a No. 3 sieve. 1 g of the powder is accurately weighed and placed in a conical flask with a stopper. 50 ml of water-saturated n-butanol is accurately added, the weight is determined, and the flask is placed in a water bath for heating and reflux extraction for 1.5 hours. After cooling, the weight is determined again, and the lost weight is made up with water-saturated n-butanol. Shake well and filter.
[0015] Preferably, the method for preparing the metabolite extract sample solution is as follows:
[0016] Accurately take 25 ml of the filtrate and place it in an evaporation dish. Evaporate to dryness. Add an appropriate amount of 50% methanol to dissolve the residue. Transfer to a 10 ml volumetric flask, add 50% methanol to the mark, shake well, and pass through a 0.22 μm microporous filter membrane. Take the filtrate to obtain the sample solution.
[0017] Preferably, the ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry detection and analysis includes:
[0018] S21. The chromatographic column, mobile phase type and gradient elution program are optimized and evaluated for chromatographic detection.
[0019] S22. Ionization is carried out by electrospray ionization for mass spectrometry detection, including ESI+ positive ion mode and ESI- negative ion mode.
[0020] Preferably, the optimized chromatographic conditions of step S21 are as follows: the chromatographic column is C18, the mobile phase is organic solvent acetonitrile and aqueous phase, the aqueous phase is 0.05 mmol / L formic acid-water, the gradient elution program is: 0-5 min, 5-15% acetonitrile; 5-20 min, 15-25% acetonitrile; 20-25 min, 25-35% acetonitrile; 25-35 min, 35-45% acetonitrile; 35-40 min, 45-55% acetonitrile; flow rate 0.4 mL / min, injection volume 2 μL;
[0021] The content of mass spectrometry detection is: before analysis, use Agilent standard mixture for accurate mass calibration, the first mass spectrum scanning range is set to 100-1700 m / z, the drying gas is nitrogen, the temperature is 325 DEG C, the flow rate is 6.8 L / min, the sheath gas temperature is 350 DEG C, the capillary voltage is 4.0 kV, and the first collision voltage is 150 V; the target mode scanning differential metabolites are analyzed by target MS, the second mass spectrum scanning range is 50-1000 m / z, the second fragmentation voltage is 10 V, 25 V, 40 V.
[0022] Preferably, the specific content of step S3 is:
[0023] S31. The original LC-MS data is pretreated, including peak extraction, identification, matching, comparison and normalization, to generate a metabolite expression matrix;
[0024] S32. Two kinds of unsupervised analysis models, principal component analysis PCA and hierarchical cluster analysis HCA, are used to identify the chemical phenotype of wild and cultivated American ginseng based on the metabolite expression matrix;
[0025] S33. In ESI + Positive ion mode and ESI - Negative ion mode, the orthogonal partial least squares discriminant analysis model is used to screen the differential components between different American ginseng varieties;
[0026] S34. The model performance is evaluated by R2, Q2 and cross-validation test, and the metabolite differences between American ginseng varieties are determined according to the variable weight value and t-test p value;
[0027] S35. The differential metabolites in American ginseng samples are identified by accurate molecular weight determination and secondary mass spectrum fragmentation analysis of molecular ion peaks, and the main metabolites are determined by comparison with standard products and mass spectrum database.
[0028] Preferably, the eight significantly different differential metabolites are ginsenoside Rg6, ginsenoside Rg2, ginsenoside F1, ginsenoside Rk2, ginsenoside R2, ginsenoside Rg10, oleanolic acid-28-O-β-D-glucopyranoside and oleanolic acid-12-ene-3,11-dione.
[0029] Preferably, the machine learning diagnosis model establishment method in step S4 is:
[0030] The machine learning model development and evaluation is performed by using Python, various classifiers such as Gaussian Bayes, logistic regression, neural network, random forest and support vector machine are used to classify American ginseng, and the American ginseng samples are classified according to the metabolites with statistical significance, the American ginseng samples are divided into a training set of 25 samples and a test set of 10 samples in a ratio of 5:2, the stability of the algorithm is verified by using 10-fold cross-validation method, and the model performance is evaluated by using accuracy, recall rate, F1 score and confusion matrix indicators.
[0031] Preferably, in step S4, the machine learning diagnosis model of the differential metabolites of wild and cultivated American ginseng is a random forest classifier.
[0032] A system for distinguishing American ginseng by integrating metabolomics and machine learning, based on the method for distinguishing American ginseng by integrating metabolomics and machine learning, comprising a metabolite extract preparation device, an LC-MS detection module, a differential metabolite characterization module and a diagnosis module.
[0033] The metabolite extract preparation device is used to prepare wild American ginseng and cultivated American ginseng samples and prepare metabolite extract sample solutions.
[0034] The LC-MS detection module is used to detect and analyze the metabolite extract sample solutions by ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry to obtain LC-MS data of the components of the wild and cultivated American ginseng samples.
[0035] The differential metabolite characterization module is used to perform multivariate data analysis on the LC-MS data of the components of the wild and cultivated American ginseng samples based on metabolomics technology and molecular network strategy, and screen differential metabolites of wild and cultivated American ginseng.
[0036] The diagnosis module is used to establish and train a machine learning diagnosis model of the differential metabolites of wild and cultivated American ginseng to distinguish wild and cultivated American ginseng.
[0037] According to the above technical solution, compared with the prior art, the present disclosure provides a method and system for distinguishing American ginseng by integrating metabolomics and machine learning, proposes to use metabolomics based on liquid chromatography-mass spectrometry LC-MS combined with machine learning to distinguish wild American ginseng and cultivated varieties, determines the differences in metabolites between wild and cultivated American ginseng, and classifies them by constructing and training an artificial intelligence diagnosis model, so as to accurately identify the origin and cultivation method of American ginseng, which is of great significance for the identification of American ginseng varieties, the tracing of the origin, and the promotion of market supervision. BRIEF DESCRIPTION OF DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only aim at the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of the provided drawings.
[0039] Figure 1 A schematic diagram of a method for distinguishing American ginseng by integrating metabolomics and machine learning provided by the present application;
[0040] Figure 2 A positive ion detection chromatogram-mass spectrum of American ginseng provided by the present application;
[0041] Figure 3 A negative ion detection chromatogram-mass spectrum of American ginseng provided by the present application;
[0042] Figure 4 American ginseng metabolic phenotypes distinguished by principal component analysis and hierarchical cluster analysis in ESI+ and ESI- modes provided by the present application;
[0043] Figure 5 A schematic diagram of OPLS-DA distinguishing American ginseng metabolic phenotypes provided by the present application;
[0044] Figure 6 A schematic diagram of different metabolites between Chinese cultivated and North American wild American ginsengs provided by the present application;
[0045] Figure 7 A schematic diagram of different metabolites between North American cultivated and North American wild American ginsengs provided by the present application;
[0046] Figure 8 A schematic diagram of different metabolites between Chinese cultivated and North American cultivated American ginsengs provided by the present application;
[0047] Figure 9 A schematic diagram of screening different metabolites of American ginseng among three different origins in a positive ion state mode provided by the present application;
[0048] Figure 10 A schematic diagram of screening different metabolites of American ginseng among three different origins in a negative ion state mode provided by the present application;
[0049] Figure 11 A schematic diagram of five kinds of machine learning distinguishing American ginsengs from different origins in a positive ion state mode provided by the present application;
[0050] Figure 12The five machine learning discriminates different producing areas of American ginseng in the positive ion state mode. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.
[0052] The embodiments of the present application disclose a method for distinguishing American ginseng by integrating metabolomics and machine learning, such as Figure 1 , comprising the following steps:
[0053] S1. Prepare wild American ginseng and cultivated American ginseng samples, and prepare a metabolite extract sample solution;
[0054] S2. Obtain LC-MS data of components of wild and cultivated American ginseng samples by detecting and analyzing the metabolite extract sample solution through ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry;
[0055] S3. Perform multivariate data analysis on the LC-MS data of components of wild and cultivated American ginseng samples based on metabolomics technology and molecular network strategy, and screen differential metabolites of wild and cultivated American ginseng;
[0056] S4. Establish and train a machine learning diagnosis model of differential metabolites of wild and cultivated American ginseng to distinguish wild and cultivated American ginseng.
[0057] In order to further implement the above technical solutions, the method for preparing wild American ginseng and cultivated American ginseng samples is as follows:
[0058] Grind the wild and cultivated American ginseng samples, pass them through a No. 3 sieve, accurately weigh 1 g of the powder, place it in a conical flask with a stopper, accurately add 50 ml of water-saturated n-butanol, weigh the weight, place it in a water bath for heating and reflux extraction for 1.5 hours, cool it down, weigh it again, make up the weight loss with water-saturated n-butanol, shake it well, and filter it.
[0059] In order to further implement the above technical solutions, the method for preparing the metabolite extract sample solution is as follows:
[0060] Accurately take 25 ml of the filtrate, place it in an evaporation dish, evaporate it to dryness, add an appropriate amount of 50% methanol to dissolve the residue, transfer it to a 10 ml volumetric flask, add 50% methanol to the calibration mark, shake it well, pass it through a 0.22 mu m microporous filter membrane, and take the filtrate to obtain the sample solution.
[0061] To further implement the above technical solutions, the ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry detection and analysis includes:
[0062] S21. Optimizing and evaluating the chromatographic column, the mobile phase type and the gradient elution program for chromatographic detection;
[0063] S22. Ionizing by electrospray ionization for mass spectrometry detection, including ESI + positive mode and ESI-negative mode.
[0064] To further implement the above technical solutions, the optimized chromatographic conditions of step S21 are as follows: the chromatographic column is C18, the mobile phase is organic solvent acetonitrile and aqueous phase, the aqueous phase is 0.05 mmol / L formic acid (mass spectrometry grade)-water, the gradient elution program is: 0-5 min, 5-15% acetonitrile; 5-20 min, 15-25% acetonitrile; 20-25 min, 25-35% acetonitrile; 25-35 min, 35-45% acetonitrile; 35-40 min, 45-55% acetonitrile; the flow rate is 0.4 mL / min, and the injection amount is 2 μL;
[0065] The content of mass spectrometry detection is as follows: before analysis, Agilent standard mass mixture is used for accurate mass calibration, the first mass spectrometry scanning range is set to 100-1700 m / z, the drying gas is nitrogen, the temperature is 325°C, the flow rate is 6.8 L / min, the sheath gas temperature is 350°C, the capillary voltage is 4.0 kV, and the first collision voltage is 150 V; the target mode is used to scan the differential metabolites for target MS analysis, the second mass spectrometry scanning range is 50-1000 m / z, and the second fragmentation voltage is 10 V, 25 V and 40 V;
[0066] The results are as follows: Figure 2 ESI + positive ion mode ion flow chart of American ginseng LC-MS, Figure 3 ESI - negative ion mode ion flow chart of American ginseng LC-MS.
[0067] To further implement the above technical solutions, the specific content of step S3 is as follows:
[0068] S31. Preprocessing the original LC-MS data, including peak extraction, identification, matching, comparison and normalization, to generate a metabolite expression matrix;
[0069] S32. Using two unsupervised analysis models of principal component analysis PCA and hierarchical cluster analysis HCA to identify the chemical phenotypes of wild and cultivated American ginseng based on the metabolite expression matrix;
[0070] S33. Identifying the chemical phenotypes of wild and cultivated American ginseng by ESI +Positive mode and ESI - In negative mode, orthogonal partial least squares discriminant analysis model was used to screen the different components between different American ginseng varieties;
[0071] S34. The performance of the model was evaluated using R2, Q2 and cross-validation test, and the differences in metabolites between American ginseng varieties were determined according to the variable weight value and the p value of t test;
[0072] S35. The differential metabolites in American ginseng samples were identified by accurate molecular weight determination and secondary mass spectrometry fragmentation analysis of molecular ion peaks, and the main metabolites were determined by comparison with standard products and mass spectrometry database;
[0073] In this embodiment, as Figure 4 ESI + and ESI - In this embodiment, as + Figure 3 A, in the positive ion state, principal component analysis was used to distinguish the metabolic phenotypes of American ginseng from different origins, and PCA showed that the quality control (QC) samples were concentrated in the scatter plot, indicating the robustness and reliability of the adopted metabolomics analysis; B, principal component analysis was used to distinguish the metabolic phenotypes of wild American ginseng, American ginseng cultivated in North America and American ginseng cultivated in China; C, hierarchical cluster analysis was used to distinguish the metabolic phenotypes of American ginseng from different origins, and HCA further showed that all wild American ginseng samples formed a unique cluster, while American ginseng cultivated in North America and China were classified into another chemical subtype; consistent with ESI + metabolome data, ESI - mode analysis also revealed different metabolic subtypes of wild, American ginseng cultivated in North America and American ginseng cultivated in China, which was confirmed by PCA and HCA, specifically: in ESI Figure 3 A, the PCA scatter plot of QC samples contains PCA scatter plot of QC samples, which shows the robust clustering and reliability of the LC-ESI-MS method, B, the PCA scatter plot of QC samples in ESI-mode, C, hierarchical cluster analysis of metabolic differences of three types of American ginseng in ESI-mode; these analyses collectively emphasize the obvious metabolic differences between wild and cultivated American ginseng;
[0074] Figure 5 In order to distinguish the chemical types of wild, American ginseng cultivated in North America and American ginseng cultivated in China by using partial least squares discriminant analysis (OPLS-DA), specifically: as Figure 4 A, in the positive ion state, OPLS-DA model was used to compare the metabolic types of American ginseng from different origins, ESI + The orthogonal partial least squares discriminant analysis (OPLS-DA) in the mode showed that there was a clear inter-group separation and intra-group clustering trend between wild and North American cultivated Panax quinquefolium samples; as shown in Figure 4 Figure B in the figure is the permutation test of the OPLS-DA model, and 200 permutation tests prove the stability and reliability of the established OPLS-DA model; as shown in Figure 4 Figure C in the figure is the SUS-PLOT screening of different origin Panax quinquefolium, Figure D is the scatter plot screening of different origin Panax quinquefolium, and Figure E is the most fluctuating differential component. Using the screening criteria of VIP greater than 1 and absolute p(corr) value greater than 0.5, 228 differential metabolites were identified, of which 40 were annotated and characterized.
[0075] Figure 6 ESI + mode, wherein Figure A shows that there is a clear separation and intra-group clustering between wild Panax quinquefolium (C) and Chinese cultivated Panax quinquefolium (N); 200 permutation tests in Figure B prove the robustness and reliability of the established OPLS-DA model; Figure C is a loading plot depicting different metabolites, showing metabolites at different levels in the periphery. The metabolite screening criteria are VIP>1, p(corr)>0.5;
[0076] Figure 7 ESI + mode, wherein Figure A is the OPLS-DA model comparing metabolites from different origins in the positive ion state, and the scatter plot shows the separation of wild Panax quinquefolium and North American cultivated Panax quinquefolium samples with intra-group clustering; 200 permutation tests in Figure B verify the robustness and reliability of the OPLS-DA model, and the replacement test in Figure B verifies the stability and reliability of the model; Figure C is a scatter plot screening different origin Panax quinquefolium, and the loading plot of differential metabolites highlights the periphery of metabolites with variable content. The metabolite screening criteria are VIP>1, p(corr)>0.5;
[0077] Figure 8 ESI + mode, wherein Figure A shows that North American (W) and Chinese (N) cultivated Panax quinquefolium samples are separated into different groups, and clustering is observed within each group; 200 permutation tests in Figure B prove the robustness and reliability of the OPLS-DA model; Figure C is a loading plot showing different metabolites, and these metabolites show metabolites at different levels in the periphery. The screening criteria are VIP>1, pcorr>0.5;
[0078] Compared with wild American ginseng (C), the levels of 34 metabolites were reduced in North American cultivated American ginseng (W), and the levels of 6 metabolites were increased in North American cultivated American ginseng (W);
[0079] Similarly, in the ESI-mode, the OPLS-DA model identified 248 different metabolites in wild American ginseng (C) and Chinese cultivated American ginseng samples, of which 39 metabolites were annotated, and compared with wild American ginseng (C), the levels of 34 metabolites were reduced in Chinese cultivated American ginseng (N), and the levels of 5 metabolites were increased;
[0080] In order to further implement the above technical solutions, the 8 significantly different differential metabolites are ginsenoside Rg6, ginsenoside Rg2, ginsenoside F1, ginsenoside Rk2, ginsenoside R2 and ginsenoside Rg10, and oleanolic acid-28-O-β-D-glucopyranoside and oleanolic acid-12-ene-3,11-dione.
[0081] As Figure 9 From the comparison of ESI+ differential metabolites among the three groups, the 8 differential metabolites of wild ginseng, North American cultivated American ginseng and Chinese cultivated American ginseng are significantly different: ginsenoside Rg6, ginsenoside Rg2, ginsenoside F1, ginsenoside Rk2, ginsenoside R2 and ginsenoside Rg10, and two oleanolic acid type triterpenes (oleanolic acid-28-O-β-D-glucopyranoside and oleanolic acid-12-ene-3,11-dione), and six ginsenosides are particularly rich in wild American ginseng compared with North American and Chinese cultivated ginseng; the expression amounts of ginsenoside Rk2, ginsenoside R2 and ginsenoside Rg6 are in the order of: wild American ginseng>North American cultivated American ginseng>Chinese cultivated American ginseng.
[0082] As Figure 10 The three groups of ESI-differential metabolites are analyzed, wherein the comparative analysis of the differential metabolites of the A figure determines three different compounds, including two ginsenosides (ginsenoside Rc and ginsenoside Re) and sucrose, the B figure shows that among the three groups, the content of Rd2 metabolite is the highest in wild American ginseng, followed by ginsenoside Re, which is also higher in wild American ginseng, and fructose is higher in North American ginseng; the C figure shows that the expression trend of ginsenoside Rd2 is: wild American ginseng>Chinese cultivated American ginseng>North American cultivated American ginseng, and the D figure is a structural description of the differential metabolites.
[0083] In order to further implement the above technical solutions, in step S4, the machine learning diagnosis model of wild and cultivated American ginseng differential metabolites is a random forest classifier.
[0084] In this embodiment, both machine learning model development and evaluation are performed using Python 3.11.4, and various classifiers such as Gaussian Bayes, logistic regression, neural network, random forest, and support vector machine are used to classify American ginseng, and the classification is performed based on metabolites with statistically significant differences. The American ginseng samples are divided into a training set of 25 samples and a test set of 10 samples in a ratio of 5:2, the stability of the algorithm is verified by 10-fold cross-validation method, and the model performance is evaluated by accuracy, recall rate, F1 score and confusion matrix indicators:
[0085]
[0086] As Figure 11 , a variety of machine learning classifiers are constructed using the expression matrix of 8 differential metabolites identified in ESI + mode to identify the source of American ginseng. The training set results show that the recognition accuracy of random forest classifier and support vector machine is 100%, which is better than Gaussian naive Bayes, logistic regression and multilayer perceptron classifier, and the recognition accuracy of the latter is 90.91%, 72.73% and 63.63%, respectively. In the validation set, the random forest classifier and the support vector machine lead with an accuracy of 100%, followed by Gaussian naive Bayes (90.91%), logistic regression (72.73%) and multilayer perceptron classifier (63.63%); misclassification is observed using Gaussian naive Bayes, logistic regression and multilayer perceptron classifier algorithms; the random forest classifier and the support vector machine have the highest diagnostic performance for American ginseng, indicating that the combination of metabolomics and machine learning can effectively identify different origins of ginseng.
[0087] In addition, as Figure 12 , a variety of machine learning classifiers are established using the expression matrix of three differential metabolites in ESI - mode to diagnose the source of American ginseng. Among them, the training set accuracy of the random forest classifier is the highest (100.0%), followed by the logistic regression classifier and the Gaussian naive Bayes classifier (95.83%), the multilayer perceptron classifier (87.5%) and the support vector machine (70.83%); In the test set, the accuracy of the random forest classifier, the logistic regression and the Gaussian naive Bayes classifier all reached 100%, among which the accuracy of the multilayer perceptron classifier was 90.91%, and the accuracy of the support vector machine was 63.64%; Except for the random forest classifier, all algorithms have misclassification problems; indicating the excellent diagnostic performance of the random forest classifier in American ginseng classification, and strengthening the practicability of the combination of metabolomics and machine learning for tracking the source of ginseng.
[0088] The application discloses a system for distinguishing American ginseng by integrating metabolomics and machine learning, and relates to a method for distinguishing American ginseng by integrating metabolomics and machine learning.
[0089] The metabolite extract preparation device is used for preparing wild American ginseng and cultivated American ginseng samples and preparing metabolite extract sample solutions.
[0090] The LC-MS detection module is used for detecting the metabolite extract sample solutions through ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry detection analysis to obtain LC-MS data of components of the wild American ginseng and the cultivated American ginseng samples.
[0091] The differential metabolite characterization module is used for performing multivariate data analysis on the LC-MS data of components of the wild American ginseng and the cultivated American ginseng samples based on metabolomics technology and molecular network strategies and screening differential metabolites of the wild American ginseng and the cultivated American ginseng.
[0092] The diagnosis module is used for establishing and training a machine learning diagnosis model of differential metabolites of the wild American ginseng and the cultivated American ginseng to distinguish the wild American ginseng and the cultivated American ginseng.
[0093] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between the embodiments can be referred to each other.
[0094] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the application. Therefore, the application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for distinguishing American ginseng by integrating metabolomics and machine learning, characterized in that, Comprise the following steps: S1. Prepare wild American ginseng and cultivated American ginseng samples, and prepare a metabolic extract sample solution; S2. Obtain LC-MS data of wild and cultivated American ginseng sample components by detecting and analyzing the metabolic extract sample solution by ultra-high performance liquid chromatography-quadrupole-time-of-flight high resolution mass spectrometry; S3. Based on metabolomics technology and molecular network strategy, perform multivariate data analysis on the LC-MS data of wild and cultivated American ginseng sample components, and screen for differential metabolites of wild and cultivated American ginseng; S4. Establish and train a machine learning diagnostic model of wild and cultivated American ginseng differential metabolites to distinguish wild and cultivated American ginseng; The chromatographic conditions are: the chromatographic column is C18, the mobile phase is organic solvent acetonitrile and aqueous phase, the aqueous phase is 0.05 mmol / L formic acid-water, and the gradient elution program is: 0~5 min, 5-15% acetonitrile; 5~20 min, 15-25% acetonitrile; 20~25 min, 25-35% acetonitrile; 25~35 min, 35~45% acetonitrile; 35~40 min, 45~55% acetonitrile; flow rate 0.4 mL / min, injection volume 2 μL; Mass spectrometry detection is performed by positive ion mode and negative ion mode ionization, the first mass spectrometry scan range is set to 100~1700 m / z, the drying gas is nitrogen, the temperature is 325℃, the flow rate is 6.8 L / min, the sheath gas temperature is 350℃, the capillary voltage is 4.0 kV, and the first collision voltage is 150V; target MS analysis is performed on the differential metabolites in target mode, the second mass spectrometry scan range is 50~1000 m / z, and the second fragmentation voltage is 10V, 25V, and 40V; 8 significantly different differential metabolites are: ginsenoside Rg6, ginsenoside Rg2, ginsenoside F1, ginsenoside Rk2, gomisin R2, and ginsenoside Rg10, and oleanolic acid-28-O-β-D-glucopyranoside and oleanolic acid-12-ene-3,11-dione.
2. The method for distinguishing American ginseng by integrating metabolomics and machine learning according to claim 1, characterized in that, The method for preparing wild American ginseng and cultivated American ginseng samples is: The wild and cultivated American ginseng samples are crushed through a No. 3 sieve, 1 g of powder is accurately weighed, placed in a conical flask with a stopper, 50 ml of water-saturated n-butanol is accurately added, the weight is determined, placed in a water bath for heating reflux extraction for 1.5 hours, cooled, the weight is determined again, the water-saturated n-butanol is added to make up for the weight loss, shaken well, and filtered.
3. The method for distinguishing American ginseng by integrating metabolomics and machine learning according to claim 1, characterized in that, The method for preparing the metabolic extract sample solution is: 25 ml of the filtrate is accurately taken, placed in an evaporation dish, evaporated to dryness, the residue is dissolved with appropriate amount of 50% methanol, transferred to a 10 ml volumetric flask, added with 50% methanol to the calibration mark, shaken well, filtered through a 0.22 μm microporous filter membrane, and the filtrate is obtained.
4. The method for distinguishing American ginseng by integrating metabolomics and machine learning according to claim 1, characterized in that, The specific content of step S3 is: S31. The original LC-MS data is pretreated, including peak extraction, identification, matching, comparison and normalization, to generate a metabolite expression matrix; S32. Two unsupervised analysis models of principal component analysis PCA and hierarchical cluster analysis HCA are used to identify the chemical phenotypes of wild and cultivated American ginseng based on the metabolite expression matrix; S33. In ESI + Positive mode and ESI - In negative mode, the different components between different varieties of American ginseng were screened by using the orthogonal partial least squares discriminant analysis model. S34. The model performance is evaluated by R2, Q2 and cross-validation test, and the difference of metabolites between American ginseng varieties is determined according to the variable weight value and the p value of t test; S35. The difference metabolites in the American ginseng sample are identified by accurate molecular weight measurement of molecular ion peaks and secondary mass spectrometry fragmentation analysis, and the main metabolites are determined by comparison with standard products and mass spectrometry database.
5. The method for distinguishing American ginseng by integrating metabolomics and machine learning according to claim 1, characterized in that, The machine learning diagnosis model establishment method in step S4 is specifically: Python is used for machine learning model development and evaluation, and multiple classifiers such as Gaussian Bayes, logistic regression, neural network, random forest and support vector machine are used to classify American ginseng, and the American ginseng samples are classified according to the statistically significant difference metabolites. According to the statistically significant difference metabolites, the American ginseng samples are divided into 25 sample training sets and 10 sample test sets in the ratio of 5:2, the stability of the algorithm is verified by 10-fold cross-validation method, and the model performance is evaluated by using accuracy, recall rate, F1 score and confusion matrix index.
6. The method for distinguishing American ginseng by integrating metabolomics and machine learning according to claim 1, characterized in that, In step S4, the machine learning diagnosis model of the difference metabolites of wild and cultivated American ginseng is a random forest classifier.
7. A system for distinguishing American ginseng by integrating metabolomics and machine learning, characterized by, A method for distinguishing American ginseng based on integrated metabolomics and machine learning according to any one of claims 1-6, comprising a metabolite extract preparation device, an LC-MS detection module, a difference metabolite characterization module and a diagnosis module; The metabolite extract preparation device is used for preparing wild American ginseng and cultivated American ginseng samples, and preparing metabolite extract sample solution; The LC-MS detection module is used for detecting the metabolite extract sample solution by ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry detection analysis, to obtain the LC-MS data of the components of wild and cultivated American ginseng samples; The difference metabolite characterization module is used for multivariate data analysis of the LC-MS data of the components of wild and cultivated American ginseng samples based on metabolomics technology and molecular network strategy, and screening of the difference metabolites of wild and cultivated American ginseng; The diagnosis module is used for establishing and training the machine learning diagnosis model of the difference metabolites of wild and cultivated American ginseng to distinguish wild and cultivated American ginseng.
Citation Information
Patent Citations
American ginseng identification platform and method for identifying American ginseng by using the same
CN111220752A
American ginseng growth age prediction method, model training method and device
CN113496309A