A method for identifying the authenticity of traditional Chinese medicine based on artificial intelligence and multi-omics technology
By integrating artificial intelligence with multi-omics technologies and combining metallomics and plant metabolomics data, a high-precision classification model was constructed, which solved the problems of time consumption and low accuracy in the process of identifying the authenticity of the Chinese medicinal herb Peucedanum praeruptorum. This enabled rapid and accurate identification of authenticity, and improved the scientific nature of quality control and efficacy evaluation.
Patent Information
- Application Number
- CN202411801541.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-12-09
AI Technical Summary
In existing technologies, the process of identifying the authenticity of the medicinal herb Peucedanum praeruptorum is time-consuming and has low accuracy. Single-dimensional detection cannot meet the quality evaluation needs after the expansion of the planting area.
By employing AI-based multi-omics technologies and combining machine learning with metallomics and plant metabolomics data, a high-precision classification model is constructed to achieve the traceability of the authenticity of traditional Chinese medicine.
It enables rapid and accurate identification of the authenticity of traditional Chinese medicine, improves the reliability and applicability of the identification results, reduces the risk of overfitting, enhances the generalization ability of the model, and provides a more accurate quality control and efficacy evaluation scheme.
Smart Images

Figure CN119649954B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medicinal material identification, and more specifically, to a method, device, medium, and program product for identifying the authenticity of traditional Chinese medicine based on artificial intelligence and multi-omics technology. Background Technology
[0002] "Authentic medicinal herbs" refers to Chinese medicinal herbs that possess distinct regional characteristics, grow in suitable environments, exhibit superior quality, and are cultivated (cultivated) and processed in a reasonable manner, resulting in quality superior to those from other producing areas. Taking *Peucedanum praeruptorum* (Qianhu) as an example, it is a traditional Chinese medicine with a long history and a prominent position in traditional Chinese medicine, serving as a key ingredient in ancient classic prescriptions. It has played a crucial role in major public health events due to its unique pharmacological effects. *Peucedanum praeruptorum* has a medicinal history of over 1500 years in my country, first recorded in the *Mingyi Bielu* (Records of Famous Physicians). It is recorded in various herbal texts throughout history and was included in the *Chinese Pharmacopoeia* in 1963. After 2005, white-flowered *Peucedanum praeruptorum* became the main variety, while purple-flowered *Peucedanum praeruptorum* was included separately in 2010. In the current 2020 edition of the *Chinese Pharmacopoeia*, *Peucedanum praeruptorum* refers to the dried root of white-flowered *Peucedanum praeruptorum*.
[0003] Peucedanum praeruptorum is distributed in many provinces and regions of my country, with the areas surrounding Tianmu Mountain in Anhui, Zhejiang, and Jiangxi being the most famous. "Xin Peucedanum praeruptorum" is particularly recognized as a genuine medicinal herb. With increasing market demand, Peucedanum praeruptorum has gradually expanded from wild cultivation to semi-wild, forest, and field cultivation, with major planting areas now extending to Ningguo in Anhui, Hangzhou in Zhejiang, Chongqing, Sichuan, and Guizhou. However, the indiscriminate expansion of planting areas has led to inconsistent quality, especially with the key quality control indicator, peucedanin B, failing to meet standards in some producing areas. Previous studies have found a close correlation between peucedanin B content and production area; therefore, the authenticity of Peucedanum praeruptorum is crucial for its safety, efficacy, and quality controllability. Traceability of different production areas is essential for improving the scientific nature of quality evaluation and strengthening supervision.
[0004] Currently, the traditional method of manually identifying the origin of Angelica dahurica using traditional techniques is not only time-consuming but also has low accuracy. Other methods rely on single-dimensional detection, such as elemental fingerprinting or secondary metabolites, to determine the authenticity of medicinal materials. However, with the continuous expansion of planting areas and the increasing inconsistency in the quality of Chinese medicinal herbs, single-dimensional detection is no longer sufficient to meet current needs. Therefore, there is an urgent need to establish a method that can accurately identify the authenticity of medicinal materials to achieve quality evaluation and scientific supervision of Chinese medicinal herbs. Summary of the Invention
[0005] This invention aims to at least solve one of the technical problems existing in the prior art. To this end, this invention provides a method, device, medium, and program product for identifying the authenticity of traditional Chinese medicine based on artificial intelligence and multi-omics technology. The method of this invention, by combining machine learning, achieves rapid capture of complex system characteristics, establishes a high-precision classification model, and realizes the traceability of the authenticity of traditional Chinese medicine.
[0006] The first aspect of this application discloses a method for identifying the authenticity of traditional Chinese medicine based on artificial intelligence and multi-omics technology, the method comprising:
[0007] 101. Obtain metallomics data and plant metabolomics data of the medicinal materials to be tested;
[0008] 102. Calculate the CPS response values of key feature elements in the metallomics data;
[0009] 103. Process the plant metabolomics data to obtain mass spectrometry data containing peak information of key chemical components in the data;
[0010] 104. Input the CPS response values of the key feature elements and the peak information of key chemical components into the authentic producing area identification model to obtain the classification result of whether it belongs to the authentic producing area.
[0011] In some embodiments, the authentic producing area includes a first authentic producing area and a second authentic producing area; in step 104, if a classification result indicating that the medicinal material belongs to an authentic producing area is obtained, the metallomics and metabolomics data of the medicinal material to be tested are input into the first authentic producing area identification model for processing, and the classification result indicating whether it belongs to the first authentic producing area is output; if the output is yes, the operation ends, and the result indicating that the medicinal material to be tested belongs to the first authentic producing area is obtained; if the output is no, the result indicating that the medicinal material to be tested does not belong to the first authentic producing area is obtained, and the metallomics and metabolomics data of the medicinal material to be tested are input into the second authentic producing area identification model for processing, and the classification result indicating whether it belongs to the second authentic producing area is output; if the output is yes, the operation ends, and the result indicating that the medicinal material to be tested belongs to the second authentic producing area is obtained; if the output is no, the result indicating that the medicinal material to be tested does not belong to the second authentic producing area is obtained.
[0012] In some embodiments, the non-authentic producing areas include a first non-authentic producing area and a second non-authentic producing area; in step 104, if a classification result indicating that the medicinal material does not belong to the authentic producing area is obtained, the metallomics and metabolomics data of the medicinal material to be tested are input into the first non-authentic producing area identification model for processing, and the classification result indicating whether it belongs to the first non-authentic producing area is output; if the output is yes, the operation ends, and the result indicating that the medicinal material to be tested belongs to the first non-authentic producing area is obtained; if the output is no, the result indicating that the medicinal material to be tested does not belong to the first non-authentic producing area is obtained, and the metallomics and metabolomics data of the medicinal material to be tested are input into the second non-authentic producing area identification model for processing, and the classification result indicating whether it belongs to the second non-authentic producing area is output; if the output is yes, the operation ends, and the result indicating that the medicinal material to be tested belongs to the second non-authentic producing area is obtained; if the output is no, the result indicating that the medicinal material to be tested does not belong to the second non-authentic producing area is obtained;
[0013] Optionally, the medicinal material to be tested includes: Peucedanum praeruptorum; when the medicinal material to be tested is Peucedanum praeruptorum, the non-authentic producing areas include Guizhou and Chongqing, with Guizhou or Chongqing being the first non-authentic producing area and Chongqing or Guizhou being the second non-authentic producing area;
[0014] Optionally, the metallomics data includes major elements and trace elements; trace elements include metallic elements and non-metallic elements.
[0015] Optionally, the metabolomics data are mass spectrometry data of secondary metabolites; the mass spectrometry data are cleaned data after reducing redundancy and highly correlated features.
[0016] In some embodiments, the method for constructing the traditional producing area identification model includes:
[0017] Obtain metallomics and plant metabolomics data of Peucedanum praeruptorum medicinal materials from both authentic and non-authentic regions, with authentic or non-authentic regions as the classification label;
[0018] The metalomics data is input into the first machine learning model to obtain the predicted classification result, which is compared with the classification label of native or non-native place. The model is optimized based on the comparison result to obtain the metalomics model.
[0019] The metabolomics data is input into a second machine learning model to obtain a predicted classification result, which is then compared with the classification labels of native or non-native origin. The model is optimized based on the comparison results to obtain a metabolomics model.
[0020] The optimal model among the metallomics and metabolomics models obtained using different algorithms is selected, and the two optimal models are combined to obtain the original product area identification model;
[0021] Optionally, the method for constructing the first or second property area identification model includes:
[0022] We obtained metallomics and metabolomics data of Peucedanum praeruptorum from the first and second producing areas as classification labels for authentic products.
[0023] The metallomics data is input into a first machine learning model to obtain a predicted classification result, which is then compared with the local classification label. The model is optimized based on the comparison result to obtain a metallomics model. The metabolomics data is input into a second machine learning model to obtain a predicted classification result, which is then compared with the local classification label. The model is optimized based on the comparison result to obtain a metabolomics model.
[0024] The optimal models of metallomics and metabolomics in the first production area are selected, and the two optimal models are aggregated to obtain the first production area identification model; the optimal models of metallomics and metabolomics in the second production area are selected, and the two optimal models are aggregated to obtain the second production area identification model.
[0025] Optionally, the method for constructing the first or second non-traditional agricultural product area includes:
[0026] Metallomics and plant metabolomics data of Angelica dahurica from the first and second non-authentic producing areas were obtained from the training set. The first and second non-authentic producing areas are used as non-authentic classification labels.
[0027] The metallomics data is input into a first machine learning model to obtain a predicted classification result, which is then compared with a non-originating region classification label. The model is then optimized based on the comparison result to obtain a metallomics model. The plant metabolomics data is input into a second machine learning model to obtain a predicted classification result, which is then compared with a non-originating region classification label. The model is then optimized based on the comparison result to obtain a metabolomics model.
[0028] The optimal metallomics and metabolomics models in the first non-authentic producing area were selected, and the two optimal models were combined to obtain the identification model for the first non-authentic producing area; the optimal metallomics and metabolomics models in the second non-authentic producing area were selected, and the two optimal models were combined to obtain the identification model for the second non-authentic producing area.
[0029] Optionally, before building the model, the training set of medicinal materials is also oversampled. The oversampling process includes: copying or amplifying minority class data points and balancing the number of samples in each class in the training data. The minority class data points include any one or more of the following: specific place of origin, content of specific components.
[0030] Optionally, the aggregation method of the aggregation model includes: bootstrapping aggregation, boosting, stacking, mixing, etc.
[0031] In some embodiments, the methods used by the first machine learning model and / or the second machine learning model include: RF, KNN, Rpa, SVM, and XGB;
[0032] Optionally, when the medicinal material to be tested is Peucedanum praeruptorum, the production area includes Zhejiang and Anhui, with Zhejiang or Anhui being the first production area and Anhui or Zhejiang being the second production area;
[0033] Optionally, the models of authentic producing areas in the metallomics model using the XGB algorithm and the models of authentic producing areas in the metabolomics model using the KNN algorithm are selected and aggregated to obtain an aggregated model of authentic producing areas.
[0034] Optionally, the key feature element includes the Ca element;
[0035] Optionally, the key characteristic elements also include: Mn and Mg.
[0036] Optionally, the peak information includes: peak 17307.
[0037] A second aspect of this application discloses a method for identifying the authenticity of traditional Chinese medicine based on artificial intelligence and multi-omics technology, the method comprising:
[0038] 201. Obtain mass spectrometry data of the CPS response values and key peak information of the key characteristic elements of the medicinal material to be tested, wherein the key characteristic elements include Ca; and the key peak information includes: peak 17307.
[0039] 202. The mass spectrometry data of the key characteristic elements of the target Ca element and the target peak information are input into the authentic producing area identification model for processing to obtain the classification result of whether it belongs to the authentic producing area.
[0040] In some embodiments, the key characteristic elements further include Mn and Mg.
[0041] Optionally, the key feature elements further include: S element and B element; the peak information includes one or more of the following: peak 16696, peak 13779, peak 19774, peak 17630;
[0042] Optionally, if the classification result of belonging to the original producing area is obtained, the Mg element and / or Mn element and / or B element and / or peak 16696 and / or peak 19774 are input into the first original producing area identification model and the second original producing area identification model to obtain the result of whether it belongs to the first original producing area and whether it belongs to the second original producing area.
[0043] Optionally, if the classification result is obtained as belonging to a non-authentic producing area, the S element and / or Mg element and / or peak 13779 and / or peak 17630 are input into the first non-authentic producing area identification model and the second non-authentic producing area identification model to obtain the result of whether it belongs to the first non-authentic producing area and whether it belongs to the second non-authentic producing area.
[0044] A third aspect of this application discloses a computer device, the device comprising: a memory and a processor; the memory being used to store a computer program; and the processor executing the computer program to implement the steps of the above-described method.
[0045] The fourth aspect of this application discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0046] The fifth aspect of this application discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0047] This application has the following beneficial effects:
[0048] 1. This application innovatively discloses a method for identifying the authenticity of traditional Chinese medicine (TCM) based on artificial intelligence and multi-omics technology. This method uses a combination of inorganic and organic substances to identify patterns, specifically for identifying the authenticity of TCM using *Pheretima as a model. Based on metallomics and metabolomics, elemental fingerprints are fused with chemical composition profiles. During the identification process, key feature elements are first used to classify whether a product belongs to an authentic producing area. Then, based on the classification results, different producing areas within authentic or non-authentic producing areas are further distinguished. For example, using *Pheretima as a model, an authentic producing area identification model is first trained using training set samples. In the application stage, the key feature elements obtained during training are directly input into the authentic producing area identification model to determine whether the product belongs to an authentic producing area. If it belongs to an authentic producing area, different key feature elements are further input into different authentic producing area identification models, and the results are output. If it belongs to a non-authentic producing area, different key feature elements are further input into different non-authentic producing area identification models, and the results are output. The identification results are accurate and reliable, the identification route is clear, and it has good practicality and applicability.
[0049] 2. In the process of building the model, this application innovatively uses oversampling technology and aggregation model. Oversampling technology improves the machine learning model's ability to identify minority class samples, providing a more accurate and reliable solution for practical applications such as quality control and efficacy evaluation of traditional Chinese medicine. Aggregation model reduces the prediction error of a single model, improves prediction accuracy, reduces the risk of overfitting, enhances the model's generalization ability, has sufficient data sources and information, and provides a more comprehensive evaluation; prediction efficiency is significantly improved; and it has good practicality and flexibility. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0051] Figure 1 This is a schematic diagram of the method flow provided in the first aspect of the present invention;
[0052] Figure 2 This is a schematic diagram of the method flow provided in the second aspect of the present invention;
[0053] Figure 3 This is a schematic diagram of a traditional Chinese medicine authenticity identification system based on artificial intelligence and multi-omics technology provided in an embodiment of the present invention;
[0054] Figure 4 This is a schematic diagram of a traditional Chinese medicine authenticity identification system based on artificial intelligence and multi-omics technology provided in another embodiment of the first aspect of the present invention;
[0055] Figure 5 This is a schematic diagram of a computer device provided in an embodiment of the present invention;
[0056] Figure 6 This is a schematic diagram of the architecture of an exemplary computing device provided in an embodiment of the present invention;
[0057] Figure 7 This is a schematic diagram of the storage medium provided in an embodiment of the present invention;
[0058] Figure 8 This is a schematic diagram of the chemometric analysis results of elemental fingerprint profiles of authentic and non-authentic producing areas provided in an embodiment of the present invention; wherein, Figure 8 A represents the PCA results analysis. Figure 8 B represents the PLSDA analysis result. Figure 8 C represents the cluster heatmap analysis result;
[0059] Figure 9 This is a schematic diagram of the chemometric analysis results of metabolomics profiles of authentic and non-authentic producing areas provided in this embodiment of the invention; wherein, Figure 9 A represents the PCA results analysis. Figure 9 B represents the PLSDA analysis result. Figure 9 C represents the cluster heatmap analysis result;
[0060] Figure 10 This is the ROC curve of the original product area established by five machine learning algorithms for element fingerprint fusion provided in this embodiment of the invention;
[0061] Figure 11 This is a schematic diagram of the correlation analysis of secondary metabolites before and after removal, provided in an embodiment of the present invention; wherein, Figure 11 A represents the state before removal. Figure 11 B represents the result after removal;
[0062] Figure 12 This is an ROC curve of an unorthodox producing area, established by fusing five machine learning algorithms with secondary metabolites, as provided in this embodiment of the invention.
[0063] Figure 13 This is an ROC curve of the original producing area established by the integrated model of metallomics and metabolomics provided in the embodiments of the present invention;
[0064] Figure 14 These are the top fifteen representative components or elements with prominent importance in the variables of the traditional producing areas provided in this embodiment of the invention; among them, Figure 14 A and B represent the importance distribution of the element model. Figure 14 C and D represent the distribution of importance of chemical components. Detailed Implementation
[0065] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0066] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0067] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0068] Figure 1 This is a schematic flowchart of a method for identifying the authenticity of traditional Chinese medicine based on artificial intelligence and multi-omics technology, provided by an embodiment of the present invention. Specifically, the method includes the following steps:
[0069] 101: Obtain metallomics and metabolomics data of the medicinal materials to be tested;
[0070] In some embodiments, the metallomics data includes major elements and trace elements; trace elements include metallic elements and non-metallic elements.
[0071] In some embodiments, the metabolomics data are mass spectrometry data of secondary metabolites; the mass spectrometry data are cleaned mass spectrometry data after reducing redundancy and highly correlated features.
[0072] 102: Calculate the CPS response values of key feature elements in the metallomics data;
[0073] In some embodiments, the CPS response value is a response value after internal standard correction and subtraction of blank background.
[0074] 103: Process the metabolomics data to obtain mass spectrometry data containing peak information of key chemical components in the data;
[0075] 104: Input the CPS response values of the key feature elements and the peak information of key chemical components into the authentic producing area identification model to obtain the classification result of whether it belongs to the authentic producing area.
[0076] In some embodiments, the method for constructing the traditional producing area identification model includes:
[0077] Obtain metallomics and metabolomics data of Peucedanum praeruptorum medicinal materials from both authentic and non-authentic regions, with authentic or non-authentic regions as the classification label;
[0078] The metalomics data is input into the first machine learning model to obtain the predicted classification result, which is compared with the classification label of native or non-native place. The model is optimized based on the comparison result to obtain the metalomics model.
[0079] The metabolomics data is input into a second machine learning model to obtain a predicted classification result, which is then compared with the classification labels of native or non-native origin. The model is optimized based on the comparison results to obtain a metabolomics model.
[0080] The optimal model among the metallomics and metabolomics models obtained using different algorithms is selected, and the two optimal models are combined to obtain the original product area identification model;
[0081] In some embodiments, the authentic producing area includes a first authentic producing area and a second authentic producing area; in step 104, if a classification result indicating that the medicinal material belongs to an authentic producing area is obtained, the metallomics and metabolomics data of the medicinal material to be tested are input into the first authentic producing area identification model for processing, and the classification result indicating whether it belongs to the first authentic producing area is output; if the output is yes, the operation ends, and the result indicating that the medicinal material to be tested belongs to the first authentic producing area is obtained; if the output is no, the result indicating that the medicinal material to be tested does not belong to the first authentic producing area is obtained, and the metallomics and metabolomics data of the medicinal material to be tested are input into the second authentic producing area identification model for processing, and the classification result indicating whether it belongs to the second authentic producing area is output; if the output is yes, the operation ends, and the result indicating that the medicinal material to be tested belongs to the second authentic producing area is obtained; if the output is no, the result indicating that the medicinal material to be tested does not belong to the second authentic producing area is obtained.
[0082] Optionally, the method for constructing the first or second property area identification model includes:
[0083] We obtained metallomics and plant metabolomics data of the medicinal material *Angelica dahurica* from the first and second producing areas, with the first and second producing areas serving as the classification labels for the authentic origin.
[0084] The metallomics data is input into a first machine learning model to obtain a predicted classification result, which is then compared with the local classification label. The model is optimized based on the comparison result to obtain a metallomics model. The metabolomics data is input into a second machine learning model to obtain a predicted classification result, which is then compared with the local classification label. The model is optimized based on the comparison result to obtain a metabolomics model.
[0085] The optimal models of metallomics and metabolomics in the first production area are selected, and the two optimal models are aggregated to obtain the first production area identification model; the optimal models of metallomics and metabolomics in the second production area are selected, and the two optimal models are aggregated to obtain the second production area identification model.
[0086] In some embodiments, the non-authentic producing areas include a first non-authentic producing area and a second non-authentic producing area; in step 104, if a classification result indicating that the medicinal material does not belong to the authentic producing area is obtained, the metallomics and metabolomics data of the medicinal material to be tested are input into the first non-authentic producing area identification model for processing, and the classification result indicating whether it belongs to the first non-authentic producing area is output; if the output is yes, the operation ends, and the result indicating that the medicinal material to be tested belongs to the first non-authentic producing area is obtained; if the output is no, the result indicating that the medicinal material to be tested does not belong to the first non-authentic producing area is obtained, and the metallomics and metabolomics data of the medicinal material to be tested are input into the second non-authentic producing area identification model for processing, and the classification result indicating whether it belongs to the second non-authentic producing area is output; if the output is yes, the operation ends, and the result indicating that the medicinal material to be tested belongs to the second non-authentic producing area is obtained; if the output is no, the result indicating that the medicinal material to be tested does not belong to the second non-authentic producing area is obtained;
[0087] Optionally, the method for constructing the first or second non-traditional agricultural product area includes:
[0088] Metallomics and metabolomics data of Peucedanum praeruptorum from the first and second non-authentic producing areas were obtained from the training set. The first and second non-authentic producing areas are used as non-authentic classification labels.
[0089] The metallomics data is input into a first machine learning model to obtain a predicted classification result, which is then compared with a non-local classification label. The model is optimized based on the comparison result to obtain a metallomics model. The metabolomics data is input into a second machine learning model to obtain a predicted classification result, which is then compared with a non-local classification label. The model is optimized based on the comparison result to obtain a metabolomics model.
[0090] The optimal metallomics and metabolomics models in the first non-authentic producing area were selected, and the two optimal models were combined to obtain the identification model for the first non-authentic producing area; the optimal metallomics and metabolomics models in the second non-authentic producing area were selected, and the two optimal models were combined to obtain the identification model for the second non-authentic producing area.
[0091] Optionally, the medicinal materials to be tested include: Peucedanum praeruptorum, Peucedanum praeruptorum var. purpurea, Saposhnikovia divaricata, etc.; when the medicinal material to be tested is Peucedanum praeruptorum, the non-authentic producing areas include Guizhou and Chongqing, with Guizhou or Chongqing being the first non-authentic producing area and Chongqing or Guizhou being the second non-authentic producing area.
[0092] Optionally, before building the model, the training set of medicinal materials is also oversampled. The oversampling process includes: copying or amplifying minority class data points and balancing the number of samples in each class in the training data. The minority class data points include any one or more of the following: specific place of origin, content of specific components.
[0093] Optionally, the aggregation method of the aggregation model includes: bootstrapping aggregation, boosting, stacking, mixing, etc.
[0094] In some embodiments, the methods used by the first machine learning model and / or the second machine learning model include: RF, KNN, Rpa, SVM, and XGB;
[0095] In some embodiments, when the medicinal material to be tested is *Peucedanum praeruptorum*, the designated producing areas include Zhejiang and Anhui, with Zhejiang or Anhui being the first producing area and Anhui or Zhejiang being the second producing area. Specifically, the models of the Guizhou producing area using the SVM algorithm in the metallomics model and the models of the Guizhou producing area using the SVM algorithm in the metabolomics model are aggregated to obtain an aggregated model for the Guizhou producing area; the models of the Anhui producing area using the RF algorithm in the metallomics model and the models of the Anhui producing area using the KNN algorithm in the metabolomics model are aggregated to obtain an aggregated model for the Anhui producing area; the models of the Zhejiang producing area using the KNN algorithm in the metallomics model and the models of the Zhejiang producing area using the XGB algorithm in the metabolomics model are aggregated to obtain an aggregated model for the Zhejiang producing area; optionally, the models of the Chongqing producing area using the XGB algorithm in the metallomics model and the models of the Chongqing producing area using the SVM algorithm in the metabolomics model are aggregated to obtain an aggregated model for the Chongqing producing area.
[0096] Optionally, the key characteristic element includes Ca; optionally, the key characteristic element further includes Mn and Mg; optionally, the peak information includes peak 17307; optionally, the key characteristic element further includes S and B; optionally, the peak information includes one or more of the following: peak 16696, peak 13779, peak 19774, and peak 17630.
[0097] In some embodiments, the second aspect of this application discloses a method for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology, such as... Figure 2 As shown, the method includes:
[0098] 201. Obtain mass spectrometry data of the CPS response values and key peak information of the key characteristic elements of the medicinal material to be tested, wherein the key characteristic elements include Ca; and the key peak information includes: peak 17307.
[0099] 202. The mass spectrometry data of the key characteristic elements of the target Ca element and the target peak information are input into the authentic producing area identification model for processing to obtain the classification result of whether it belongs to the authentic producing area.
[0100] In some embodiments, the key characteristic elements further include Mn and Mg.
[0101] Optionally, the key feature elements further include: S element and B element; the peak information includes one or more of the following: peak 16696, peak 13779, peak 19774, peak 17630;
[0102] Optionally, if the classification result of belonging to the original producing area is obtained, the elements Mg and / or Mn and / or B and / or peak 16696 and / or peak 19774 are input into the first original producing area identification model and the second original producing area identification model to obtain the result of whether it belongs to the first original producing area and whether it belongs to the second original producing area. Specifically, when the first original producing area is Anhui and the second original producing area is Zhejiang, the elements Mn and peak 16696 are input into the first original producing area identification model to obtain the result of whether it belongs to the first original producing area; the elements Mg, B and peak 19774 are input into the second original producing area identification model to obtain the result of whether it belongs to the second original producing area.
[0103] Optionally, if the classification result indicates a non-authentic producing area, the S element and / or Mg element and / or peak 13779 and / or peak 17630 are input into the first non-authentic producing area identification model and the second non-authentic producing area identification model to obtain the result of whether it belongs to the first non-authentic producing area and whether it belongs to the second non-authentic producing area. Specifically, when the first non-authentic producing area is Guizhou and the second non-authentic producing area is Chongqing, the S element, Mg element, and peak 13779 are input into the first non-authentic producing area identification model to obtain the result of whether it belongs to the first non-authentic producing area; the Mg element and peak 17630 are input into the second non-authentic producing area identification model to obtain the result of whether it belongs to the second non-authentic producing area.
[0104] In some embodiments, if the result is that the medicinal material comes from a traditional producing area, then the medicinal material is considered to have authenticity.
[0105] Figure 5 This is a schematic diagram of a computer device provided in an embodiment of the present invention, such as... Figure 5 As shown, the device may include: one or more processors and one or more memories; wherein the memories store computer-readable code that, when run by the one or more processors, can perform the methods described above.
[0106] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, operations, and logic block diagrams disclosed in this embodiment. The general-purpose processor can be a microprocessor or any conventional processor, and can be based on an x86 or ARM architecture.
[0107] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0108] For example, the method or apparatus according to embodiments of this disclosure can also be used by means of Figure 6 The architecture of the computing device 3000 shown is used for implementation. For example... Figure 6 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the methods provided in this disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 6 The architecture shown is merely exemplary and can be omitted as needed when implementing different devices. Figure 6 One or more components in the computing device shown.
[0109] This invention also includes a computer-readable storage medium, such as... Figure 7 The diagram illustrates a storage medium provided in an embodiment of the present invention. The computer storage medium 4020 stores computer-readable instructions 4010. When the computer-readable instructions 4010 are executed by a processor, the method described above according to embodiments of the present disclosure can be performed. The computer-readable storage medium in the embodiments of the present disclosure can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Synchronous Link Dynamic Random Access Memory (SLDRAM), and Direct Memory Bus Random Access Memory (DR RAM). It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0110] This disclosure also provides a computer program product or system, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0111] In some embodiments, this embodiment also discloses a system for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology, such as... Figure 3 As shown, the system includes:
[0112] The first data acquisition module 301 is used or configured to acquire metallomics data and metabolomics data of the medicinal material to be tested;
[0113] The first data processing module 302 is used or configured to calculate the CPS response values of key feature elements in the metallomics data;
[0114] The second data processing module 303 is used or configured to process the metabolomics data to obtain mass spectrometry data containing peak information of key chemical components in the data.
[0115] The first classification result output module 304 is used or configured to input the CPS response value of the key feature element and the peak information of the key chemical components into the authentic producing area identification model to obtain the classification result of whether it belongs to the authentic producing area.
[0116] In some embodiments, the classification result output module 304 specifically includes: the authentic producing area includes a first authentic producing area and a second authentic producing area; if a classification result indicating that the herb belongs to an authentic producing area is obtained, the metallomics data and metabolomics data of the herb to be tested are input into the first authentic producing area identification model for processing, and the classification result indicating whether it belongs to the first authentic producing area is output; if the output is yes, the operation ends, and the result indicating that the herb to be tested belongs to the first authentic producing area is obtained; if the output is no, the result indicating that the herb to be tested does not belong to the first authentic producing area is obtained, and the metallomics data and metabolomics data of the herb to be tested are input into the second authentic producing area identification model for processing, and the classification result indicating whether it belongs to the second authentic producing area is output; if the output is yes, the operation ends, and the result indicating that the herb to be tested belongs to the second authentic producing area is obtained; if the output is no, the result indicating that the herb to be tested does not belong to the second authentic producing area is obtained. The non-authentic producing areas include a first non-authentic producing area and a second non-authentic producing area. In step 304, if a classification result indicating the medicinal material does not belong to the authentic producing area is obtained, the metallomics and metabolomics data of the medicinal material to be tested are input into the first non-authentic producing area identification model for processing, and the classification result indicating whether it belongs to the first non-authentic producing area is output. If the output is yes, the process ends, and the result indicating the medicinal material to be tested belongs to the first non-authentic producing area is obtained. If the output is no, the result indicating the medicinal material to be tested does not belong to the first non-authentic producing area is obtained. Then, the metallomics and metabolomics data of the medicinal material to be tested are input into the second non-authentic producing area identification model for processing, and the classification result indicating whether it belongs to the second non-authentic producing area is output. If the output is yes, the process ends, and the result indicating the medicinal material to be tested belongs to the second non-authentic producing area is obtained. If the output is no, the result indicating the medicinal material to be tested does not belong to the second non-authentic producing area is obtained.
[0117] In some embodiments, another embodiment also discloses a system for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technologies, such as... Figure 4 As shown, the system includes:
[0118] The second data acquisition module 401 is used or configured to acquire mass spectrometry data of the CPS response values and key peak information of the key characteristic elements of the medicinal material to be tested, wherein the key characteristic elements include Mg; and the key peak information includes any one or more of the following: peak 16696, peak 13779, peak 19774, and peak 17630.
[0119] The second classification result output module 402 is used or configured to input the mass spectrometry data of the key characteristic elements of the target Mg element and the target peak information into the authentic producing area identification model for processing, so as to obtain the classification result of whether it belongs to the authentic producing area. Specific Implementation
[0121] This study, using Peucedanum praeruptorum as a model traditional Chinese medicine, innovatively integrates non-targeted and targeted metallomics, plant metabolomics, and machine learning techniques to construct a highly efficient and precise analytical strategy, addressing the challenges of authentic pedigree and quality identification. By comprehensively analyzing the elemental fingerprint and secondary metabolite profiles of Peucedanum praeruptorum, combined with intelligent model construction, rapid and accurate identification of its authentic pedigree was achieved. This research provides new ideas, methods, and tools for the quality evaluation of Peucedanum praeruptorum, and sets a benchmark for the improvement and scientific supervision of the overall quality evaluation system for traditional Chinese medicine.
[0122] 1. Materials and Methods
[0123] 1.1 Materials and Reagents
[0124] 1.1.1 Instruments: Inductively Coupled Plasma Mass Spectrometer (Shimadzu Corporation, Japan, Model 2030LF); Microwave Digester (CEM Corporation, USA, Model MARS6); Electronic Balance (METTLER TOLEDO, Model XSE205DU); Ultrapure Water System (Milli-Q Advantage A10, Millipore, USA); Acid Removal Heater (Beijing Donghang Science & Technology Instrument Co., Ltd., Model BHW-C); Ultra-High Performance Liquid Chromatography (Thermo Scientific, Model Vanquish); Orbitrap Fusion Lumos Tribrid Mass Spectrometer (Thermo Scientific), equipped with Xcalibur 3.2 software and Compound Discover 2.0 software (Thermo Scientific).
[0125] 1.1.2 Reagents: Chromatographic grade nitric acid (Sigma-Aldrich, USA); single-element standard solutions of Al (aluminum), Ag (silver), As (arsenic), Ba (barium), B (boron), Be (beryllium), Ca (calcium), Cu (copper), Cd (cadmium), Co (cobalt), Cr (chromium), Fe (iron), Hg (mercury), K (potassium), Mg (magnesium), Mn (manganese), Mo (molybdenum), Na (sodium), Ni (nickel), Pb (lead), Rb (rubidium), Sr (strontium), Sb (antimony), Se (selenium), Sn (tin), Ti (titanium), Tl (thallium), V (vanadium), and Zn (zinc) were purchased from the National Institute of Metrology, China; the tuning solution was a mixed standard solution of Li, Y, Ce, Tl, and Co (1 μg / L); the internal standard solution was a mixed internal standard solution containing Ge 100 μg / mL (Agilent Technologies, USA). Company, batch number 5188-6525; chromatographic methanol (DiKMA, mass spectrometry grade), chromatographic acetonitrile (DiKMA, mass spectrometry grade), mass spectrometry grade formic acid (Thermo Fisher Scientific, mass spectrometry grade, batch number A117-50), and deionized water was purified by Milli-Q.
[0126] 1.1.3 Samples: A total of 73 batches of Angelica dahurica (white-flowered angelica) were collected in this study, mainly from major producing areas, medicinal material markets, medicinal material companies, and pharmacies. The main producing areas were Zhejiang, Anhui, Chongqing, and Guizhou, with some samples originating from the authentic producing areas of Zhejiang and Anhui. All samples were identified as Angelica dahurica (white-flowered angelica) by Chief Pharmacist Jin Hongyu of the China National Institutes for Food and Drug Control. Peucedanum praeruptorum Dried roots of Dunn. Samples were stored under cool conditions in the laboratory of the Institute of Traditional Chinese Medicine and Ethnic Medicine, China National Institutes for Food and Drug Control.
[0127] 1.2 Metallomics Analysis Method: After pulverizing the sample and passing it through a No. 4 sieve, accurately weigh 0.5 g of the sample to be tested and place it in a microwave digestion vessel. Add 8.0 mL of nitric acid, set up the apparatus according to the operating procedure, and digest (heat to 120 ℃ for 3 min and hold for 3 min, heat to 150 ℃ for 3 min for 2 min, heat to 200 ℃ for 12 min for 2 min). After digestion, cool to below 60 ℃, place the digestion vessel on a 100 ℃ heating plate, and perform acid removal until no brownish-yellow acid fumes are emitted. Remove the digestion vessel, let it cool, and transfer the digestion solution to a 50 mL volumetric flask. Wash the digestion vessel three times with a small amount of water, combine the washings in the volumetric flask, dilute to the mark with water, and shake well to obtain the test solution. Prepare a reagent blank solution simultaneously using the same method.
[0128] This study employed a collision cell mode. During ICP-MS operation, the high-frequency power was 1.20 kW, the carrier gas (high-purity argon) flow rate was 0.70 L / min, the plasma gas flow rate was 9.0 L / min, the auxiliary gas flow rate was 1.1 L / min, the peristaltic pump speed was 0.3 r / s, the sampling depth was 7 mm, the cell gas flow rate was 6.0 mL / min, the cell voltage was -21 V, the energy filter was 7.0 V, and the nebulizer temperature was 5 °C. Semi-quantitative analysis included elements and isotopes such as… 107 Ag、 27 Al、 75 As、 197 Au、 11 B 138 Ba、 9 Be、 79 Br、 44 Ca, 114 Cd, 140 Ce、 35 Cl、 59 Co、 52 Cr 133 Cs、 63 Cu, 164 Dy、 166 Er、 153 Eu、 56 Fe、 69 Ga, 158 Gd, 180 Hf, 202 Hg, 165 Ho、 127 I, 193 Ir、 39 K, 139 La、 7 Li, 24 Mg 55 Mn, 95 Mo、 23 Na、 93 Nb, 142 Nd, 60 Ni、 192 Os、 31 P, 208 Pb, 108 Pd, 141 Pr, 195 Pt, 85 Rb、 187 Re、 102 Ru、 34 S, 121 Sb、 78 Se、 28 Si、 152 Sm, 118 Sn、88 Sr. 181 Ta、 130 Te、 232 Th、 47 Ti、 205 Tl、 169 Tm、 238 U、 51 V. 184 W, 89 Y、 174 Yb、 66 Zn, 90 Accurately measure an appropriate amount of internal standard solution (Zr) and place it in a volumetric flask. Dilute to the mark with 5% nitric acid solution (v / v). 72 Ge 6 Li, 45 Sc、 115 In、 209 Bi、 175 Lu、 103 Rh、 159 A Tb concentration of 500 ng / mL was used to obtain the internal standard solution. Considering the principles of stability and similar mass numbers, a suitable internal standard was selected. The sample tubes were then sequentially inserted into the standard solution and the sample solution for determination. The CPS response values for each element were obtained through internal standard correction and blank background subtraction.
[0129] 1.3 Metabolite Profile Analysis Method: Take about 0.5g of Peucedanum praeruptorum powder (passed through a No. 3 sieve), accurately weigh it, place it in a stoppered conical flask, accurately add 25ml of methanol, seal tightly, weigh, sonicate (power 250W, frequency 33kHz) for 30 minutes, cool, shake well, and filter to obtain the sample solution.
[0130] Chromatographic conditions: The column was an ACCQUITY UHPLC HSS T3 (2.1 × 100 mm, 1.8 μm), and the mobile phase was acetonitrile (A)-0.1% formic acid (B). The gradient elution program was as follows: 0–3 min, 90%–70% B; 3–7 min, 70%–50% B; 7–20 min, 50%–40% B; 20–30 min, 40%–10% B; 30–32 min, 10% B; 32–35 min, 10%–90%. The flow rate was 0.3 mL / min; the column temperature was 30 ℃; and the injection volume was 2 μL.
[0131] Mass spectrometry conditions: Heated electrospray ionization (H-ESI) was used in positive ion mode. The nebulizing gas was high-purity nitrogen (Ar), and the collision gas was high-purity helium (He). The sheath gas flow rate was 45 L / min, the auxiliary gas flow rate was 15 L / min, the spray voltage was 3.5 kV in positive ion mode, the capillary temperature was 350 °C, and the auxiliary gas temperature was 320 °C. The primary scan mode was MS OT, with a mass range of m / z 70–1000, a resolution of 120,000, and an injection time of 15 ms. The secondary mass spectrometry detector was Orbitrap (dd-MS2 OT HCD) with a resolution of 30,000. The collision energies (CE) were 15, 30, and 50 eV.
[0132] 1.4 General Data Processing Methods: The multi-element CPS values obtained from semi-quantitative sampling were initially standardized by dividing the CPS by the sample size. For different production area distributions, the R language chemometrics package was used for initial normalization, followed by PCA, OPLS-DA, and cluster heatmap analysis. General chemometric characteristics were obtained through data dimensionality reduction.
[0133] The plant metabolomics processing method was as follows: the collected raw files were input into QI software (Waters Corporation) for preprocessing, including peak extraction and peak matching, to initially obtain the ion response intensity of each sample. R language chemometrics packages were used for initial normalization, followed by PCA, OPLS-DA, and cluster heatmap analysis. General chemometric characteristics were obtained through data dimensionality reduction.
[0134] 1.5 Machine Learning Modeling Methods
[0135] 1.5.1 Dataset Partitioning: In this study based on semi-quantitative elemental CPS response data and comprehensive metabolomics secondary metabolite data from 73 batches of Peucedanum praeruptorum, a carefully planned data utilization strategy was employed. This study explicitly uses non-targeted metallogroup data, metabolomics data, and combinations of both as the foundation for constructing the machine learning model. Given the large volume of secondary metabolite data, correlation removal was performed before dataset partitioning to reduce metabolite correlation features and improve computational efficiency. To ensure the model's generalization ability and reduce the risk of overfitting, a scientific data partitioning method was adopted, randomly and evenly dividing the data into a training set (70%) and a test set (30%). The training set was fully utilized for training and optimizing the model, while the test set served as an independent validation set for objectively evaluating key performance indicators such as accuracy and error. This optimization strategy not only preserved the original intent but also further enhanced the rigor and reliability of the research through clear data partitioning and scientific evaluation methods.
[0136] 1.5.2 Selection of Machine Learning Methods: Building upon the solid foundation of our research group's previous work, this study will fully utilize the powerful capabilities of R and Python languages, focusing on a binary classification algorithm framework and employing a data oversampling strategy to handle data imbalance. We then carefully selected and integrated six efficient and complementary machine learning algorithms based on the mlr3 framework in R: kNN (K-Nearest Neighbors), RF (Random Forest), rpart (Recursive Segmentation Tree), Stacking (Stacked Ensemble), SVM (Support Vector Machine), and XGBoost (Extreme Gradient Boosting). By employing an aggregation model to integrate the optimal models of metallomics and metabolomics, the core objective of this study is to optimize and construct an advanced model that accurately identifies key factors related to the authenticity and origin of Peucedanum praeruptorum. Through in-depth analysis and application of this model, this study aims to further explore and reveal the complex scientific laws behind the quality of Peucedanum praeruptorum, providing strong data support and scientific basis for the quality evaluation, origin traceability, and authenticity certification of Chinese medicinal materials. This optimization scheme, while retaining the original content, emphasizes the diversity and complementarity of algorithm selection, as well as the scientific and forward-looking nature of model construction.
[0137] 1.5.3 Classification Model Construction and Evaluation: In constructing and evaluating the classification model, this study employs a 5-fold cross-validation strategy, an effective model evaluation method that can more comprehensively test the model's generalization ability. This study selects accuracy (ACC) and area under the curve (AUC) as core evaluation metrics, both jointly measuring model performance. In particular, AUC is given primary importance during model parameter optimization because it can more comprehensively reflect the model's classification ability under different thresholds. To ensure the reliability and stability of model evaluation, this study further strengthens the experimental design by repeating the modeling process for 10 cycles. This measure not only helps reduce the impact of random errors on the results but also obtains more robust model performance evaluation results through multiple iterations. In each cycle, this study strictly follows the 5-fold cross-validation process to ensure full utilization of data and fairness in model evaluation.
[0138] 1.5.4 Variable Importance Analysis: In this study, the model output is interpreted by assigning importance values to features. The `fastshap` package is used to calculate variable importance, and the `shapviz` package is used to visualize variable importance. This study introduces SHAP into the interpretation of machine learning models to calculate the contribution of each feature to the model's prediction results. Through SHAP values, the impact of features on the model's prediction results can be visually observed, and the prediction results for individual data points can be explained, helping to understand why the model makes a certain prediction.
[0139] 1.6 Statistical Methods: This study efficiently integrates multiple professional tools such as Origin 2024 (OriginLab), Prism 8 (GraphPad Software), and Excel 2019 (Microsoft) to optimize the entire process of data processing, statistical analysis, and graphical display.
[0140] 2. Results
[0141] 2.1 General Characteristics of Elemental Fingerprint Profiles in Peucedanum praeruptorum: This study used a semi-quantitative method to monitor 66 macro and micro elements in Peucedanum praeruptorum. All micro elements basically covered both metallic and non-metallic elements. PCA, PLSDA, and cluster heatmap analyses were performed on all elements using R packages, covering geographical origin factors and focusing on distinguishing the characteristics of traditional and non-traditional producing areas. The results are as follows: Figure 8 As shown, some samples from the traditional producing area and the non-traditional producing area showed separation, but the feature differentiation effect did not reach the ideal state. Figure 8 A). Further PLSDA analysis revealed a certain separation trend between authentic and non-authentic local products. Figure 8 B). Regarding locality, cluster heatmap analysis results showed that no significant characteristic clustering effect was observed ( Figure 8 C). Preliminary results of elemental fingerprint profile studies indicate that traditional chemometric methods are insufficient to identify overall factor characteristics, thus hindering the revelation of local characteristics. Therefore, it is necessary to explore more suitable data identification methods.
[0142] 2.3 Profile Analysis of Small Molecule Metabolites: In this study, high-resolution mass spectrometry was used to collect the m / z and response values of metabolites from multiple batches of Peucedanum praeruptorum medicinal materials from Anhui, Zhejiang, Chongqing, Guizhou and other producing areas. After preliminary standardization and normalization, PCA, OPLSDA and cluster analysis were used to further observe the profile and characteristics of secondary metabolites. Figure 9 Meanwhile, this study also incorporated regionality into the metabolite profile to observe whether changes in metabolites could reflect regionality and non-regionality. PCA analysis results showed ( Figure 9 A), there is also overlap between the distances between traditional and non-traditional producing areas. OPLSDA analysis results show ( Figure 9 B), the trend between authentic and non-authentic samples is relatively obvious, but some sample areas still overlap. Cluster analysis results show ( Figure 9 C) Clustering results from authentic producing areas and non-authentic producing areas showed overlap. In summary, this study provides a preliminary understanding of the variation profile of secondary metabolites in Peucedanum praeruptorum samples. While authenticity factors were observed, relatively few characteristic factors were identified, and further research is needed to explore more effective methods to identify factor specificity.
[0143] 2.4 Machine learning for integrating metallomics and metabolomics in the classification of Qianhu geography
[0144] Preliminary chemometric analysis results from non-targeted omics indicate that exploring more suitable feature classification methods is essential. Based on these preliminary chemometric analysis results, to more accurately discover and reveal the quality connotation information of *Peucedanum praeruptorum* under multiple factors, this study uses machine learning from artificial intelligence. Using *Peucedanum praeruptorum* elemental fingerprints, collected metabolite characteristics, and their combination as a basic database, the study develops and optimizes algorithms to achieve classification modeling under different factors, thus revealing the key characteristics of *Peucedanum praeruptorum* more accurately and clearly. Previous research revealed that conventional experimental data on traditional Chinese medicine, due to their limited data volume, makes it difficult to obtain suitable models using appropriate algorithms. When dealing with imbalanced data, classification algorithm training often leads to poor predictive performance. The model significantly favors the majority class, ignoring minority class samples that are crucial for many practical applications. This makes the model impractical when facing real-world problems containing rare but high-priority events. To address this issue, oversampling provides a method to rebalance classes before model training. By replicating minority class data points, oversampling balances the training data, thus preventing the algorithm from ignoring important but few-numbered classes. Although this method carries the risk of overfitting, it effectively offsets the adverse effects of imbalanced learning, enabling machine learning models to handle critical use cases. In this study, we innovatively employ a data oversampling method to effectively address the imbalanced sample size. Measurements are standardized and normalized to serve as the machine learning dataset, which is then randomly divided into training and test sets with a 7:3 ratio. Five machine learning algorithms—RF, KNN, RPA, SVM, and XGB—are used for model building and optimization, while AUC and ACC are used to evaluate model performance. The samples cover the four major production areas of Anhui, Zhejiang, Guizhou, and Chongqing. Furthermore, the study includes the identification of authentic producing areas, primarily Anhui and Zhejiang. In this research, data from these two regions are merged to form the initial authentic producing areas.
[0145] 2.4.1 Machine Learning Classification Modeling Based on Metalomics: This study follows an established machine learning modeling framework and explores the application potential of the response values of 66 elements obtained from semi-quantitative analysis after internal standard correction in the field of non-targeted metalomics. By constructing a preliminary model and evaluating its performance, this study obtained detailed results as shown in Table 1. In the model evaluation stage, this study pays particular attention to the AUC value of the test set, while also considering the ACC index for comprehensive evaluation. The data in Table 1 reveals significant differences among the various regional origin models. Specifically, for traditional high-quality producing areas, the AUC of the model on the test set ranges from 0.77 to 0.96, showing good overall performance. Further analysis of the ROC curve results of all algorithms is shown below. Figure 10 As shown, the closer the ROC curve is to the upper left corner, the better the classifier's performance, meaning a high true positive rate while maintaining a low false positive rate. The characteristics of the ROC curve suggest it can help select the optimal classification threshold. Combined with the ACC results of the test set, this study found that for authentic producing areas, the XGB algorithm, with its powerful ensemble learning ability and robustness, is the suitable algorithm for building the optimal model. In summary, this study explored the construction and performance evaluation of different authentic metalomics models through a combination of semi-quantitative analysis and machine learning modeling. The significant differences between models from different producing areas not only reveal the influence of regional elemental content characteristics on model construction but also provide a strong basis for subsequent customized model development for different producing areas.
[0146] Table 1 Evaluation results of the Qianhu locality classification model based on elemental fingerprinting
[0147]
[0148] 2.4.2 Metabolomics-Based Machine Learning Classification Modeling: This study used high-resolution mass spectrometry to obtain mass spectrometry data of secondary metabolites from different producing areas of *Angelica dahurica*. Based on the distribution profile characteristics of these secondary metabolites, five machine learning algorithms were used to establish a classification model for *Angelica dahurica* from different producing areas and based on its geographical origin. Considering the complexity of secondary metabolites and the sensitivity of mass spectrometry detection, this study first cleaned the detected molecular ion peaks. Using compound information with response values above 1×10⁵ as the base dataset, molecular ion peaks close to the mass spectrometry baseline and with high accuracy were removed. Furthermore, to further optimize the base dataset and improve its quality, this study specifically performed correlation removal processing on complex metabolite information. The core purpose of this step was to reduce redundant and highly correlated features through careful selection, thereby improving the performance, interpretability, and computational efficiency of the machine learning model, reducing the risk of overfitting, and simultaneously reducing noise and storage requirements. The analysis results before and after correlation removal are shown below. Figure 11As shown in the figure. By comparing the correlation characteristics before and after removal, this study found that the correlation between secondary metabolites was lower after removal, which significantly improved the quality of the data.
[0149] This study follows an established machine learning modeling framework, using cleaned secondary metabolomics data as input. Through meticulous dataset partitioning, algorithm selection, and model parameter tuning, diverse algorithmic models for regional origin classification were constructed. The preliminary model construction results and evaluation details based on non-targeted metabolomics are shown in Table 2. In model evaluation, this study primarily relied on the AUC value of the test set, supplemented by ROC curve characteristic analysis and accuracy (ACC) for comprehensive consideration. Table 2 data shows that the AUC value of the regional origin model remained stable in a high range of 0.90 to 0.96. Further in-depth analysis of the ROC curves of the models constructed by all algorithms (e.g.) Figure 12 As shown in the figure, and by combining the accuracy (ACC) of the test set for horizontal comparison, this study found that for traditional producing areas, the optimal model is the KNN algorithm.
[0150] Table 2 Evaluation results of the Peucedanum praeruptorum locality classification model based on secondary metabolites
[0151]
[0152] 2.4.3 Establishment of a Machine Learning Aggregation Model Based on Metallomics and Metabolomics: Based on the classification model results from both metallomics and metabolomics, this study found that although data oversampling significantly improved model accuracy, the AUC and ACC of the metallomics-based model were low for some producing areas, such as Guizhou. Furthermore, the stability of the various metabolomics-based algorithms was poor. To further improve the accuracy and stability of predictions, reduce the risk of misjudgment, and increase the model's applicability to different scenarios, this study attempted to adopt an aggregation model approach, utilizing the differences between different models to improve the overall accuracy and stability of the prediction system. Based on the high-accuracy models obtained from five algorithms, metallomics and metabolomics were aggregated in different authentic producing areas. This further improved the overall model's predictive effect and obtained more reliable prediction results. More notably, this strategy combined inorganic and organic substances in *Peucedanum praeruptorum* for the first time, providing insight into the authenticity of *Peucedanum praeruptorum* at the overall material level. The establishment, optimization, and evaluation results of the aggregation model are shown in Table 3. Comparing the classification model results of metallomics and metabolomics alone (Tables 2 and 3), this study found that the aggregation model significantly improved the AUC and ACC of authenticity. This indicates that the aggregation model improves accuracy, enhances predictive robustness, and thus improves the model's generalization ability. A similar trend was observed when comparing with other producing regions, combining the ROC curve characteristics of different producing regions (…). Figure 13The ROC values for different producing areas are all close to the left, and the overfitting phenomenon between models is small, further verifying the effectiveness and stability of the aggregation model. The aggregation model results in this study show that the strategy of combining inorganic and organic methods fully utilizes the differences between different models. By integrating information from metallomics and metabolomics, effective aggregation of information from different producing areas can be achieved, thereby comprehensively improving the accuracy and stability of the overall prediction system and realizing the evaluation of the authenticity of Peucedanum praeruptorum.
[0153] Table 3 Evaluation results of the convergence model based on metallomics and metabolomics
[0154] 3. Discussion
[0155] 3.1 A Novel Strategy for Identifying the Origin and Authenticity of Traditional Chinese Medicine (TCM): The quality of TCM herbs is closely related to their origin. Different origins result in variations in growth environment, soil conditions, and climate, directly impacting the chemical composition and metal element content. Metallomics, a relatively new research direction, is used to study the types, content, and distribution patterns of metal elements in TCM herbs. Metal elements play a crucial role in TCM, serving not only as fundamental chemical components but also participating in pharmacological effects. Metallomics analysis can reveal the characteristic spectra of metal elements in TCM, providing vital information for origin identification. Plant TCM metabolomics studies the types, content, and variation patterns of small-molecule metabolites in TCM herbs. Metabolites are chemical substances produced during the growth, development, and metabolism of TCM herbs, reflecting their physiological state and metabolic pathways. Metabolomics analysis can obtain metabolic fingerprints of TCM herbs, further revealing their origin characteristics. A novel strategy for identifying the origin of traditional Chinese medicine (TCM) using a machine learning-integrated metallomics and metabolomics model is an innovative and scientific approach. It combines the advantages of metallomics, metabolomics, and artificial intelligence, encompassing both inorganic and organic compounds to reveal the characteristics of TCM origins more comprehensively, deeply, and intelligently. Therefore, through integrated metallomics and metabolomics analysis, the origin of TCM can be identified more accurately, providing a scientific basis for the quality control and efficacy evaluation of TCM materials. This innovative strategy has broad application prospects in TCM quality control, efficacy evaluation, and the protection of TCM resources. This study perfectly applies this strategy to the medicinal herb *Peucedanum praeruptorum*, integrating information on 66 trace elements and abundant secondary metabolites. Utilizing the intelligent algorithm of the machine learning aggregation model, it achieves the traceability and authenticity identification of *Peucedanum praeruptorum*. Therefore, this strategy is comprehensive, accurate, and scientific in the traceability and authenticity identification of *Peucedanum praeruptorum*, reflecting the quality connotation and scientific principles of *Peucedanum praeruptorum*.
[0156] 3.2 Importance of Inorganic Elements and Chemical Composition in the Identification of the Origin of *Phyllostachys edulis*: In the fields of machine learning and deep learning, model interpretability plays a crucial role. Although complex models such as deep neural networks and ensemble models (e.g., XGBoost, LightGBM) perform well in predictive accuracy, they are often like "black boxes," with their internal decision-making logic difficult to understand. To address this challenge, SHAP (Shapley Additive Explanations) has emerged as an effective tool to reveal and interpret the model's output by assigning importance values to each feature. SHAP is used to calculate the contribution of each feature to the model's prediction results. This study further employs Shapley additive interpretation, calculating all possible permutations among the variables and then calculating the average contribution of each variable. In this study, the fastshap package is used to calculate variable importance, and the shapviz package is used for visualizing variable importance. In the interpretation results of the original producing area prediction model ( Figure 14 This study found that Ca and peak 17307 had the greatest impact on the model, and the constant elements Mg and Mn were also among the features with significant influence. Figure 14 B and Figure 14 D further reveals visualizations of the SHAP values of 15 high-impact features.
[0157] Furthermore, in this study, the `fastshap` package was used to calculate variable importance, and the `shapviz` package was used to visualize variable importance. The calculated results for the Anhui production area show that Mn has the greatest impact on the model's predictions in the elemental model. Among the chemical components of the Anhui production area, peak 16696 has the greatest impact on the model's predictions. In the interpretation of the Guizhou production area model, S and Mg are the two most influential constant elements, while the characteristic importance of peak 13779 is much greater than other peaks. In the Zhejiang production area, B, Mg, and peak 19774 have a significant impact on the model's predictions. In the Chongqing production area, Mg and peak 17630 are the variable features with the greatest impact on the model. A horizontal comparison of the Guizhou, Anhui, Zhejiang, and Chongqing production areas reveals distinct characteristics in the elemental and chemical composition summary figures, which effectively reflects the characteristic variability among the samples from different production areas.
[0158] 3.3 Advantages of oversampling technology for classifying the origin and authenticity of Angelica dahurica
[0159] The training process of classification algorithms faces significant challenges when dealing with datasets containing imbalanced classes. This applies not only to general classification problems but also to the field of traditional Chinese medicine (TCM) testing. Because the majority class samples (such as common TCM ingredients or types) far outnumber the minority class samples (such as rare or specific TCM ingredients, varieties, or quality grades), models are prone to severe bias towards the majority class during training, neglecting those minority class samples that, while few in number, are crucial for practical applications such as TCM quality control and efficacy evaluation. Take Peucedanum praeruptorum (Qianhu) as an example. As a TCM herb, its efficacy can be affected by factors such as quality, component content, and origin. However, in actual TCM testing, the large number of majority class samples (such as common qualified Peucedanum praeruptorum samples) and the scarcity of minority class samples (such as Peucedanum praeruptorum samples from specific origins or with specific component contents) can lead to poor performance of the trained model in identifying these minority class samples. To effectively address this challenge, oversampling techniques can also be applied to the field of TCM testing. By replicating or amplifying minority class data points (such as samples of Peucedanum praeruptorum from specific origins or containing specific components), the number of samples in each class in the training data tends to be balanced. This prevents the algorithm from easily ignoring important but scarce classes during training, allowing it to learn the features and information in the data more comprehensively and improving the model's ability to identify minority class samples. However, oversampling also carries the risk of overfitting. In the field of traditional Chinese medicine (TCM) testing, this may lead to the model overfitting to minority class samples in the training data, resulting in poor performance in practical applications. To mitigate this risk, reasonable parameter settings and cross-validation can be employed. For example, model performance can be optimized by adjusting the oversampling ratio and selecting appropriate machine learning algorithms and parameters. In summary, applying oversampling techniques to TCM testing can significantly offset the negative impact of imbalanced learning, improve the machine learning model's ability to identify minority class samples, and provide more accurate and reliable solutions for practical applications such as TCM quality control and efficacy evaluation. Simultaneously, this also provides new ideas and methods for data processing and the application of machine learning algorithms in the field of TCM testing.
[0160] 3.4 Advantages of the Aggregate Model in Predicting the Origin and Authenticity of Angelica dahurica: The aggregate model has significant advantages in predicting the origin and authenticity of Angelica dahurica. These advantages are mainly reflected in the following aspects: (1) Improved prediction accuracy. By integrating the results of multiple prediction models, the aggregate model can combine the advantages of different models, reduce the prediction error of a single model, and thus improve the accuracy of prediction. In the prediction of the origin and authenticity of Angelica dahurica, different models may make predictions based on different features (such as climate, soil, ecological environment, etc.). The aggregate model can comprehensively consider these features and give more accurate prediction results. (2) Enhanced model generalization ability. By combining multiple models, the aggregate model can reduce the risk of overfitting a single model to specific data and improve the generalization ability of the model. This means that when facing new and unknown Angelica dahurica origin data, the aggregate model can give more reliable prediction results. This is of great significance for the planting, procurement and quality control of Chinese medicinal materials. (3) Full utilization of multiple data sources and information. The origin and authenticity of Chinese medicinal materials are affected by a variety of factors, including climate, soil, ecological environment, cultivation techniques, etc. The aggregation model can integrate information from different data sources, such as meteorological data, soil data, remote sensing images, and expert judgment based on traditional knowledge and experience, so as to more comprehensively assess the origin and authenticity of Angelica dahurica. (4) Improve prediction efficiency. The aggregation model can significantly improve prediction efficiency by processing multiple prediction models in parallel. This is of great significance for the rapid response and decision-making of the Chinese medicinal materials industry. For example, in the process of planting Chinese medicinal materials, timely and accurate prediction of the origin and authenticity of Angelica dahurica helps guide growers to select suitable planting areas and cultivation techniques, thereby improving the yield and quality of Chinese medicinal materials. (5) Adaptability and flexibility. The aggregation model has high adaptability and flexibility. It can select different prediction models to combine according to actual needs and data characteristics, and adjust the combination method and weight. This enables the aggregation model to adapt to different application scenarios and data distributions, and provide more personalized solutions. For example, when predicting the origin and authenticity of Angelica dahurica for a specific region or a specific variety, the aggregation model can select more suitable prediction models to combine according to the climate, soil and other characteristics of the region.
[0161] In conclusion, the aggregation model has many advantages in predicting the origin and authenticity of Peucedanum praeruptorum, making it an indispensable tool in the Chinese medicinal materials industry. By applying the aggregation model, this study can more accurately and efficiently predict the origin and authenticity of Peucedanum praeruptorum, providing strong support for the cultivation, procurement, and quality control of Chinese medicinal materials.
[0162] 3.5 The Correlation Between the Origin and Quality of Peucedanum praeruptorum: Based on the in-depth analysis of 528 batches of Peucedanum praeruptorum, this study not only confirmed the significant regional differences in the content and pass rate of A and B components of Peucedanum praeruptorum, but also revealed the close relationship between the quality of Peucedanum praeruptorum and the characteristics of specific production areas. To further explore the scientific mechanism behind this complex relationship, this study innovatively integrated a multi-component perspective, elevating the analysis of inorganic and organic matter to a multi-level material level, and cleverly combined metallomics and metabolomics technologies, opening a new chapter in the study of the origin and quality authenticity of Peucedanum praeruptorum.
[0163] By precisely integrating chemometrics and machine learning methods, this study successfully constructed a dual evaluation system based on elemental fingerprints and secondary metabolite profiles, comprehensively elucidating how these two dimensions synergistically characterize the quality attributes of Peucedanum praeruptorum. At the elemental fingerprint analysis level, this study found that the abundance of macroelements such as K, Na, Ca, and Al is not only regulated by the soil environment but also profoundly reflects the unique enrichment mechanism of Peucedanum praeruptorum. This discovery highlights the crucial role of metallomics in revealing the characteristics of traditional Chinese medicine.
[0164] Meanwhile, the application of non-targeted metabolomics enabled this study to deeply analyze the diverse profiles of small organic molecules in Peucedanum praeruptorum, particularly confirming the dominant role of coumarins. The abundance of coumarins in Peucedanum praeruptorum is a key characteristic of this plant belonging to the Apiaceae family. Faced with these structurally complex secondary metabolites that are difficult to capture using traditional analysis, this study innovatively employed machine learning techniques to effectively mine and extract key features, providing strong technical support for the quality identification and authenticity evaluation of Peucedanum praeruptorum at the small organic molecule level.
[0165] In summary, this study not only preliminarily elucidates the intrinsic connection between the authentic quality of Peucedanum praeruptorum, but also successfully verifies the innovative strategy of integrating metallomics and metabolomics in identifying the authenticity of traditional Chinese medicine. This contributes significantly to the improvement and scientific development of the quality evaluation system for traditional Chinese medicine. In the future, this strategy is expected to be widely applied to more varieties of traditional Chinese medicine, promoting the development of the traditional Chinese medicine industry towards a more precise, efficient, and sustainable direction.
[0166] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0167] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0168] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0169] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0170] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0171] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0172] The exemplary embodiments of this disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of this disclosure, and such modifications should fall within the scope of this disclosure.
Claims
1. A method for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology, characterized in that, The method includes:
101. Obtain metallomics and metabolomics data of the medicinal materials to be tested; 102. Calculate the CPS response values of key feature elements in the metallomics data; 103. Process the metabolomics data to obtain mass spectrometry data containing peak information of key chemical components in the data; 104. Input the CPS response values of the key characteristic elements and the peak information of key chemical components into the authentic producing area identification model to obtain a classification result of whether it belongs to an authentic producing area; the authentic producing area includes the first authentic producing area and the second authentic producing area; if the classification result of belonging to an authentic producing area is obtained in 104, input the metallomics data and metabolomics data of the medicinal material to be tested into the first authentic producing area identification model for processing, and output the classification result of whether it belongs to the first authentic producing area; if the output is yes, the operation ends, and the result of the medicinal material to be tested belonging to the first authentic producing area is obtained; if the output is no, the result of the medicinal material to be tested not belonging to the first authentic producing area is obtained, and input the metallomics data and metabolomics data of the medicinal material to be tested into the second authentic producing area identification model for processing, and output the classification result of whether it belongs to the second authentic producing area; if the output is yes, the operation ends, and the result of the medicinal material to be tested belonging to the second authentic producing area is obtained; if the output is no, the result of the medicinal material to be tested not belonging to the second authentic producing area is obtained; Non-authentic producing areas include a first non-authentic producing area and a second non-authentic producing area. In step 104, if a classification result indicating the medicinal material does not belong to the authentic producing area is obtained, the metallomics and metabolomics data of the medicinal material to be tested are input into the first non-authentic producing area identification model for processing, and the classification result indicating whether it belongs to the first non-authentic producing area is output. If the output is yes, the process ends, and the result indicating the medicinal material to be tested belongs to the first non-authentic producing area is obtained. If the output is no, the result indicating the medicinal material to be tested does not belong to the first non-authentic producing area is obtained. Then, the metallomics and metabolomics data of the medicinal material to be tested are input into the second non-authentic producing area identification model for processing, and the classification result indicating whether it belongs to the second non-authentic producing area is output. If the output is yes, the process ends, and the result indicating the medicinal material to be tested belongs to the second non-authentic producing area is obtained. If the output is no, the result indicating the medicinal material to be tested does not belong to the second non-authentic producing area is obtained.
2. The method for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology according to claim 1, characterized in that, When the medicinal material to be tested is Peucedanum praeruptorum, the non-authentic producing areas include Guizhou and Chongqing. Guizhou or Chongqing is the first non-authentic producing area, and Chongqing or Guizhou is the second non-authentic producing area.
3. The method for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology according to claim 1, characterized in that, The metallomics data includes major elements and trace elements; trace elements include metallic elements and non-metallic elements.
4. The method for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology according to claim 1, characterized in that, The method for constructing the traditional product area identification model includes: Obtain metallomics and metabolomics data of Peucedanum praeruptorum medicinal materials from both authentic and non-authentic regions, with authentic or non-authentic regions as the classification label; The metalomics data is input into the first machine learning model to obtain the predicted classification result, which is compared with the classification label of native or non-native place. The model is optimized based on the comparison result to obtain the metalomics model. The metabolomics data is input into a second machine learning model to obtain a predicted classification result, which is then compared with the classification labels of native or non-native origin. The model is optimized based on the comparison results to obtain a metabolomics model. The optimal model among the metallomics and metabolomics models obtained using different algorithms is selected, and the two optimal models are combined to obtain the original product area identification model.
5. The method for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology according to claim 1, characterized in that, The methods for constructing the first or second property area identification model include: We obtained metallomics and metabolomics data of Peucedanum praeruptorum from the first and second producing areas as classification labels for authentic products. The metallomics data is input into a first machine learning model to obtain a predicted classification result, which is then compared with the local classification label. The model is optimized based on the comparison result to obtain a metallomics model. The metabolomics data is input into a second machine learning model to obtain a predicted classification result, which is then compared with the local classification label. The model is optimized based on the comparison result to obtain a metabolomics model. The optimal models of metallomics and metabolomics in the first producing area were selected, and the two optimal models were combined to obtain the first producing area identification model; the optimal models of metallomics and metabolomics in the second producing area were selected, and the two optimal models were combined to obtain the second producing area identification model.
6. The method for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology according to claim 1, characterized in that, The construction methods for the first non-traditional agricultural product area identification model or the second non-traditional agricultural product area identification model include: Metallomics and metabolomics data of Peucedanum praeruptorum from the first and second non-authentic producing areas were obtained from the training set. The first and second non-authentic producing areas are used as non-authentic classification labels. The metallomics data is input into a first machine learning model to obtain a predicted classification result, which is then compared with a non-local classification label. The model is optimized based on the comparison result to obtain a metallomics model. The metabolomics data is input into a second machine learning model to obtain a predicted classification result, which is then compared with a non-local classification label. The model is optimized based on the comparison result to obtain a metabolomics model. The optimal models of metallomics and metabolomics in the first non-authentic producing area were selected, and the two optimal models were combined to obtain the identification model of the first non-authentic producing area; the optimal models of metallomics and metabolomics in the second non-authentic producing area were selected, and the two optimal models were combined to obtain the identification model of the second non-authentic producing area.
7. The method for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology according to claim 4, characterized in that, The methods used by the first machine learning model and / or the second machine learning model include: RF, KNN, Rpa, SVM, and XGB.
8. The method for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology according to claim 1, characterized in that, When the medicinal material to be tested is Peucedanum praeruptorum, the production areas include Zhejiang and Anhui. Zhejiang or Anhui is the first production area, and Anhui or Zhejiang is the second production area.
9. The method for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology according to claim 7, characterized in that, The models of authentic producing areas in the metallomics model using the XGB algorithm and the models of authentic producing areas in the metabolomics model using the KNN algorithm were selected and aggregated to obtain the aggregated model of authentic producing areas.
10. The method for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology according to claim 1, characterized in that, The key characteristic element includes Ca.
11. The method for identifying the geographical origin of traditional Chinese medicine based on artificial intelligence and multi-omics technology according to claim 10, characterized in that, The key characteristic elements also include: Mn and Mg.
12. A computer device, characterized in that, The device includes: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1-11.
13. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1-11.
14. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1-11.
Citation Information
Patent Citations
Traditional Chinese medicine genuine nature discrimination method and device based on data enhancement
CN115423042A
Chinese Hongkong oyster producing area tracing method based on metabolites and mineral elements
CN117686577A