Method and system for identifying medicinal snake bile based on metabolomics and machine learning
By combining liquid chromatography-mass spectrometry with machine learning, differential metabolites in snake bile are screened and a classification model is constructed. This solves the problem of strong subjectivity in traditional identification methods and achieves efficient and reliable identification between medicinal and non-medicinal snake bile, which is suitable for the quality control of traditional Chinese medicine preparations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI INST OF FOOD & DRUG INSPECTION (ANHUI NAT AGRI & SIDELINE PROCESSED FOOD QUALITY SUPERVISION & INSPECTION CENT)
- Filing Date
- 2026-03-26
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies are insufficient to effectively distinguish between medicinal snake bile and non-medicinal snake bile. Traditional identification methods are highly subjective and difficult to quantify. Furthermore, adulteration with python bile seriously affects the quality and stability of traditional Chinese medicine preparations.
The metabolite information of snake bile was obtained by liquid chromatography-mass spectrometry. Combined with multivariate statistical analysis and machine learning methods, differential metabolites were screened, a classification model was constructed, and the automatic identification of medicinal snake bile and non-medicinal snake bile was realized.
It significantly improves the accuracy and reliability of snake bile identification, provides a scientific identification method, ensures the quality of traditional Chinese medicine, is applicable to the identification of closely related snake species, and realizes a complete automated process from sample detection to result output.
Smart Images

Figure CN122282991A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of traditional Chinese medicine identification, and more specifically, to a method and system for identifying medicinal snake bile based on metabolomics and machine learning. Background Technology
[0002] Snake bile is a valuable and commonly used traditional Chinese medicine, first recorded in the Liang Dynasty's *Mingyi Bielu*, and subsequently documented in various herbal classics such as *Zhenglei Bencao* and *Bencao Gangmu*. Snake bile has immense medicinal value and has been widely used since ancient times, remaining prevalent in modern clinical practice. Snake bile comes from many sources, with varying distributions and usage methods across the country. Furthermore, there is currently no national quality standard; the only existing description in the 2025 edition of the *Chinese Pharmacopoeia* is that "snake bile is the bile of various snakes belonging to the Elapidae, Colubridae, or Viperidae families." Currently, the state has approved numerous traditional Chinese medicine preparations using snake bile as the main ingredient, including over 30 varieties such as Snake Bile and Fritillaria Liquid, Bezoar Snake Bile and Fritillaria Liquid, Snake Bile and Tangerine Peel Oral Liquid, Three Snake Bile and Fritillaria Syrup, and Snake Bile and Fritillaria Loquat Paste, with as many as 356 approval numbers. However, due to high market demand and resource scarcity, adulteration frequently occurs. Currently, python bile is similar to snake bile in morphology and some chemical components. Furthermore, because pythons are easy to farm, python bile is often adulterated with animal bile in the market, making it impossible to guarantee the authenticity of snake bile medicinal materials. This seriously affects the quality stability and clinical efficacy of related traditional Chinese medicines. Traditional methods for identifying snake bile mainly rely on morphological identification and experience, which are highly subjective and difficult to quantify and standardize.
[0003] High-resolution liquid chromatography-mass spectrometry (LC-HRMS) is widely used in the analysis of components in traditional Chinese medicine. This technique can capture information on thousands of metabolites in a sample without bias. However, the resulting high-dimensional, massive amounts of data are difficult to interpret directly. Quantitative analysis of only a few indicator components cannot comprehensively reflect the overall quality characteristics of snake bile, nor can it effectively distinguish bile from closely related snake species. Therefore, there is an urgent need for a data mining method that can extract essential features from complex mass spectrometry data and achieve rapid, objective classification. Summary of the Invention
[0004] To solve the above-mentioned technical problems, the present invention provides the following technical solution: A method for identifying medicinal snake bile based on metabolomics and machine learning, comprising: S1. Obtain metabolic extracts of the samples, including medicinal snake bile samples and non-medicinal snake bile samples, and use liquid chromatography-mass spectrometry to detect them and obtain mass spectrometry data of all samples. S2, preprocess the mass spectrometry data, extract the intensity information of metabolite characteristic peaks, and construct the metabolite characteristic matrix of all samples; S3. Based on the metabolite feature matrix, a multivariate statistical analysis method is used to screen and distinguish the differential metabolites between medicinal snake bile and non-medicinal snake bile. S4. Based on the characteristic peak intensity data of the differential metabolites, construct and train a machine learning classification model; S5. After the mass spectrometry data of the snake bile sample to be tested undergoes the same preprocessing as in S2, it is input into the machine learning classification model. Based on the output of the machine learning classification model, it is determined whether the source is medicinal snake bile or non-medicinal snake bile.
[0005] This invention combines liquid chromatography-mass spectrometry metabolomics with machine learning methods to systematically reveal the metabolite differences between medicinal and non-medicinal snake bile. Furthermore, it utilizes an artificial intelligence diagnostic model for classification, significantly improving the accuracy and reliability of snake bile identification and providing an efficient data-driven solution for medicinal material traceability and quality control. This method overcomes the limitations of traditional identification methods, which are highly subjective and difficult to quantify, achieving a shift from experience-based judgment to objective data-driven identification.
[0006] Preferably, the liquid chromatography-mass spectrometry (LC-MS) technique is ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry (UHPLC-QQS-MS), and its chromatographic conditions include: A C18 column was used, with mobile phase A being methanol solution and mobile phase B being 10 mmol·L⁻¹. -1 Ammonium acetate solution; gradient elution program: 0-5 min, 30%-50% A; 5-20 min, 50%-90% A; 20-25 min, 90% A; 25-26 min, 90%-30% A; 26-30 min, 30% A.
[0007] This invention utilizes ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry (UHPLC-quadrupole-time-of-flight), which boasts advantages such as high separation efficiency, high resolution, and high mass accuracy, to achieve efficient separation and precise detection of complex metabolites in snake bile.
[0008] Preferably, the preprocessing of the mass spectrometry data in S2 includes: Peak detection, retention time alignment, and metabolite identification were performed on the mass spectrometry data. Common metabolite characteristic peaks were extracted from all samples to construct the original data matrix. The original data matrix is standardized to obtain a standardized data matrix.
[0009] This invention provides a systematic preprocessing method for mass spectrometry data. Peak detection identifies metabolite signals in each sample, retention time alignment corrects for chromatographic drift between samples, and normalization eliminates systematic errors caused by instrument response or sample concentration differences. The preprocessing yields a metabolite feature matrix that accurately reflects the biological differences between samples. This raw data matrix is further standardized to eliminate differences in the dimensions and absolute intensities of different metabolite features.
[0010] As a preferred option, the multivariate statistical analysis methods in S3 include: Principal component analysis (PCA) unsupervised model was used to identify the separation trend and clustering of principal components in the samples, and orthogonal partial least squares discriminant analysis model was used to screen differential variables. Metabolites with statistical significance were screened based on the criteria of variable weight value VIP>1 and statistical test value P<0.05.
[0011] This invention employs an unsupervised principal component analysis model to reduce the dimensionality and visualize samples without any prior classification information. It visually demonstrates the natural clustering and separation trends of medicinal and non-medicinal snake bile samples, allowing for an overall assessment of significant differences between the two. Orthogonal partial least squares discriminant analysis (PLS) is used as a supervised discriminant analysis method to extract classification-related variables, effectively identifying the differential variables that contribute most to distinguishing the two classes of samples. Differential metabolites with statistical significance and biological relevance are selected based on the criteria of VIP>1 and P<0.05.
[0012] Preferably, the differential metabolites include bile acid metabolites, specifically taurocholic acid, taurodeoxycholic acid, and tauroursodeoxycholic acid, which have different content distribution characteristics in medicinal snake bile and non-medicinal snake bile.
[0013] This invention clarifies the key chemical basis for distinguishing between medicinal snake bile and non-medicinal snake bile. The main difference lies in bile acid metabolites. Taurocholic acid (TCA), taurodeoxycholic acid (TDCA), and tauroursodeoxycholic acid (TUDCA) are relatively enriched in medicinal snake bile, while the remaining bile acid components are present in higher amounts in non-medicinal snake bile.
[0014] As a preferred option, S4 specifically includes: Construct at least one machine learning classification model, including logistic regression, neural network, random forest, support vector machine, adaptive boosting, and gradient boosting; The optimal machine learning classification model is selected by dividing the samples into training and test sets, evaluating the performance of the machine learning classification model using cross-validation, and selecting the model based on at least one of the following evaluation metrics: accuracy, recall, F1 score, and confusion matrix.
[0015] This invention divides samples into independent training and test sets, effectively evaluating the model's generalization ability and avoiding overfitting. Cross-validation is used to fully utilize data under limited sample conditions, stably evaluating model performance. Multiple evaluation metrics, including accuracy, recall, F1 score, and confusion matrix, are employed to comprehensively measure the model's classification performance, ensuring that the selected optimal model achieves the best diagnostic results.
[0016] Preferably, the machine learning classification model is a neural network classifier, whose classification accuracy in both the training and test sets reaches a preset accuracy threshold.
[0017] Through this invention, a neural network classifier is obtained through training, evaluation, and selection, becoming the most suitable machine learning classification model for this method. The neural network possesses strong nonlinear fitting capabilities and can effectively learn the complex mapping relationship between differential metabolites and sample categories, thus performing optimally in the classification task of this invention. Using a preset accuracy threshold as an optimal condition ensures the stability and reliability of the selected model, meeting the accuracy requirements for identification in practical applications.
[0018] This invention also discloses a medicinal snake bile identification system based on metabolomics and machine learning, which includes: The mass spectrometry detection unit is used to detect metabolic extracts of samples using liquid chromatography-mass spectrometry (LC-MS) to obtain mass spectrometry data for all samples, including medicinal snake bile samples and non-medicinal snake bile samples. The data processing unit is used to preprocess the mass spectrometry data, extract metabolite characteristic peak intensity information, and construct a metabolite characteristic matrix for all samples. The differential analysis unit is used to screen and distinguish differential metabolites between medicinal snake bile and non-medicinal snake bile based on the metabolite feature matrix and using multivariate statistical analysis methods. The model building unit is used to build and train a machine learning classification model based on the characteristic peak intensity data of the differential metabolites. The discrimination output unit is used to input the mass spectrometry data of the snake bile sample to be tested, which has been preprocessed by the data processing unit, into the machine learning classification model, and to determine whether the source is medicinal snake bile or non-medicinal snake bile based on the output of the machine learning classification model.
[0019] This invention integrates functional units such as mass spectrometry detection, data processing, difference analysis, model building, and discriminant output, forming a complete automated process from sample detection to result output. The collaborative work of each unit enables efficient and accurate automated identification of medicinal snake bile, facilitating its widespread application in quality inspection institutions and manufacturing enterprises, and demonstrating a high level of integration and automation.
[0020] In summary, the present invention has the following beneficial effects: First, this invention obtains comprehensive metabolite information of snake bile through liquid chromatography-mass spectrometry, combines multivariate statistical analysis to screen differential metabolites, and uses machine learning models for automatic classification, avoiding the subjectivity of traditional experience-based identification and significantly improving the accuracy and reliability of identification. Secondly, this invention not only achieves an effective distinction between medicinal and non-medicinal snake bile, but also uses chemometric methods such as principal component analysis and orthogonal partial least squares discriminant analysis to screen out the key metabolites that cause the differences between the two types of bile, providing a material basis for the scientific validity of the identification method. Third, the panoramic analysis approach based on metabolomics in this invention can address the challenges of identifying snake bile with its complex composition and subtle differences between different snake species. It is particularly suitable for identifying the origin of snake species that are closely related and difficult to distinguish using traditional methods. Fourth, the identification system provided by this invention realizes a complete automated process from sample detection to result output, which is convenient for promotion and application in actual production and is of great significance for ensuring the quality of raw materials for snake bile-based traditional Chinese medicine and maintaining the market order of traditional Chinese medicine. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the overall process of the method for identifying medicinal snake bile based on metabolomics and machine learning in this embodiment.
[0022] Figure 2 This is the total ion chromatogram of medicinal snake bile and non-medicinal snake bile in negative ion mode in this embodiment.
[0023] Figure 3 This shows the distribution of medicinal snake bile and non-medicinal snake bile samples on the PC1-PC2 score map in this embodiment.
[0024] Figure 4 This is a schematic diagram and a substitution test diagram showing the differentiation of snake bile metabolic phenotypes using the OPLS-DA model in this embodiment.
[0025] Figure 5 This is the volcano plot analysis result of the differential metabolites in this embodiment.
[0026] Figure 6This is a schematic diagram illustrating the six machine learning methods used in this embodiment to distinguish between medicinal and non-medicinal snake bile. Detailed Implementation
[0027] To further understand the content of this invention, the invention will be described in detail with reference to the embodiments. It should be understood that the embodiments are merely illustrative and not limiting of the invention.
[0028] Example 1 like Figure 1 As shown, this embodiment provides a method for identifying medicinal snake bile based on metabolomics and machine learning, which includes: S1. Obtain metabolic extracts of the samples, including medicinal snake bile samples and non-medicinal snake bile samples, and use liquid chromatography-mass spectrometry to detect them and obtain mass spectrometry data of all samples. S2, preprocess the mass spectrometry data, extract the intensity information of metabolite characteristic peaks, and construct the metabolite characteristic matrix of all samples; S3. Based on the metabolite feature matrix, a multivariate statistical analysis method is used to screen and distinguish the differential metabolites between medicinal snake bile and non-medicinal snake bile. S4. Based on the characteristic peak intensity data of the differential metabolites, construct and train a machine learning classification model; S5. After the mass spectrometry data of the snake bile sample to be tested undergoes the same preprocessing as in S2, it is input into the machine learning classification model. Based on the output of the machine learning classification model, it is determined whether the source is medicinal snake bile or non-medicinal snake bile.
[0029] Specifically: S1. Obtain metabolic extracts of the samples, including medicinal snake bile samples and non-medicinal snake bile samples, and use liquid chromatography-mass spectrometry to detect them and obtain mass spectrometry data of all samples.
[0030] In this step, snake bile samples from different sources, verified by authoritative experts, are first collected. Medicinal snake bile samples come from various snake species in the families Elapidae, Viperidae, or Colubridae, while non-medicinal snake bile samples mainly come from pythons in the family Pythonidae. After sample collection, a standardized sample pretreatment is performed: an appropriate amount of snake bile sample is placed in a centrifuge tube, an extraction solvent (such as methanol-water solution) is added, the sample is shaken well and weighed, then subjected to ultrasonic extraction. After cooling, the sample is weighed again, and the weight loss is replenished with the extraction solvent. The sample is then shaken well, filtered through a microporous membrane, and the filtrate is collected to obtain the metabolic extract sample solution.
[0031] The prepared metabolic extract sample solutions were detected using liquid chromatography-mass spectrometry (LC-MS). During the detection process, metabolites in the sample were separated by a chromatographic column and then entered the mass spectrometer. The mass spectrometer acquired data in a specific scanning mode to obtain raw mass spectrometry data for all samples. The raw mass spectrometry data includes information such as the mass-to-charge ratio, retention time, and corresponding ionic strength of thousands of metabolites in each sample.
[0032] S2, preprocess the mass spectrometry data, extract the intensity information of metabolite characteristic peaks, and construct the metabolite characteristic matrix of all samples.
[0033] In this step, the acquired raw mass spectrometry data is imported into data processing software for preprocessing. The preprocessing process includes: peak detection to identify metabolite signal peaks in each sample; retention time alignment to correct retention time shifts between different samples caused by chromatographic drift; and normalization to eliminate systematic errors caused by instrument response fluctuations or differences in sample concentrations.
[0034] Through the above preprocessing, common metabolite characteristic peaks are extracted from all samples, and an original data matrix is constructed. The rows of the original data matrix represent samples, the columns represent metabolite characteristics, and the matrix elements are the intensity values of each characteristic peak. This original data matrix contains the original metabolite abundance information of the samples.
[0035] To further eliminate differences in dimensions and orders of magnitude among different metabolites, the original data matrix was standardized. The standardization process involved calculating the arithmetic mean and standard deviation of the intensity of all samples for each metabolite feature column, and then performing a standard normal transformation on each matrix element to obtain the standardized data matrix. In the standardized data matrix, the mean of each metabolite feature is 0, and the standard deviation is 1, making metabolite features of different dimensions and orders of magnitude comparable and facilitating subsequent multivariate statistical analysis.
[0036] S3. Based on the metabolite feature matrix, multivariate statistical analysis is used to screen and distinguish the differential metabolites between medicinal snake bile and non-medicinal snake bile.
[0037] In this step, principal component analysis (PCA), an unsupervised pattern recognition method, is first used to analyze the standardized data matrix. PCA performs a linear transformation on the original data, projecting high-dimensional data into a low-dimensional space, achieving dimensionality reduction and visualization while preserving the main information of the original data. In the principal component score plot, each sample point represents a sample, and the distribution of sample points reflects the similarity and differences between samples. By observing the distribution of medicinal snake bile samples and non-medicinal snake bile samples in the principal component space, it is possible to intuitively determine whether there is a significant natural separation trend and clustering characteristics between the two.
[0038] Based on the confirmed significant differences between the two types of samples, a supervised discriminant analysis method, orthogonal partial least squares discriminant analysis, was further employed. Orthogonal partial least squares discriminant analysis can extract variable information relevant to sample classification to the greatest extent possible, establishing a correlation model between metabolite characteristics and sample categories. Using this model, the variable projection importance value for each metabolite characteristic is calculated, reflecting the degree to which each metabolite contributes to distinguishing the two types of samples.
[0039] Metabolites with statistical significance were screened based on the criteria of variable weight value VIP>1 and statistical test value P<0.05. Metabolites that meet this dual criterion are both statistically significant and biologically relevant, and can serve as potential markers to distinguish between medicinal snake bile and non-medicinal snake bile.
[0040] S4. Based on the characteristic peak intensity data of the differential metabolites, construct and train a machine learning classification model.
[0041] In this step, firstly, based on the differentially expressed metabolites selected in step S3, their characteristic peak intensity data are extracted across all samples to construct the input feature matrix for the machine learning classification model. Simultaneously, each sample is assigned a category label (medicinal snake bile or non-medicinal snake bile).
[0042] The sample set is divided into a training set and a test set according to a preset ratio. The training set is used for model construction and parameter optimization, while the test set is used to evaluate the model's generalization ability and final performance. Simultaneously, cross-validation is used to further divide the training set, and training and validation are repeated multiple times during model training to fully utilize the limited sample data, stably evaluate model performance, and avoid overfitting.
[0043] Various types of machine learning classification models are constructed, including but not limited to logistic regression, neural networks, random forests, support vector machines, adaptive boosting, and gradient boosting. These classifiers have different algorithmic principles and characteristics: logistic regression is suitable for linearly separable problems; neural networks have strong nonlinear fitting capabilities; random forests improve generalization performance through ensemble learning; support vector machines handle nonlinear problems through kernel functions; and adaptive boosting and gradient boosting improve classification performance through iterative optimization.
[0044] During model training, at least one evaluation metric from precision, recall, F1 score, and confusion matrix is used to comprehensively evaluate the performance of each classification model. Precision reflects the overall correctness of classification, recall reflects the ability to identify the positive class, the F1 score is the harmonic mean of precision and recall, and the confusion matrix visually displays the classification status of each category. By comparing the performance of different models on the training set cross-validation and the test set, the machine learning classification model with the best performance is selected as the final diagnostic model.
[0045] S5. After the mass spectrometry data of the snake bile sample to be tested undergoes the same preprocessing as in S2, it is input into the machine learning classification model. Based on the output of the machine learning classification model, it is determined whether the source is medicinal snake bile or non-medicinal snake bile.
[0046] In this step, for the unknown snake bile sample to be tested, the same sample pretreatment method and liquid chromatography-mass spectrometry detection conditions as in step S1 are used to obtain its mass spectrometry data. Then, following the same pretreatment procedure as in step S2, peak detection, retention time alignment, and metabolite identification are performed on the raw mass spectrometry data, and characteristic peak intensity data corresponding to the differential metabolites screened in step S3 are extracted.
[0047] The processed sample data is input into the optimal machine learning classification model trained in step S4. The model automatically calculates and outputs the classification result according to the learned classification rules. Then, based on the output classification result, it is determined whether the sample comes from medicinal snake gallbladder or non-medicinal snake gallbladder.
[0048] Through the steps described above, this embodiment combines liquid chromatography-mass spectrometry metabolomics technology with machine learning methods to systematically reveal the metabolite differences between medicinal and non-medicinal snake bile. Furthermore, it utilizes an artificial intelligence diagnostic model for classification, significantly improving the accuracy and reliability of snake bile identification and providing an efficient data-driven solution for medicinal material traceability and quality control. This method overcomes the limitations of traditional identification methods, which are highly subjective and difficult to quantify, achieving a shift from experience-based judgment to objective data-driven identification.
[0049] In this embodiment, the liquid chromatography-mass spectrometry (LC-MS) technique is ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry (UHPLC-QPC-TOF-MS), and its chromatographic conditions include: A C18 column was used. Mobile phase A was methanol solution and mobile phase B was 10 mmol·L⁻¹ ammonium acetate solution. The gradient elution program was 0–5 min, 30%–50% A; 5–20 min, 50%–90% A; 20–25 min, 90% A; 25–26 min, 90%–30% A; 26–30 min, 30% A.
[0050] In this embodiment, ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry (UHPLC-quadrupole-time-of-flight) is used, which has the advantages of high separation efficiency, high resolution, and high quality accuracy, to achieve efficient separation and accurate detection of complex metabolites in snake bile.
[0051] In this embodiment, the preprocessing of mass spectrometry data in S2 includes: Peak detection, retention time alignment, and metabolite identification were performed on the mass spectrometry data. Common metabolite characteristic peaks were extracted from all samples to construct the original data matrix. The original data matrix is standardized to obtain a standardized data matrix.
[0052] This embodiment systematically preprocesses mass spectrometry data, identifies metabolite signals in each sample through peak detection, corrects chromatographic drift between different samples through retention time alignment, and eliminates systematic errors caused by instrument response or sample concentration differences through normalization. The preprocessing yields a metabolite feature matrix that accurately reflects the biological differences between samples. This raw data matrix is then further standardized to eliminate differences in the dimensions and absolute intensity of different metabolite features.
[0053] In this embodiment, the multivariate statistical analysis method in S3 includes: Principal component analysis (PCA) unsupervised model was used to identify the separation trend and clustering of principal components in the samples, and orthogonal partial least squares discriminant analysis model was used to screen differential variables. Metabolites with statistical significance were screened based on the criteria of variable weight value VIP>1 and statistical test value P<0.05.
[0054] This embodiment employs an unsupervised principal component analysis model to reduce the dimensionality and visualize the samples without any prior classification information. It visually demonstrates the natural clustering and separation trends of medicinal and non-medicinal snake bile samples, allowing for an overall assessment of whether significant differences exist between the two. Orthogonal partial least squares discriminant analysis (OLS) is used as a supervised discriminant analysis method to extract classification-related variables, effectively identifying the differential variables that contribute most to distinguishing the two classes of samples. Based on the criteria of VIP>1 and P<0.05, differentially expressed metabolites with statistical significance and biological relevance are selected.
[0055] In this embodiment, the differential metabolites include bile acid metabolites, specifically taurocholic acid, taurodeoxycholic acid, and tauroursodeoxycholic acid, which have different content distribution characteristics in medicinal snake bile and non-medicinal snake bile.
[0056] This embodiment clarifies the key chemical basis for distinguishing between medicinal snake bile and non-medicinal snake bile. The main difference lies in bile acid metabolites. Taurocholic acid (TCA), taurodeoxycholic acid (TDCA), and tauroursodeoxycholic acid (TUDCA) are relatively enriched in medicinal snake bile, while the remaining bile acid components are present in higher amounts in non-medicinal snake bile.
[0057] In this embodiment, S4 specifically includes: Construct at least one machine learning classification model, including logistic regression, neural network, random forest, support vector machine, adaptive boosting, and gradient boosting; The optimal machine learning classification model is selected by dividing the samples into training and test sets, evaluating the performance of the machine learning classification model using cross-validation, and selecting the model based on at least one of the following evaluation metrics: accuracy, recall, F1 score, and confusion matrix.
[0058] This embodiment divides the samples into independent training and test sets to effectively evaluate the model's generalization ability and avoid overfitting. Cross-validation is used to fully utilize the data under limited sample conditions and stably evaluate model performance. Multiple evaluation metrics, including accuracy, recall, F1 score, and confusion matrix, are employed to comprehensively measure the model's classification performance, ensuring that the selected optimal model has the best diagnostic effect.
[0059] In this embodiment, the machine learning classification model is preferably a neural network classifier, whose classification accuracy in both the training set and the test set reaches a preset accuracy threshold.
[0060] In this embodiment, a neural network classifier was obtained through training, evaluation, and selection, becoming the most suitable machine learning classification model for this method. The neural network possesses strong nonlinear fitting capabilities and can effectively learn the complex mapping relationship between differential metabolites and sample categories, thus performing optimally in the classification task of this invention. Using a preset accuracy threshold as an optimal condition ensures the stability and reliability of the selected model, meeting the accuracy requirements for identification in practical applications.
[0061] Example 2 This embodiment provides a medicinal snake bile identification system based on metabolomics and machine learning, which includes: The mass spectrometry detection unit is used to detect metabolic extracts of samples using liquid chromatography-mass spectrometry (LC-MS) to obtain mass spectrometry data for all samples, including medicinal snake bile samples and non-medicinal snake bile samples. The data processing unit is used to preprocess the mass spectrometry data, extract metabolite characteristic peak intensity information, and construct a metabolite characteristic matrix for all samples. The differential analysis unit is used to screen and distinguish differential metabolites between medicinal snake bile and non-medicinal snake bile based on the metabolite feature matrix and using multivariate statistical analysis methods. The model building unit is used to build and train a machine learning classification model based on the characteristic peak intensity data of the differential metabolites. The discrimination output unit is used to input the mass spectrometry data of the snake bile sample to be tested, which has been preprocessed by the data processing unit, into the machine learning classification model, and to determine whether the source is medicinal snake bile or non-medicinal snake bile based on the output of the machine learning classification model.
[0062] This embodiment integrates functional units such as mass spectrometry detection, data processing, differential analysis, model building, and discrimination output to form a complete automated process from sample detection to result output. The collaborative work of each unit enables efficient and accurate automated identification of medicinal snake bile, facilitating its widespread application in quality inspection institutions and manufacturing enterprises, and demonstrating a high level of integration and automation.
[0063] Example 3 This embodiment is a practical experimental verification of the identification method for medicinal snake bile based on metabolomics and machine learning. The specific operation is as follows: 1. Sample collection and preparation Snake bile samples, confirmed by morphological and DNA barcoding technologies, were collected, including non-medicinal snake bile (including python bile) and medicinal snake bile (Viperidae, Elapidae, Colubridae).
[0064] Metabolic extracts were prepared from all samples using a uniform method: 0.5 mL of medicinal and non-medicinal snake bile samples were placed in a 2 mL centrifuge tube, 1 mL of 50% methanol was precisely added, the mixture was shaken well, the weight was measured, the mixture was sonicated for 30 min, the weight was measured again, the weight loss was replenished with 50% methanol, the mixture was filtered through a 0.22 μm microporous membrane, and the filtrate was collected to obtain the test solution.
[0065] 2. Mass spectrometry data acquisition The test solution was detected using an ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry system.
[0066] The chromatographic conditions were as follows: Thermo Acclaim™ RSLC 120 C18 column (2.2 μm, 2.1 × 100 mm); mobile phase was methanol solution (A) and 10 mmol·L⁻¹. -1 Ammonium acetate solution (B); gradient elution program: 0–5 min, 30%–50% A; 5–20 min, 50%–90% A; 20–25 min, 90% A; 25–26 min, 90%–30% A; 26–30 min, 30% A; column temperature: 35 ℃; flow rate: 0.3 mL·min -1 Injection volume: 1 μL.
[0067] The mass spectrometry conditions were as follows: negative ion (ESI-) mode, first-level full scan, capillary voltage of 3 kV, cone voltage of 40 V; ion source temperature of 120 ℃, desolvation gas temperature of 350 ℃; desolvation gas flow rate of 500 L / h; collision gas of argon; collision voltage of 50~90 V; scan time of 0.2 s; mass scan range of 100~1500 m / z, with real-time calibration of the acquisition process using Leucineenkephalin; the mass of the generated reference fragment was 554.2615 m / z (ESI(-)).
[0068] Through the above tests, we can obtain the following results: Figure 2 The negative ion detection chromatograms of all samples are shown.
[0069] 3. Data Preprocessing The acquired raw mass spectrometry data were imported into data processing software for peak detection, retention time alignment, and normalization. This included baseline filtering, peak extraction, deconvolution, peak alignment, peak screening, compound identification, and normalization. Common metabolite characteristic peaks were extracted from all samples to construct the original data matrix. After preprocessing, over a thousand common metabolite characteristic peaks were extracted from all samples. The original data matrix was then standardized to obtain a standardized data matrix.
[0070] 4. Screening for differentially expressed metabolites Principal component analysis (PCA) was used to perform unsupervised dimensionality reduction on the standardized data matrix. The PCA score graph is shown below. Figure 3 As shown, the variance contribution rates of the first principal component and the second principal component are 13.64% and 5.35%, respectively. The score plot shows that medicinal snake bile samples and non-medicinal snake bile samples exhibit a clear separation trend in the principal component space. Medicinal snake bile samples are mainly distributed in the negative PC1 region and show an aggregated distribution; while non-medicinal snake bile samples are mainly located in the positive PC1 region and also form obvious clusters, indicating that there are significant differences in the metabolic profiles of the two.
[0071] Further construct an orthogonal partial least squares discriminant analysis model, such as Figure 4As shown, a partial least squares discriminant analysis (PLS-DA) model was constructed for medicinal snake bile and non-medicinal snake bile under negative ion conditions. The permutation test results showed that the model had high reliability: the original model's explained rate (R²Y = 0.988) and predictive ability (Q² = 0.975) were significantly better than the permuted model, and the permuted regression line showed a negative slope trend, indicating that the model did not have the risk of overfitting and could effectively resolve the metabolic differences between the two types of snake bile samples. Based on this PLS-DA model, the variable projected importance (VIP) values of each metabolite under negative ion conditions were further calculated and derived to assess the role and contribution of each variable in distinguishing medicinal and non-medicinal snake bile. Loading analysis indicated that certain bile acids and phospholipid metabolites had high loadings on PC1, and were key chemical substances leading to the distinction.
[0072] like Figure 5 As shown, based on the PLS-DA analysis results in negative ion mode, volcano plot analysis was performed on the screened differential metabolites. The results showed that compared with medicinal snake bile, 67 metabolites were significantly upregulated, 319 metabolites were significantly downregulated, and 811 metabolites showed no significant difference. Further screening for statistically significant differential metabolites using VIP>1 and P<0.05 as the standard, and comparison with standard controls and mass spectrometry databases, identified bile acid metabolites as the main differential components. A scatter plot with VIP value on the x-axis and peak response value on the y-axis revealed that taurocholic acid (TCA), taurodeoxycholic acid (TDCA), and tauroursodeoxycholic acid (TUDCA) were relatively enriched in medicinal snake bile; the remaining bile acid components were present at higher levels in non-medicinal snake bile.
[0073] 5. Machine Learning Model Building and Training like Figure 5As shown, various machine learning classifiers, including logistic regression, neural networks, random forests, and support vector machines, were built and evaluated using Python to classify snake bile based on statistically significant differential metabolites. A total of 162 samples were included in the experiment, randomly divided into a training set of 132 samples and a test set of 30 samples at an 8:2 ratio. 10-fold cross-validation was used to evaluate model stability, and accuracy, recall, F1 score, and confusion matrix were used as performance metrics. A machine learning classification model was constructed using the expression matrices of 22 differential metabolites identified under ESI negative ion mode to distinguish the source of snake bile. Training set results showed that the neural network model achieved 100% classification accuracy, outperforming random forest (97.4%), logistic regression (96.7%), support vector machine (99.3%), adaptive boosting (97.0%), and gradient boosting (97.4%). In independent test set validation, the neural network model maintained 100% accuracy, while the other classifiers showed varying degrees of misclassification. The above results demonstrate that the neural network classifier exhibits optimal diagnostic performance in determining the origin of snake bile. 6. Method Validation To verify the universality and stability of this method, medicinal snake bile samples and non-medicinal snake bile samples were collected from different origins for validation. The results showed that all medicinal and non-medicinal snake bile samples were correctly identified, with an overall high accuracy rate. This result confirms that the method of this invention has good applicability and stability for snake bile samples from different origins.
[0074] The above experiments verify that the method for identifying medicinal snake bile based on metabolomics and machine learning provided by this invention can effectively and accurately distinguish between medicinal snake bile and non-medicinal snake bile, and has important application value.
[0075] It is readily understood that those skilled in the art can combine, split, or reorganize the embodiments provided in this application to obtain other embodiments, all of which do not exceed the protection scope of this application.
[0076] The present invention and its embodiments have been described above illustratively. This description is not restrictive, and the embodiments shown are only part of the embodiments of the present invention. The actual structure is not limited thereto. Therefore, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the present invention, they should all fall within the protection scope of the present invention.
Claims
1. A method for identifying medicinal snake bile based on metabolomics and machine learning, comprising: S1. Obtain metabolic extracts of the samples, including medicinal snake bile samples and non-medicinal snake bile samples, and use liquid chromatography-mass spectrometry to detect them and obtain mass spectrometry data of all samples. S2, preprocess the mass spectrometry data, extract the intensity information of metabolite characteristic peaks, and construct the metabolite characteristic matrix of all samples; S3. Based on the metabolite feature matrix, a multivariate statistical analysis method is used to screen and distinguish the differential metabolites between medicinal snake bile and non-medicinal snake bile. S4. Based on the characteristic peak intensity data of the differential metabolites, construct and train a machine learning classification model; S5. After the mass spectrometry data of the snake bile sample to be tested undergoes the same preprocessing as in S2, it is input into the machine learning classification model. Based on the output of the machine learning classification model, it is determined whether the source is medicinal snake bile or non-medicinal snake bile.
2. The method for identifying medicinal snake bile based on metabolomics and machine learning according to claim 1, wherein, The liquid chromatography-mass spectrometry (LC-MS) technique is ultra-high performance liquid chromatography-quadrupole-time-of-flight high-resolution mass spectrometry (UHPLC-QMS), and its chromatographic conditions include: A C18 column was used, with mobile phase A being methanol solution and mobile phase B being 10 mmol·L⁻¹. -1 Ammonium acetate solution; gradient elution program: 0-5 min, 30%-50% A; 5-20 min, 50%-90% A; 20-25 min, 90% A; 25-26 min, 90%-30% A; 26-30 min, 30% A.
3. The method for identifying medicinal snake bile based on metabolomics and machine learning according to claim 1, wherein, The preprocessing of mass spectrometry data described in S2 includes: Peak detection, retention time alignment, and metabolite identification were performed on the mass spectrometry data. Common metabolite characteristic peaks were extracted from all samples to construct the original data matrix. The original data matrix is standardized to obtain a standardized data matrix.
4. The method for identifying medicinal snake bile based on metabolomics and machine learning according to claim 1, wherein, The multivariate statistical analysis methods described in S3 include: Principal component analysis (PCA) unsupervised model was used to identify the separation trend and clustering of principal components in the samples, and orthogonal partial least squares discriminant analysis model was used to screen differential variables. Metabolites with statistical significance were screened based on a variable weight value VIP > 1 and a statistical test value P < 0.
05.
5. The method for identifying medicinal snake bile based on metabolomics and machine learning according to claim 1, wherein: The differential metabolites include bile acid metabolites, specifically taurocholic acid, taurodeoxycholic acid, and tauroursodeoxycholic acid, which have different content distribution characteristics in medicinal snake bile and non-medicinal snake bile.
6. The method for identifying medicinal snake bile based on metabolomics and machine learning according to claim 1, wherein, S4 specifically includes: Construct at least one machine learning classification model, including logistic regression, neural network, random forest, support vector machine, adaptive boosting, and gradient boosting; The optimal machine learning classification model is selected by dividing the samples into training and test sets, evaluating the performance of the machine learning classification model using cross-validation, and selecting the model based on at least one of the following evaluation metrics: accuracy, recall, F1 score, and confusion matrix.
7. The method for identifying medicinal snake bile based on metabolomics and machine learning according to claim 6, wherein: The machine learning classification model is preferably a neural network classifier, whose classification accuracy in both the training and test sets reaches a preset accuracy threshold.
8. A system for identifying medicinal snake bile based on metabolomics and machine learning, comprising: The mass spectrometry detection unit is used to detect metabolic extracts of samples using liquid chromatography-mass spectrometry (LC-MS) to obtain mass spectrometry data for all samples, including medicinal snake bile samples and non-medicinal snake bile samples. The data processing unit is used to preprocess the mass spectrometry data, extract metabolite characteristic peak intensity information, and construct a metabolite characteristic matrix for all samples. The differential analysis unit is used to screen and distinguish differential metabolites between medicinal snake bile and non-medicinal snake bile based on the metabolite feature matrix and using multivariate statistical analysis methods. The model building unit is used to build and train a machine learning classification model based on the characteristic peak intensity data of the differential metabolites. The discrimination output unit is used to input the mass spectrometry data of the snake bile sample to be tested, which has been preprocessed by the data processing unit, into the machine learning classification model, and to determine whether the source is medicinal snake bile or non-medicinal snake bile based on the output of the machine learning classification model.