Method for identifying traditional chinese medicine polysaccharide based on multispectral fusion and machine learning

By employing multispectral fusion and machine learning methods, the problems of speed, non-destructive nature, and accuracy in the identification of polysaccharides from traditional Chinese medicine were solved. A robust and interpretable identification model was constructed, which is suitable for on-site quality control and high-throughput detection of polysaccharides from traditional Chinese medicine.

CN122238252APending Publication Date: 2026-06-19CHINESE MEDICINE GUANGDONG LABORATORY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINESE MEDICINE GUANGDONG LABORATORY
Filing Date
2026-05-20
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing technologies are insufficient for rapid, non-destructive, and efficient identification of polysaccharides from traditional Chinese medicines. Furthermore, single-spectral techniques suffer from overlapping peaks and unclear features when identifying polysaccharides from traditional Chinese medicines with similar chemical structures. Machine learning models lack interpretability and fail to fully utilize the complementarity of multispectral methods.

Method used

We employ a multispectral fusion and machine learning approach. By acquiring infrared and Raman spectral data, we perform standardized preprocessing and feature extraction, fuse subsets of infrared and Raman features, construct a cross-validation model, and use SHAP interpretability analysis to identify key spectral features.

Benefits of technology

It enables rapid, non-destructive, and accurate identification of polysaccharides from traditional Chinese medicine, improves the robustness and interpretability of the identification model, and is suitable for on-site quality control and high-throughput detection of polysaccharides from traditional Chinese medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122238252A_ABST
    Figure CN122238252A_ABST
Patent Text Reader

Abstract

This invention provides a method for identifying polysaccharides from traditional Chinese medicine (TCM) based on multispectral fusion and machine learning, applicable to the rapid and non-destructive identification of polysaccharides from different types and origins. The method first uses infrared and Raman spectroscopy to acquire and construct a raw spectral dataset of TCM polysaccharide samples. The dataset is then preprocessed using standardization to eliminate noise, baseline drift, and other non-sample-related interferences. Effective spectral information is then filtered through feature extraction, followed by feature-level complementary fusion of the infrared and Raman spectral features. Based on the fused feature set, a machine learning classification model is constructed and trained using a cross-validation strategy to achieve accurate identification of the type and origin of TCM polysaccharides. Furthermore, the SHAP interpretability analysis tool is used to analyze the key spectral feature bands upon which the model's discrimination decisions rely, and to analyze the correlation between these key bands and specific functional groups and chemical bond vibrations in the TCM polysaccharide molecules, providing chemical mechanism support for the identification results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of traditional Chinese medicine quality analysis and spectral identification technology. Specifically, it relates to an intelligent identification method for polysaccharides in traditional Chinese medicine based on multispectral fusion and machine learning. In particular, it relates to the fusion of complementary feature-level data of infrared and Raman spectra, the construction of a classification model combined with cross-validation model parameter optimization, and its interpretability analysis technology based on SHAP. This method can achieve rapid, non-destructive, and accurate identification of polysaccharides from different types and origins of traditional Chinese medicine. Background Technology

[0002] Polysaccharides from traditional Chinese medicine (TCM) are an important class of bioactive substances, and their differences in type and structure directly affect the efficacy and quality uniformity of TCM. Therefore, establishing accurate, efficient, rapid, and non-destructive methods for identifying TCM polysaccharides is of great significance for ensuring the clinical efficacy of TCM, evaluating its authenticity, and formulating quality control standards.

[0003] Currently, the identification of polysaccharides from traditional Chinese medicine (TCM) mainly relies on traditional analytical methods such as high-performance liquid chromatography (HPLC) and mass spectrometry (MS). While these methods are highly accurate, they typically require complex extraction and derivatization processes, which are time-consuming and destructive, making them unsuitable for rapid, non-destructive, and high-throughput on-site quality control. Spectroscopic analysis techniques, especially infrared and Raman spectroscopy, have been widely applied in TCM quality control research due to their advantages of speed, non-destructiveness, and ease of operation. However, existing studies mostly focus on the analysis of raw TCM materials and have not yet conducted multispectral fusion identification studies on TCM polysaccharides. The few related studies only focus on single spectral techniques, and are mostly infrared spectroscopy. A mature Raman spectroscopy detection method for TCM polysaccharides has not yet been developed. If the Raman acquisition method for raw TCM powder is directly used, problems such as weak signals, low characteristic peak resolution, and susceptibility to laser ablation are likely to occur. Even increasing laser power is unlikely to effectively improve signal quality and may even exacerbate sample damage, leading to poor repeatability of spectral data. In addition, single-spectral techniques provide limited information dimensions. When identifying polysaccharides from traditional Chinese medicines with highly similar chemical structures, there are problems such as overlapping spectral peaks and unclear characteristics, resulting in insufficient identification specificity and robustness.

[0004] In recent years, machine learning technology has been widely applied to spectral data analysis, enabling the discovery of potential patterns from complex spectra and significantly improving the automation and objectivity of identification. However, existing methods are mostly based on single spectral sources to build models, failing to fully utilize the complementarity of different spectral techniques: infrared spectroscopy can reflect the vibrational information of molecular groups, while Raman spectroscopy can reflect the vibrational information of the molecular skeleton. Combining the two can achieve a comprehensive characterization of polysaccharide structural features, but this technological advantage has not yet been fully utilized in the field of traditional Chinese medicine polysaccharide identification. In addition, machine learning models often lack interpretability, making it difficult for analysts to understand which specific spectral features the model uses for discrimination. This hinders the in-depth application and credibility establishment of this technology in the quality analysis and standardization research of traditional Chinese medicine that requires mechanistic support.

[0005] In summary, optimizing the Raman spectroscopy acquisition scheme for polysaccharides from traditional Chinese medicine (TCM), establishing a standardized spectral preprocessing workflow, and effectively integrating complementary information from multiple spectral sources such as infrared and Raman to construct a more robust and accurate identification model, while ensuring that the identification process is mechanistically sound and results are interpretable, and clarifying the core basis for identification conclusions, are crucial technical challenges that urgently need to be addressed in this field. This will facilitate the transition of intelligent identification technology for TCM polysaccharides from theory to practical application. Summary of the Invention

[0006] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: A method for identifying polysaccharides in traditional Chinese medicine based on multispectral fusion and machine learning includes the following steps: (1) Obtain standard spectral data of polysaccharide samples of Chinese medicine of different types and origins, including infrared spectral data and Raman spectral data; preheat the instrument before collection, collect air background data before infrared spectral collection to eliminate environmental interference, and collect Raman spectral data in the dark. (2) The infrared spectral data and Raman spectral data are respectively subjected to standardization preprocessing and feature extraction to obtain the corresponding infrared feature subsets and Raman feature subsets; the preprocessing operations are performed in the following order: smoothing → baseline correction → normalization → multivariate scattering correction (MSC); (3) The infrared feature subset and the Raman feature subset are fused complementaryly to generate a multispectral fusion feature set; the fusion method is feature-level fusion, which splices the infrared and Raman feature vectors of the same sample along the feature dimension column direction, and the fused feature vector is the sum of the dimensions of the original infrared and Raman feature vectors; (4) Based on the multispectral fusion features, the machine learning classification model is trained and the hyperparameters are optimized using the 5-fold cross-validation strategy. The training set and the test set are randomly divided in a 7:3 ratio to obtain the identification model of Chinese medicine polysaccharides. The stratified random division ensures that the sample ratio of each type / origin of Chinese medicine in the training set and the test set is consistent, thereby reducing the impact of sample imbalance on the model and obtaining the identification model of Chinese medicine polysaccharides. (5) Use the polysaccharide identification model of traditional Chinese medicine obtained in step (4) to identify the multispectral fusion feature set of the polysaccharide extract of the sample to be tested: Based on the infrared spectral data and Raman spectral data of the polysaccharide sample of traditional Chinese medicine, perform the feature extraction and fusion processing of steps (2) to (3) in sequence, and input the obtained fusion features into the polysaccharide identification model of traditional Chinese medicine, and output the category discrimination result of the sample; use independent test set data to evaluate the performance of the model, and use the evaluation result as the basis for determining whether the model is usable; (6) An interpretable artificial intelligence model is used to analyze the decision-making mechanism of the traditional Chinese medicine polysaccharide identification model, identify key spectral feature bands that significantly contribute to the classification results, and the key bands are related to the vibration of functional groups / chemical bonds of polysaccharide molecules, providing a basis for discrimination at the chemical mechanism level.

[0007] Preferably, in step (1), an infrared spectral data is acquired using a Fourier transform infrared spectrometer equipped with an attenuated total reflectance (ATR) accessory, with 16 scans and a resolution of 4 cm⁻¹. -1 The spectral acquisition range is 650~4000 cm⁻¹ -1 Ultimately, fingerprints were extracted from an area of ​​650-1800 cm. -1 For subsequent analysis; Raman spectral data were acquired using a laser confocal microscopy Raman spectrometer. The laser wavelength was 532 nm, the laser power was 50 mW, the integration time was 2 seconds, the number of integrations was 10, and the resolution was 2 cm⁻¹. -1 The spectral acquisition range is 200~1800 cm⁻¹ -1 Ultimately, 350-1500 cm sections were cut. -1 For subsequent analysis; Raman spectroscopy acquisition adopts a pretreatment method of polysaccharide dissolution and drying to form a film, which avoids laser burning of the sample and improves the Raman signal detection effect.

[0008] Preferably, in step (2), the data preprocessing includes at least one or more of the following operations: Savitzky-Golay (SG) spectral smoothing, adaptive iterative reweighted penalized least squares baseline correction, normalization, and MSC processing. Infrared data is processed using the built-in functions of the acquisition software, and Raman data is processed using R language.

[0009] Preferably, in step (2), the spectral feature extraction adopts the principal component analysis method, with the cumulative variance contribution rate not less than 95% as the standard, and the infrared spectrum and Raman spectrum data are reduced in dimension to extract representative features; before the principal component analysis, the spectral data are Z-score standardized to eliminate the difference in dimensions and intensity, and the principal components are selected by sorting the feature values ​​from large to small, and noise features with low variance contribution rate are deleted.

[0010] Preferably, in step (3), the spectral data fusion method is feature-level fusion, specifically: the infrared feature vector (i.e., infrared PCA feature vector) and the Raman feature vector (i.e., Raman PCA feature vector) corresponding to the same Chinese herbal polysaccharide sample are concatenated along the feature dimension to form the fused feature vector of the sample, thereby constituting the fused feature dataset; the specific concatenation process is: the 1×m dimension infrared PCA feature vector (m represents the number of infrared principal component features obtained after the principal component feature extraction in step (2), which is the infrared core feature dimension determined by the cumulative variance contribution rate not less than 95%) corresponding to the same Chinese herbal polysaccharide sample is concatenated along the feature dimension to form the fused feature vector of the sample, thereby constituting the fused feature dataset; the specific concatenation process is: the 1×m dimension infrared PCA feature vector (m represents the number of infrared principal component features obtained after the principal component feature extraction in step (2), which is the infrared core feature dimension determined by the cumulative variance contribution rate not less than 95%) and the 1×n dimension Raman PCA feature vector (n represents the number of infrared principal component features obtained after the principal component feature extraction in step (2)) are concatenated along the feature dimension to form the fused feature vector of the sample, which is ... The Raman principal component features obtained after component feature extraction and screening are determined by the standard that the cumulative variance contribution rate is not less than 95% (Raman core feature dimension). They are directly spliced ​​along the column direction of the feature dimension to form a 1×(m+n) dimension fusion feature vector of the sample. This fusion feature vector contains the core information of the molecular group vibration of Chinese herbal polysaccharide represented by the infrared PCA feature vector and the core information of the molecular skeleton vibration of Chinese herbal polysaccharide represented by the Raman PCA feature vector. Then, the 1×(m+n) dimension fusion feature vectors of all Chinese herbal polysaccharide samples are arranged in the row direction to construct an N×(m+n) dimension feature-level fusion vector matrix of all samples (N is the total number of samples), which is the fusion feature dataset of the present invention.

[0011] Preferably, in step (4), the machine learning classification model includes at least one of the following algorithms: K-Nearest Neighbors (KNN), Linear Discriminant Analysis (LDA), Random Forest (RF), Logistic Regression (LR), Support Vector Machine (SVM), Gradient Boosting (XGBoost), Gaussian Naive Bayes (GNB), and Decision Tree Analysis (DT).

[0012] KNN, a distance-based local classification model, determines the class of a sample by voting among its neighboring classes, thus adapting to the local feature patterns of spectral data. For hyperparameter optimization of the KNN model on spectral data, candidate parameters were selected: the number of nearest neighbors (3, 5, 7, 9, 11), equal weight / distance weighting strategies, and Euclidean distance / Manhattan distance. Stratified sampling ensured a balanced sample distribution, and 20 iterations of random search were used to optimize the hyperparameters, validating the adaptation effect of different distance metrics on spectral features.

[0013] LDA addresses the classification problem of high-dimensional spectral data by finding the optimal projection direction that minimizes intra-class variance and maximizes inter-class variance. Parameter optimization focuses on the solver and shrinkage coefficient: the solver is selected from Singular Value Decomposition (SVD), Least Squares QR Decomposition (LSQR), and Eigenvalue Decomposition (eigen) to adapt to different matrix decomposition scenarios; the shrinkage coefficient is set to None, Auto, and gradient values ​​from 0.1 to 0.9, using covariance matrix regularization to alleviate matrix singularity issues in high-dimensional, small-sample scenarios and improve model stability. Hyperparameter optimization is achieved through 20 iterations of random search.

[0014] RF ensemble learning uses majority voting from multiple independent decision trees for classification, and Bootstrap sampling combined with random feature selection reduces overfitting. Parameter optimization aims to control complexity and improve generalization ability: the number of trees (n_estimators) is set to 50~300, the maximum tree depth (max_depth) is set to 5~20 or unlimited, the minimum number of samples for node splits (min_samples_split) is set to 2~15, the minimum number of samples for leaf nodes (min_samples_leaf) is set to 1~8, and the number of features for each split (max_features) adopts three strategies: square root (sqrt), logarithm (log2), and full (None), and compares the bootstrap sampling on / off state; multi-dimensional hyperparameter optimization is completed through 50 iterations of random search.

[0015] LR maps linear regression results to the 0-1 interval using the Sigmoid function, determining the category based on probability values, thus adapting to linear classification of high-dimensional fused data. Parameter optimization aims to enhance regularization and ensure convergence: the regularization strength (C) is set to 0.001-100, the regularization type (penalty) includes L1 and L2, the solver uses lbfgs (quasi-Newton optimization algorithm), liblinear (coordinate descent method), and saga (stochastic average gradient descent method), and the maximum number of iterations (max_iter) is set to 2000; the optimal parameters of the linear model under high-dimensional spectral data are obtained through 30 iterations of random search.

[0016] SVM maps low-dimensional spectral data to high-dimensional data using kernel functions, seeking the optimal classification hyperplane to fit the nonlinear distribution of the spectral data. Parameter optimization aims to balance the fitting and generalization capabilities of SVM: the penalty coefficient (C) is set to 0.001~100; the kernel width (gamma) is set to scale / auto (automatic scaling) and 0.001~1 (manual gradient); the RBF (radial basis function) is selected to fit the nonlinear distribution of the spectral data; the optimal combination of C and gamma is explored through 30 iterations of random search to address the parameter sensitivity issue of SVM.

[0017] XGBoost gradient boosting ensemble learning is a serial decision tree ensemble that iteratively trains multiple decision trees, with each tree learning the residuals / errors of the previous tree, gradually accumulating and optimizing the model loss. It features built-in regularization, pruning, and automatic handling of missing values, balancing fitting accuracy and generalization ability, and can efficiently handle nonlinear and complex classification tasks with high-dimensional spectral data. Parameter optimization aims for lightweight design and strong regularization: the maximum tree depth (max_depth) is set to 3, 5, or 7; the learning rate (learning_rate) is set to 0.01~0.2; and the subsample and feature sampling ratios (colsample_bytree) are set to 0.6~1.0. Gradient search is performed on the split gain threshold (gamma), L1 regularization (reg_alpha), and L2 regularization (reg_lambda), and the number of trees (n_estimators) is controlled between 50 and 150. Parameter optimization is completed through 25 random searches, improving model generalization ability while ensuring training efficiency.

[0018] GNB is based on Bayes' theorem and the feature independence assumption, assuming that the spectral features follow a Gaussian distribution, and calculates the posterior probability to achieve classification. To address the problem of minimal variance in the spectral bands, the variance smoothing coefficient (var_smoothing) is optimized to 10. -9 ~10 -5 To avoid zero-variance calculation errors, optimization was completed with 20 random searches, and Naive Bayes was used as the benchmark model to verify the effectiveness of complex models.

[0019] DT constructs a tree model by recursively partitioning the feature space, using information gain / Gini coefficient as the splitting criterion to adapt to the hierarchical partitioning of spectral features. Parameter optimization focuses on addressing overfitting in a single decision tree: the splitting criteria are gini coefficient and entropy; the gradient ranges for maximum tree depth, node splits, and minimum number of samples in leaf nodes are consistent with RF; and the number of features for each split employs three strategies: sqrt, log2, and None. Thirty random searches are used for optimization, serving as a benchmark model to verify the advantages of the ensemble method and simultaneously determining the optimal complexity of a single tree.

[0020] Preferably, in step (4), the average classification accuracy of the 5-fold cross-validation strategy is used as the evaluation index to train and optimize the hyperparameters of the machine learning classification model, so as to further improve the overfitting phenomenon of the model, enhance the generalization ability, and finally obtain the optimal parameter combination that fits the spectral data. The specific implementation method is as follows: the training set after hierarchical random partitioning is randomly divided into 5 subsets. Each time, 4 subsets are selected as training subsets and 1 subset is selected as validation subsets to train and validate the model. The cycle is repeated 5 times so that each subset is used as a validation subset once. The average accuracy of the 5 validations is taken as the validation accuracy of the model under the parameter combination. Different parameter combinations are searched through a random network. The parameter combination with the highest average accuracy of 5-fold cross-validation is selected as the optimal parameters of the model. Finally, the model is obtained by training the entire training set based on the optimal parameters.

[0021] The optimal parameter combination for the identification model of polysaccharides from different types of traditional Chinese medicine is as follows: KNN: Weights: uniform (all neighbors have the same weight); Number of neighbors (n_neighbors): 5; Metric: Manhattan (using Manhattan distance).

[0022] LDA: Solver: SVD is suitable for high-dimensional data scenarios. Shrinkage: None, no additional regularization is required.

[0023] RF: Number of decision trees (n_estimators): 300; Minimum number of samples for node split (min_samples_split): 10; Minimum number of samples for leaf nodes (min_samples_leaf): 4; Maximum number of features (max_features): log2; Maximum tree depth (max_depth): None (no depth limit). Bootstrap: False (each tree is trained using all the original data).

[0024] LR: Solver: lbfgs (optimization loss function algorithm); Regularization type: L2; Regularization strength (C): 1.0.

[0025] SVM: Kernel: RBF; Kernel coefficient (gamma): 0.001; Penalty coefficient (C): 10.

[0026] XGBoost: Subsample ratio: 0.6, i.e., using 60% of the samples; L2 regularization coefficient (eg_lambda): 1.5; No L1 regularization coefficient (reg_alpha); Number of decision trees (n_estimators): 50; Maximum tree depth (max_depth): 3; Learning rate (learning_rate): 0.1; Split threshold (gamma): 0.2; Column sampling ratio (colsample_bytree): 0.6, i.e., using 60% of the features.

[0027] GNB: Variance smoothing term (var_smoothing): 10 -9 The smallest values ​​have almost no effect on the original variance distribution.

[0028] DT: Minimum number of samples for node split (min_samples_split): 5; Minimum number of samples for leaf nodes (min_samples_leaf): 1; Maximum number of features (max_features): None (use all features); Maximum tree depth (max_depth): None (no depth limit); Split criterion (criterion): entropy (information entropy).

[0029] The optimal parameter combinations for each model for identifying ginseng polysaccharides from different origins are as follows: KNN: Weights: distance (samples that are closer to each other have a higher weight); Number of nearest neighbors (n_neighbors): 3; Metric: Manhattan (using Manhattan distance).

[0030] LDA: Solver: SVD is suitable for high-dimensional data scenarios. Shrinkage: None, no additional regularization is required.

[0031] RF: Number of decision trees (n_estimators): 300; Minimum number of samples for node split (min_samples_split): 5; Minimum number of samples for leaf nodes (min_samples_leaf): 2; Maximum number of features (max_features): sqrt; Maximum tree depth (max_depth): 20. Bootstrap: True (uses sampling with replacement to generate subsets, enhancing model diversity).

[0032] LR: Solver: liblinear (an optimization algorithm suitable for small sample and high-dimensional data); Regularization type: L1, L1 regularization ratio: 0.9 (i.e., L1 regularization accounts for 90%, L2 regularization accounts for 10%); Regularization strength (C): 1.0.

[0033] SVM: Kernel: RBF; Kernel coefficient (gamma): scale (automatically calculated based on feature variance); Penalty coefficient (C): 100.

[0034] XGBoost: Subsample ratio: 0.8, meaning 80% of the samples are used for training; L2 regularization coefficient (reg_lambda): 2; No L1 regularization (reg_alpha); Number of decision trees (n_estimators): 150; Maximum tree depth (max_depth): 7; Learning rate (learning_rate): 0.05; Split threshold (gamma): 0; Column sampling ratio (colsample_bytree): 0.6, meaning 60% of the features are used.

[0035] GNB: Variance smoothing term (var_smoothing): 10 -9 The smallest values ​​have almost no effect on the original variance distribution.

[0036] DT: Minimum number of samples for node split (min_samples_split): 5; Minimum number of samples for leaf nodes (min_samples_leaf): 1; Maximum number of features (max_features): sqrt; Maximum tree depth (max_depth): None (no depth limit); Split criterion (criterion): entropy (information entropy).

[0037] Preferably, in step (5), the trained Chinese herbal polysaccharide identification model is evaluated using stratified, randomly partitioned, independent test set data. The evaluation metrics include at least one of the following: accuracy, precision, recall, F1 score, and Matthews correlation coefficient (MCC), calculated using the following formulas. All evaluation metrics of the multispectral fusion model are superior to those of the model constructed using a single spectrum, demonstrating the effectiveness of multispectral fusion.

[0038] ; TP, FP, TN, and FN represent true positive, false positive, true negative, and false negative, respectively.

[0039] Preferably, in step (6), the interpretable artificial intelligence method is an analysis model based on the SHAP framework, which is used to quantify the contribution of each spectral feature to the model output and identify key discrimination bands; the feature contribution is analyzed through SHAP beehive diagram and wavenumber importance spectrum, and the core principal component features are the key basis for model decision-making.

[0040] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention provides a method for identifying the types and origins of polysaccharides in traditional Chinese medicine based on multispectral fusion and machine learning. It can identify 14 types of polysaccharides from traditional Chinese medicine and pinpoint the origins of ginseng polysaccharides from Jilin, Liaoning, and Heilongjiang provinces, with good identification results. 2. This invention provides a Raman spectroscopy acquisition preprocessing method for polysaccharide dissolution, drying, and film formation. The method involves preparing polysaccharides into commonly used working concentration solutions, drying them into films, and then detecting them. This method can specifically address the problems of weak Raman signals, indistinct characteristic peaks, and susceptibility to laser burns in traditional Chinese medicine polysaccharide powder samples, thereby improving the quality of spectral data detection and enhancing the practicality of Raman spectroscopy detection. 3. This invention effectively improves the model's identification accuracy by combining steps of optimizing polysaccharide spectral preprocessing and selecting optimal spectral fusion strategies; at the same time, it introduces the SHAP interpretability analysis model to identify key spectral features that make significant contributions to the discrimination results, thereby providing interpretable chemical evidence for the identification conclusions; 4. The method of the present invention is fast and non-destructive throughout the process, requiring no complex derivatization of Chinese herbal polysaccharide samples. The operation process is simple and suitable for on-site quality control and high-throughput detection of Chinese herbal polysaccharides. Attached Figure Description

[0041] Figure 1 For analysis of the flowchart; Figure 2 Images showing the properties of polysaccharides from 14 kinds of traditional Chinese medicine; Figure 3 The average infrared spectrum of polysaccharides from 14 kinds of traditional Chinese medicines; Figure 4 Comparison of Raman signals detected by dry membranes of polysaccharide powder and polysaccharide solution; Figure 5 The average Raman spectra of polysaccharides from 14 kinds of traditional Chinese medicines; Figure 6 Interpretation of variance for infrared principal component analysis of polysaccharides from 14 kinds of traditional Chinese medicines; Figure 7 A variance interpretation plot for Raman principal component analysis of polysaccharides from 14 kinds of traditional Chinese medicines; Figure 8 Confusion matrix diagram of classification models for polysaccharides of 14 kinds of traditional Chinese medicines; Figure 9SHAP bee colony graph showing the contribution of polysaccharide classification model features of 14 kinds of traditional Chinese medicines; Figure 10 The top three principal component loading plots based on infrared and Raman importance obtained from the global SHAP analysis of polysaccharides from 14 kinds of traditional Chinese medicines; Figure 11 SHAP analysis was performed on key spectral characteristic bands for the classification of polysaccharides from 14 traditional Chinese medicines. Figure 12 Average infrared spectra of ginseng polysaccharides from different origins; Figure 13 Average Raman spectra of ginseng polysaccharides from different origins; Figure 14 Interpretive plot of variance for infrared principal component analysis of ginseng polysaccharides from different origins; Figure 15 Interpretive plot of variance for Raman principal component analysis of ginseng polysaccharides from different origins; Figure 16 Confusion matrix diagram for classification models of ginseng polysaccharides from different origins; Figure 17 SHAP bee colony diagram showing the contribution of ginseng polysaccharide classification model features from different origins; Figure 18 The top three principal component loading plots based on infrared and Raman importance obtained from the global SHAP analysis of ginseng polysaccharides from different origins; Figure 19 Key spectral characteristic bands for SHAP analysis of ginseng polysaccharides from different origins were used to classify them. Detailed Implementation

[0042] Figure 1 A flowchart of the method for identifying polysaccharides in traditional Chinese medicine based on multispectral fusion and machine learning according to the present invention is provided. The invention is further described below with reference to specific embodiments, and the advantages and features of the invention will become clearer with the description. However, these embodiments are merely exemplary and do not constitute any limitation on the scope of protection defined by the claims of the present invention.

[0043] Example 1. Model Building 1. Construction of Spectral Dataset (1) Extraction of polysaccharides from traditional Chinese medicine In this embodiment, infrared spectroscopy and Raman spectroscopy were fused and combined with machine learning to identify 14 kinds of polysaccharides from traditional Chinese medicine. The specific steps are as follows: Experimental Materials: The experiment used various traditional Chinese medicine (TCM) granules produced by Jiangsu Jiangyin Tianjiang Pharmaceutical Co., Ltd., all with maltodextrin as the excipient. The sample composition included 5 batches of ginseng, poria cocos, platycodon grandiflorus, astragalus membranaceus, ophiopogon japonicus, gastrodia elata, and codonopsis pilosula granules; 3 batches of angelica dahurica, codonopsis pilosula, and polygonatum sibiricum granules; 2 batches of wolfberry, American ginseng, and yam granules; and 1 batch of polygonatum odoratum granules, totaling 14 kinds of TCM. Polysaccharides were extracted from the above TCM granules to obtain different TCM polysaccharide samples, which were named according to the Chinese pinyin abbreviations of the above TCMs (RS, FL, JG, HQ, MD, TM, TZS, BZ, DS, HJ, GQZ, XYS, SY, YZ, as detailed in Table 1). The control excipient maltodextrin was represented by MYHJ. Figure 2 These are morphological diagrams of different polysaccharides from traditional Chinese medicine.

[0044] Extraction method of polysaccharides from traditional Chinese medicine: Weigh 20g of the traditional Chinese medicine formula granules and place them in a beaker. Add 10 times the volume of boiling water and stir thoroughly for 1 hour with a magnetic stirrer. Add 95% ethanol until the ethanol volume reaches 80%. Let stand at 4℃ for 12 hours, centrifuge at 3000r / min for 10 minutes to collect the precipitate, wash three times with anhydrous ethanol, evaporate the residual ethanol from the precipitate, dissolve it in ultrapure water, and further dialyze using a 1000Da dialysis bag for 24 hours. Then freeze-dry to obtain different polysaccharide powders from traditional Chinese medicine. The yield and polysaccharide content (determined by phenol-sulfuric acid method) of the extracted polysaccharides are shown in Table 1. The polysaccharide powders are white / pale yellow cotton-like and have good water solubility.

[0045] Table 1. Information on polysaccharides in different Chinese herbal medicine granules .

[0046] (2) Spectral acquisition 1) Infrared Spectroscopy Acquisition: The infrared spectra of polysaccharides from traditional Chinese medicine were acquired using a Spectrum two Fourier transform infrared spectrometer equipped with an ATR top plate, achieving a resolution of 4 cm⁻¹. -1 Before data acquisition, the instrument was fully preheated, and background air sampling was performed before acquiring data for each polysaccharide to eliminate environmental interference. The dried polysaccharide powder was placed on an ATR crystal, and the acquisition scan was set to 16 times, with a sampling range of 650–4000 cm⁻¹. -1 Ultimately, fingerprints were harvested and preserved in areas ranging from 650 to 1800 cm. -1 For subsequent analysis, 45 spectra were collected for each type of Chinese herbal polysaccharide, resulting in a total of 675 infrared spectral standard data. Figure 3 The average infrared spectrum of different Chinese herbal polysaccharides.

[0047] 2) Raman spectroscopy acquisition: The Raman signals of the polysaccharides from traditional Chinese medicine were acquired using a HORIBA laser confocal microscopy spectrometer. This instrument is equipped with a 532nm laser, a charge-coupled device (CCD) detector, an inverted microscope, and an automatically moving stage. The grating specification is 600g / mm, and the spectral resolution is 2 cm⁻¹. -1 Before testing, the instrument is fully preheated. This invention acquires polysaccharide spectra using a polysaccharide dissolution and drying film-forming method (preparing polysaccharides into a solution to form a uniformly distributed, flat film layer), which improves the signal-to-noise ratio and enhances the sample's resistance to photodamage. Conventional polysaccharide powder samples are loose, rough, and uneven, easily burned by heat, and can only be acquired for extended periods at low laser power, resulting in low signal-to-noise ratios and easy degradation. The solution film-forming method significantly improves thermal conductivity and resistance to photodamage, allowing for rapid and stable acquisition at higher laser power, resulting in higher signal-to-noise ratios, superior spectral quality, and higher detection efficiency. This method, to a certain extent, solves the technical problems of easy burning, weak signal, and poor repeatability in Raman detection of traditional Chinese medicine polysaccharides. (Detection comparison example follows.) Figure 4 (The figures show the comparison results of polysaccharide detection in 10 batches of ginseng medicinal materials. Only the raw data has been baseline corrected.) a) Polysaccharide powder detected at 25 mW for 3 seconds with a cumulative total of 15 detections; b) Polysaccharide solution detected at 25 mW for 2 seconds with a cumulative total of 10 detections; c) Polysaccharide solution detected at 50 mW for 2 seconds with a cumulative total of 10 detections. (Direct detection of polysaccharide powder at 50 mW power would completely burn it, resulting in no effective signal. Detection of the solution after film formation avoids thermal damage and yields an effective Raman signal.)

[0048] The specific detection method is as follows: Each polysaccharide is prepared as a 2 mg / mL solution. 3 μL of each solution is placed on a foil-wrapped glass slide and allowed to air dry for 5 minutes. After film formation, the slide is placed on a loading stage. The sample is focused using a 50× objective lens with a laser power of 50 mW, an integration time of 2 seconds, and 10 consecutive scans, acquiring images from 200 to 1800 cm⁻¹. -1 Raman spectral signals within the specified range were obtained. The entire acquisition process was conducted in darkness to avoid environmental interference. At least three points were randomly selected for each sample, and the average value was taken to reduce measurement bias caused by random cosmic rays and local sample inhomogeneities. The intercept range was 350–1500 cm⁻¹. -1 For subsequent analysis, 45 spectra were collected for each type of Chinese herbal polysaccharide, and a total of 675 Raman spectral standard data were obtained. Figure 5 The average Raman spectra of polysaccharides from different traditional Chinese medicines are shown.

[0049] (3) Preprocessing of infrared and Raman spectroscopic data: In order to eliminate interference from non-samples themselves, unify data scale, and make subsequent analysis more accurate and stable, the infrared and Raman spectroscopic data were smoothed, baseline corrected, normalized, and subjected to MSC (Multiplicative Scatter Correction). Infrared data was automatically baseline corrected, smoothed using 5-point SG, and normalized using the built-in functions of the acquisition software. Raman data was processed using R language, specifically: SG smoothing with a 15-point filter window and second-order polynomial fitting; adaptive iterative reweighted penalized least squares baseline correction with a smoothing parameter of 800 and a maximum number of iterations of 50. After the above preprocessing, the model's accuracy on the test set reached a maximum of 97.5%. Further MSC processing was performed on the infrared and Raman spectral data distributions using R language. MSC correction significantly eliminated interference between different batches of the same variety, making the sample clusters more compact. It also weakened spectral differences caused by non-chemical components, highlighting the spectral characteristics corresponding to polysaccharides, making the differences between different categories of samples more obvious. After MSC processing, the model's accuracy on the test set was further improved to 100%.

[0050] (4) Feature extraction from infrared and Raman datasets Feature extraction for the infrared and Raman datasets was performed using principal component analysis (PCA), with a cumulative variance contribution rate of at least 95% as the criterion.

[0051] The specific steps of the PCA algorithm process are as follows: 1) Data input: The preprocessed spectral data matrix is ​​X, with dimensions n×p (n is the number of samples, p is the number of wavenumber points); 2) Standardization: X is Z-score standardized to eliminate dimensional and intensity differences between different wavenumber points, ensuring that all features participate in the calculation equally. The calculation formula is: ; Among them, X (i,j) Let μ be the intensity value of the j-th wavenumber point of the i-th sample in the original spectral data matrix. j Let σ be the mean of the j-th wavenumber point. j Let be the standard deviation of the j-th wavenumber point.

[0052] 3) Covariance Matrix Calculation: Calculate the covariance matrix C of the standardized data, with dimensions p×p. The calculation formula is: ; in, For the standardized data matrix X stdThe transpose of the matrix has dimensions p×n; n is the total number of samples, and n-1 is the unbiased estimate of the covariance; where the elements of the covariance matrix C (i,j) This represents the correlation strength between the i-th wavenumber point and the j-th wavenumber point. 4) Eigenvalue and eigenvector solving: Solve for the eigenvalues ​​λ (λ1≥λ2≥…≥λ) of the covariance matrix C using linear algebra methods. p (≥0) and the corresponding eigenvectors v1, v2, ..., v p (Each feature vector has a dimension of p×1). The eigenvalue λ represents the variance information carried by the corresponding feature vector, and the feature vector v represents the projection direction of the data in that direction.

[0053] 5) Principal Component Selection: Sort the eigenvalues ​​from largest to smallest, and select the top k eigenvectors with a cumulative variance contribution rate ≥ 95% as the principal component loading matrix V (dimension p × k). The calculation formula is as follows: ; Where λi is the i-th eigenvalue of the covariance matrix C (sorted from largest to smallest, λ1≥λ2≥…≥λp≥0), representing the original data variance information carried by the corresponding eigenvector; This represents the sum of the first k eigenvalues, i.e., the total variance retained by the first k principal components, where k is the final number of principal components selected. This represents the sum of all eigenvalues, i.e., the total variance of the original standardized data.

[0054] 6) Principal Component Calculation: The final principal component matrix PC (with dimensions n×k) is obtained by multiplying the standardized data with the loading matrix. The calculation formula is as follows: ; Each column of PC corresponds to a principal component (PC1, PC2, ..., PC). k Each row corresponds to the score of a sample on the principal component; V is the principal component loading matrix, which is composed of the eigenvectors v1, v2, ..., vk corresponding to the first k largest eigenvalues, with a dimension of p×k, where vi is the i-th eigenvector (dimension p×1) and represents the projection direction of the i-th principal component.

[0055] Six infrared principal components were finally extracted, and the characteristics of each infrared principal component (PC) are as follows: The variance contribution of the infrared PC1 was 60.77%, mainly explaining the variance in the 1250–1350 cm⁻¹ range. -1 1450~1600 cm -1 Wavenumber range (load values ​​> 0.0493 and > 0.05 can be considered as bands with some contribution; load values ​​> 0.07 are usually bands with strong contributions). The variance contribution of infrared PC2 was 20.50%, mainly explaining the variance in the 800–1000 cm⁻¹ range. -1 Wavenumber range (load value > 0.0611); The variance contribution of the infrared PC3 was 5.57%, mainly explaining the variance in the 1000–1200 cm⁻¹ range. -1 Wavenumber range (load value > 0.0734); The variance contribution of the infrared PC4 was 4.51%, mainly explaining the variance in the 1150–1200 cm⁻¹ range. -1 1300~1400 cm -1 Wavenumber range (load value > 0.0691); The PC5 variance contribution rate in the infrared region was 2.45%, primarily explaining the variance in the 1750–1800 cm⁻¹ range. -1 Wavenumber range (load value > 0.0583); The PC6 variance contribution rate in the infrared region was 1.78%, primarily explaining the variance in the 1100–1200 cm⁻¹ range. -1 1750~1800 cm -1 The wavenumber range (load value > 0.0477) is shown in Table 2 for a more detailed interpretation of the wavenumbers. The maximum variance of the six infrared principal component features is preserved at 95.6%.

[0056] Twenty-nine Raman principal component features were extracted, of which the first ten principal component features include: The Raman PC1 variance contribution rate was 43.73%, mainly explaining the variance in the 400–500 cm⁻¹ range. -1 950~1050 cm -1 1200~1300cm -1 1400~1500 cm -1 Wavenumber range (load value > 0.0509); The Raman PC2 variance contribution rate was 14.88%, mainly explaining the variance in the 500–600 cm⁻¹ range. -1 650~700 cm -1 850~950 cm -1 Wavenumber range (load value > 0.0603); The Raman PC3 variance contribution rate was 7.86%, mainly explaining the variance in the 500–650 cm⁻¹ range. -1 800~850 cm -1 Wavenumber range (load value > 0.0703); The Raman PC4 variance contribution rate was 7.45%, mainly explaining the variance in the 450–500 cm⁻¹ range. -1 750~850 cm -1 1050~1100cm -1Wavenumber range (load value > 0.0676); The Raman PC5 variance contribution rate was 5.43%, mainly explaining the variance in the 450–500 cm⁻¹ range. -1 900~1000 cm -1 1100~1200cm -1 Wavenumber range (load value > 0.0650); The Raman PC6 variance contribution rate was 3.85%, mainly explaining the variance in the 650–700 cm⁻¹ range. -1 800~850 cm -1 900–950 cm -1 Wavenumber range (load value > 0.0660); The Raman PC7 variance contribution rate was 2.34%, mainly explaining the variance in the 450–500 cm⁻¹ range. -1 800~900 cm -1 950~1000 cm -1 Wavenumber range (load value > 0.0577); The Raman PC8 variance contribution rate was 1.36%, mainly explaining the variance in the 450–500 cm⁻¹ range. -1 650–800 cm -1 900~950 cm -1 Wavenumber range (load value > 0.0636); The Raman PC9 variance contribution rate was 0.87%, mainly explaining the variance in the 450–500 cm⁻¹ range. -1 650~700 cm -1 Wavenumber range (load value > 0.0572); The Raman PC10 variance contribution rate was 0.65%, mainly explaining the variance in the 650–700 cm⁻¹ range. -1 1150~1200 cm -1 The wavenumber interval (loading value > 0.0574), and the corresponding wavenumbers are more specific as shown in Table 3. The 29 Raman principal component features ultimately retained a maximum variance of 95.0%.

[0057] The variance contribution rate of the infrared principal components is obtained as follows: Figure 6 The variance contribution rate of the Raman principal components is as follows: Figure 7 The feature vectors corresponding to the high feature values ​​extracted are all linear combinations of the original spectral wavenumbers, representing the core feature information of the spectrum.

[0058] Table 2. Interpreted wavenumbers of principal components in infrared spectroscopy for different types of polysaccharides. .

[0059] Table 3. Wavenumbers of principal components interpreted by Raman spectroscopy for different types of polysaccharides. .

[0060] (5) Fusion of infrared and Raman principal component eigenvectors Subsequently, the infrared and Raman principal component feature vectors of the same sample are concatenated to generate a fused feature dataset. The specific method of feature fusion is as follows: the infrared and Raman principal component feature vectors of the same sample are directly concatenated along the feature dimension (column direction). Based on the extracted infrared feature vector being 1×6 and the Raman feature vector being 1×29, the fused feature vector is 1×35. The fused feature vectors of all samples are arranged in the row direction, forming a 675×35 fused feature dataset.

[0061] (6) Dataset partitioning The dataset was stratified and randomly divided, with 70% serving as the model training set (473 samples) and 30% as the model test set (202 samples). The final analysis revealed that the Linear Discriminant Analysis (LDA), Random Forest (RF), and Logistic Regression (LR) models achieved 100% accuracy on the test set. Compared to analysis using a basic data fusion strategy (where only the LDA model achieved 100% accuracy on the test set), feature set data fusion further improved model accuracy while significantly reducing the dimensionality of spectral features and computational resource consumption. The basic data fusion strategy, however, directly concatenates raw infrared and Raman data to construct a large dataset without feature extraction, resulting in high dimensionality, redundancy, and increased computational cost.

[0062] 2. Model Building Multiple machine learning classification models were employed: K-Nearest Neighbors (KNN), Linear Discriminant Analysis (LDA), Random Forest (RF), Logistic Regression (LR), Support Vector Machine (SVM), Gradient Boosting (XGBoost), Gaussian Naive Bayes (GNB), and Decision Tree Analysis (DT); a total of 8 machine learning algorithms across 6 categories. The average classification accuracy of the 5-fold cross-validation strategy was used as the evaluation metric for training and hyperparameter optimization of the machine learning classification model. The stratified, randomly partitioned training set (473 samples) was randomly divided into 5 subsets. Four subsets were selected as the training subset and one subset as the validation subset in each iteration. This process was repeated 5 times, ensuring each subset served as a validation subset once. The average accuracy of the 5 validation iterations was taken as the model's validation accuracy for that parameter combination. A random network search was used to traverse different parameter combinations, and the parameter combination with the highest average accuracy of 5-fold cross-validation was selected as the optimal parameters for the model. Finally, the model was trained using the entire training set based on the optimal parameters to obtain the final model.

[0063] Model structure and learning mechanism: (1) Random Forest (RF): In this invention, the random forest is a parallel tree-based ensemble learning model. The overall structure is a multi-decision tree parallel structure without hierarchical relationships. The core classification unit consists of 50-300 independently trained classification decision trees. There is no information interaction between individual decision trees; they are trained and predicted independently. The final classification result is determined by an absolute majority voting mechanism based on the prediction results of all decision trees (i.e., the category with the highest percentage of votes is the final predicted category; if the votes are tied, the category with the highest weight during the decision tree training process is selected). The classification errors of individual decision trees are mutually offset and weakened through the ensemble effect, making the overall model's generalization ability and anti-overfitting ability significantly better than that of individual decision trees, thus meeting the classification requirements of high-dimensional features fused from multispectral fusion of traditional Chinese medicine polysaccharides. The specific structure and operating logic of the model are as follows (including basic default parameters and the optimization range of this invention): 1) Sample Sampling: The default method is Bootstrap sampling, which involves extracting a subset of samples with replacement from the original training set that is the same size as the original dataset. These subsets are used for training each decision tree, and the samples that are not extracted can be used for model evaluation without an additional validation set. In order to optimize model performance, the Bootstrap sampling switch status (on / off) is included in the hyperparameter optimization range, and the effects of different sampling methods are compared simultaneously.

[0064] 2) Feature selection: By default, when splitting a node, each decision tree randomly selects a feature with the number of square roots (sqrt) from all features as a candidate splitting feature to avoid a single tree over-relying on strong features and improve the diversity of the ensemble. In this invention, the feature selection strategy is extended to include the three methods of square root (sqrt), log (log2), and all features (None) in the hyperparameter optimization range, and the optimal strategy that is suitable for the target scenario is determined through optimization.

[0065] 3) Decision tree growth: The basic default is that there is no maximum depth limit for a single decision tree (max_depth=None), the minimum number of samples for node splitting (min_samples_split) is 2, and the minimum number of samples for leaf nodes (min_samples_leaf) is 1, until the leaf node samples are of pure class or there are not enough samples to split; In this invention, optimization ranges are set for the above key parameters, where the maximum tree depth (max_depth) is set to 5~20 or unlimited, the minimum number of samples for node splitting (min_samples_split) is set to 2~15, and the minimum number of samples for leaf nodes (min_samples_leaf) is set to 1~8, which is used to control model complexity and improve generalization ability.

[0066] 4) Parameter baseline: The default random forest contains 100 decision trees (n_estimators=100), which is a general setting that balances training efficiency and model performance; this invention sets the optimization range of the number of trees (n_estimators) to 50~300 to further adapt to the classification requirements of high-dimensional features of multispectral fusion of Chinese herbal polysaccharides.

[0067] This invention aims to control model complexity and improve generalization ability. It completes the optimization of the above-mentioned multi-dimensional hyperparameters through 50 iterations of random search, and finally determines the optimal hyperparameter combination that is suitable for the classification of high-dimensional features of multispectral fusion of Chinese herbal polysaccharides.

[0068] RF ensemble learning uses majority voting from multiple independent decision trees for classification, and Bootstrap sampling combined with random feature selection reduces overfitting. Parameter optimization aims to control complexity and improve generalization ability: the number of trees (n_estimators) is set to 50~300, the maximum tree depth (max_depth) is set to 5~20 or unlimited, the minimum number of samples for node splits (min_samples_split) is set to 2~15, the minimum number of samples for leaf nodes (min_samples_leaf) is set to 1~8, and the number of features for each split (max_features) adopts three strategies: square root (sqrt), logarithm (log2), and full (None), and compares the bootstrap sampling on / off state; multi-dimensional hyperparameter optimization is completed through 50 iterations of random search.

[0069] Parameter settings for the classification of polysaccharides in traditional Chinese medicine In this invention, the parameter combination is optimized according to the difficulty of the identification task. For different types of polysaccharide classification tasks, 14 different Chinese medicine polysaccharides need to be completely classified, that is, divided into 14 categories. The classification labels in this embodiment are the pinyin abbreviations of the 14 Chinese medicine polysaccharides (RS, FL, JG, etc.).

[0070] 1) Modify the feature selection to log2 (logarithm to base 2). After PCA dimensionality reduction, the data still retains dozens of principal component features. The number of candidate features selected by log2 is less than that of sqrt, which can further focus on core features, filter redundant features corresponding to spectral noise, reduce feature dependence of individual trees, and improve ensemble diversity. 2) Increase the number of decision trees (n_estimators) to 300. Different polysaccharides have high dimensionality of spectral features. More decision trees can reduce ensemble variance, avoid misjudgments caused by local spectral noise in a single tree, and enhance the model's ability to capture complex feature patterns. 3) Increase the node splitting and leaf node thresholds (min_samples_split=10, min_samples_leaf=4), while keeping max_depth=None. The feature patterns of different categories are complex, and no depth limit ensures the model captures multi-level features; while increasing the splitting and leaf node thresholds limits the decision tree from overlearning irrelevant features such as spectral baseline fluctuations and instrument noise, avoiding overfitting. Compared to the default parameters, the model accuracy increased by 0.42%.

[0071] (2) Linear Discriminant Analysis (LDA): Linear Discriminant Analysis (LDA) solves the classification problem of high-dimensional spectral data by finding the optimal projection direction that minimizes within-class variance and maximizes between-class variance. Parameter optimization focuses on the solver and shrinkage coefficient: the solver is selected from Singular Value Decomposition (SVD), Least Squares QR Decomposition (LSQR), and Eigenvalue Decomposition (Eigen) to adapt to different matrix decomposition scenarios; the shrinkage coefficient is set to None, Auto, and gradient values ​​from 0.1 to 0.9, which alleviates the matrix singularity problem under high-dimensional small sample conditions through covariance matrix regularization, thus improving model stability. Hyperparameter optimization is completed using 20 iterations of random search.

[0072] In this invention, linear discriminant analysis is a supervised linear dimensionality reduction and classification integrated model. Its core structure consists of a feature projection layer and a class discrimination layer. This two-layer structure achieves dimensionality reduction and classification of high-dimensional multispectral fusion features of traditional Chinese medicine polysaccharides, solving the matrix singularity problem of spectral data under high-dimensional small sample conditions. It also adapts to the linear classification requirements of multispectral fusion features. The specific structure and operating logic are as follows: 1) Feature projection layer: The core is to find the optimal linear projection matrix W, which is p×k dimension (p is the fusion feature dimension m+n, k is the feature dimension after dimensionality reduction, k≤number of categories-1). The solution of the projection matrix has the dual objectives of minimizing intra-class variance and maximizing inter-class variance, so that after the high-dimensional fusion features are projected, the feature points of the same type of Chinese herbal polysaccharide samples have the strongest clustering and the feature points of the different types of samples have the greatest separation.

[0073] The specific solution process is as follows: First, calculate the intra-class scatter matrix Sw and inter-class scatter matrix Sb of each category of samples in the fusion feature training set. Then, by solving the generalized eigenvalue problem SbW=λSwW, obtain the eigenvectors corresponding to the eigenvalues ​​λ. Select the eigenvectors corresponding to the first k largest eigenvalues ​​to form the projection matrix W. Finally, map the high-dimensional fusion feature X (n×p) to the low-dimensional discriminant feature X' (n×k) through X'=XW.

[0074] 2) Classification Layer: Based on the dimensionality-reduced discriminant features X', a Bayesian linear discriminant function is constructed, and an independent discriminant function is established for each category of traditional Chinese medicine polysaccharides: ; Where, μ i S is the low-dimensional feature mean of the i-th class of samples. i Let P(ω) be the low-dimensional feature variance of the i-th class of samples. i Let be the prior probability of the i-th class sample. Substitute the low-dimensional discriminant features of the sample to be tested into all class discriminant functions, and select the class with the largest function value as the predicted class of the sample to be tested.

[0075] To address the issues of high-dimensional small samples and matrix singularity in multispectral fusion features of traditional Chinese medicine polysaccharides, hyperparameter optimization and structural improvements were implemented in the model: the Eigen eigenvalue solver was selected, and covariance matrix regularization was combined with the addition of a small regularization term αI (I being the identity matrix) to the intra-class scatter matrix Sw, transforming the singular matrix into a non-singular matrix and resolving the problem of matrix invertibility under high-dimensional features; the shrinkage coefficient was set to 0.1–0.9, and the intra-class scatter matrix Sw was corrected through shrinkage estimation, reducing the impact of high-dimensional spectral noise on matrix solving and improving the stability of the projection matrix W and the model's generalization ability.

[0076] (3) Logistic Regression (LR): Logistic Regression (LR) maps the linear regression results to the 0-1 interval using the Sigmoid function, determining the category based on probability values, thus adapting to linear classification of high-dimensional fused data. Parameter optimization aims to enhance regularization and ensure convergence: the regularization strength (C) is set to 0.001-100, the regularization type (penalty) includes L1 and L2, the solver uses lbfgs (quasi-Newton optimization algorithm), liblinear (coordinate descent method), and saga (stochastic average gradient descent method), and the maximum number of iterations (max_iter) is set to 2000; the optimal parameters of the linear model under high-dimensional spectral data are obtained through 30 iterations of random search.

[0077] In this invention, logistic regression is a supervised linear probabilistic classification model. Its core structure consists of a linear regression layer and a probability mapping layer. This two-layer structure maps the multispectral fusion features of high-dimensional traditional Chinese medicine polysaccharides into class probabilities, achieving accurate classification of multiple categories of traditional Chinese medicine polysaccharides. This adapts to the linear classification requirements of high-dimensional fusion features. Simultaneously, regularization optimization addresses the overfitting problem under high-dimensional data. The specific structure and operating logic are as follows: 1) Linear Regression Layer: The core is to construct a linear regression function, which performs a linear weighted summation on the m+n-dimensional multispectral fusion features X to obtain the linear predicted value z. ; Where w is the feature weight vector (dimension 1×(m+n)), each weight value corresponds to a fusion feature, representing the contribution of the spectral feature to the classification result of Chinese herbal polysaccharides, and b is the bias term, which is used to correct the systematic error of the linear model. In order to meet the multi-class classification requirements of Chinese herbal polysaccharides, a one-to-many (O v R) strategy is adopted to decompose the multi-class classification problem into multiple binary classification problems, and train an independent feature weight vector w and bias term b for each Chinese herbal polysaccharide category.

[0078] 2) Probability Mapping Layer: A Sigmoid non-linear activation function is introduced to map the continuous value z output by the linear regression layer to the interval 0~1, obtaining the posterior probability P that the test sample belongs to a certain category of traditional Chinese medicine polysaccharides. .

[0079] The Sigmoid function is an S-shaped curve. When z→+∞, P→1, indicating that the probability of the sample belonging to this category is extremely high. When z→-∞, P→0, indicating that the probability of the sample belonging to this category is extremely low. For multi-class classification results, the category with the largest posterior probability P is selected as the final predicted category of the sample to be tested. The model undergoes hyperparameter optimization and convergence improvement: L1 / L2 regularization terms are introduced to constrain the feature weight vector w, constructing a regularized loss function. The regularization term penalizes excessively large feature weight values, preventing the model from overfitting noisy features in high-dimensional spectra. L1 regularization is suitable for feature selection, compressing the weights of irrelevant spectral features to 0, while L2 regularization is suitable for weight smoothing, reducing the weight proportion of individual strong spectral features and improving model stability. In this invention, L2 regularization is preferred, adapting to the complementary characteristics of multispectral fusion features. The regularization strength C is the reciprocal of λ (C=1 / λ), ranging from 0.001 to 100. A larger C value indicates weaker regularization constraints and a higher degree of model fit to the training set, while a smaller C value indicates stronger regularization constraints and a stronger generalization ability. In this invention, the optimal regularization strength C=10, achieving the best balance between fitting accuracy and generalization ability. The number of model iterations (max_iter) was increased to 1000, and the optimal solution of the loss function was obtained by using the stochastic gradient descent (SGD) optimization algorithm. Through multiple rounds of iteration, the feature weight vector w was made to converge to a stable value, avoiding the model non-convergence problem caused by insufficient iterations under high-dimensional data, and ensuring the classification accuracy of the model on the multispectral fusion features of Chinese herbal polysaccharides.

[0080] (4) The K-Nearest Neighbors (KNN) algorithm is a local classification algorithm based on distance metrics. It determines the category of the sample to be tested by voting on the category of the sample's neighborhood, and adapts to the local feature patterns of spectral data. For the hyperparameter optimization of the KNN model for spectral data, the candidate parameters are the number of nearest neighbors 3, 5, 7, 9, 11, equal weight / distance weighting strategy, and Euclidean distance / Manhattan distance. The balanced distribution of samples is ensured by stratified sampling, and the hyperparameter optimization is completed by 20 iterations of random search. The adaptation effect of different distance metrics on spectral features is verified.

[0081] (5) Support Vector Machine (SVM) maps low-dimensional spectral data to high-dimensional data through kernel functions to find the optimal classification hyperplane and adapt to the nonlinear distribution of spectral data. Parameter optimization aims to balance the fitting and generalization ability of SVM: the penalty coefficient (C) is set to 0.001~100; the kernel function width (gamma) is set to scale / auto (automatic scaling) and 0.001~1 (manual gradient), and RBF (radial basis function) is selected to adapt to the nonlinear distribution of spectral data; the optimal combination of C and gamma is explored through 30 iterations of random search to solve the parameter sensitivity problem of SVM.

[0082] (6) Gradient Boosting (XGBoost) Gradient boosting ensemble learning is a type of serial decision tree ensemble. It iteratively trains multiple decision trees, with each tree learning the residual / error of the previous tree, gradually accumulating and optimizing the model loss. It has built-in regularization, pruning, and automatic handling of missing values, balancing fitting accuracy and generalization ability, and can efficiently handle nonlinear and complex classification tasks of high-dimensional spectral data. Parameter optimization aims at lightweight and strong regularization: the maximum tree depth (max_depth) is set to 3, 5, and 7, the learning rate (learning_rate) is set to 0.01~0.2, and the sample sampling ratio (subsample) and feature sampling ratio (colsample_bytree) are set to 0.6~1.0. Gradient search is performed on the split gain threshold (gamma), L1 regularization (reg_alpha), and L2 regularization (reg_lambda), and the number of trees (n_estimators) is controlled between 50 and 150. Parameter optimization is completed through 25 random searches, improving the model's generalization ability while ensuring training efficiency.

[0083] (7) Gaussian Naive Bayes (GNB) is based on Bayes' theorem and the feature independence assumption. It assumes that the spectral features follow a Gaussian distribution and calculates the posterior probability to achieve classification. To address the problem of minimal variance in spectral bands, the variance smoothing coefficient (var_smoothing) is optimized to 10. -9 ~10 -5 To avoid zero-variance calculation errors, optimization was completed with 20 random searches, and Naive Bayes was used as the benchmark model to verify the effectiveness of complex models.

[0084] (8) Decision Tree Analysis (DT) constructs a tree model by recursively partitioning the feature space, using information gain / Gini coefficient as the splitting criterion to adapt to the hierarchical partitioning of spectral features. Parameter optimization focuses on solving the overfitting of a single decision tree: the splitting criteria are gini (Gini coefficient) and entropy (information entropy), and the gradient range of the maximum tree depth, node splitting, and minimum number of samples in leaf nodes is consistent with RF. The number of features for each split adopts three strategies: sqrt, log2, and None. Optimization is achieved through 30 random searches, which serve as a benchmark model to verify the advantages of the ensemble method and determine the optimal complexity of a single tree.

[0085] 3. Model Performance Evaluation A classification model was built on the training set, and its performance was evaluated on the test set. In this example, the model performance evaluation samples consisted of 675 spectra of 14 kinds of Chinese herbal polysaccharides: ginseng (RS), poria cocos (FL), platycodon grandiflorus (JG), astragalus membranaceus (HQ), ophiopogon japonicus (MD), gastrodia elata (TM), codonopsis pilosula (TZS), angelica dahurica (BZ), codonopsis pilosula (DS), polygonatum sibiricum (HJ), wolfberry (GQZ), American ginseng (XYS), yam (SY), and polygonatum odoratum (YZ). Each type of polysaccharide had 45 spectra, which were randomly divided into a training set of 473 spectra (30-34 spectra of each of the 14 polysaccharides) and a test set of 202 spectra (14-15 spectra of each of the 14 polysaccharides) using a 7:3 stratified random partitioning. After preprocessing and feature extraction of the spectral data, the models were trained, tested, and analyzed. Ultimately, the identification accuracy of the RF, LDA, and LR models all reached 100%, and all 14 Chinese herbal polysaccharides were completely classified by the models. Compared with single-spectral modeling, the infrared spectroscopy model achieved an accuracy of 95.1%, but only 7 kinds of Chinese herbal polysaccharides (Angelica dahurica, Poria cocos, Astragalus membranaceus, Platycodon grandiflorus, Dioscorea opposita, Pseudostellaria heterophylla, and Polygonatum odoratum) were completely classified, while the remaining 7 kinds of Chinese herbal polysaccharides were not completely classified, with identification accuracy ranging from 78.6% to 92.9%. The Raman spectroscopy model achieved an accuracy of 99.0%, and 12 kinds of Chinese herbal polysaccharides (Codonopsis pilosula, Poria cocos, Lycium barbarum, Polygonatum sibiricum, Astragalus membranaceus, Platycodon grandiflorus, Ophiopogon japonicus, Panax ginseng, Dioscorea opposita, Gastrodia elata, Pseudostellaria heterophylla, and Polygonatum odoratum) were completely classified, but Angelica dahurica and Panax quinquefolius polysaccharides were not completely classified, with identification accuracy of 92.3% for both. The multispectral fusion strategy improved the model performance. Figure 8 The confusion matrix of the optimal model is shown, demonstrating that all test samples were accurately classified.

[0086] Model Evaluation: By combining infrared and Raman spectroscopy with machine learning, the model achieved accurate classification of different Chinese herbal polysaccharides on the test set. The test set accuracy of the RF, LDA, and LR models all reached 100%. The accuracy, precision, recall, F1 score, and Matthews correlation coefficient (MCC) of the models were all 1, indicating that the models have excellent discrimination ability and reliability.

[0087] 4. Model Analysis To reveal the basis for model identification, this invention employs the SHAP (SHapley Additive exPlanations) framework to quantify and analyze the decision-making mechanism of the optimal model. The core principle is based on the Shapley value principle in game theory, quantifying the contribution of each input feature (such as infrared / Raman principal component features) to the model's prediction results, thereby revealing the model's core identification criteria. The specific algorithm, combined with the code implementation logic, is as follows: SHAP analysis, based on conditional expectation decomposition and the Shapley value fair allocation criterion, decomposes the model output f(x) into: ; in, =E[f(X)] (the model's average prediction of the background dataset, i.e., the baseline value); Let be the SHAP value of the i-th feature, representing the contribution of this feature to the prediction result relative to the baseline value (positive contribution increases the prediction value, negative contribution decreases the prediction value); M is the feature dimension (e.g., the number of principal components). Differentiated SHAP interpreters (TreeExplainer / LinearExplainer / KernelExplainer) were adapted for different model types (tree model / linear model / general model), and the background dataset sampling strategy was optimized to address the sensitivity of TreeExplainer to background data. The specific steps are: loading the trained optimal classification model and fused feature dataset; constructing the corresponding SHAP interpreter according to the model type, selecting the training set as the background baseline data; calculating the SHAP value of each spectral feature to the prediction result, quantifying the feature contribution; ranking the features based on the average absolute SHAP value; mapping the core contributing features to the original spectral wavenumber interval to determine the model identification criteria. The mapping formula is as follows: ; Where: X pc W represents the high-contribution principal component vectors; W is the PCA loading matrix (eigenvector matrix), whose column vectors are the principal component directions obtained after eigenvalue decomposition of the original data covariance matrix; W TX is the transpose of the PCA loading matrix, used in the inverse transform to project the principal component scores back to the original spectral space; μ is the original data mean vector; X recon This is the restored original spectral signal.

[0088] like Figure 9 The SHAP beehive diagram is shown below: PC is a comprehensive spectral feature vector obtained by dimensionality reduction extraction of preprocessed Chinese medicine polysaccharide spectral data through the PCA algorithm. It is a linear combination of all wavenumber points of the original spectrum. Its core function is to eliminate data redundancy and noise while preserving the original spectral structure information to the maximum extent, and to extract the core spectral features that can characterize the structural differences of different Chinese medicine polysaccharides. The specific calculation method is calculated according to the formulas in the feature extraction and data partitioning section above.

[0089] After spectral feature fusion, the SHAP bee colony diagram shows: The principal components PC3, PC5, and PC2 of the infrared spectrum are the core features for model decision-making (they are orthogonal projections onto the eigenvectors corresponding to the 3rd, 5th, and 2nd largest eigenvalues ​​of the infrared covariance matrix, respectively, with variance contribution rates of 5.57%, 2.45%, and 20.50%, and are linear combinations of the 581 wavenumber points of the original infrared spectrum). PC3 mainly interprets the spectral range of 1000–1200 cm⁻¹. -1 Wavenumber range, PC5 mainly interprets 1700~1800cm -1 Wavenumber range, PC2 mainly interprets 900~1000 cm -1 The load diagram for the wavenumber range (peak positive load > 0.05) is shown below. Figure 10 In point A, the SHAP bee colony diagram reveals that the values ​​of these three infrared principal components are stably negatively correlated with the model predictions. The principal components PC3, PC4, and PC2 of the Raman spectrum are the core features for model decision-making (they are orthogonal projections onto the eigenvectors corresponding to the 3rd, 4th, and 2nd largest eigenvalues ​​of the Raman covariance matrix, respectively, with variance contribution rates of 7.86%, 7.45%, and 14.88%, respectively, and are linear combinations of the 673 wavenumber points of the original Raman spectrum). PC3 mainly interprets the 500–550 cm⁻¹ range. -1 600~700 cm -1 and 800~850 cm -1 Wavenumber range, PC4 mainly interprets 450~500 cm -1 800~850 cm -1 and 1100~1150 cm -1 Interval wavenumber and PC2 primarily interpret the 500–600 cm⁻¹ range. -1 650~700 cm -1 and 850~950 cm -1The load diagram for interval wavenumbers (positive load > 0.05) is shown below. Figure 10 By using the SHAP beehive diagram, it can be found that the high values ​​of these three Raman principal components positively drive the prediction results. The principal components of the two types of spectra form complementary decision logics and jointly determine the classification prediction results of the model.

[0090] At the same time, we further obtained key spectral features that contribute significantly to the classification results, such as Figure 11 The infrared and Raman wavenumber spectra, obtained after SHAP analysis of the three optimal models (RF, LDA, and LR), respectively, represent the importance of the classification features. Analysis was conducted on the core wavenumbers covered by all three models, specifically the infrared spectrum from 1780 to 1795 cm⁻¹. -1 The region is attributed to the C=O stretching vibration, a characteristic peak of the carboxyl group (-COOH) of uronic acid residues in polysaccharides; 1100~1200 cm⁻¹ -1 The region is attributed to the stretching vibration of the COC glycosidic bond and the bending vibration of the COH glycosidic bond, which are characteristic peaks of polysaccharide glycosidic bond types; while the Raman peaks at 860~861cm... -1 The region is attributed to skeletal vibrations of the pyranose ring and is a characteristic peak in polysaccharide-monosaccharide composition; 1100–1150 cm⁻¹ -1 The region is attributed to COC glycosidic bond vibrations. Focusing on the SHAP analysis of the three models, the aforementioned infrared and Raman core wavenumbers correspond to differences in polysaccharide structure (uronic acid content, monosaccharide composition, and glycosidic bond type). The limitations of a single spectrum allow for polysaccharide identification based on only partial differences. However, a spectral fusion strategy complements and refines this information, improving the accuracy of the model analysis. Furthermore, the key wavenumbers clarify the chemical mechanism of the model classification.

[0091] Example 2. Model for Identifying the Origin of Traditional Chinese Medicine In this embodiment, infrared spectroscopy and Raman spectroscopy are further fused and combined with machine learning to identify the origin of traditional Chinese medicine at the polysaccharide level. The specific steps are as follows: Experimental materials: The Chinese medicinal materials used in the experiment were ginseng from three different origins, sourced from Jiangyin Tianjiang Pharmaceutical Co., Ltd., namely Jilin (6 batches), Liaoning (5 batches), and Heilongjiang (3 batches). Detailed information is shown in Table 2. Polysaccharide extraction was performed on the ginseng from the above three different origins to obtain ginseng polysaccharide samples from different origins, which were named according to the pinyin abbreviations of the above place names (JL, LN, HLJ, respectively).

[0092] Extraction method of ginseng polysaccharides: Weigh 100g of sliced ​​ginseng and place it in a round-bottom flask. Add 15 times the volume of ultrapure water and reflux at 100℃ for 1.5h. Extract twice and combine the extracts. Centrifuge at 3000r / min for 15min and collect the supernatant. Concentrate by rotary evaporation to 1mg / mL. Add 95% ethanol until the ethanol volume reaches 80%. Let stand at 4℃ for 12h. Centrifuge at 3000r / min for 10min and collect the precipitate. Wash three times with anhydrous ethanol. After evaporating the residual ethanol from the precipitate, dissolve it in ultrapure water. Further dialyze using a 1000Da dialysis bag for 24h. Then freeze-dry to obtain different Chinese herbal polysaccharide powders. The yield and polysaccharide content (determined by phenol-sulfuric acid method) of ginseng polysaccharides from different origins are shown in Table 4. The polysaccharides appear as white cotton-like particles and have good water solubility.

[0093] Table 4. Information on ginseng medicinal materials from different producing areas .

[0094] Infrared spectral acquisition: The same infrared spectral acquisition method as in Example 1 was used. 45 infrared spectra were collected from Heilongjiang, 60 from Jilin, and 50 from Liaoning, resulting in a total of 155 standard infrared spectral data points. Figure 12 The average infrared spectrum of ginseng polysaccharides from different origins.

[0095] Raman spectroscopy acquisition: The same Raman spectroscopy acquisition method as in Example 1 was used. 45 spectra were collected from Heilongjiang, 60 from Jilin, and 50 from Liaoning, resulting in a total of 155 Raman spectral standard data. Figure 13 The average Raman spectra of ginseng polysaccharides from different origins.

[0096] Infrared and Raman spectral data preprocessing: The same infrared and Raman spectral data preprocessing methods as in Example 1 were used. In this example, the samples for model performance evaluation consisted of 155 spectra of ginseng polysaccharides from three different origins: Jilin (JL), Liaoning (LN), and Heilongjiang (HLJ). There were 60 spectra from Jilin, 50 from Liaoning, and 45 from Heilongjiang. These spectra were randomly divided in a 7:3 stratified ratio into a training set of 108 spectra (42, 35, and 31 spectra from each of the three origins) and a test set of 47 spectra. Spectra (18, 15, and 14 spectra of polysaccharides from 3 origins each) were used. After only SG smoothing, baseline correction, and normalization, the highest model accuracy of the spectral fusion samples in the test set was 91.5%, with identification accuracies of 94.4%, 93.3%, and 85.7% for Jilin, Liaoning, and Heilongjiang, respectively. After MSC correction, the highest model accuracy in the test set increased to 95.7%, with identification accuracies of 100%, 100%, and 85.7% for Jilin, Liaoning, and Heilongjiang, respectively.

[0097] Data partitioning and feature extraction: The same data partitioning and feature extraction methods as in Example 1 were used. Seven infrared principal component (PC) features were finally extracted. The variance contribution rate of infrared PC1 was 34.99%, mainly explaining the variance of the 700–900 cm⁻¹ region. -1 Wavenumber interval (loadings > 0.0633); the principal component PC2 variance contribution rate was 25.21%, mainly explaining the 950–1000 cm⁻¹ range. -1 1700–1800cm -1 Wavenumber interval (loadings > 0.0692); Principal component PC3 variance contribution rate was 13.48%, mainly explaining the 1200–1300 cm⁻¹ range. -1 1700–1800 cm -1 Wavenumber range (load value > 0.0803); Infrared PC4 variance contribution rate is 10.59%, mainly explaining the 1300–1400 cm⁻¹ range. -1 Wavenumber range (load value > 0.0695); Infrared PC5 variance contribution rate is 6.05%, mainly explaining the 1100–1200 cm⁻¹ range. -1 Wavenumber range (load value > 0.0720); Infrared PC6 variance contribution rate is 3.01%, mainly explaining the 1200–1400 cm⁻¹ range. -1 Wavenumber range (load value > 0.0696); the infrared PC7 variance contribution rate is 2.23%, mainly explaining the 900–1000 cm⁻¹ range. -1 1650–1750 cm -1 Wavenumber intervals (loading values ​​> 0.0649), more detailed interpretation of wavenumbers is shown in Table 5. The cumulative variance contribution rate of the 7 infrared principal components is 96.1%, and the PCA variance interpretation plot is shown below. Figure 14 Fifteen Raman principal component (PC) features were extracted. (Taking the first 10 as an example), the Raman PC1 variance contribution rate was 42.33%, mainly explaining the variance of 1200–1400 cm⁻¹. -1 In the wavenumber range (loading > 0.0531), the Raman PC2 variance contribution rate was 24.66%, primarily explaining the 900–1100 cm⁻¹ range. -1 Wavenumber range (loading > 0.0612); Raman PC3 variance contribution rate was 7.15%, mainly explaining the 750–850 cm⁻¹ range. -1 Wavenumber range (loading > 0.0605); Raman PC4 variance contribution rate was 6.58%, mainly explaining the 400–550 cm⁻¹ range. -1 1050–1100 cm -1 Wavenumber range (loading > 0.0648); Raman PC5 variance contribution rate was 3.02%, mainly explaining the 850–900 cm⁻¹ range. -1 1100–1200 cm -1Wavenumber range (loading > 0.0611); Raman PC6 variance contribution rate was 2.47%, mainly explaining the 400–600 cm⁻¹ range. -1 800–850 cm -1 1150–1200 cm -1 Wavenumber range (loading > 0.0623); Raman PC7 variance contribution rate was 1.84%, mainly explaining the 450–500 cm⁻¹ range. -1 1450–1550 cm -1 Wavenumber range (loading > 0.0665); Raman PC8 variance contribution rate was 1.68%, mainly explaining the 400–500 cm⁻¹ range. -1 800–900 cm -1 Wavenumber range (loading > 0.0646); Raman PC9 variance contribution rate was 1.39%, mainly explaining the 400–450 cm⁻¹ range. -1 1100–1200 cm -1 Wavenumber range (loading > 0.0606); Raman PC10 variance contribution rate was 1.14%, mainly explaining the 400–550 cm⁻¹ range. -1 1100–1200 cm -1 Wavenumber interval (loading value > 0.0634); a more detailed interpretation of the wavenumbers is shown in Table 6. The cumulative variance contribution rate of the 15 Raman principal component features is 95.3%. The PCA variance interpretation plot is shown below. Figure 15 The extracted infrared and Raman principal component features were fused to obtain a fused dataset, which was used for subsequent model analysis. The highest accuracy of the model test set was 95.7%. Compared with the highest accuracy of 93.1% of the model test set obtained by directly fusing primary spectral data, feature-level fusion further improved the model accuracy.

[0098] Table 5. Interpreted wavenumbers of principal components in infrared spectroscopy for polysaccharides from different origins .

[0099] Table 6. Explanation wavenumbers of principal components of polysaccharides from different origins using Raman spectroscopy. .

[0100] Model building: The same model building method as in Example 1 was used.

[0101] Parameter settings for the task of classifying the origin of polysaccharides from traditional Chinese medicine The task involved classifying ginseng polysaccharides from three different origins: Heilongjiang, Jilin, and Liaoning. 1) Similarly, modify the feature selection to log2 (logarithm to base 2) for feature selection; 2) Adjust the number of decision trees (n_estimators) to 300. The differences in origin are subtle and the feature patterns are relatively simple. 300 decision trees ensure the ability to identify subtle features. 3) Moderately increase the split threshold (min_samples_split=5), the minimum number of samples per leaf node (min_samples_leaf=2), and strictly limit the maximum depth (max_depth=20), and enable bootstrap sampling. The feature signals from geographical differences are weak; retaining the default leaf node values ​​prevents key subtle features from being merged. Limiting the maximum depth prevents the model from learning from irrelevant noise such as instrument fluctuations and environmental interference. Compared to the default parameters, the model accuracy increased by 2.86%.

[0102] After spectral fusion of infrared and Raman spectroscopy, the accuracy of the model on the test set was improved compared to the infrared spectroscopy model alone (91.0% accuracy, 88.9% accuracy in Jilin, 85.7% accuracy in Liaoning, and 100% accuracy in Heilongjiang) and the Raman spectroscopy model alone (89.1% accuracy, 100% accuracy in Jilin, 86.7% accuracy in Liaoning, and 78.6% accuracy in Heilongjiang). The overall accuracy reached 95.7%, validating the effectiveness of feature fusion. The confusion matrix of the optimal model is as follows: Figure 16 As shown, although complete identification of polysaccharides from all three producing regions was not achieved, 100% accurate identification of ginseng polysaccharides from Jilin and Liaoning was achieved. During dataset construction, ginseng polysaccharides from Jilin, Liaoning, and Heilongjiang were clearly distinguished as independent classification labels. However, some samples of ginseng polysaccharides from Heilongjiang were misclassified as originating from Jilin. This phenomenon is attributed to the relatively small sample size of Heilongjiang samples in the current dataset. Increasing the sample size in the future is expected to further improve the model's generalization ability.

[0103] Model Evaluation: The optimal model for identifying ginseng from different origins was determined by spectral fusion of polysaccharide infrared spectroscopy and Raman spectroscopy combined with machine learning. The accuracy, precision, recall, F1 score, and Matthews correlation coefficient (MCC) of this optimal model on the test set were 95.74%, 96.17%, 95.74%, 95.69%, and 93.80%, respectively, indicating that the model has good discrimination performance and reliability.

[0104] Model Analysis: SHAP was also used to analyze the origin identification model to clarify its decision-making basis. For example... Figure 17As shown in the SHAP beehive diagram, after spectral feature fusion, the infrared spectral model is based on principal components PC5, PC6, and PC2 (orthogonal projections onto the eigenvectors corresponding to the 5th, 6th, and 2nd largest eigenvalues ​​of the infrared covariance matrix, respectively, with variance contribution rates of 6.05%, 3.01%, and 25.21%, respectively, representing a linear combination of the 581 wavenumber points of the original infrared spectrum). PC5 mainly interprets the spectral range of 1100–1200 cm⁻¹. -1 Wavenumber range, PC6 mainly interprets 1200~1400 cm -1 Wavenumber range, PC2 mainly interprets 950~1000cm -1 1700~1800 cm -1 The load diagram for the wavenumber range (peak positive load > 0.05) is shown below. Figure 18 In model A, the Raman spectroscopy model is centered on principal components PC2, PC11, and PC13 (orthogonal projections onto the eigenvectors corresponding to the 2nd, 11th, and 13th largest eigenvalues ​​of the Raman covariance matrix, respectively, with variance contribution rates of 24.66%, 0.79%, and 0.49%, respectively, representing a linear combination of the 673 wavenumber points of the original Raman spectrum). PC2 primarily interprets the spectral range of 900–1100 cm⁻¹. -1 Wavenumber range, PC11 mainly interprets 400~500 cm -1 500~550cm -1 900~950 cm -1 and 1150~1200 cm -1 Wavenumber range, PC13 mainly interprets 500~550 cm -1 750~800 cm -1 1100~1150 cm -1 and 1350~1400 cm -1 The load diagram for the wavenumber range (peak positive load > 0.05) is shown below. Figure 18 In point B, the SHAP beehive diagram reveals that the values ​​of infrared PC5 and PC2, and Raman PC11 and PC13, are stably negatively correlated with the model predictions, while high values ​​of infrared PC6 and Raman PC2 positively drive the prediction results. The fusion of the two types of spectra, through the positive and negative synergistic regulation of different principal components, demonstrates complementarity, further validating the scientific value of spectral fusion. Further analysis of the key wavenumber spectra of the optimal model RF classification, such as... Figure 19 The infrared value is 1676 cm. -1 The region is attributed to the hydrogen bond vibration of the carboxyl group -COOH of the uronic acid residue; 1122 cm -1 The region is attributed to COC glycosidic bond vibrations; for the Raman spectrum at 466 cm⁻¹ -1 The region is attributed to the vibration of the pyranose ring; 801–876 cm⁻¹-1 The region is attributed to respiratory and deformation vibrations of the sugar ring; 1128–1154 cm -1 The region is attributed to the stretching vibration of COC glycosidic bonds. Compared to different types of Chinese herbal polysaccharides, the structural differences between polysaccharides from different origins of the same Chinese herbal medicine are subtle. The important wavenumbers of the infrared and Raman classification features mentioned above are the vibrations of polysaccharide rings and glycosidic bonds. These subtle differences are effectively captured by multispectral fusion and machine learning models and transformed into a basis for determining the origin.

[0105] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific implementations of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles and spirit disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for identifying polysaccharides in traditional Chinese medicine based on multispectral fusion and machine learning, characterized in that, Includes the following steps: (1) Obtain standard spectral data of polysaccharide extracts from traditional Chinese medicine samples, wherein the standard spectral data includes infrared spectral data and Raman spectral data; (2) The infrared spectral data and Raman spectral data are preprocessed and feature extracted respectively to obtain the corresponding infrared feature subsets and Raman feature subsets; (3) The infrared feature subset and the Raman feature subset are fused to generate a multispectral fusion feature set, and the dataset is divided. (4) Based on the multispectral fusion features, the machine learning classification model is trained and the hyperparameters are optimized using the 5-fold cross-validation strategy to obtain the identification model of Chinese herbal polysaccharides; (5) Use the polysaccharide identification model of traditional Chinese medicine obtained in step (4) to identify the multispectral fusion feature set of the polysaccharide extract of the drug sample to be tested.

2. The method according to claim 1, characterized in that, The preprocessing in step (2) includes one or more of the following operations: Savitzky-Golay smoothing, baseline correction, normalization, and multivariate scattering correction; the feature extraction adopts the principal component analysis method, and the data dimensionality is reduced and representative spectral features are extracted with the cumulative variance contribution rate not less than 95% as the standard.

3. The method according to claim 1, characterized in that, The data fusion in step (3) is a feature-level fusion. By splicing the infrared feature vector and the Raman feature vector of the sample, the fusion feature vector of the sample is formed, thereby forming a fusion feature dataset. The feature fusion of the infrared spectral data and the Raman spectral data is a complementary fusion. The infrared feature subset reflects the vibration information of the molecular groups of traditional Chinese medicine polysaccharides, and the Raman feature subset reflects the vibration information of the molecular skeleton of traditional Chinese medicine polysaccharides.

4. The method according to claim 1, characterized in that, The machine learning classification model used in step (4) is selected from one of the following: K-nearest neighbors algorithm, linear discriminant analysis, random forest, logistic regression, support vector machine, gradient boosting, Gaussian Naive Bayes, decision tree analysis; the hyperparameter optimization is achieved through random network search optimization method.

5. The method according to claim 4, characterized in that, The machine learning classification model is a random forest classification model, which consists of 50 to 300 independently trained decision trees. The final classification result is determined by a majority voting mechanism. Sample sampling adopts the bootstrap sampling method. When splitting a node, each decision tree randomly selects features with the square root of the total number of features from all features as candidate split features. There is no maximum depth limit for a single decision tree. The minimum number of samples for node splitting is 2, and the minimum number of samples for leaf nodes is 1, until the leaf node samples are of pure class or there are not enough samples to split.

6. The method according to claim 5, characterized in that, When the identification of the Chinese herbal polysaccharide is type identification, the feature selection in the random forest classification model is log2, the decision tree is increased to 300, the node split threshold is increased to 10, the leaf node threshold is 4, and the maximum depth is None. Alternatively, when the identification of the Chinese herbal polysaccharide is place of origin identification, the feature selection in the random forest classification model is log2, the decision tree is 200, the split threshold is 5, the leaf node is 1, and the maximum depth is 5.

7. The method according to claim 1, characterized in that, Step (5) also includes evaluating the performance of the traditional Chinese medicine polysaccharide identification model using stratified randomized test set data. The evaluation index includes at least one of the following: accuracy, precision, recall, F1 score, and Matthews correlation coefficient.

8. The method according to any one of claims 1-7, characterized in that, The method further includes the following steps: (6) Using an interpretable artificial intelligence analysis method based on SHAP values, the decision-making mechanism of the traditional Chinese medicine polysaccharide identification model is analyzed, the SHAP values ​​of each principal component and original wavenumber are calculated, and the high contribution wavenumbers of SHAP are matched with the known functional group vibration intervals in combination with the classical vibrational assignments of infrared and Raman spectra to determine the key wavebands.

9. A system for identifying polysaccharides in traditional Chinese medicine, characterized in that, include: The spectral acquisition module is used to acquire infrared spectral data and Raman spectral data of traditional Chinese medicine polysaccharide samples using an infrared spectrometer and a Raman spectrometer. The data processing module is used to execute steps (2)-(3) of the method as described in claim 1, and to complete the preprocessing, feature extraction and multispectral feature fusion of spectral data; The model identification module is used to store and execute the identification model of traditional Chinese medicine polysaccharides trained by the method described in claim 1, receive fused features and output sample category discrimination results; An interpretability analysis module is used to perform step (6) of the method as described in claim 8, analyze the model decision mechanism and identify key spectral feature bands.

10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the method for identifying polysaccharides of traditional Chinese medicine based on multispectral fusion and machine learning as described in claim 1.