Method for rapid classification and multi-component quantitative analysis of bighead atractylodes rhizome from different producing areas
By using near-infrared spectroscopy and chemometrics, PLS-DA and PLSR models were constructed, solving the problems of Atractylodes macrocephala origin identification and component quantification. This enabled rapid and accurate classification and multi-component quantitative analysis of Atractylodes macrocephala, which is suitable for quality control and industrial application of Chinese medicinal materials.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI UNIVERSITY OF TRADITIONAL CHINESE MEDICINE
- Filing Date
- 2025-11-28
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies make it difficult to quickly and accurately identify and quantify the chemical components of Atractylodes macrocephala from different origins, resulting in uncontrollable differences in efficacy.
Near-infrared spectroscopy combined with chemometrics was used to construct a PLS-DA model for origin classification and a PLSR model for component quantification. HPLC was used for component quantification, and IPLS characteristic wavelength screening and spectral preprocessing were combined to improve the accuracy and stability of the models.
It enables efficient identification of Atractylodes macrocephala from different origins and rapid quantitative analysis of various chemical components, providing a reference for the quality control of Atractylodes macrocephala. It is suitable for quality monitoring of batch samples and improves the efficiency of identification of origin and component analysis of Chinese medicinal materials.
Smart Images

Figure CN121978225A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Atractylodes macrocephala origin classification and component quantification technology, and in particular to a method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins. Background Technology
[0002] Atractylodes macrocephala Koidz, a plant of the Asteraceae family, is used medicinally for its dried rhizome, which has the effects of invigorating the spleen and replenishing qi, as well as drying dampness and promoting diuresis. With the increasing demand for Atractylodes macrocephala in the traditional Chinese medicine industry, and the continuous expansion and relocation of its production areas, there are significant differences in its appearance and chemical composition due to variations in climate, soil, and other natural conditions. Furthermore, the content of components in Atractylodes macrocephala varies significantly by region due to different production areas, thus affecting its efficacy. Therefore, establishing rapid and accurate methods for identifying the origin of Atractylodes macrocephala and for quantitative analysis of its chemical components is of great significance for ensuring the quality and clinical efficacy of Atractylodes macrocephala.
[0003] Traditional methods for identifying Chinese medicinal herbs include morphological examination, microscopic identification, and chemical identification. However, these methods are limited by the subjective reliance of the assessor and are not easily standardized. Modern methods often use high-performance liquid chromatography (HPLC), gas chromatography (GC), and thin-layer chromatography (TLC) to combine information such as peak shape, retention time, and peak area of chemical components in Chinese medicinal herbs, enabling the differentiation of the origin of Chinese medicinal herbs and the quantitative determination of chemical components. However, in chromatographic analysis, gradient conditions are complex to determine, sample integrity cannot be restored, and batch processing is time-consuming. In contrast, near-infrared spectroscopy (NIR) has the advantages of being rapid, green, non-destructive, and economical, and is becoming increasingly popular in the field of Chinese medicinal herbs. Near-infrared spectroscopy is a spectrum between visible light (VIS) and mid-infrared light (MIR), with a wavenumber range of approximately 12,000-4,000 cm⁻¹. -1 The NIR spectrum primarily consists of overtone and combination frequency absorptions of the XH vibrations (X=C, N, O), and can also detect overtone signals from some carbon-containing functional groups (C=O, CO), containing molecular structural information for most organic compounds. Chemometrics can transform the complex signals of NIR into interpretable qualitative and quantitative outputs. In discrimination and classification tasks, methods such as partial least squares discriminant analysis (PLS-DA) and support vector machines (SVM) are widely used for authenticity identification or origin tracing. In content prediction, partial least squares regression (PLSR) is robust to multicollinearity and is one of the commonly used regression methods for mapping chemical component concentrations from NIR spectral data. Combining NIR with chemometrics allows for the extraction of key information from complex spectra, enabling rapid identification of the origin of traditional Chinese medicine and prediction of its intrinsic components.
[0004] Existing research on Atractylodes macrocephala mainly focuses on the separation and determination of atractylone, lactone components and polysaccharides, while the determination of phenolic acids, fatty acids and other volatile components in Atractylodes macrocephala is relatively limited.
[0005] Therefore, a rapid qualitative and quantitative analysis method for Atractylodes macrocephala from different origins is needed, based on a combination of NIR and HPLC methods and chemometrics. On one hand, a PLS-DA model was established based on NIR spectral data to distinguish Atractylodes macrocephala from different origins; on the other hand, a PLSR model was used to predict the content of 11 chemical components (neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, atractylenolide III, atractylenolide II, atractylenolide I, β-eudesminol, linoleic acid, β-elemene, and oleic acid) in Atractylodes macrocephala. The model performed well, providing a reference for rapid content determination and comprehensive quality control of Atractylodes macrocephala, and offering an important tool for manufacturers, consumers, and regulatory agencies to help ensure a safe and sustainable market for Atractylodes macrocephala. Summary of the Invention
[0006] To address the aforementioned problems, the present invention aims to provide a method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins, thereby solving the problem of the lack of efficient origin identification and accurate prediction of multi-component content in existing Atractylodes macrocephala methods.
[0007] This invention provides a method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins, comprising the following steps:
[0008] S1. Collect Atractylodes macrocephala samples from multiple production areas and perform sample preprocessing;
[0009] S2. Based on HPLC testing, quantitative analysis of multiple components contained in Atractylodes macrocephala samples was performed to obtain the actual content of each component in the Atractylodes macrocephala samples;
[0010] S3. NIR spectral data of Atractylodes macrocephala samples acquired based on NIR spectroscopy;
[0011] S4. Based on the actual content of Atractylodes macrocephala samples and NIR spectral data, a training set was constructed to train the origin classification model PLS-DA and the component quantification model PLSR.
[0012] S5. Based on the trained PLS-DA and PLSR models, the origin classification and component quantitative analysis of Atractylodes macrocephala were performed.
[0013] In one embodiment of the present invention, the sample preprocessing in S1 includes:
[0014] S11. Remove impurities from the surface of the Atractylodes macrocephala sample;
[0015] S12. Place the treated Atractylodes macrocephala samples in a 35℃ environment and dry them to constant weight.
[0016] S13. Crush the dried sample with a pulverizer and pass it through a No. 3 sieve;
[0017] S14. Store the sieved Atractylodes macrocephala sample powder in a desiccator for later use.
[0018] In one embodiment of the present invention, step S2 includes the following steps:
[0019] S21. Prepare mixed standard solutions and Atractylodes macrocephala sample solutions with different components;
[0020] S22. Set HPLC detection parameters;
[0021] S23. Inject mixed standard solutions and test solutions of different components separately. Calculate the actual content of each component in the Atractylodes macrocephala sample by using the linear relationship between the concentration and peak area of the mixed standard solution, and obtain HPLC data.
[0022] In one embodiment of the present invention, step S3 includes the following steps:
[0023] S31. Set the parameters of the NIR spectrometer and collect near-infrared NIR spectral data of the Atractylodes macrocephala sample with the gold mirror background inside;
[0024] S32. Preprocess the near-infrared (NIR) spectral data to obtain the preprocessed near-infrared (NIR) spectral data.
[0025] S33. Perform characteristic wavelength screening on the preprocessed near-infrared (NIR) spectral data to obtain the screened NIR spectral data.
[0026] In one embodiment of the present invention, when training the PLS-DA model, NIR spectral data is used as the independent variable and the sample origin category is used as the dependent variable; when training the PLSR model, NIR spectral data is used as the independent variable and the actual content of each component tested by HPLC is used as the dependent variable.
[0027] In one embodiment of the present invention, the near-infrared (NIR) spectral data preprocessing method in S32 includes: SG smoothing, first derivative, second derivative, multivariate scattering correction, standard normal variable transformation and / or data centering.
[0028] In one embodiment of the present invention, the feature wavelength screening in S33 includes: dividing continuous feature variables into multiple sub-wavelength intervals based on IPLS, establishing partial least squares models of different combinations of sub-wavelength intervals and evaluating their performance, screening out the key sub-wavelength intervals that contribute the most to the model as effective features, and screening NIR spectral data maps with effective features.
[0029] In one embodiment of the present invention, qualitative and quantitative evaluation of the PLS-DA model and the PLSR model is also included.
[0030] The qualitative evaluation includes cross-validation accuracy, precision, accuracy, recall, and F1 score.
[0031] Quantitative evaluation includes residual prediction bias, corrected coefficient of determination, prediction coefficient of determination, cross-validation coefficient of determination, corrected root mean square error, prediction root mean square error, and cross-validation root mean square error.
[0032] In one embodiment of the present invention, the qualitative evaluation formula for cross-validation accuracy is expressed as:
[0033]
[0034] Where k is the cross-validation fold number, Acc i Let be the accuracy in the i-th cross-validation. This is the cross-validation accuracy.
[0035] In one embodiment of the present invention, the quantitative evaluation formula for the residual prediction bias is expressed as follows:
[0036]
[0037] Where RPD is the residual prediction bias, SD is the standard deviation of the measured values, and RMSEP is the root mean square error of the prediction.
[0038] In one embodiment of the present invention, the Atractylodes macrocephala sample components include neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, atractylodes macrocephala lactone III, atractylodes macrocephala lactone II, atractylodes macrocephala lactone I, β-eucalyptol, atractylone, linoleic acid, β-elemene and oleic acid.
[0039] The beneficial effects of this invention are:
[0040] This invention utilizes near-infrared spectroscopy combined with chemometrics to achieve efficient identification of Atractylodes macrocephala from different origins, and also constructs rapid quantitative models for various active components, demonstrating significant application potential and advantages in the quality evaluation of Atractylodes macrocephala. The method is simple, non-destructive, and rapid, serving as an effective alternative to HPLC methods. It is suitable for quality monitoring of batches of Atractylodes macrocephala samples and provides a valuable research approach for the identification of origin and component analysis of other traditional Chinese medicinal materials. Attached Figure Description
[0041] Figure 1 This is a spectral diagram of the origin of Atractylodes macrocephala in this invention;
[0042] Figure 2 This is the PLS-DA score chart of the present invention;
[0043] Figure 3 The spectra after different pretreatments according to the present invention;
[0044] Figure 4 The spectrum is the result of MSC+FD preprocessing according to this invention;
[0045] Figure 5 This invention provides IPLS characteristic wavelength screening and LVS line graphs.
[0046] Figure 6 This is a confusion matrix diagram of the training and test sets for this invention;
[0047] Figure 7 This is an HPLC result chromatogram of the mixed standard and sample of this invention;
[0048] Figure 8 This is a scatter plot of the PLSR model of 11 components of Atractylodes macrocephala in this invention;
[0049] Figure 9 This invention provides IPLSR screening bands and LVS line graphs for 11 components of Atractylodes macrocephala. Detailed Implementation
[0050] Embodiments of the present invention are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar symbols denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0051] Example 1:
[0052] This embodiment discloses a method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins. The method includes:
[0053] S1. Collect Atractylodes macrocephala samples from multiple production areas and perform sample preprocessing.
[0054] As shown in Table 1, rhizome samples of Atractylodes macrocephala were collected from 12 production areas in different provinces. They were purchased from different local pharmacies and Chinese raw material markets and were randomly numbered in sequence.
[0055] Table 1. Origin and Harvesting Schedule of Atractylodes macrocephala
[0056]
[0057] The samples were then pretreated, including removing surface impurities from the collected Atractylodes macrocephala samples and drying them at 35°C to constant weight. 26-31 samples were taken from each origin. The dried samples were then pulverized using a pulverizer, and all samples were passed through a No. 3 sieve. The sieved Atractylodes macrocephala powder was stored in a desiccator for later use.
[0058] S2. Based on HPLC testing, quantitative analysis of multiple components contained in Atractylodes macrocephala samples was performed to obtain the actual content of each component in the Atractylodes macrocephala samples.
[0059] First, 11 chemical reagents were prepared according to the main components of Atractylodes macrocephala, including neochlorogenic acid (purity ≥98%), chlorogenic acid (purity ≥98%), cryptochlorogenic acid (purity ≥98%), atractylodes lactone III (purity ≥98%), atractylodes lactone II (purity ≥98%), atractylodes lactone I (purity ≥98%), β-eucalyptol (purity ≥98%), atractylone (purity ≥98%), linoleic acid (purity ≥98%), β-elemene (purity ≥98%), and oleic acid (purity ≥98%). Also included were HPLC-grade acetonitrile, methanol, analytical reagent-grade formic acid, and ultrapure water for testing.
[0060] Using the above reagents, a mixed standard solution of 11 components was accurately weighed and prepared. The mixed standard solution was diluted to different concentrations using a serial dilution method (e.g., neochlorogenic acid 1.85-59 μg / ml; chlorogenic acid 3.47-111 μg / ml; cryptochlorogenic acid 1.29-41 μg / ml; atractylodes lactone III 12.28-393 μg / ml; atractylodes lactone II 9.79-313 μg / ml; atractylodes lactone I 11.03-353 μg / ml; β-eucalyptol 44.25-1416 μg / ml; atractylone 65.81-2106 μg / ml; linoleic acid 5.19-166 μg / ml; β-elemene 3.25-104 μg / ml; oleic acid 6.5-208 μg / ml).
[0061] Simultaneously, to prepare a sample solution of Atractylodes macrocephala, take 0.5g of Atractylodes macrocephala powder that has passed through a No. 3 sieve, extract it with 5mL of methanol using ultrasonic extraction (40kHz, 240W) for 30min, and filter the extract through a 0.45μm nylon membrane to obtain the sample solution of Atractylodes macrocephala.
[0062] Then, the HPLC detection parameters were set. HPLC detection was performed using an Agilent G7180A high-performance liquid chromatograph (Agilent Technologies Co., Ltd., Romania). The parameters included: an Agilent Eclipse Plus C18 column (250 mm × 4.6 mm, 5 μm); mobile phase A: acetonitrile; mobile phase B: 0.1% formic acid. The sample was scanned at all wavelengths using HPLC-DAD. The detection wavelengths were as follows: neochlorogenic acid, chlorogenic acid, and cryptochlorogenic acid at 325 nm; atractylodes lactone III, atractylodes lactone II, and atractylone at 220 nm; and atractylodes lactone I at 27 nm. 6 nm, β-eudesmol, linoleic acid, β-elemene, and oleic acid at 205 nm; flow rate 1 mL / min; injection volume 20 μL; column temperature 40 ℃; gradient elution (0-25 min, 7%-13% A; 25-27 min, 13%-55% A; 27-38 min, 55%-60% A; 38-55 min, 60% A; 55-56 min, 60%-75% A; 56-76 min, 75%-90% A; 76-77 min, 90%-100% A; 77-80 min, 100% A; 80-82 min, 100%-7% A; 82-85 min, 7% A).
[0063] By injecting mixed standard solutions and test solutions of different components separately, the actual content of each component in the Atractylodes macrocephala sample can be calculated by the linear relationship between the concentration and peak area of the mixed standard solution. The actual content of each component is used as one of the standard data for training the PLSR model.
[0064] S3. NIR spectral data of Atractylodes macrocephala samples collected based on NIR spectroscopy.
[0065] Spectroscopic acquisition was performed using a laboratory-grade near-infrared spectrometer (Thermo Fisher Scientific Inc., USA) equipped with an indium gallium arsenide diode array detector and a diffuse reflectance probe. 26-31 samples of Atractylodes macrocephala powder from each origin were collected, totaling 344 samples. Each sample weighed approximately 10g and was spread evenly in a quartz sample cup. Near-infrared (NIR) spectral data were acquired after one background scan using an integrated gold mirror. The sampling method was integrating sphere diffuse reflectance, with a wavenumber range of 10000-4000 cm⁻¹. -1 8.0cm resolution -1 , scanned 82 times.
[0066] The differences in the NIR and near-infrared spectra of Atractylodes macrocephala samples are usually very slight. Therefore, appropriate spectral preprocessing methods are used to eliminate various interfering factors, including noise, background interference, baseline drift, and scattering effects. Commonly used preprocessing techniques include SG smoothing, first derivative (1stD), second derivative (2ndD), multivariate scattering correction (MSC), standard normal variable transformation (SNV), and data centering.
[0067] Among these methods, SG smoothing reduces spectral noise through local weighting while preserving the original spectral features as much as possible; first and second derivatives effectively eliminate baseline drift and improve spectral resolution; MSC and SNV aim to eliminate spectral differences caused by diffuse scattering from the sample surface; data centering eliminates systematic bias by subtracting the mean of each variable to make the average of the data zero. This embodiment utilizes the above preprocessing methods and their combinations, comprehensively selecting the optimal preprocessing method to improve the predictive power and stability of the model.
[0068] Furthermore, near-infrared spectral variables often exhibit strong correlations and redundant information, and effective features are often concentrated in specific wavelength ranges. Using Interval Partial Least Squares (IPLS), continuous feature variables (such as spectral wavelength points) are divided into multiple sub-wavelength ranges. By establishing partial least squares models with different range combinations and evaluating their performance, the key ranges that contribute most to the model are selected as effective features. NIR spectral data maps are then selected based on these effective features. IPLS extracts continuous bands rather than discrete wavelength points, enhancing the model's interpretability of spectral features.
[0069] S4. Based on the actual content of Atractylodes macrocephala samples and NIR spectral data, a training set was constructed to train the origin classification model PLS-DA and the component quantification model PLSR.
[0070] In terms of classification and regression model selection, the origin classification model PLS-DA and the component quantification model PLSR were constructed. In this embodiment, the PLS-DA model was used to distinguish the 12 origins of Atractylodes macrocephala, and the PLSR model was used for regression prediction of 11 components.
[0071] Training sets were constructed based on the actual content of Atractylodes macrocephala samples and NIR spectral data. For PLS-DA model training, NIR spectral data was used as the independent variable and sample origin category as the dependent variable; for PLSR model training, NIR spectral data was used as the independent variable and the actual content of each component tested by HPLC was used as the dependent variable.
[0072] After collecting the near-infrared (NIR) spectral data of Atractylodes macrocephala samples, the dataset was divided to improve the representativeness of the training set. In this embodiment, the training set and the test set were divided in a 7:3 ratio.
[0073] For the origin classification model, stratified sampling was used, and 344 Atractylodes macrocephala samples from 12 origins were divided into 241 training sets and 103 test sets.
[0074] For the component quantification model, a random partitioning method was used to divide 120 Atractylodes macrocephala samples into 84 as the training set and 36 as the test set.
[0075] S5. Based on the trained PLS-DA and PLSR models, the origin classification and component quantitative analysis of Atractylodes macrocephala were performed.
[0076] The PLS-DA and PLSR models were evaluated. In the evaluation of the origin qualitative model, the following indicators were used to comprehensively measure the performance: cross-validation accuracy, precision, accuracy, recall, and F1 score.
[0077] Cross-validation accuracy refers to the average accuracy calculated through cross-validation. It is a core indicator for evaluating the overall generalization ability of the origin qualitative model. The calculation formula is as follows:
[0078]
[0079] Where k is the cross-validation fold number, Acc i Let be the accuracy in the i-th cross-validation.
[0080] In the evaluation of the component quantitative model, the following indicator is used: RPD (Residual Prediction Deviation). (Corrected coefficient of determination) (Prediction Determination Coefficient) (Cross-validation determination coefficient), RMSEC (corrected root mean square error), RMSEP (predicted root mean square error), and RMSECV (cross-validation root mean square error).
[0081] RPD (Root Mean Square Error) is the ratio of the standard deviation of measured values to the root mean square error of prediction. It reflects whether the model can effectively capture differences between samples and is a core indicator for evaluating the performance of quantitative component models. The calculation formula is as follows:
[0082]
[0083] Where SD is the standard deviation of the measured value and RMSEP is the root mean square error of the prediction.
[0084] The higher the RPD value, the higher the model's prediction accuracy. In the quantitative analysis of traditional Chinese medicine, an RPD value of 2-3 is generally considered to indicate good model quality, which can meet general quantitative needs. An RPD value exceeding 3 is considered to indicate excellent model quality, suitable for pharmacopoeia standards or scientific research analyses requiring high precision.
[0085] Near-infrared (NIR) spectral data from different origins, such as Figure 1 As shown, Figure 1 The figures show the raw NIR spectra of Atractylodes macrocephala from 12 different origins. As can be seen from the figures, the NIR signals of the Atractylodes macrocephala samples are mainly concentrated in the following wavenumber range: 8400-8300 cm⁻¹. -1 6800-6700cm -1 5800-5600cm -1 5200-5100cm -1 4800-4700cm -1 4500-4200cm -1 Combining spectroscopic knowledge, 8400-8300cm -1 The absorption peaks in the vicinity are mostly attributed to the second overtone of the CH stretching vibration, around 6800-6700 cm⁻¹. -1 The absorption peak usually corresponds to the first overtone of the OH stretching vibration, 5800-5600 cm⁻¹. -1 The broad absorption peaks in the vicinity are typically associated with the first overtone of CH stretching, 5200-5100 cm⁻¹ -1 The second overtone of the common OH bending and stretching vibration combination band or CO stretching vibration is found at 4800-4700cm. -1 In the adjacent bands, a combination of CH and CC / C=C skeleton vibrations is visible, along with related skeleton bands, at 4500-4200 cm⁻¹. -1 Combinations of CH or CH3 groups are frequently observed nearby. Because the spectral curves of different Atractylodes macrocephala products overlap significantly and exhibit largely consistent overall trends, distinguishing differences in origin solely by visual inspection is extremely challenging. Therefore, further analysis using near-infrared spectroscopy combined with chemometrics is necessary.
[0086] PLS-DA score maps were constructed by analyzing the near-infrared spectra of Atractylodes macrocephala from 12 origins to visually present the spatial distribution characteristics and inter-group separation of different samples. Figure 2 As shown in the PLS-DA score plot, samples that deviate excessively from the interior of the ellipse are typically outliers. The results show that most samples fall within the 95% confidence ellipse. Each group exhibits varying degrees of clustering and separation trends in the two-dimensional space formed by the first and second latent variables, with no obvious outliers. The first latent variable explains 86% of the model's score, while the second latent variable explains 12.1%, both supporting the PLS-DA model's ability to discriminate between groups.
[0087] To optimize the preprocessing methods, stratified sampling was used to divide the dataset. Based on cross-validation accuracy, several preprocessing methods for near-infrared spectra in the PLS-DA model were compared, as shown in Table 2. The spectra after different preprocessing methods are shown in Table 2. Figure 3In contrast, such as Figure 4 As shown, the optimal spectral preprocessing method is MSC+FD. Compared with the original data, the cross-validation accuracy of the model improved from 93.33% to 95.83%.
[0088] Table 2. Model evaluation of different preprocessing methods
[0089]
[0090] For the selection of characteristic bands, the IPLS (Interval Partial Least Squares) method was used to divide the optimal preprocessed spectrum into 30 intervals, each with 26 wavelengths. The cross-validation accuracy of each interval against the model classification was calculated, and intervals with higher accuracy were selected to further determine key bands. With the premise of improving the overall prediction accuracy of the model, the band that ultimately contributed the most to the model was determined to be 7205-7012 cm⁻¹. -1 6001-5608cm -1 5400-4003cm -1 ( Figure 5 (A and B in the image) It can be observed that the screened bands are all located near several peaks with high absorption values in the near-infrared spectrum. These peaks cover multiple groups such as CH, OH, and C-O in the Atractylodes macrocephala samples. Therefore, it can be inferred that there are significant differences in the chemical composition content among Atractylodes macrocephala from different origins.
[0091] In the PLS-DA model, the choice of the number of latent variables (LVs) directly affects the model's classification performance, stability, and generalization ability. PLS-DA reduces dimensionality and maximizes inter-group differences by extracting LVs. However, too many LVs may lead to overfitting of the model to noise or irrelevant variables in the training data, thus reducing classification accuracy. Therefore, methods such as cross-validation are needed to determine the LVs that optimize the model's predictive ability. Through 10-fold cross-validation analysis of the LVs used to select bands for model construction, the optimal LVs for the model were ultimately determined to be 26. Figure 5 (C in the middle).
[0092] Initially, the model's prediction accuracy increased with increasing LVs. The cross-validation accuracy reached its peak when the LVs were 26. As the LVs exceeded 26, the model's accuracy gradually plateaued and showed a downward trend. To further validate the reasonableness of this LVs, 200 permutation tests were performed on the 26LVs model. Figure 5 In the figure, R² (coefficient of determination) represents the model’s fit to the data, while Q² (cross-validation coefficient of determination) is used to evaluate the model’s predictive ability. In this test, the intercepts of R² and Q² were 0.089 and -0.302, respectively, and were much lower than the highest point, which indicates that the model did not overfit.
[0093] It is worth noting that the 26LVs determined in this embodiment is slightly higher than the 3-20 range commonly used in PLS-DA models. However, this is closely related to the spectral characteristics of the samples. In actual measurements, since the overall similarity of the spectra of different categories of samples is high, their differences are mostly manifested as subtle absorption peak shifts, absorbance changes in specific bands, or small fluctuations in the baseline. These signals, which are crucial for classification, are often quite weak. Therefore, the selection of 26LVs is precisely to extract and utilize these key information more comprehensively, ensuring that the model can effectively resolve complex spectral features. Based on the above verification results, it can be seen that the selection of this LVs is reasonable and necessary.
[0094] The results show that after optimal preprocessing combined with IPLS wavelength selection, the cross-validation accuracy of the final model increased from 95.83% to 96.67%, the training set accuracy increased from 96.68% to 98.34%, and the test set accuracy remained at 98.06%, compared to the model that was only preprocessed. This result demonstrates the effectiveness of wavelength selection in improving the model's generalization ability. Figure 6 The confusion matrix for the training and test sets is shown in Figure 1 (A represents the training set, and B represents the test set). Among the 103 test set samples, only 2 samples were misclassified. One sample misclassified the origin of S3 as S4, and the other sample misclassified the origin of S1 as S6. This may be because the component content of some Atractylodes macrocephala samples from different origins is relatively similar, making it difficult for the model to completely distinguish them during feature extraction. Even with the above situation, the overall error rate of the model is still less than 2%, indicating that the model has excellent classification and discrimination capabilities.
[0095] HPLC method validation: The HPLC method in this embodiment was validated through linearity, precision, stability, repeatability, and accuracy (Table 3). RC of the calibration curve. 2 The values were all above 0.999, indicating a good linear relationship between the concentration and peak area of the 11 components. The relative standard deviations (RSDs) of the precision, stability, and repeatability of each compound were all below 2%. The RSDs of the average recoveries of the 11 compounds were all below 4%. These results demonstrate that the developed HPLC method can meet the requirements of quantitative analysis. The chromatograms of the mixed standard and the test sample under these conditions are shown below. Figure 7 As shown (A is the mixed standard, B is the sample. Peaks 1-11 are neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, atractylodes lactone III, atractylodes lactone II, atractylodes lactone I, β-eucalyptol, atractylone, linoleic acid, β-elemene, and oleic acid, respectively).
[0096] Table 3. Linear relationship of 11 components in Atractylodes macrocephala.
[0097]
[0098] The pretreatment method was optimized, and combined with the HPLC content determination results, a dataset of 120 Atractylodes macrocephala samples was randomly divided using RPD and [other methods]. Based on this, several preprocessing methods for near-infrared spectra in the PLSR model of 11 chemical components in Atractylodes macrocephala were compared. The model evaluation of different preprocessing methods is shown in Tables 5-1 to 5-11. Finally, the optimal preprocessing methods for the 11 components were determined as shown in Table 4. Compared with the original data, these preprocessing methods all improved the performance of the model to a certain extent.
[0099] Table 4. Optimal pretreatment methods for the 11 components of Atractylodes macrocephala.
[0100]
[0101] Table 5-1 Model evaluation of the chlorogenic acid pretreatment method for Atractylodes macrocephala
[0102]
[0103] Table 5-2 Model Evaluation of Atractylodes macrocephala chlorogenic acid Pretreatment Method
[0104]
[0105] Table 5-3 Model evaluation of the cryptochlorogenic acid pretreatment method for Atractylodes macrocephala
[0106]
[0107] Table 5-4 Model evaluation of the atractylodes lactone III pretreatment method
[0108]
[0109] Table 5-5 Model evaluation of the atractylodes lactone II pretreatment method
[0110]
[0111] Table 5-6 Model evaluation of the atractylodes lactone I pretreatment method
[0112]
[0113] Table 5-7 Model evaluation of the Atractylodes macrocephala β-eucalyptol pretreatment method
[0114]
[0115] Table 5-8 Model Evaluation of Atractylodes macrocephala and Atractylodes lancea Pretreatment Method
[0116]
[0117] Table 5-9 Model Evaluation of Linoleic Acid Pretreatment Methods for Atractylodes macrocephala
[0118]
[0119] Table 5-10 Model Evaluation of Atractylodes macrocephala β-elemene Pretreatment Method
[0120]
[0121] Table 5-11 Model Evaluation of Atractylodes macrocephala Oleic Acid Pretreatment Method
[0122]
[0123] For the selection of characteristic bands, the IPLS (Interval Partial Least Squares) method was used to screen characteristic bands from the preprocessed spectra of the 11 components. By comparing the RMSECV of different bands, bands with lower RMSECV were selected, thus identifying key bands while improving the overall model performance. 10-fold cross-validation was used to determine the optimal LVs of the model, ultimately yielding the optimal bands and evaluation metrics for the 11 components, as shown in Table 6. Figure 8 As shown (A is the mixed standard, B is the sample. Peaks 1-11 are neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, atractylodes lactone III, atractylodes lactone II, atractylodes lactone I, β-eucalyptol, atractylone, linoleic acid, β-elemene, and oleic acid, respectively).
[0124] After characteristic band screening, the model performance of all 11 components improved. Among them, the phenolic acid components neochlorogenic acid, chlorogenic acid, and cryptochlorogenic acid contain phenolic hydroxyl groups (-OH) and carboxyl groups (-COOH), and the screened bands may be related to the vibrations of these groups; the lactone components atractylone III, atractylone II, and atractylone I have lactone rings containing C=O and CO, and their characteristic bands may be related to the stretching vibrations of C=O and CO in the lactone ring; the volatile oil components β-eudesmol, atractylone, and β-elemene are all sesquiterpenoids, containing hydroxyl groups. Structures such as OH, ketone C=O, and C=C may correspond to the characteristic bands of β-cineole (hydroxyl OH vibration), atractylone (ketone C=O vibration), and β-elemene (C=C vibration). Among fatty acid components, linoleic acid and oleic acid both contain carboxyl-COOH and C=C double bonds. Linoleic acid is a polyunsaturated fatty acid, while oleic acid is a monounsaturated fatty acid; their characteristic bands may correspond to the number of double bonds and the presence of carboxyl-COOH groups. The analysis of the potential correlation between these characteristic bands and the structures of chemical components is based on IPLS screening results and inferences derived from the fundamental principles of chemical group vibrations and infrared spectroscopy. The specificity and universality of the screened bands, as well as the correspondence between characteristic bands and specific functional groups, require further investigation.
[0125] Table 6. Characteristic bands corresponding to the 11 components of Atractylodes macrocephala.
[0126]
[0127] The calibration results, under optimal spectral preprocessing and selection of the most suitable wavelength range, show the correlation between the HPLC values and near-infrared spectral predictions of the 11 components as follows: Figure 7 As shown in Table 6, it can be seen that among the PLSR models of different chemical components of Atractylodes macrocephala, β-cineole and atractylone have the strongest predictive ability. β-cineole has an RPD of 4.2366, a training set R² of 0.9646, and a test set R² of 0.9530. Atractylone has an RPD of 4.1316, a training set R² of 0.9569, and a test set R² of 0.9520, indicating excellent model predictive ability and suitability for high-precision analysis in practical quantitative measurements. β-elemene has an RPD of 2.4778 compared to the other two volatile oil components, indicating good model predictive ability. For lactone components, the RPDs of atractylodes lactone I, II, and III are between 3 and 3.5, indicating strong model predictive ability. For phenolic acid components, the RPDs of neochlorogenic acid, chlorogenic acid, and cryptochlorogenic acid are between 2 and 2.5, indicating good model predictive ability. Among the fatty acid components, linoleic acid had an RPD of 2.2143, while oleic acid had an RPD of 2.7123, indicating that oleic acid had a higher predictive ability than linoleic acid. Overall, these chemical components all possess a certain degree of quantitative modeling capability.
[0128] Table 7 lists the statistical values of 11 chemical components of Atractylodes macrocephala in the PLSR model. Comparison reveals that β-cineole and atractylone have higher contents and larger standard deviations and chemical value fluctuations, which may be one of the important factors affecting the predictive performance of the PLSR model. Specifically, higher contents and larger chemical value fluctuations result in smaller errors during modeling, providing more comprehensive feature patterns. The model can capture more spectral response differences at different content levels during training, thus helping to improve the model's prediction accuracy for unknown samples. Other components, such as phenolic acids and fatty acids, have lower contents and relatively narrower chemical value fluctuations, which may limit the available feature information during modeling. Therefore, their predictive performance is slightly inferior to β-cineole and atractylone. This provides direction for subsequent model optimization for different components; for components with narrow content ranges, expanding the content range of the modeling samples can be considered to further improve the model's predictive ability.
[0129] Table 7 Statistical values of 11 chemical components of Atractylodes macrocephala.
[0130]
[0131] This embodiment uses a Fourier transform near-infrared spectrometer to collect spectral data of Atractylodes macrocephala samples, and combines the PLS-DA method to classify and distinguish Atractylodes macrocephala from 12 origins. A near-infrared quantitative analysis model for 11 components of Atractylodes macrocephala is constructed using the PLSR method. In the PLS-DA model construction, after optimal spectral preprocessing and characteristic band screening, the model achieved a classification accuracy of 98.06% for Atractylodes macrocephala from the 12 origins, demonstrating excellent discrimination ability. It is worth mentioning that during the sample collection process, multiple classification modeling verifications were conducted for different combinations of origins. When the number of origins to be classified was small, such as 3-6, the model's classification accuracy reached 100%. Therefore, in practical applications, by controlling the number of origins to be identified, near-infrared spectroscopy technology can be used for rapid and accurate identification of the origin or authenticity of Chinese medicinal materials. Figure 9 As shown (AK represents neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, atractylodes lactone III, atractylodes lactone II, atractylodes lactone I, β-cineole, atractylone, linoleic acid, β-elemene, and oleic acid, respectively), for the quantitative analysis of 11 chemical components in Atractylodes macrocephala, the PLSR model showed that β-cineole and atractylone were the most prominent among the volatile oil components, with RPD values both higher than 4, and R² values between the training and test sets both greater than 0.95. Further analysis revealed that compared with other chemical components, β-cineole and atractylone exhibited higher peak area response values in HPLC, and their content varied significantly among different samples of Atractylodes macrocephala. This phenomenon indicates that in quantitative modeling based on HPLC combined with near-infrared spectroscopy, the prediction accuracy of the constructed model is usually better when the content level of the target component is high and the chemical value fluctuation range is large. The RPD values of the remaining components were all higher than 2, which also met the corresponding quantitative analysis requirements. This result verifies the feasibility of near-infrared spectroscopy for rapid quantitative analysis of multiple components in Atractylodes macrocephala.
[0132] In summary, near-infrared spectroscopy combined with chemometrics not only achieves efficient identification of Atractylodes macrocephala from different origins but also constructs rapid quantitative models for various active components, demonstrating significant application potential and advantages in the quality evaluation of Atractylodes macrocephala. This method is simple to operate, non-destructive, and rapid, serving as an effective alternative to HPLC. It is suitable for quality monitoring of batches of Atractylodes macrocephala samples and provides a valuable research approach for the identification of origins and component analysis of other Chinese medicinal materials. In future applications within the Chinese medicine industry, samples from different harvesting periods and processing techniques can be incorporated to enhance the adaptability of near-infrared spectroscopy to the industry. Furthermore, with the widespread availability of portable near-infrared spectrometers, their miniaturization, low energy consumption, and real-time detection capabilities make field monitoring of Chinese medicinal materials, quality control in storage and logistics, and enforcement against counterfeiting possible. In the future, near-infrared spectroscopy is expected to become a standardized tool for the quality evaluation of Chinese medicinal materials, promoting the high-quality development of the Chinese medicine industry towards intelligence and digitalization.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins, characterized in that, include: S1. Collect Atractylodes macrocephala samples from multiple production areas and perform sample preprocessing; S2. Based on HPLC testing, quantitative analysis of multiple components contained in Atractylodes macrocephala samples was performed to obtain the actual content of each component in the Atractylodes macrocephala samples; S3. NIR spectral data of Atractylodes macrocephala samples acquired based on NIR spectroscopy; S4. Based on the actual content of Atractylodes macrocephala samples and NIR spectral data, a training set was constructed to train the origin classification model PLS-DA and the component quantification model PLSR. S5. Based on the trained PLS-DA and PLSR models, the origin classification and component quantitative analysis of Atractylodes macrocephala were performed.
2. The method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins according to claim 1, characterized in that, The sample preprocessing in S1 includes: S11. Remove impurities from the surface of the Atractylodes macrocephala sample; S12. Place the treated Atractylodes macrocephala samples in a 35℃ environment and dry them to constant weight. S13. Crush the dried sample with a pulverizer and pass it through a No. 3 sieve; S14. Store the sieved Atractylodes macrocephala sample powder in a desiccator for later use.
3. The method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins according to claim 1, characterized in that, S2 includes the following steps: S21. Prepare mixed standard solutions and Atractylodes macrocephala sample solutions with different components; S22. Set HPLC detection parameters; S23. Inject mixed standard solutions and test solutions of different components separately, calculate the actual content of each component in the Atractylodes macrocephala sample, and obtain HPLC data.
4. The method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins according to claim 1, characterized in that, S3 includes the following steps: S31. Set the parameters of the NIR spectrometer and collect near-infrared NIR spectral data of the Atractylodes macrocephala sample with the gold mirror background inside; S32. Preprocess the near-infrared (NIR) spectral data to obtain the preprocessed near-infrared (NIR) spectral data. S33. Perform characteristic wavelength screening on the preprocessed near-infrared (NIR) spectral data to obtain the screened NIR spectral data.
5. The method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins according to claim 1, characterized in that: When training the PLS-DA model, NIR spectral data map is used as the independent variable and sample origin category is used as the dependent variable. When training the PLSR model, NIR spectral data is used as the independent variable, and the actual content of each component tested by HPLC is used as the dependent variable.
6. The method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins according to claim 4, characterized in that, The preprocessing methods for near-infrared (NIR) spectral data in S32 include: SG smoothing, first derivative, second derivative, multivariate scattering correction, standard normal variable transformation, and / or data centering.
7. The method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins according to claim 4, characterized in that, The characteristic wavelength screening in S33 includes: Based on IPLS, continuous feature variables are divided into multiple sub-wavelength intervals. Partial least squares models with different combinations of sub-wavelength intervals are established and their performance is evaluated. The key sub-wavelength intervals that contribute the most to the model are selected as effective features, and NIR spectral data maps are selected based on the effective features.
8. The method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins according to claim 1, characterized in that, It also includes qualitative and quantitative assessments of origin and components for the PLS-DA and PLSR models; The qualitative assessment of origin includes cross-validation accuracy, precision, accuracy, recall, and F1 score. The quantitative assessment of components includes residual prediction bias, corrected coefficient of determination, prediction coefficient of determination, cross-validation coefficient of determination, corrected root mean square error, prediction root mean square error, and cross-validation root mean square error.
9. The method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins according to claim 8, characterized in that, The qualitative evaluation formula for cross-validation accuracy is expressed as follows: Where k is the cross-validation fold number, Acc i Let be the accuracy in the i-th cross-validation. This is the cross-validation accuracy.
10. The method for rapid classification and multi-component quantitative analysis of Atractylodes macrocephala from different origins according to claim 8, characterized in that, The formula for quantitatively evaluating the remaining prediction bias is expressed as follows: Where RPD is the residual prediction bias, SD is the standard deviation of the measured values, and RMSEP is the root mean square error of the prediction.