A method, device and computer program for rapid identification of the storage period of warm turmeric

CN122814529APending Publication Date: 2026-09-25WENZHOU MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611028541.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而近红外光谱采集时样品的颗粒度会导致谱图产生差异,从而影响对光谱数据的利用

Benefits of technology

[0024]本发明的方法填补了现有技术中对温郁金储存期快速识别的空白,利用FT-NIR方法,实现了不同储存期温郁金的快速、高精度的鉴别及成分含量估算。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122814529A_ABST
    Figure CN122814529A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of quick identification method, equipment and computer program of storage period of warm root of rubia, including the near infrared spectrum data of different storage period warm root sample is preprocessed using different preprocessing method, and the content of index component is constructed with the spectrum data after preprocessing and content prediction model;The accuracy evaluation result of content prediction model is used to determine the candidate preprocessing method;The spectrum data processed using different candidate preprocessing method is used to train and construct the classification model of the storage period of warm root of rubia;The storage period of unknown warm root sample is quickly identified using the content prediction model and classification model.The method of the present application fills the blank of the quick identification of the storage period of warm root of rubia in the prior art, realizes the quick, high-precision identification estimation of different storage period warm root of rubia using FT-NIR method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of quality evaluation technology of Chinese medicinal materials, specifically involving a method, device and computer program for rapid identification of the storage period of Curcuma longa. Background Technology

[0002] Curcuma longa has a long history of medicinal use, and high-quality medicinal materials are the cornerstone of ensuring the efficacy of clinical medications. Quality evaluation is a crucial and challenging aspect of the development of traditional Chinese medicine. There are two main models for evaluating the quality of Chinese medicinal materials: subjective and objective. Subjective evaluation often relies on the appearance, color, and odor of the herbs; objective evaluation uses pharmacopoeia testing standards or content determination methods. Experience-based identification of Chinese medicinal materials is still in use. This involves judging the quality of herbs by observing their overall or characteristic characteristics, using methods such as sight and smell. This method is simple, fast, intuitive, and highly distinctive. Through continuous accumulation, it still holds a place in the quality evaluation of Chinese medicinal materials. However, traditional identification methods suffer from strong subjectivity and difficulty in standardization.

[0003] Currently, techniques such as spectrophotometry, electronic nose, and near-infrared spectroscopy (FT-NIR) are used for the evaluation of Chinese medicinal materials. However, different Chinese medicinal materials have different components and properties, making existing techniques incompatible. FT-NIR obtains absorption spectra by irradiating samples with a continuous wavelength infrared light source. It comprehensively reflects information in the sample, has a fast detection speed, minimal sample loss, simple operation, and high sensitivity. However, the particle size of the sample during near-infrared spectroscopy acquisition can cause differences in the spectrum, thus affecting the utilization of spectral data. Current spectral detection of Curcuma longa mainly focuses on the qualitative identification of its authenticity. Most research on the quality control of Curcuma longa is based on different origins and harvesting periods; the correlation between the quality of Curcuma longa and storage time has not been explored. Summary of the Invention

[0004] The present invention aims to provide a method, device and computer program for rapid identification of the storage period of Curcuma longa.

[0005] To achieve the above technical objectives, the present invention adopts the following solution:

[0006] A rapid method for identifying the storage period of Curcuma longa, the method comprising:

[0007] Samples of Curcuma aromatica from different storage periods were obtained, the content of indicative components in each sample was determined, and near-infrared spectral data of the samples were collected; the indicative components are chemically stable, specific active ingredients of Curcuma aromatica.

[0008] The near-infrared spectral data are preprocessed using different preprocessing methods. The spectral data processed by different preprocessing methods are combined with the content data of the indicative components. Based on regression analysis, content prediction models are constructed for each indicative component. The accuracy of the content prediction models is evaluated using preset model evaluation indicators to determine the optimal content prediction model for each indicative component. The preprocessing method corresponding to the optimal content prediction model is then used as a candidate preprocessing method.

[0009] Using the spectral data processed by each candidate preprocessing method as input, and the storage period of Curcuma longa samples as classification label, a storage period prediction model of Curcuma longa was constructed using machine learning methods, and the optimal model was selected as the classification model for Curcuma longa samples.

[0010] An unknown Curcuma longa sample was obtained. The spectral data of the unknown Curcuma longa sample was collected and preprocessed using the same processing method as that used for modeling the Curcuma longa sample. The data was then input into the corresponding classification model to identify the storage period of the unknown Curcuma longa sample.

[0011] In some embodiments of the present invention, the indicative components are turmeric dione, furandiene, gemmaconone, and β-elemene.

[0012] In some embodiments of the present invention, after pulverizing Curcuma longa, the dried powder is used for spectral acquisition and determination of the content of indicative components;

[0013] The unknown Curcuma longa sample was processed in the same way to obtain dry powder, and then spectral data were collected.

[0014] In some embodiments of the present invention, an FT-NIR spectrometer is used in a diffuse reflectance manner in the range of 10000~4000 cm⁻¹. -1 Near-infrared spectral data of Curcuma aromatica samples were collected within the spectral range.

[0015] In some embodiments of the present invention, full-spectrum data are used as independent variables and the measured content of each index component is used as dependent variables to construct content prediction models for each index component.

[0016] In some embodiments of the present invention, a content prediction model for each index component is constructed based on the partial least squares method.

[0017] In some embodiments of the present invention, a classification model for Curcuma longa samples with different storage periods is constructed based on an artificial neural network (ANN).

[0018] In some embodiments of the present invention, after acquiring spectral data, the spectral data is divided into three parallel groups for preprocessing:

[0019] The spectral data after standard normal distribution and convolution smoothing were used to construct a prediction model for turmeric dione content;

[0020] Spectral data processed by multivariate scattering correction, convolution smoothing, and first derivative were used to construct a furandiene content prediction model.

[0021] The spectral data after convolution smoothing were used to construct prediction models for the contents of gemmaconone and β-elemene, respectively.

[0022] The present invention further provides a computer device including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0023] The present invention further provides a computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0024] The method of this invention fills the gap in the prior art for rapid identification of the storage period of Curcuma longa. By using the FT-NIR method, it achieves rapid and high-precision identification and component content estimation of Curcuma longa at different storage periods. Attached Figure Description

[0025] Figure 1 The images show the original near-infrared spectra of Curcuma longa at different storage periods. The left image shows the original near-infrared spectra of a single sample at different storage periods (one sample is taken from each storage period); the right image shows the original near-infrared spectra of all samples at different storage periods.

[0026] Figure 2 Scatter plot of the prediction model for the components of Curcuma longa at different storage periods.

[0027] Figure 3 The confusion matrix and ROC curve of classification models constructed using different methods on the original spectral data are shown.

[0028] Figure 4 Confusion matrices and ROC curves of classification models constructed using different methods for spectral data preprocessed using the MSC+SG+1d method.

[0029] Figure 5 The confusion matrix and ROC curve of classification models constructed using different methods for spectral data preprocessed based on the SNV+SG method are shown.

[0030] Figure 6 The confusion matrix and ROC curve of classification models constructed using different methods for spectral data preprocessed based on the SG method are shown.

[0031] Figures 3 to 6In the diagram: 2018-2019 represents Curcuma longa stored for 5-6 years, 2020-2021 represents Curcuma longa stored for 3-4 years, and 2022-2023 represents Curcuma longa stored for 1-2 years; each figure, from top to bottom, shows the classification model constructed using DT, KNN, SVM, and ANN methods. Detailed Implementation

[0032] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0033] The reagents and instruments involved in the examples are as follows:

[0034] Curcuma dione (Shanghai Yuanye Biotechnology Co., Ltd., purity ≥ 99.8%, batch number: D18GB171806);

[0035] Furandiene (Shanghai Yuanye Biotechnology Co., Ltd., purity ≥ 99.8%, batch number: J23HB188661);

[0036] Gemmazone (Shanghai Yuanye Biotechnology Co., Ltd., purity ≥ 99.8%, batch number: D02D11S132892);

[0037] β-elemene (Shanghai Yuanye Biotechnology Co., Ltd., purity ≥ 98%, batch number: M20IB215514);

[0038] Methanol (chromatographic grade, Macklin Laboratories, USA, lot number: C15895751);

[0039] Acetonitrile (chromatographic grade, Spectrum Corporation, USA);

[0040] Near-infrared spectrometer (AntarisIl FT-NIR, Thermo Fisher Scientific, USA);

[0041] High-speed multi-functional pulverizer (FTT-2500 T, Shanghai Naise Machinery Co., Ltd.);

[0042] Gas chromatography (8890)-mass spectrometry (7000 D) system (Agilent Technologies, USA);

[0043] Chromatographic column (HP-5MS, Agilent Technologies, USA);

[0044] Ultrasonic cleaner (JP-100 S, Shenzhen Jiemeng Cleaning Equipment Co., Ltd., power 500 W, frequency 40 kHz);

[0045] Ultrapure water system (Milli-Q, Merck GmbH, Germany).

[0046] In this embodiment, 60 batches of Curcuma aromatica (origin: Ruian, Zhejiang) from different storage periods (2018-2023) were used as the research object to illustrate the technical solution of the present invention. The Curcuma aromatica obtained were samples in the form of dried tuberous roots, which were crushed and passed through a No. 5 sieve to obtain dried powder for subsequent use.

[0047] Example 1

[0048] This embodiment is based on HPLC determination of the changes in the content of indicative components of Curcuma longa during different storage periods.

[0049] By analyzing the components of dried Curcuma longa powder, this embodiment selected turmeric dione, furandiene, gemmaconone, and β-elemene, which have relatively stable chemical structures (ensuring stability during content determination) and are specific effective components of Curcuma longa, as indicator components of Curcuma longa for content determination and subsequent content prediction.

[0050] The content of the indicative components in Curcuma longa was determined by HPLC as follows:

[0051] Preparation of mixed reference solutions: Weigh out the reference standards of turmericdione, furandiene, gemimazone, and β-elemene, respectively, and dilute to volume with methanol to prepare single-standard stock solutions of turmericdione, furandiene, gemimazone, and β-elemene. Dilute the above single-standard stock solutions separately to prepare mixed reference solutions with a content of turmericdione of 100.20 μg / mL, furandiene of 25.08 μg / mL, gemimazone of 25.10 μg / mL, and β-elemene of 1.01 μg / mL.

[0052] Preparation of the test solution: Accurately weigh 1 g of dried Curcuma aromatica powder, place it in a 10 mL stoppered conical flask, accurately add 5 mL of methanol, shake well, seal, and ultrasonically extract for 45 min (power 500 W, frequency 40 kHz). Let it stand at room temperature, make up the weight with methanol, and filter through a 0.22 μm organic phase microporous membrane to obtain the test solution.

[0053] HPLC parameters:

[0054] Column: EC-C 18 Chromatographic column (4.6 mm × 150 mm, 4 µm);

[0055] Mobile phase: water (A) - acetonitrile (B), gradient elution (0–10 min, 50% B; 10–22 min, 50%–78% B; 22–36 min, 78%–83% B; 36–37 min, 83%–95% B; 37–40 min, 95% B; 40–41 min, 95%–50% B; 41–50 min, 50% B) see Table 1;

[0056] Detection wavelength: 201 nm;

[0057] Flow rate: 1.0 mL / min;

[0058] Column temperature: 30 ℃;

[0059] Injection volume: 20 μL.

[0060] Table 1. HPLC Elution Procedure for Curcuma longa

[0061]

[0062] Example 2

[0063] This embodiment specifically illustrates the method for acquiring and preprocessing spectral data of Curcuma longa, and for constructing an index component prediction model using the preprocessed spectral data.

[0064] Near-infrared spectral information of 60 batches of dried Curcuma aromatica powder samples was acquired using a near-infrared spectrometer with diffuse reflectance. The sample powder was poured into quartz dishes and gently pressed to achieve a uniform packing density. The spectral range was 10000–4000 cm⁻¹. -1 The resolution is 16 cm. -1 Each spectrum was scanned 32 times, and each sample was scanned in triplicate.

[0065] The particle size of the sample during near-infrared spectroscopy acquisition can cause differences in the spectra. In this embodiment, the near-infrared spectra of all samples showed similar characteristics. Figure 1 Furthermore, considering that the original spectra are highly susceptible to instrument background noise and baseline drift, this embodiment preprocesses the original spectra of 60 batches of samples to improve the accuracy of discrimination.

[0066] To screen for the optimal spectral preprocessing method, this embodiment investigated 17 near-infrared spectral preprocessing methods: multivariate scattering correction (MSC), standard normal distribution (SNV), convolution smoothing (SG), first derivative (1d), second derivative (2d), MSC+1d, MSC+2d, MSC+SG, MSC+SG+1d, MSC+SD+2d, SNV+1d, SNV+2d, SNV+SG, SNV+SG+1d, SNV+SG+2d, SG+1d, and SG+2d. The reliability of each method was evaluated by the accuracy of predicting the content of different indicative components. Specifically, after processing the spectral data using the above different preprocessing methods, a partial least squares (PLS) model was constructed to predict the content of each indicative component. Specifically, the full-spectrum data was used as the independent variable, and the measured content of each indicative component was used as the dependent variable. A PLS model was constructed for each indicative component separately.

[0067] The constructed PLS model is evaluated for accuracy using preset evaluation metrics to determine the reliability of the preprocessing method. In this embodiment, the accuracy of the PLS model is assessed by calibrating the coefficient of determination (R²). 2 C ), prediction certainty coefficient (R) 2 The metrics are evaluated using the following methods: root mean square error of calibration (RMSEC), root mean square error of prediction (RMSEP), RMSEP / RMSEC ratio, and relative percentage deviation (RPD). The reasonable ranges for each metric are as follows: R 2 C > 0.80, R 2 p>0.80, RMSEP / RMSEC values ​​are between 0.80 and 1.2, and the smaller the RMSEC and RMSEP values, the better.

[0068] The results are shown in Tables 2-5, where RAW refers to the raw spectral data without preprocessing.

[0069] Table 2. Prediction results of turmeric dione by different spectral preprocessing methods

[0070]

[0071] Table 3. Prediction results of furandiene by different spectral preprocessing methods

[0072]

[0073] Table 4. Prediction results of different spectral preprocessing methods for gemmaconazole

[0074]

[0075] Table 5. Prediction results of β-elemene using different spectral preprocessing methods

[0076]

[0077] It can be seen that the SNV+SG pretreatment method is suitable for the analysis of curcuminone; the MSC+SG+1d pretreatment method is suitable for the analysis of furandiene; and the SG pretreatment method is suitable for the analysis of gemmaconone and β-elemene.

[0078] In this embodiment, the PLS model corresponding to the optimal results in Tables 2-5 is selected as the prediction model for the corresponding component. Therefore, after collecting the spectral data, the spectral data is divided into three parallel groups for preprocessing. The spectral data preprocessed by SNV+SG is used to construct the turmeric dione content prediction model, the spectral data preprocessed by MSC+SG+1d is used to construct the furandiene content prediction model, and the spectral data preprocessed by SG is used to construct the gemmaconone and β-elemene content prediction models, respectively.

[0079] like Figure 2 As shown, Curcuma dione R 2 c = 0.9184, R 2 p = 0.8053; furandiene R 2 c = 0.8632, R 2 p = 0.7812; Gemmaconazole R 2 c = 0.8051, R 2 p = 0.8150; β-elemene R 2 c = 0.8278, R 2 p = 0.8483, which shows that the constructed PLS model has good predictive performance for all four indicative components.

[0080] Example 3

[0081] This embodiment uses machine learning methods to classify Curcuma aromatica at different storage periods based on preprocessed spectral data.

[0082] The machine learning methods tested in this embodiment include decision tree (DT), artificial neural network (ANN), support vector machine (SVM), and K-nearest neighbor classification (KNN).

[0083] DT, SVM, KNN, and ANN models were developed using MATLAB 2023a software, and 10-fold cross-validation was used to test the accuracy of the classification models. Clustering results were visualized using a confusion matrix. The classification models were mainly evaluated using accuracy (Acc%), precision (Pr%), recall (Re%), and F1 score (F1). The overall performance of each classification model was represented by an operating characteristic curve (ROC curve), with a larger area under the ROC curve indicating higher model performance.

[0084] Based on the spectral data preprocessed by SNV+SG, spectral data preprocessed by MSC+SG+1d, and spectral data preprocessed by SG selected in Example 2, classification models were constructed using different machine learning methods, and the performance evaluation results are shown in Table 6.

[0085] Table 6 Performance Evaluation of Different Classification Models

[0086]

[0087] It can be seen that the accuracies of the DT, ANN, SVM, and KNN models in the original near-infrared spectrum are 57.27%, 91.67%, 72.21%, and 91.67%, respectively. After processing with SNV+SG preprocessing, the accuracy of the above four models is significantly improved, with accuracies of 83.91%, 96.67%, 97.78%, and 96.12%, respectively. After processing with MSC+SG+1d preprocessing, the accuracy of the four models is 82.77%, 97.78%, 96.16%, and 98.89%, respectively. However, after processing with SG preprocessing, the accuracy improvement of the above four models, except for the ANN model, is too small, with accuracies of 60.60%, 92.24%, 60.60%, and 66.76%, respectively. Therefore, the ANN model is used to construct the classification model.

Claims

1. A rapid identification method for the storage period of Curcuma aromatica, characterized in that, The method includes: Samples of Curcuma aromatica from different storage periods were obtained, the content of indicative components in each sample was determined, and near-infrared spectral data of the samples were collected; the indicative components are chemically stable, specific active ingredients of Curcuma aromatica. The near-infrared spectral data are preprocessed using different preprocessing methods. The spectral data processed by different preprocessing methods are combined with the content data of the indicative components. Based on regression analysis, content prediction models are constructed for each indicative component. The accuracy of the content prediction models is evaluated using preset model evaluation indicators to determine the optimal content prediction model for each indicative component. The preprocessing method corresponding to the optimal content prediction model is then used as a candidate preprocessing method. Using the spectral data processed by each candidate preprocessing method as input, and the storage period of Curcuma longa samples as classification label, a storage period prediction model of Curcuma longa was constructed using machine learning methods, and the optimal model was selected as the classification model for Curcuma longa samples. An unknown Curcuma longa sample was obtained. The spectral data of the unknown Curcuma longa sample was collected and preprocessed using the same processing method as that used for modeling the Curcuma longa sample. The data was then input into the corresponding classification model to identify the storage period of the unknown Curcuma longa sample.

2. The method according to claim 1, characterized in that, The indicative components are turmeric dione, furandiene, gemmaconone, and β-elemene.

3. The method according to claim 1, characterized in that, After pulverizing Curcuma longa, the dried powder was used for spectral acquisition and determination of the content of indicative components. The unknown Curcuma longa sample was processed in the same way to obtain dry powder, and then spectral data were collected.

4. The method according to claim 1, characterized in that, Using an FT-NIR spectrometer in the diffuse reflectance mode, in the range of 10000–4000 cm⁻¹ -1 Near-infrared spectral data of Curcuma aromatica samples were collected within the spectral range.

5. The method according to claim 1, characterized in that, Using full-spectrum data as independent variables and the measured content of each indicative component as dependent variables, content prediction models were constructed for each indicative component.

6. The method according to claim 1, characterized in that, A content prediction model for each indicative component was constructed based on partial least squares method.

7. The method according to claim 1, characterized in that, A classification model for Curcuma longa samples with different storage periods was constructed based on an artificial neural network (ANN).

8. The method according to claim 2, characterized in that, After acquiring the spectral data, the spectral data was divided into three parallel groups for preprocessing: The spectral data after standard normal distribution and convolution smoothing were used to construct a prediction model for turmeric dione content; Spectral data processed by multivariate scattering correction, convolution smoothing, and first derivative were used to construct a furandiene content prediction model. The spectral data after convolution smoothing were used to construct prediction models for the contents of gemmaconone and β-elemene, respectively.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 8.