Method for detecting chemical component content of andrographis paniculata based on hyperspectral imaging technology
Through hyperspectral imaging technology and stoichiometric methods, a regression model of diterpenoid lactone compounds in truncao is established, which solves the problems of high detection cost, long time and high reagent consumption in the existing technology, and achieves fast, accurate and non-destructive chemical composition detection of truncao.
Patent Information
- Application Number
- CN202410031871.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-11
AI Technical Summary
When the prior art detects the content of four main diterpene lactone compounds in the piercing heart, there are problems such as high detection cost, long time, and large reagent consumption, which is difficult to meet the needs of efficient, green and sustainable development.
Hyperspectral imaging technology combined with stoichiometric methods was used to establish a regression model of diterpenoid lactone compounds in the cystole. Sample data were obtained through hyperspectral imaging equipment, content values were measured by liquid chromatography, and partial least squares regression, backpropagation neural network and other models were used for prediction.
It realizes fast, accurate, lossless and pollution-free detection, simplifies the operation process, reduces the detection cost, and improves the detection efficiency. It is suitable for the chemical composition analysis of different germplasm pierced heartworms.
Smart Images

Figure CN120293623A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of content detection, and particularly relates to a method for detecting the content of chemical components of Andrographis paniculata based on hyperspectral imaging technology. Background Art
[0002] Andrographis paniculata is the dried aerial part of the plant Andrographis paniculata (Burm.f.) Nees of the Acanthaceae family. Diterpenoid lactone components are the main active components of Andrographis paniculata, with pharmacological effects such as anti-inflammatory, anti-infective, anti-tumor, and heart protection. Andrographolide, neoandrographolide, deoxyandrographolide, and dehydrated andrographolide account for more than 75% of the total diterpenoid lactones and are the quantitative components specified in the Chinese Pharmacopoeia (2020 Edition). At present, the content determination of the above four active substances mainly uses high-performance liquid chromatography and ultra-high performance liquid chromatography. Although these detection methods have good selectivity and high sensitivity, they have disadvantages such as high detection cost, long time, and large consumption of reagents. Therefore, the current means and methods for detecting the quality of Andrographis paniculata in the market cannot meet the concept of efficient, green, and sustainable development of society. Summary of the Invention
[0003] The present invention provides a method for detecting the content of chemical components in Andrographis paniculata based on hyperspectral technology, and the method includes the following steps:
[0004] Step (1): Obtain Andrographis paniculata samples;
[0005] Step (2): Measure and obtain the hyperspectral data of the Andrographis paniculata samples;
[0006] Step (3): Measure and obtain the content values of the chemical components in the Andrographis paniculata samples;
[0007] Step (4): Use the hyperspectral data of the Andrographis paniculata samples in step (2) and the content values of the chemical components in the Andrographis paniculata samples in step (3) to train a model;
[0008] Step (5): Measure the hyperspectral data of the to-be-detected Andrographis paniculata samples, input them into the trained model, and obtain the content values of the chemical components in the to-be-detected Andrographis paniculata samples.
[0009] According to the embodiments of the present invention, the chemical components in the Andrographis paniculata samples are diterpenoid lactone compounds, preferably andrographolide compounds; further preferably, the andrographolide compounds are selected from any one or more of the following: andrographolide, neoandrographolide, deoxyandrographolide, dehydrated andrographolide.
[0010] According to the embodiments of the present invention, in step (1), the Andrographis paniculata samples are Andrographis paniculata samples of different strains.
[0011] According to an embodiment of the present invention, step (2) specifically is: Grind the andrographis paniculata sample, sieve it (for example, through a 100-mesh sieve), evenly distribute it on the surface of the container, and use a hyperspectral imaging device to obtain the hyperspectral data of each andrographis paniculata sample.
[0012] According to an embodiment of the present invention, the hyperspectral imaging device is a HySpex series imaging spectrometer. According to an embodiment of the present invention, the hyperspectral imaging device includes a visible-near infrared camera (spectral range 410nm - 990nm) and a short-wave infrared camera (spectral range 950nm - 2500nm). According to an embodiment of the present invention, the camera lens is equipped with a tungsten halogen lamp. According to an embodiment of the present invention, the spectral resolution of the camera lens is 6nm. According to an embodiment of the present invention, the distance between the camera lens and the sample is 25cm. According to an embodiment of the present invention, the platform moving speed is 1.5nm / s.
[0013] According to an embodiment of the present invention, in step (3), liquid chromatography (such as high performance liquid chromatography HPLC, ultra-high performance liquid chromatography UPLC) is used to measure and obtain the content values of chemical components in the andrographis paniculata sample.
[0014] According to an embodiment of the present invention, step (3) specifically includes the following steps:
[0015] (a) Grind the andrographis paniculata sample, sieve it (for example, through a 100-mesh sieve), and perform solvent extraction to obtain a test solution;
[0016] (b) Use liquid chromatography for detection to obtain the content values of chemical components in the test solution.
[0017] According to an embodiment of the present invention, in step (a), the extraction solvent is 40% methanol.
[0018] According to an embodiment of the present invention, in step (b), the liquid chromatography detection conditions are:
[0019] The chromatographic column is a reversed-phase C 18 chromatographic column (chromatographic column size, for example, 100mm×2.1mm, 1.8μm);
[0020] The mobile phase consists of phase A and phase B; phase A is water, and phase B is acetonitrile; gradient elution;
[0021] Preferably, the elution gradient is: 0 - 7.5min, 20%B → 25%B; 7.5 - 9min, 25%B → 28%B; 9 - 15min, 28%B → 32%B; 15 - 16min, 32%B → 35%B; 16 - 17min, 35%B → 90%B; 17 - 19min, 90%B; 19 - 20min, 90%B → 20%B; 20 - 22min, 20%B;
[0022] Preferably, the flow rate is: 0.3 mL / min;
[0023] Preferably, the column temperature is: 35 °C;
[0024] Preferably, the injection volume is: 1.0 μL,
[0025] Preferably, the detection wavelength is: 205 nm.
[0026] According to an embodiment of the present invention, in step (4), the model is selected from: partial least squares regression (PLSR), back propagation neural network (BPNN), random forest regression (RFR).
[0027] According to an embodiment of the present invention, when the chemical component to be detected is andrographolide or neoandrographolide, the model is a PLSR model or a BPNN model; preferably, it is a PLSR model.
[0028] According to an embodiment of the present invention, when the chemical component to be detected is deoxyandrographolide or dehydrated andrographolide, the model is a BPNN model or a PLSR model; preferably, it is a BPNN model.
[0029] According to an embodiment of the present invention, when the chemical component to be detected is a mixture of four compounds including andrographolide, neoandrographolide, deoxyandrographolide, and dehydrated andrographolide, the model is a PLSR model or a BPNN model; preferably, it is a PLSR model.
[0030] According to an embodiment of the present invention, in step (2), optionally, it further includes: preprocessing the hyperspectral data of the obtained andrographis paniculata samples; preferably, the preprocessing method is selected from: multiplicative scatter correction (MSC), first derivative (D1), second derivative (D2), Savitzky-Golay (SG) smoothing.
[0031] According to an embodiment of the present invention, when the chemical component to be detected is andrographolide and / or neoandrographolide and / or dehydrated andrographolide, preprocessing is performed based on the hyperspectral data of the obtained andrographis paniculata samples.
[0032] According to an embodiment of the present invention, when the chemical component to be detected is andrographolide, SG preprocessing is performed on the hyperspectral data of the obtained andrographis paniculata samples, and the model is selected from PLSR models.
[0033] According to an embodiment of the present invention, when the chemical component to be detected is neoandrographolide, the hyperspectral data of the andrographis paniculata sample obtained is preprocessed by MSC, and the model is selected from the PLSR model.
[0034] According to an embodiment of the present invention, when the chemical component to be detected is dehydroandrographolide, the hyperspectral data of the andrographis paniculata sample obtained is preprocessed by MSC, and the model is selected from the BPNN model.
[0035] According to an embodiment of the present invention, in step (2), optionally, it further includes: based on the hyperspectral data of the andrographis paniculata sample obtained, screening the characteristic bands of the chemical component (SPA screening).
[0036] According to an embodiment of the present invention, the characteristic bands of andrographolide are as follows:
[0037] 524nm, 660nm, 676nm, 746nm, 806nm, 876nm, 958nm, 1450nm, 1690nm, 1897nm, 1930nm, 1995nm, 2039nm, 2142nm, 2202nm, 2306nm, 2382nm, 2502nm, 2513nm.
[0038] According to an embodiment of the present invention, the characteristic bands of neoandrographolide are as follows: 432nm, 475nm, 503nm, 530nm, 622nm, 643nm, 660nm, 676nm, 687nm, 698nm, 757nm, 811nm, 952nm, 990nm, 1314nm, 1336nm, 1357nm, 1374nm, 1385nm, 1390nm, 1516nm, 1586nm, 1674nm, 1717nm, 1794nm, 1870nm, 1897nm, 2044nm, 2110nm, 2246nm, 2273nm, 2306nm, 2371nm, 2502nm.
[0039] According to an embodiment of the present invention, the characteristic bands of deoxyandrographolide are as follows:
[0040] 421nm, 427nm, 443nm, 448nm, 459nm, 481nm, 508nm, 530nm, 551nm, 622nm, 643nm, 806nm, 876nm, 903nm, 936nm, 1434nm, 1597nm, 1897nm, 1924nm, 1995nm, 2055nm, 2208nm, 2246nm, 2273nm, 2306nm, 2377nm, 2480nm, 2491nm, 2497nm。
[0041] According to the embodiments of the present invention, the characteristic bands of dehydroandrographolide are as follows:
[0042] 421nm, 465nm, 508nm, 540nm, 551nm, 584nm, 622nm, 643nm, 708nm, 844nm, 876nm, 903nm, 998nm, 1145nm, 1336nm, 1347nm, 1352nm, 1363nm, 1385nm, 1390nm, 1401nm, 1407nm, 1434nm, 1663nm, 1695nm, 1717nm, 1744nm, 1804nm, 1843nm, 1864nm, 1875nm, 1924nm, 1995nm, 2028nm, 2050nm, 2110nm, 2170nm, 2208nm, 2246nm, 2300nm, 2317nm, 2344nm, 2377nm, 2404nm, 2431nm, 2491nm, 2502nm, 2508nm, 2513nm。
[0043] According to the embodiments of the present invention, the characteristic bands of the total content of four andrographolide compounds (andrographolide, neoandrographolide, deoxyandrographolide, dehydroandrographolide) are as follows:
[0044] 416nm, 421nm, 432nm, 448nm, 535nm, 622nm, 643nm, 952nm, 990nm, 1477nm, 1695nm, 1870nm, 1924nm, 2050nm, 2148nm, 2246nm, 2311nm, 2502nm, 2513nm。
[0045] According to the embodiments of the present invention, when the chemical component to be detected is andrographolide, the hyperspectral data of the obtained andrographis paniculata sample is subjected to SG pretreatment, and then the characteristic bands are screened (SPA screening), and the model is selected from the PLSR model.
[0046] According to the embodiments of the present invention, when the chemical component to be detected is neoandrographolide, the hyperspectral data of the andrographis paniculata samples obtained are subjected to special MSC preprocessing, and then the characteristic bands are screened (SPA screening), and the model is selected from the PLSR model.
[0047] According to the embodiments of the present invention, when the chemical component to be detected is deoxyandrographolide, the characteristic bands of the hyperspectral data of the andrographis paniculata samples obtained are screened (SPA screening), and the model is selected from the BPNN model.
[0048] According to the embodiments of the present invention, when the chemical component to be detected is dehydrated deoxyandrographolide, the hyperspectral data of the andrographis paniculata samples obtained are subjected to MSC preprocessing, and then the characteristic bands are screened (SPA screening), and the model is selected from the BPNN model.
[0049] According to the embodiments of the present invention, when the chemical component to be detected is the total content of four andrographolide compounds, the characteristic bands of the hyperspectral data of the andrographis paniculata samples obtained are screened (SPA screening), and the model is selected from the PLSR model.
[0050] Beneficial effects
[0051] The present invention provides a method for detecting the content of chemical components in andrographis paniculata. The method can quickly and accurately detect the content of chemical components (such as andrographolide, neoandrographolide, deoxyandrographolide, and dehydrated andrographolide) in andrographis paniculata of different germplasms in a convenient, fast, non-destructive, and pollution-free manner, avoiding problems such as cumbersome detection operations, long time, high cost, and large consumption of reagents, providing a new idea for realizing the rapid non-destructive detection of the quality of andrographis paniculata, and having the beneficial effects of high efficiency, non-destructiveness, high throughput, greenness, and environmental protection.
[0052] The present invention uses hyperspectral imaging technology, combined with chemometric methods, to establish regression models for andrographolide, neoandrographolide, deoxyandrographolide, dehydrated andrographolide, and the total content of the above four diterpenoid lactone compounds in andrographis paniculata, providing a new method for the rapid and accurate detection of the quality of andrographis paniculata. Description of the drawings
[0053] Figure 1 It is a box plot of the chemical component contents of andrographis paniculata samples of different germplasms; among them, Figure 1 a. Andrographolide content diagram; Figure 1 b. Neoandrographolide content diagram; Figure 1 c. Deoxyandrographolide content diagram; Figure 1 d. Dehydrated andrographolide content diagram; Figure 1 e. Total content diagram of andrographolide, deoxyandrographolide, deoxyandrographolide, and dehydrated andrographolide; the broken line in the figure is the connecting line of the average values of the active ingredient contents.
[0054] Figure 2 It is the average spectral curve, original spectral curve and preprocessed spectral curve diagram of andrographis paniculata samples; among them, Figure 2 a. Average spectrum; Figure 2 b. Original spectrum; Figure 2 c. D1 preprocessing; Figure 2 d. D2 preprocessing; Figure 2 e. SG preprocessing; Figure 2 f. MSC preprocessing.
[0055] Figure 3 It is the scatter plot of the reference values and predicted values of the best prediction model for the content of andrographis paniculata compounds established based on the full wavelength; among them, Figure 3 a. Andrographolide; Figure 3 b. Neoandrographolide; Figure 3 c. Deoxyandrographolide; Figure 3 d. Dehydroandrographolide; Figure 3 e. Total content.
[0056] Figure 4 It is the characteristic wavelength bands screened by SPA; among them, Figure 4 a. Andrographolide; Figure 4 b. Deoxyandrographolide; Figure 4 c. Deoxyandrographolide; Figure 4 d. Dehydroandrographolide; Figure 4 e. Total content.
[0057] Figure 5 It is the scatter plot of the reference values and predicted values of the best prediction model for the content of andrographis paniculata compounds established based on the characteristic wavelengths; among them, Figure 5 a. Andrographolide; Figure 5 b. Neoandrographolide; Figure 5 c. Deoxyandrographolide; Figure 5 d. Dehydroandrographolide; Figure 5 e. Total content. Detailed implementation manners
[0058] The technical solutions of the present invention will be further described in detail below in conjunction with specific embodiments. It should be understood that the following embodiments are only used to illustrate and explain the present invention exemplarily, and should not be construed as limiting the protection scope of the present invention. All technologies implemented based on the above content of the present invention are covered within the scope of protection intended by the present invention.
[0059] Unless otherwise specified, the raw materials and reagents used in the following embodiments are all commercially available products, or can be prepared by known methods.
[0060] 1 Experimental part
[0061] 1.1 Andrographis paniculata sample
[0062] All samples were collected from Zhanjiang, Guangdong in September 2022, including 6 strains and 209 batches, including A-1 (30 batches), A-2 (50 batches), A-3 (31 batches), A-4 (68 batches), A-5 (19 batches) and A-6 (11 batches). All samples were identified as dried stems and leaves of Andrographis paniculata (Burm.f.) Nees, a plant of the Acanthaceae family. After collection, the samples were stored at room temperature and ground into powder, sieved through a 100-mesh sieve, and the powder was sealed in a polyethylene bag and refrigerated at 4°C. Each powder sample was used for hyperspectral data collection, and the samples were randomly divided into a training set and a prediction set in a ratio of 7:3 for subsequent modeling analysis.
[0063] 1.2 Main instruments and reagents
[0064] I-Class ACQUITY UPLC TM Ultra-high performance liquid chromatography (Waters, USA); reference substances andrographolide, neoandrographolide, deoxyandrographolide, and dehydroandrographolide (98%, Beijing BetterKang Biopharmaceutical Technology Co., Ltd.).
[0065] 1.3 Hyperspectral Imaging System
[0066] The hyperspectral imaging equipment is a HySpex series hyperspectral imaging spectrometer (Norsk Elektro OptikkA / S, Norway), which mainly consists of two hyperspectral cameras, two 150w / 12v halogen tungsten lamps (H-LAM Norsk Elektro Optikk, Norway), a CCD detector, a mobile platform, and the instrument's own computer and built-in software. The hyperspectral cameras are SN0605 VNIR visible-near infrared camera (Norsk Elektro Optikk, Norway) with a spectral range of 410-990nm and N3124 SWIR short-wave infrared camera (Norsk Elektro Optikk, Norway) with a spectral range of 950-2500nm. The spectral resolution of the camera lens is 5nm, and the distance from the sample is 25cm. The platform moving speed is 1.5nm / s.
[0067] 1.4 Hyperspectral image acquisition, correction, and data extraction of regions of interest
[0068] Grind the Andrographis paniculata sample into powder and pass it through a No. 4 sieve. Take about 2g of powder in a 35mm×12mm disposable petri dish and oscillate to make the powder surface flat and evenly distributed. Take 5 to 10 samples in turn and place them on a black horizontal moving platform, and place a Teflon whiteboard at the same time, scan, and obtain the Andrographis paniculata powder spectrum data. The hyperspectral acquisition and imaging system needs to be preheated before image acquisition. During the scanning process, the SN0605 VNIR visible-near infrared camera lens integration time is 4000us, and the frame time is 19000; the N3124 SWIR short-wave infrared camera lens integration time is 5400us, and the frame time is 49535.
[0069] In order to eliminate the influence of factors such as uneven light source distribution, unstable illumination and camera lens dark current on sample data during image acquisition, the original hyperspectral image needs to be RAD corrected using the instrument's built-in RAD correction software. Then, black and white correction is performed, and the correction formula is shown in (1):
[0070]
[0071] Where R is the corrected spectral data, Rraw is the original spectral data, Rw is the white reference data obtained from a white board with a reflectivity of 99%, and Rd is the dark reference data obtained by turning off the light and blocking the camera lens.
[0072] After correction, ENVI 5.3 software was used to extract the region of interest, and the average relative reflectance in the region of interest was the original spectral data of the sample.
[0073] 1.5 UPLC content determination
[0074] The UPLC method was used to determine the content of andrographolide, neoandrographolide, deoxyandrographolide, and dehydroandrographolide in the Andrographis paniculata sample. Accurately weigh 50 mg of Andrographis paniculata sample powder (andrographis paniculata sample powder under "1.4"), add 1.5 mL of 40% methanol solution to extract andrographolide, neoandrographolide, deoxyandrographolide, and dehydroandrographolide to obtain a test solution. Finally, the test solution was quantitatively detected using the UPLC method.
[0075] 1.6 Spectral data preprocessing
[0076] In order to reduce the errors caused by background, noise and other factors and improve the prediction ability and stability of the model, four methods, including multiplicative scatter correction (MSC), first-order derivative (D1), second-order derivative (D2) and SG smoothing (Savitzky-Golay, SG), were used to preprocess the spectral data.
[0077] 1.7 Model establishment
[0078] Three models, namely partial least squares regression (PLSR), back propagation neural network (BPNN), and random forest regression (RFR), were used to predict the content of Andrographis paniculata samples.
[0079] BPNN is a commonly used artificial neural network model, consisting of three or more neurons such as the input layer, hidden layer, and output layer. During the analysis, each neuron receives the input from the neurons in the previous layer and converts it into an output through an activation function. Through multiple iterative trainings, the connection weights between neurons are adjusted to make the actual value as close as possible to the predicted value. PLSR is a multivariate regression method used to establish a relationship model between input variables (X) and output variables (Y). It is applicable when there are multicollinearity or high-dimensional problems among input variables, and when there is a correlation among output variables. RFR is an ensemble learning method that establishes a regression model by combining multiple decision trees. It combines the simplicity of decision trees and the advantages of ensemble learning, and can effectively handle regression problems. RFR can handle high-dimensional data and a large number of samples, and has good robustness for data with multicollinearity.
[0080] 1.8 Model evaluation
[0081] The coefficient of determination (R 2 ), root mean squares error of calibration (RMSEC), root mean squares error of prediction (RMSEP), and residual predictive deviation (RPD) of the training set and prediction set were used to evaluate the performance of the regression prediction model. The closer R 2 is to 1, the smaller the RMSE value, and the larger the RPD value, the better the performance of the model. If 0.60 < R 2 < 0.80 and 1.5 < RPD < 2.5, it indicates that the model can be used for prediction; if 0.81 < R 2 < 0.90 and 2.51 < RPD < 3.0, it indicates that the model has good prediction performance; if R 2>0.90, RPD > 3.0 indicates that the model has excellent predictive ability.
[0082] 1.9 Data processing and analysis software
[0083] The image correction software is the HySpex RAD software of the hyperspectral imaging system (Norsk Elektro Optikk, Norway), the black and white calibration and region of interest extraction software is ENVI 5.3, and the spectral data preprocessing and prediction model construction software is Matlab 2020a.
[0084] 2 Results and discussion
[0085] 2.1 UPLC content analysis of the chemical components of Andrographis paniculata
[0086] The contents of Andrographis paniculata samples of different germplasms were determined, and the results are as Figure 1 shown. Among the four diterpenoid lactone compounds, the content of andrographolide is the highest, with a content range of 28.16 mg / g - 62.19 mg / g ( Figure 1 a); the content of dehydroandrographolide is the lowest, with a content range of 1.01 mg / g - 4.76 mg / g ( Figure 1 d); the content ranges of neoandrographolide and deoxyandrographolide are 2.56 mg / g - 8.09 mg / g ( Figure 1 b) and 1.03 mg / g - 10.25 mg / g ( Figure 1 c), respectively.
[0087] 2.2 Hyperspectral curve analysis
[0088] The average value of the original spectral data of Andrographis paniculata samples of different germplasms (the hyperspectral data of Andrographis paniculata powder under "1.4") was calculated, and the average spectral curve graph was drawn (as Figure 2a). Generally speaking, the average spectral curve change trends of Andrographis paniculata with different germplasms are relatively similar. The spectral curve forms an absorption peak at 550 nm, shows an upward trend within 670 - 1300 nm, especially within 670 - 960 nm where the reflectance increases sharply, and shows a fluctuating downward trend within the range of 1300 - 2500 nm with obvious absorption peaks and absorption valleys. Attributing the different peaks, the absorption peak near 550 nm corresponds to the fifth overtone of O - H; the absorption near 840 nm is related to the third overtone of C–H in C=C - H; the absorption near 970 nm may be caused by the second overtone of O - H; the absorption peak at 1120 nm and the absorption in the region of 1300 - 1400 nm are related to the second stretching overtone of C - H; the absorption at 1470 nm corresponds to the first overtone of O - H; the absorption at 1600 - 1700 nm is related to the first overtone of C - H. The absorption within 1950 - 2000 nm is related to the second stretching overtone of C=O in the ester bond; the absorption peak at 2000 nm is related to the stretching vibration and bending vibration of O - H, the absorption peak at 2200 nm is a combined absorption peak of C - H and C - O; the absorption near 2370 nm corresponds to the second overtone of C - H. The positions of the absorption peaks of Andrographis paniculata with different germplasms have relatively similar characteristics, but there are differences in absorption intensity, indicating that the types of chemical components of Andrographis paniculata with different germplasms are not very different, but the contents are different.
[0089] Plot the original spectral data and the pre - processed spectral data, and the results are shown in Figure 2 b - Figure 2 f. As Figure 2 shown in b, the original spectral map lines overlap, the baseline drifts severely, and the absorption peaks are difficult to observe directly. After pre - processing with D1 and D2, the spectral curve baseline is basically horizontal, and the absorption peak characteristics are obvious, especially after pre - processing with D2, but at the same time, noise is introduced ( Figure 2 c, Figure 2 d). Compared with Figure 2 b, Figure 2 the spectral curve in e is smoother, Figure 2 and the spectral curve spacing in f is significantly reduced.
[0090] 2.3 Prediction and evaluation of the contents of four diterpene lactone compounds in Andrographis paniculata
[0091] 2.3.1 Prediction of the contents of four diterpene lactone compounds in Andrographis paniculata based on full - wavelength data
[0092] Using the original hyperspectral data of Andrographis paniculata samples and combining with the UPLC content determination results, PLSR, BPNN, and RFR models for andrographolide, neoandrographolide, deoxyandrographolide, dehydrated andrographolide, and total content were established respectively. The results are shown in Table 1. The PLSR models for andrographolide, neoandrographolide, and total content have the best performance, followed by the BPNN models; for deoxyandrographolide and dehydrated andrographolide, the BPNN models have the best performance, followed by the PLSR models.
[0093] Table 1 Prediction results of different models for the full-band original spectral data of Andrographis paniculata samples
[0094]
[0095]
[0096] Then, the spectral data pretreated by D1, D2, SG, and MSC were combined with the content determination results to establish PLSR and BPNN models. The results are shown in Table 2.
[0097] In the prediction of andrographolide content, compared with the BPNN model, the PLSR model has better prediction performance, with higher R 2 and RPD values and lower RMSE values. Among them, the R 2 of the training set and prediction set of the SG-PLSR model are 0.841 and 0.829 respectively, the RPD value is 2.36, and the scatter plot of the prediction results is as shown in Figure 3 a. The slope and R 2 are both greater than 0.8, indicating that the SG-PLSR model can be used for the prediction of andrographolide content. In the prediction of neoandrographolide content, compared with the BPNN model, the PLSR model also shows better prediction performance. Among them, the R 2 of the training set and prediction set of the MSC-PLSR model are 0.895 and 0.783 respectively, the RPD value is 2.25, and the scatter plot of the prediction results is as shown in Figure 3 b. The slope is 0.933 and R 2 is 0.806, indicating that the MSC-PLSR model can be used for the prediction of neoandrographolide content. In the prediction of deoxyandrographolide content, the R 2 of the training set and prediction set of most models are both greater than 0.9, the RMSE values are all close to 0, and the RPD values are all greater than 3.0. Among them, the Raw Data-BPNN model has the best prediction performance, and the scatter plot of the prediction results is as shown in Figure 3 c. The slope is close to 1 and R 2 is greater than 0.9, indicating that the predicted value of the model is close to the reference value. In the prediction of dehydrated andrographolide content, most models have good prediction ability. Among them, the R of the training set and prediction set of the MSC-BPNN model2 is the largest, the RMSE value is the smallest, and the RPD value is the largest. The scatter plot of the prediction results of this model is as shown in Figure 3 d, the slope is close to 1, and the R 2 is greater than 0.9, indicating that the difference between the predicted value and the reference value is small. In the prediction of the total content, the R values of the training set and the prediction set of the Raw Data-PLSR, SG-PLSR, and D1-PLSR models 2 are all greater than 0.8, and the RPD values are all greater than 2.5. Among them, the R values of the training set and the prediction set of the Raw Data-PLSR model 2 are the largest, the RMSE value is the smallest, and the RPD value is the largest. The scatter plot of the prediction results of this model is as shown in Figure 3 e, the slope is 0.969, and the R 2 is 0.867.
[0098] Table 2 Prediction results of PLSR and BPNN models established after preprocessing the full-band spectral data of andrographis paniculata samples by different methods
[0099]
[0100] 2.3.2 Prediction of the contents of four diterpenoid lactones in andrographis paniculata based on characteristic wavelength data
[0101] According to the prediction results of the full wavelength, SPA is further combined to select characteristic variables. The specific models selected are: the SG-PLSR model for predicting andrographolide, the MSC-PLSR model for predicting neoandrographolide, the Raw Data-BPNN model for predicting deoxyandrographolide, the MSC-BPNN model for predicting dehydrated andrographolide, and the Raw Data-PLSR model for predicting the total content.
[0102] The characteristic wavelengths screened by SPA are as shown in Figure 4 a- Figure 4 e. Most of these characteristic wavelengths are distributed near the absorption peaks and absorption valleys. The number of characteristic wavelengths screened for andrographolide, neoandrographolide, deoxyandrographolide, dehydrated andrographolide, and total content are 19, 34, 29, 49, and 19 respectively, and the number of wavelengths is reduced to 4.80%, 8.59%, 7.32%, 12.37%, and 4.80% of the full wavelength respectively. Andrographolide, neoandrographolide, deoxyandrographolide, and dehydrated andrographolide have similar structures and all contain -C=C-, -COOR, and -OH functional groups. Combining Figure 2It can be seen that the characteristic wavelengths selected by SPA are closely related to the structures of the four diterpenoid lactone compounds. The characteristic wavelengths selected near 410 - 670 nm, 900 - 1000 nm, and 1470 nm may be related to the -OH of the four diterpenoid lactone compounds. The characteristic wavelengths selected near 800 - 900 nm may be related to the -C=C- of the four diterpenoid lactone compounds. The characteristic wavelengths selected in the regions of 1950 - 2000 nm and 2200 nm may be related to the -COOR of the four diterpenoid lactone compounds. The characteristic wavelengths selected near 1120 nm, 1300 - 1400 nm, 1600 - 1700 nm, and 2370 nm may be related to the -CH, -CH2, and -CH3 of the four diterpenoid lactone compounds. Therefore, the functional group information of Andrographis paniculata can be analyzed through its spectral information, and then the quality of Andrographis paniculata can be analyzed.
[0103] After screening by SPA and modeling, the prediction results are shown in Table 3. The R values of the training set and prediction set of the SG - SPA - PLSR model for andrographolide and the MSC - SPA - PLSR model for neoandrographolide 2 are both greater than 0.6, and the RPD values are both greater than 1.5. The scatter plots of the prediction results are as shown in Figure 5 a and Figure 5 b. The slopes and R 2 are both greater than 0.75, indicating that after extracting the characteristic wavelengths by SPA, not only the complexity of the content prediction models for andrographolide and neoandrographolide is reduced, but also the content prediction is achieved under the condition of limited wavelength numbers. In the content prediction of deoxyandrographolide, compared with the full - wavelength Raw Data - BPNN model, the Raw Data - SPA - BPNN model established after extracting the characteristic wavelengths has an RPD value increased by 4.42%. The scatter plot of the prediction result is as shown in Figure 5 c. The slope is close to 1, and the R 2 is greater than 0.9, and the prediction performance of the model is good. In the content prediction of dehydrated andrographolide, compared with the full - wavelength SG - BPNN model, the SG - SPA - BPNN model established after screening the characteristic wavelengths has an RPD value increased by 7.25%. The scatter plot of the prediction result is as shown in Figure 5 d. The slopes and R 2 are both greater than 0.9, and the predicted values of the model are close to the reference values. Therefore, in the content prediction of deoxyandrographolide and dehydrated andrographolide, after extracting the characteristic wavelengths by SPA, not only the modeling process is simplified, but also the irrelevant variables are removed, and the accuracy of the model is improved. In the prediction of the total content, the R values of the training set and prediction set of the Raw Data - SPA - PLSR model 2 are in the range of 0.6 - 0.8, the RPD value is greater than 2.0, and the scatter plot of the prediction result is as shown in Figure 5As shown in e, the slope is greater than 0.8, indicating that the model can be used for the prediction of the total content of andrographolide, neoandrographolide, deoxyandrographolide and dehydrated andrographolide.
[0104] Therefore, the characteristic wavelengths screened by SPA can effectively characterize the spectral information of the four diterpenoid lactone compounds, simplify the model, and provide a method reference for the development of a dedicated miniaturized hyperspectral device for the quality detection of Andrographis paniculata in the future.
[0105] Table 3 Prediction results of different combined models for the spectral data of Andrographis paniculata samples after SPA screening
[0106]
[0107]
[0108] 3 Conclusions
[0109] The best models for andrographolide, neoandrographolide, deoxyandrographolide, dehydrated andrographolide and total content established in the present invention are SG-PLSR, MSC-PLSR, Raw Data-SPA-BPNN, MSC-SPA-BPNN and RawData-PLSR respectively. The detection method based on hyperspectral imaging technology provided by the present invention can quickly and accurately detect the chemical component content of Andrographis paniculata, avoiding problems such as cumbersome detection operations, long time, high cost, and large consumption of reagents, and providing a new idea for the rapid and non-destructive detection of the quality of Andrographis paniculata.
[0110] The above describes the embodiments of the present invention. However, the present invention is not limited to the above embodiments. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for detecting the content of chemical components in Andrographis paniculata based on hyperspectral technology, the method comprising the following steps: Step (1): Obtain Andrographis paniculata samples; Step (2): Measure and obtain the hyperspectral data of the Andrographis paniculata samples; Step (3): Measure and obtain the content values of the chemical components in the Andrographis paniculata samples; Step (4): Use the hyperspectral data of the Andrographis paniculata samples in step (2) and the content values of the chemical components in the Andrographis paniculata samples in step (3) to train a model; Step (5): Measure the hyperspectral data of the Andrographis paniculata samples to be tested, input it into the trained model, and obtain the content values of the chemical components in the Andrographis paniculata samples to be tested.
2. The method according to claim 1, wherein the chemical components in the Andrographis paniculata samples are diterpene lactone compounds, preferably andrographolide compounds; further preferably, the andrographolide compounds are selected from any one or more of the following: andrographolide, neoandrographolide, deoxyandrographolide, dehydrated andrographolide.
3. The method according to claim 1 or 2, characterized in that In step (1), the Andrographis paniculata samples are Andrographis paniculata samples of different strains; Preferably, step (2) is specifically: Grind the Andrographis paniculata samples, sieve them, evenly distribute them on the surface of the container, and use a hyperspectral imaging device to obtain the hyperspectral data of each Andrographis paniculata sample.
4. The method according to any one of claims 1 to 3, characterized in that, In step (3), liquid chromatography is used to measure and obtain the content values of the chemical components in the Andrographis paniculata samples; Preferably, step (3) specifically comprises the following steps: (a) Grind the Andrographis paniculata samples, sieve them, and perform solvent extraction to obtain a test solution; (b) Use liquid chromatography for detection to obtain the content values of the chemical components in the test solution.
5. The method according to any one of claims 1 to 4, characterized in that, In step (4), the model is selected from: partial least squares regression PLSR, backpropagation neural network BPNN, random forest regression RFR; Preferably, when the chemical component to be detected is andrographolide or neoandrographolide, the model is a PLSR model, a BPNN model; preferably a PLSR model; Preferably, when the chemical component to be detected is deoxyandrographolide or dehydrated andrographolide, the model is a BPNN model, a PLSR model; preferably a BPNN model; Preferably, when the chemical component to be detected is four compounds including andrographolide, neoandrographolide, deoxyandrographolide, and dehydrated andrographolide, the model is a PLSR model, a BPNN model; preferably a PLSR model.
6. The method according to any one of claims 1-5, characterized in that, In step (2), optionally further comprising: preprocessing the obtained hyperspectral data of the Andrographis paniculata samples; Preferably, the preprocessing method is selected from: multiplicative scatter correction MSC, first derivative D1, second derivative D2, Savitzky-Golay smoothing SG.
7. The method according to claim 6, wherein When the chemical component to be detected is andrographolide and / or neoandrographolide and / or dehydrated andrographolide, preprocessing is performed based on the obtained hyperspectral data of the Andrographis paniculata samples; Preferably, when the chemical component to be detected is andrographolide, perform SG preprocessing on the obtained hyperspectral data of the Andrographis paniculata samples, and the model is selected from PLSR models; Preferably, when the chemical component to be detected is neoandrographolide, perform MSC preprocessing on the obtained hyperspectral data of the Andrographis paniculata samples, and the model is selected from PLSR models; Preferably, when the chemical component to be detected is dehydroandrographolide, the hyperspectral data of the andrographis paniculata sample obtained is preprocessed by MSC, and the model is selected from the BPNN model.
8. The method according to any one of claims 1 to 7, characterized in that, In step (2), optionally further included: based on the hyperspectral data of the andrographis paniculata sample obtained, screening the characteristic bands of the chemical components.
9. The method according to claim 8, characterized in that, The characteristic bands of andrographolide are as follows: 524nm, 660nm, 676nm, 746nm, 806nm, 876nm, 958nm, 1450nm, 1690nm, 1897nm, 1930nm, 1995nm, 2039nm, 2142nm, 2202nm, 2306nm, 2382nm, 2502nm, 2513nm; Preferably, the characteristic bands of neoandrographolide are as follows: 432nm, 475nm, 503nm, 530nm, 622nm, 643nm, 660nm, 676nm, 687nm, 698nm, 757nm, 811nm, 952nm, 990nm, 1314nm, 1336nm, 1357nm, 1374nm, 1385nm, 1390nm, 1516nm, 1586nm, 1674nm, 1717nm, 1794nm, 1870nm, 1897nm, 2044nm, 2110nm, 2246nm, 2273nm, 2306nm, 2371nm, 2502nm; Preferably, the characteristic bands of deoxyandrographolide are as follows: 421nm, 427nm, 443nm, 448nm, 459nm, 481nm, 508nm, 530nm, 551nm, 622nm, 643nm, 806nm, 876nm, 903nm, 936nm, 1434nm, 1597nm, 1897nm, 1924nm, 1995nm, 2055nm, 2208nm, 2246nm, 2273nm, 2306nm, 2377nm, 2480nm, 2491nm, 2497nm; Preferably, the characteristic wavelengths of dehydroandrographolide are as follows: 421 nm, 465 nm, 508 nm, 540 nm, 551 nm, 584 nm, 622 nm, 643 nm, 708 nm, 844 nm, 876 nm, 903 nm, 998 nm, 1145 nm, 1336 nm, 1347 nm, 1352 nm, 1363 nm, 1385 nm, 1390 nm, 1401 nm, 1407 nm, 1434 nm, 1663 nm, 1695 nm, 1717 nm, 1744 nm, 1804 nm, 1843 nm, 1864 nm, 1875 nm, 1924 nm, 1995 nm, 2028 nm, 2050 nm, 2110 nm, 2170 nm, 2208 nm, 2246 nm, 2300 nm, 2317 nm, 2344 nm, 2377 nm, 2404 nm, 2431 nm, 2491 nm, 2502 nm, 2508 nm, 2513 nm; Preferably, the characteristic wavelengths of the total content of four andrographolide compounds (andrographolide, neoandrographolide, deoxyandrographolide, dehydroandrographolide) are as follows: 416 nm, 421 nm, 432 nm, 448 nm, 535 nm, 622 nm, 643 nm, 952 nm, 990 nm, 1477 nm, 1695 nm, 1870 nm, 1924 nm, 2050 nm, 2148 nm, 2246 nm, 2311 nm, 2502 nm, 2513 nm.
10. The method according to any one of claims 1-9, characterized in that when the chemical component to be detected is andrographolide, the hyperspectral data of the obtained andrographis paniculata sample is preprocessed by SG, and then the characteristic wavelengths are screened, and the model is selected from the PLSR model; Preferably, when the chemical component to be detected is neoandrographolide, the hyperspectral data of the obtained andrographis paniculata sample is preprocessed by MSC, and then the characteristic wavelengths are screened, and the model is selected from the PLSR model; According to the embodiments of the present invention, when the chemical component to be detected is deoxyandrographolide, the characteristic wavelengths of the hyperspectral data of the obtained andrographis paniculata sample are screened, and the model is selected from the BPNN model; Preferably, when the chemical component to be detected is dehydrooxyandrographolide, the hyperspectral data of the obtained andrographis paniculata sample is preprocessed by MSC, and then the characteristic wavelengths are screened, and the model is selected from the BPNN model; Preferably, when the chemical component to be detected is four andrographolide compounds, the characteristic wavelengths of the hyperspectral data of the obtained andrographis paniculata sample are screened, and the model is selected from the PLSR model.