Lithocarpus litseifolius producing area distinguishing and chemical component content detecting method and device
By building a model using hyperspectral imaging technology and chemometric methods, the problems of rapid, no-complex-pretreatment-required, and environmentally friendly origin identification and chemical component content detection of Ligusticum chuanxiong were solved, thus achieving efficient origin identification and chemical component content detection.
Patent Information
- Application Number
- CN202510908730.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies make it difficult to achieve rapid, environmentally friendly identification of the origin of Ligusticum chuanxiong and detection of chemical component content, without the need for complex pretreatment, especially large-scale detection in on-site scenarios.
Hyperspectral imaging technology is combined with chemometric methods to construct an origin classification model and a content prediction model. By training the model with hyperspectral data, the origin identification and rapid detection of chemical component content can be achieved.
It achieves rapid origin identification and chemical composition content detection without complex preprocessing, improves detection efficiency, is environmentally friendly, and is suitable for large-scale detection in on-site scenarios.
Smart Images

Figure CN120801208A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of origin discrimination and content detection of Lithocarpus phyllochrous, and particularly relates to a method and device for origin discrimination and content detection of Lithocarpus phyllochrous. BACKGROUND
[0002] Lithocarpus phyllochrous is a Fagaceae plant, also known as Gantian or Tiantian, and has high medicinal and edible values. Lithocarpus phyllochrous mainly contains phloridzin and 3-hydroxyphloridzin chemical components, wherein, phloridzin is a key substance for exerting various biological activities, and 3-hydroxyphloridzin can improve insulin resistance and has a positive effect on the prevention and treatment of diabetes. However, the contents of phloridzin and 3-hydroxyphloridzin in Lithocarpus phyllochrous from different origins are quite different. Therefore, it is necessary to discriminate the origin and detect the content of chemical components of Lithocarpus phyllochrous.
[0003] At present, the traditional quality evaluation method of Lithocarpus phyllochrous mainly includes sensory evaluation and instrument analysis. However, the sensory evaluation is highly subjective, and although the instrument analysis such as liquid chromatography and liquid chromatography-mass spectrometry can solve the problem of highly subjective sensory evaluation, it needs complex pretreatment, takes a long time, has low efficiency, is not environmentally friendly, and is difficult to realize large-scale detection based on a field scene. Therefore, how to provide a method and device for origin discrimination and content detection of Lithocarpus phyllochrous, which is rapid, does not need complex pretreatment, is environmentally friendly, can realize large-scale detection in a field scene, has become a technical problem to be solved in the field. SUMMARY
[0004] The purpose of the present application is to provide a method and device for origin discrimination and content detection of Lithocarpus phyllochrous, which can be rapidly detected without complex pretreatment, improve the efficiency of origin discrimination and content detection of Lithocarpus phyllochrous, and is environmentally friendly and can realize large-scale detection in a field scene.
[0005] To achieve the above-mentioned purpose, the present application provides the following solutions.
[0006] In a first aspect, the present application provides a method for origin discrimination and content detection of Lithocarpus phyllochrous, which comprises the following steps:
[0007] obtaining origin information, hyperspectral data and content information of each chemical component of a Lithocarpus phyllochrous sample.
[0008] constructing an origin classification model, and training the origin classification model by using the hyperspectral data and the origin information of the Lithocarpus phyllochrous sample to obtain a trained origin classification model.
[0009] constructing a content prediction model, and training the content prediction model by using the hyperspectral data of the L. hypoglaucum samples and the content information of the chemical components, to obtain a trained content prediction model.
[0010] acquiring hyperspectral data of a to-be-tested L. hypoglaucum sample, inputting the hyperspectral data of the to-be-tested L. hypoglaucum sample into the trained origin classification model, to obtain an origin discrimination result of the to-be-tested L. hypoglaucum sample; and inputting the hyperspectral data of the to-be-tested L. hypoglaucum sample into the trained content prediction model, to obtain a chemical component content prediction result of the to-be-tested L. hypoglaucum sample.
[0011] In a second aspect, the present application provides a L. hypoglaucum origin discrimination and chemical component content detection system, which comprises the following functional modules:
[0012] a data acquisition module, configured to acquire origin information, hyperspectral data, and content information of chemical components of L. hypoglaucum samples.
[0013] an origin classification model construction and training module, configured to construct an origin classification model, and train the origin classification model by using the hyperspectral data of the L. hypoglaucum samples and the origin information, to obtain a trained origin classification model.
[0014] a content prediction model construction and training module, configured to construct a content prediction model, and train the content prediction model by using the hyperspectral data of the L. hypoglaucum samples and the content information of the chemical components, to obtain a trained content prediction model.
[0015] an origin discrimination and content prediction module, configured to acquire hyperspectral data of a to-be-tested L. hypoglaucum sample, input the hyperspectral data of the to-be-tested L. hypoglaucum sample into the trained origin classification model, to obtain an origin discrimination result of the to-be-tested L. hypoglaucum sample; and input the hyperspectral data of the to-be-tested L. hypoglaucum sample into the trained content prediction model, to obtain a chemical component content prediction result of the to-be-tested L. hypoglaucum sample.
[0016] In a third aspect, the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the L. hypoglaucum origin discrimination and chemical component content detection method according to any one of the above.
[0017] In a fourth aspect, the present application provides a computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the steps of the method for identifying the origin of L. suberifolium and detecting the content of chemical components according to any one of the preceding embodiments.
[0018] In a fifth aspect, the present application provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method for identifying the origin of L. suberifolium and detecting the content of chemical components according to any one of the preceding embodiments.
[0019] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0020] The present application provides a method and device for identifying the origin of L. suberifolium and detecting the content of chemical components. The present application uses hyperspectral imaging technology to replace traditional sensory evaluation and instrument analysis methods. The origin information, hyperspectral data and content information of each chemical component of the L. suberifolium sample are combined to form multi-dimensional comprehensive data of origin, spectrum and content. Then, taking the hyperspectral data as the main data, the hyperspectral data and the origin information are combined to train the origin classification model, so that the origin classification model fully learns the hyperspectral information of the L. suberifolium sample from different origins. The hyperspectral data and the content information of each chemical component are combined to train the content prediction model, so that the content prediction model fully learns the hyperspectral information and the content information of each chemical component. After obtaining the trained origin classification model and the trained content prediction model, only the hyperspectral data of the to-be-tested L. suberifolium sample is needed, which is input into the trained origin classification model and the trained content prediction model, respectively, so that the origin identification result and the chemical component content prediction result can be quickly obtained. Without complex pretreatment process, the origin of L. suberifolium and the content of chemical components can be simultaneously and efficiently detected, the efficiency of identifying the origin of L. suberifolium and detecting the content of chemical components is effectively improved, the environment is friendly, large-scale detection in a field scene can be realized, and the problems of traditional methods, such as complex pretreatment, long time consumption, low efficiency, unfriendliness to the environment and difficulty in realizing large-scale detection based on a field scene, are solved. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0022] Figure 1 An application environment diagram of a method for identifying the origin of L. suberifolium and detecting the content of chemical components provided by an embodiment of the present application is shown.
[0023] Figure 2 A flow chart of a method for identifying the origin of Ligusticum chuanxiong and detecting the content of chemical components provided in one embodiment of the present application.
[0024] Figure 3 This is a box plot of the phlorizin content in Zingiber officinale from different origins provided in one embodiment of the present application.
[0025] Figure 4 This is a box plot of the 3-hydroxyphlorizin content in Zingiber officinale from different origins provided in an embodiment of the present application.
[0026] Figure 5 This is the original spectral curve of samples of Ligusticum chuanxiong from different origins provided in one embodiment of the present application.
[0027] Figure 6 This is an average spectral curve diagram of samples of Ligusticum chuanxiong from different origins provided in one embodiment of the present application.
[0028] Figure 7 This is a schematic diagram of the confusion matrix of the prediction results of the SD-SVM classification model of Ligusticum chuanxiong based on SPA provided in one embodiment of the present application.
[0029] Figure 8 A schematic diagram of the confusion matrix of the prediction results of the SD-SVM classification model of Ligusticum chuanxiong based on CARS provided in one embodiment of the present application.
[0030] Figure 9 This is a graph showing the change in reflectivity of characteristic wavelengths of CARS screening based on phlorizin as a function of the number of wavelengths provided in one embodiment of the present application.
[0031] Figure 10 This is a graph showing the change in reflectivity of characteristic wavelengths of CARS screening based on 3-hydroxyphlorizin as a function of the number of wavelengths provided in one embodiment of the present application.
[0032] Figure 11 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0034] The related art research shows that the content of flavonoids, polyphenols and triterpenes in Lithocarpus polystachyus Rehder is rich, and it has the effects of reducing blood sugar, anti-allergy, anti-oxidation, improving memory, etc. Among them, phlorizin is the key substance for it to play multiple biological activities, and it has been widely used in the research and development of blood sugar lowering drugs, and is also an ingredient index for quality evaluation of Lithocarpus polystachyus Rehder. The related art research also shows that 3-hydroxyphlorizin can improve insulin resistance and has a positive effect on the prevention and treatment of diabetes. However, Lithocarpus polystachyus Rehder grows in the wild and is rarely artificially cultivated. The contents of phlorizin and 3-hydroxyphlorizin in Lithocarpus polystachyus Rehder from different producing areas differ greatly, and the quality is difficult to control. At present, the quality evaluation method of Lithocarpus polystachyus Rehder mainly includes sensory evaluation and instrument analysis. These traditional methods have the disadvantages of high cost and long time-consuming, and it is difficult to meet the demand of on-site rapid detection in the market scene. Therefore, it is urgent to develop a quality evaluation method with high efficiency, accuracy, high throughput and environmental friendliness.
[0035] Hyperspectral imaging technology is a new technology that combines traditional two-dimensional imaging technology and spectral technology to obtain a data cube. It has the characteristics of high resolution, rich band information and strong spatial recognition. Chemometrics is a powerful data mining and processing method, which can solve the problem of multi-dimension and complexity of modern instrument data. By integrating hyperspectral imaging technology and chemometrics methods, it has been successfully applied in the field of intelligent quality detection such as agricultural product origin traceability, adulteration identification and effective component content prediction. For example, Wang Zixuan et al. combined hyperspectral imaging technology with deep learning model to predict the content of total soluble solids of mulberry under different postharvest storage temperature conditions; Wang Danyun et al. realized the identification of watermelon from different producing areas based on hyperspectral imaging technology and self-adaptive enhancement network algorithm. At present, there is no reported literature on the spectral research of rapid discrimination of Lithocarpus polystachyus Rehder producing area and effective component content detection. Therefore, the present application aims to collect Lithocarpus polystachyus Rehder samples from four different producing areas of Jiangxi, Guizhou, Hunan and Yunnan, obtain their hyperspectral data, compare the performance of different combinations of preprocessing methods, characteristic wavelength screening algorithms and modeling methods, establish the optimal producing area discrimination and content prediction model, and realize the quality evaluation of Lithocarpus polystachyus Rehder that meets the market demand, and realize the accurate discrimination of Lithocarpus polystachyus Rehder producing area and the detection of chemical component content.
[0036] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below in combination with the drawings and specific embodiments.
[0037] The Lithocarpus polystachyus Rehder producing area discrimination and chemical component content detection method provided by the embodiments of the present application can be applied to, for example, Figure 1The application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data required by the server 104 to process. The data storage system can be set up separately, or integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the origin information, hyperspectral data and content information of each chemical component of the Lithocarpus polystachyus sample and the hyperspectral data of the to-be-tested Lithocarpus polystachyus sample to the server 104. After receiving the origin information, hyperspectral data and content information of each chemical component of the Lithocarpus polystachyus sample and the hyperspectral data of the to-be-tested Lithocarpus polystachyus sample, the server 104 obtains the origin information, hyperspectral data and content information of each chemical component of the Lithocarpus polystachyus sample; construct the origin classification model, and train the origin classification model using the hyperspectral data and origin information of the Lithocarpus polystachyus sample; construct the content prediction model, and train the content prediction model using the hyperspectral data and content information of each chemical component of the Lithocarpus polystachyus sample; input the hyperspectral data of the to-be-tested Lithocarpus polystachyus sample into the trained origin classification model to obtain the origin discrimination result; at the same time, input the hyperspectral data of the to-be-tested Lithocarpus polystachyus sample into the trained content prediction model to obtain the chemical component content prediction result. The server 104 can feed back the obtained origin discrimination result and chemical component content prediction result to the terminal 102. In addition, in some embodiments, the Lithocarpus polystachyus origin discrimination and chemical component content detection method can also be realized by the server 104 or the terminal 102 alone, such as the terminal 102 can directly perform origin discrimination and chemical component content detection processing on the origin information, hyperspectral data and content information of each chemical component of the Lithocarpus polystachyus sample and the hyperspectral data of the to-be-tested Lithocarpus polystachyus sample, or the server 104 can obtain the origin information, hyperspectral data and content information of each chemical component of the Lithocarpus polystachyus sample and the hyperspectral data of the to-be-tested Lithocarpus polystachyus sample from the data storage system, and perform origin discrimination and chemical component content detection processing on the origin information, hyperspectral data and content information of each chemical component of the Lithocarpus polystachyus sample and the hyperspectral data of the to-be-tested Lithocarpus polystachyus sample.
[0038] Among them, the terminal 102 can be but not limited to various desktop computers, notebook computers, smart phones, tablet computers and Internet of Things devices. The server 104 can be realized by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0039] In an exemplary embodiment, as Figure 2As shown, a method for identifying the origin of Lithocarpus polystachyus and detecting the content of chemical components is provided. The method is executed by a computer device, which can be a terminal or a server, or both. In the embodiments of the present application, the method is applied to Figure 1 The server 104 in the method is taken as an example for illustration, which includes the following steps S1 to S4.
[0040] S1: Obtain the origin information, hyperspectral data, and content information of each chemical component of the Lithocarpus polystachyus sample.
[0041] S2: Construct an origin classification model, and train the origin classification model using the hyperspectral data and origin information of the Lithocarpus polystachyus sample to obtain a trained origin classification model.
[0042] S3: Construct a content prediction model, and train the content prediction model using the hyperspectral data and content information of each chemical component of the Lithocarpus polystachyus sample to obtain a trained content prediction model.
[0043] S4: Obtain the hyperspectral data of the to-be-tested Lithocarpus polystachyus sample, and input the hyperspectral data of the to-be-tested Lithocarpus polystachyus sample into the trained origin classification model to obtain the origin identification result of the to-be-tested Lithocarpus polystachyus sample; at the same time, input the hyperspectral data of the to-be-tested Lithocarpus polystachyus sample into the trained content prediction model to obtain the content prediction result of the chemical components of the to-be-tested Lithocarpus polystachyus sample.
[0044] It should be noted that the Lithocarpus polystachyus sample in step S1 of the present embodiment refers to the training sample used in the model training stage, which is mainly used to train the origin classification model and the content prediction model using the origin information, hyperspectral data, and content information of each chemical component of the Lithocarpus polystachyus sample, so as to obtain the trained origin classification model and the trained content prediction model. Unlike the Lithocarpus polystachyus sample in step S1, the to-be-tested Lithocarpus polystachyus sample in step S4 refers to the target object to be detected in the model application stage, that is, after obtaining the trained origin classification model and the trained content prediction model, the hyperspectral data of the to-be-tested Lithocarpus polystachyus sample is input into the trained origin classification model and the trained content prediction model, respectively, so as to obtain the origin identification result and the content prediction result of the chemical components of the to-be-tested Lithocarpus polystachyus sample.
[0045] In the present embodiment, step S1 obtains the origin information, hyperspectral data, and content information of each chemical component of the Lithocarpus polystachyus sample, which specifically includes the following steps:
[0046] S11: Obtain a plurality of Lithocarpus polystachyus samples from different origins, and determine the corresponding origin information of each Lithocarpus polystachyus sample.
[0047] S12: Collecting the hyperspectral image of each of the Lithocarpus polystachyus sample to obtain original spectral data.
[0048] S13: Carrying out data correction and black-white correction on the original hyperspectral data respectively to obtain corrected spectral data.
[0049] S14: Extracting a region of interest from the corrected spectral data and determining the average relative reflectance in each of the regions of interest as the hyperspectral data of the Lithocarpus polystachyus sample.
[0050] S15: Carrying out chemical component content determination on each of the Lithocarpus polystachyus samples to obtain the content information of each chemical component of each of the Lithocarpus polystachyus samples.
[0051] In the embodiment, the differences in chemical component content of Lithocarpus polystachyus from different producing areas mainly reflect in phloridzin and 3-hydroxyphloridzin, and the edible value and medicinal value of phloridzin and 3-hydroxyphloridzin are the highest. Therefore, the chemical component content detection of Lithocarpus polystachyus in the embodiment mainly includes the most core phloridzin and 3-hydroxyphloridzin, and other components can be ignored.
[0052] After the step S1 of obtaining the producing area information, the hyperspectral data and the content information of each chemical component of the Lithocarpus polystachyus sample, the Lithocarpus polystachyus producing area discrimination and chemical component content detection method in the embodiment further includes a pretreatment step, which is specifically as follows:
[0053] The hyperspectral data of the Lithocarpus polystachyus sample is pretreated by using a SD (Second-order Derivative, second-order derivative) pretreatment method, a MSC (Multiple Scattering Correction, multiple scattering correction) pretreatment method or a SNV (Standard Normalized Variate, standard normal variate) pretreatment method to obtain pretreated hyperspectral data; the pretreated hyperspectral data is used as the hyperspectral data used when training the producing area classification model and the content prediction model.
[0054] It should be noted that, the same as the pretreatment process of the hyperspectral data of the Lithocarpus polystachyus sample, the hyperspectral data of the Lithocarpus polystachyus sample to be detected in the embodiment also needs to be pretreated and the characteristic wavelength for modeling is selected.
[0055] In this embodiment, the origin classification model is preferably an SD-CARS-SVM model. The SD-CARS-SVM model refers to a SVM (Support Vector Machine) model based on an SD preprocessing method and a CARS (Competitive Adaptive Reweighted Sampling) algorithm. The SD preprocessing method is used to perform SD processing on the hyperspectral data of the L. subcostata samples to obtain preprocessed hyperspectral data. The CARS algorithm is used to perform feature wavelength screening according to the preprocessed hyperspectral data and the origin information. The SVM model is used to classify the origins of the L. subcostata samples and output origin discrimination results.
[0056] In this embodiment, the content prediction model includes an SD-CARS-PLSR model and an MSC-CARS-PLSR model.
[0057] In this embodiment, the SD-CARS-PLSR model refers to a PLSR (Partial Least Squares Regression) model based on an SD preprocessing method and a CARS algorithm. The SD preprocessing method is used to perform SD processing on the hyperspectral data of the L. subcostata samples to obtain preprocessed hyperspectral data. The CARS algorithm is used to perform feature wavelength screening according to the preprocessed hyperspectral data and the content information of each chemical component. The PLSR model is used to predict the content of phlorizin in the L. subcostata samples and output a phlorizin content prediction result.
[0058] In this embodiment, the MSC-CARS-PLSR model refers to a PLSR model based on an MSC preprocessing method and a CARS algorithm. The MSC preprocessing method is used to perform MSC processing on the hyperspectral data of the L. subcostata samples to obtain preprocessed hyperspectral data. The CARS algorithm is used to perform feature wavelength screening according to the preprocessed hyperspectral data and the content information of each chemical component. The PLSR model is used to predict the content of 3-hydroxyphlorizin in the L. subcostata samples and output a 3-hydroxyphlorizin content prediction result.
[0059] In order to make the technical solutions of the present embodiment clearer, the implementation process of the technical solutions of the present embodiment will be described in the form of examples as follows.
[0060] In this embodiment, when collecting samples of Lithocarpus phyllochrous from different producing areas, considering that Lithocarpus phyllochrous is mainly distributed in Jiangxi, Guizhou, Hunan and Yunnan, etc., the collected samples of Lithocarpus phyllochrous (taking Lithocarpus phyllochrous leaves as an example) are from Jiangxi (15 samples), Guizhou (19 samples), Hunan (47 samples) and Yunnan (12 samples), a total of 93 samples. After collection, the Lithocarpus phyllochrous samples are ground into powder, sieved through a 100-mesh sieve, sealed in a polyethylene bag, and stored in a 4°C refrigerator for standby.
[0061] In this embodiment, when determining the content of the active ingredient, the instruments and reagents used include an I-Class type ACQUITY UPLC TM ultra-high performance liquid chromatography system (Waters Corporation, USA). Phlorizin, 3-hydroxyphlorizin (98%, Beijing Betin Kang Biomedicine Technology Co., Ltd.). Distilled water (Watsons); methanol, acetonitrile are chromatographically pure.
[0062] In this embodiment, when preparing the test solution, 20.0 mg of powder is weighed into a 2 mL centrifuge tube, 1.5 mL of 80% methanol is added, and the weight is weighed. Ultrasonic extraction for 30 min (power 250 W, frequency 4 kHz), let it cool to room temperature, weigh again, supplement the lost mass with 80% methanol, shake well. Centrifuge at 12000 r / min for 10 min, filter through a 0.22 μm microporous filter, and take the filtrate to obtain.
[0063] In this embodiment, when setting the chromatographic conditions, an ACQUITY UPLC BEH C18 chromatographic column (2.1 mm x 100 mm, 1.8 μm) is used, the column temperature is 40°C, the injection volume is 1 μL, the flow rate is 0.6 mL·min-1, and the mobile phase is 0.1% formic acid acetonitrile (A) and 0.1% formic acid water (B). Elution gradient: 0-1 min, 5%-25% A; 1-3.5 min, 25%-40% A; 3.5-4.5 min, 40%-60% A; 4.5-5 min, 60%-5% A; 5-7 min, 5% A.
[0064] In the present embodiment, the HySpex series hyperspectral imaging spectrometer (VIS-NIR-HIS, HySpex VNIR-1800 / HySpex SWIR384, Norsk Elektro Optikk, Oslo, Norway) was used for image acquisition during hyperspectral image acquisition and spectral data extraction. During acquisition, a black spot plate containing about 2 g of sample powder was placed on a horizontal moving platform, and a white plate was placed behind the sample. To ensure that the picture is not distorted, the distance between the sample and the lens is 32 cm, and the moving speed of the conveyor belt is 2.5 mm / s; the integration time of the VNIR lens is 4000 μs, the frame time is 19000, and the spectral range is 410-990 nm; the integration time of the SWIR lens is 4500 μs, the frame time is 46928, and the spectral range is 950-2500 nm; the spectral resolution of the two lenses is 6 nm. Subsequently, the original spectral data correction and black and white correction were performed, wherein the original spectral data correction was completed using the same software as the spectrometer (HySpex RAD Norsk Elektro Optikk, Norway), and the black and white correction was completed using the ENVI 5.3 software. Finally, the ENVI 5.3 software was used for region of interest extraction, and the average relative reflectance in the region of interest was the hyperspectral data of the L. subcostata sample.
[0065] In order to reduce errors caused by background, noise and other factors and improve the performance of the model, the present embodiment uses MSC pretreatment method, SD pretreatment method and SNV pretreatment method for pretreatment of spectral data. Among them, MSC is mainly used to eliminate the baseline shift and intensity change of diffuse reflectance spectrum caused by sample surface scattering and optical path change, which can effectively enhance the correlation between spectrum and chemical information, and is especially suitable for spectral correction of solid or particulate samples. SD is a pretreatment method based on mathematical differentiation. By processing the spectral data through SD, the baseline drift and background interference can be effectively eliminated, and the subtle features in the spectrum can be enhanced, thereby improving the spectral resolution. SNV is a zero-mean normalization method. The core idea is to center the mean and normalize the standard deviation of each spectrum, eliminate the spectral intensity change caused by scattering effect, and is suitable for eliminating the non-chemical related changes caused by sample preparation, measurement conditions or instrument differences, thereby enhancing the comparability and consistency of spectral data.
[0066] In this example, Partial Least Squares-Discriminant Analysis (PLS-DA), SVM and k-Nearest Neighbor for Classification (KNN-Class) were used to classify and identify Lithocarpus litseifolius from different origins; PLSR and k-Nearest Neighbor for Regression (KNN-Reg) were used to predict the content of chemical components in leaves. Among them, PLSR and PLS-DA are two multivariate modeling methods based on partial least squares algorithm. PLSR extracts orthogonal latent variables between spectral data matrix X and response variable matrix Y through iteration, with the goal of maximizing covariance, to establish the linear relationship between spectral characteristics and quantitative values. PLS-DA, as a variant of PLSR, is dedicated to classification scenarios, which needs to convert discrete category variables Y into binary encoding matrix, and through iteration to construct hidden space to maximize the difference between categories, and then generate discriminant function for sample classification decision. k-Nearest Neighbor (KNN) is a supervised learning method based on instances, which is widely used in classification and regression tasks. In the classification task, the algorithm calculates the distance measure of the sample to be tested and all training samples, selects the k samples with the shortest distance, and determines the predicted category by majority voting. For regression problems, the predicted value is usually the mean or weighted mean of the k nearest neighbor target variables, in order to reduce noise interference. Since KNN model belongs to non-parametric and lazy learning method, it is more sensitive to high-dimensional data, and sensitive to feature scale difference and unbalanced distribution. SVM model is a supervised learning algorithm based on statistical learning theory, and its core goal is to realize high generalization ability of classification decision by constructing maximum interval hyperplane. The performance of SVM model depends on the selection of kernel function and the regulation of regularization parameter, which balances the fitting accuracy and structural complexity of the model to the training data, and shows strong robustness in small sample, high-dimensional data and nonlinear scene.
[0067] In the selection of characteristic wavelength screening method, the successive projection algorithm (SPA) and CARS algorithm are used for characteristic wavelength screening. The SPA algorithm extracts the most representative characteristic variables by eliminating redundant information and multicollinearity. The core idea is based on the iterative process of orthogonal projection of vectors, and the variables that are linearly independent of each other are gradually screened out through projection operation, so as to construct a characteristic subset with strong explanatory ability and low correlation. The CARS algorithm combines Monte Carlo sampling and exponential decay function mechanism, and screens out the most effective characteristic variables through iterative competition. In each iteration, the CARS algorithm uses the absolute value of the regression coefficient in the PLS regression model as the weight to adaptively reweight the characteristic variables, and gradually eliminates unimportant variables through the exponential decay function, and finally retains the characteristics with the largest contribution to the model prediction.
[0068] In the model evaluation, the accuracy (Accuracy) is used to evaluate the performance of the classification model, and the calculation formula is as follows:
[0069]
[0070] Among them, TP is the number of true positive samples, TN is the number of true negative samples, FP is the number of false positive samples, and FN is the number of false negative samples.
[0071] The coefficient of determination (R 2 ), root mean square error (RMSE) and residual prediction error (RPD) are also used to evaluate the performance of the regression model. Among them, R 2 The closer to 1, the smaller the RMSE, the larger the RPD, and the better the performance of the model. If 0.61 < R 2 <0.80, 2.0 < RPD < 2.5, indicating that the model can be used for prediction; if 0.81 < R 2 <0.90, 2.5 < RPD < 3.0, indicating that the model has good prediction performance; if R 2 > 0.90, RPD > 3.0, indicating that the model has excellent prediction ability.
[0072] In the embodiment, the contents of phlorizin and 3-hydroxyphlorizin in samples from Jiangxi, Guizhou, Hunan and Yunnan are compared and analyzed, as shown in Figure 3 and Figure 4 , the broken line in Figure 3 and Figure 4 is the connecting line of the average content of chemical components. According to Figure 3 and Figure 4It can be found that the content of phloridzin in all samples from different producing areas is significantly higher than that of 3-hydroxyphloridzin. In addition, the content of the same component also has significant differences among different producing areas. The average content of phloridzin is 100.28 mg / g, among which the average content of samples from Yunnan is the highest, reaching 140.31 mg / g, while the average content of samples from Hunan, Guizhou and Jiangxi is lower than 100.28 mg / g, being 97.26 mg / g, 95.80 mg / g and 83.38 mg / g, respectively. Among 3-hydroxyphloridzin, the average content of samples from Jiangxi (6.40 mg / g) is significantly higher than that of samples from Yunnan (0.80 mg / g), Hunan (0.67 mg / g) and Guizhou (0.34 mg / g). These differences provide a scientific basis for the provenance tracing and quality evaluation of Lithocarpus phyllochlamys.
[0073] Figure 5 and Figure 6 The original spectrum curve and the average spectrum curve of Lithocarpus phyllochlamys samples from different producing areas are shown in FIGS. 1 and 2, respectively, from which Figure 5 and Figure 6 It can be seen from FIGS. 1 and 2 that the spectral curves of samples from different producing areas have similar trends, and all show obvious absorption peaks and absorption valleys at the same wavelength positions, indicating that Lithocarpus phyllochlamys from different producing areas has high similarity in chemical composition. However, there are significant differences in reflectivity at specific absorption peaks (such as 589 nm, 632 nm, 1204 nm, 1313 nm, 2218 nm). This may be due to the differences in soil conditions and climate in different producing areas, leading to differences in the content of certain compounds, and thus causing differences in spectral reflectivity, which provides a theoretical basis for subsequent hyperspectral provenance identification. In addition, the absorption in the 1000-1200 nm band may be closely related to the second overtone of C-H bond, and the absorption in the 1300-1400 nm region and near 2000 nm wavelength may be related to the stretching and bending vibration of O-H bond, which provides support for the content prediction of phloridzin and 3-hydroxyphloridzin in Lithocarpus phyllochlamys.
[0074] In the provenance identification of Lithocarpus phyllochlamys, PLS-DA, SVM and KNN-Class algorithms combined with SD, MSC and SNV pretreatment methods were used to establish the provenance classification model for the classification of Lithocarpus phyllochlamys. The provenance classification results of each provenance classification model are shown in Table 1.
[0075] Table 1 Provenance classification results based on full wavelength spectral data
[0076]
[0077] As can be seen from Table 1, in the KNN-Class model, the prediction set accuracy is significantly lower than the training set accuracy, whether based on original spectral data or based on preprocessed data modeling, and the model is over-fitted. This may be due to the inconsistent sample quantity of different origins, and KNN is sensitive to class balance, and the sample of the small class is submerged by the majority class. The PLS-DA model and the SVM model are global models, which can better handle unbalanced data and high-dimensional features. Even if the original spectral data is modeled, the prediction set accuracy is 96.4% and 92.9%, respectively. After preprocessing by the SD preprocessing method, the prediction set accuracy of the two models is improved to 100%, and after MSC and SNV preprocessing, the prediction set accuracy remains unchanged. This may be because the SD preprocessing method removes the baseline drift while enhancing the subtle information in the spectrum, making the spectra of different origins more easily distinguishable. While the MSC preprocessing method and the SNV preprocessing method can remove the scattering effect in the spectrum, some useful information is lost, and the model performance does not improve. Therefore, the SD-PLS-DA model (i.e., the combination of the SD preprocessing method and the PLS-DA model) and the SD-SVM model (i.e., the combination of the SD preprocessing method and the SVM model) are selected for subsequent analysis.
[0078] Hyperspectral data is huge, which contains a large amount of redundant information, and needs to screen the characteristic wavelengths to reduce the data dimension and simplify the model. In this embodiment, two characteristic wavelength screening methods, SPA and CARS, are used, and the origin classification results are shown in Table 2.
[0079] Table 2 Origin classification results based on characteristic wavelength spectral data
[0080]
[0081] As can be seen from Table 2, SPA only screens out 14 wavelengths, which is 3.55% of the full wavelength. CARS screens out 71 wavelengths, which is 18.2% of the full wavelength. Whether SPA or CARS is used, the performance of the PLS-DA model is significantly reduced, while the performance of the SVM model can achieve almost the same effect as the full wavelength. Figure 7 and Figure 8The confusion matrix of SD-SPA-SVM model (i.e. the combination of SD pretreatment method, SPA feature wavelength screening method and SVM model) and SD-CARS-SVM model (i.e. the combination of SD pretreatment method, CARS feature wavelength screening method and SVM model) on the prediction set is shown respectively. Among them, the SD-SPA-SVM model has 1 misjudgment in discriminating Jiangxi and Hunan samples, while the SD-CARS-SVM model achieves 100% accurate classification in all test samples. This shows that the feature wavelengths selected by CARS algorithm retain the spectral characteristics of origin specificity more effectively. However, considering the detection efficiency and cost in the actual application process, the SD-SPA-SVM model can be applied to solve the origin discrimination problem of L. subcostata within a certain error range.
[0082] In this embodiment, for the content prediction and evaluation of phlorizin and 3-hydroxyphlorizin in L. subcostata, firstly, based on the content prediction of full wavelength data, PLSR, KNN-Reg algorithm and SD, MSC, SNV pretreatment methods were applied to establish the content prediction model of phlorizin and 3-hydroxyphlorizin in L. subcostata samples. The content prediction results are shown in Table 3.
[0083] Table 3 Content prediction results of phlorizin and 3-hydroxyphlorizin based on full wavelength spectral data
[0084]
[0085] Overall, the prediction performance of KNN-Reg model is lower than that of PLSR model. In high-dimensional space, KNN-Reg model is easily affected by dimension disaster, and it is also difficult to find meaningful neighbors in small sample scenarios, resulting in poor prediction performance. While PLSR model can effectively handle small sample and high-dimensional data problems by reducing dimension and extracting important information, and the prediction result is better. The three pretreatment methods of SD, MSC and SNV can effectively handle errors caused by physical factors such as background and noise, and improve model accuracy. For phlorizin, the PLSR model established after SD treatment has the best prediction performance, and the R 2 of training set and prediction set are 0.88 and 0.85 respectively, and the RPD value is improved to 2.61. For 3-hydroxyphlorizin, the prediction effect of MSC-PLSR model is the best, and the R 2 of training set and prediction set are 0.89 and 0.81 respectively, and the RPD value is 2.31, which is improved by 26.23%.
[0086] In this embodiment, for the content prediction based on feature wavelength data, the SD-PLSR model of phlorizin and the MSC-PLSR model of 3-hydroxyphlorizin are respectively reduced in dimension, and the content prediction results are shown in Table 4.
[0087] Table 4 Prediction results of phloridzin and 3-hydroxyphloridzin content based on characteristic wavelength spectral data
[0088]
[0089] As shown in Table 4, CARS is a very effective characteristic wavelength screening method. For phloridzin, the model established based on the CARS screened characteristic wavelengths, the training set and the prediction set R 2 respectively increased to 0.97 and 0.93, and the RPD value was 3.78, which was about 44.83% higher than the performance of the full wavelength model.
[0090] Figure 9 Figure 4 is a graph of reflectance as a function of wavelength number based on CARS screening of characteristic wavelengths of phloridzin, Figure 10 Figure 5 is a graph of reflectance as a function of wavelength number based on CARS screening of characteristic wavelengths of 3-hydroxyphloridzin, from Figure 9 It can be seen that the 33 characteristic wavelengths screened are mainly distributed in the 410-670 nm, 973-1095 nm and 1646-2491 nm intervals. Among them, the dense distribution in the SWIR region may be related to the characteristic absorption of the phenolic hydroxyl group and the glycoside bond in the molecular structure of phloridzin, especially near 2100 nm (O-H) and 2300 nm (C-O-C). For 3-hydroxyphloridzin, the model was established based on the 19 characteristic wavelengths screened by CARS, and the training set and the prediction set R 2 respectively 0.90 and 0.81, and the RPD value was 2.27. This indicates that the characteristic wavelengths screened by CARS fully retain the spectral response information related to the component, so that the performance of the model can still be comparable to that of the model based on the full wavelength in the case of limited wavelength number. From Figure 10 It can be seen that these wavelengths are mainly concentrated in the 442-914 nm and 1237-2000 nm intervals. Among them, wavelengths such as 1237 nm, 1488 nm and 1602 nm are related to O-H primary frequency doubling and C-H stretching vibration. Overall, the CARS algorithm can accurately identify the characteristic waveband related to the structure of the target compound through the step-by-step shrinkage process of variable weights, laying a foundation for the development of portable hyperspectral detection equipment.
[0091] Lithocarpus phuthalanus has both medicinal and edible values, and has a long history of application. However, the traditional quality evaluation method does not have the characteristics of rapid, accurate, and environmentally friendly. In view of this problem, the present embodiment adopts hyperspectral imaging technology combined with chemometrics to establish the origin classification model and the effective component content prediction model of Lithocarpus phuthalanus, and to provide a new method for its quality control and evaluation. In the discrimination of the origin, three models of SVM model, PLS-DA model and KNN-Class model are used. The results show that when modeling based on full wavelength, the classification performance of SVM model is the best, PLS-DA model is the second, and KNN-Class model is the weakest. When modeling based on characteristic wavelength, SVM model still maintains high accuracy, among which the accuracy of SD-CARS-SVM model in the prediction set reaches 100%, and PLS-DA model appears overfitting phenomenon. The excellent performance of SVM model is not only related to its regularization mechanism to inhibit overfitting; it may also be attributed to the strong generalization ability of its kernel function to nonlinear data; in addition, the global optimization characteristics based on convex optimization of SVM model further ensure the stability and reliability of the model. The related technical research also verifies the advantages: in the discrimination of black tea quality grade, the SVM model based on radial basis kernel function achieves the optimal discrimination accuracy of 95.28%; in the adulteration detection of buckwheat flour, the accuracy of CARS-SVM model is 100%, which is significantly better than that of PLS-DA model; in the maturity classification of rapeseed, the accuracy of the SVM model fused with multi-band optimization reaches 97.86%, which is also better than other methods.
[0092] In terms of content prediction, the optimal prediction models of phlorizin and 3-hydroxyphlorizin are SD-CARS-PLSR model and MSC-CARS-PLSR model, respectively. Obviously, the optimal pretreatment methods of the two compounds are different. This may be due to the difference in their contents. The average content of phlorizin in Lithocarpus phuthalanus is 100.28 mg / g, while the average content of 3-hydroxyphlorizin is only 1.54 mg / g, which is much lower than that of phlorizin. SD is a baseline correction method, which can effectively eliminate baseline drift and enhance the subtle spectral characteristics such as shoulder peak and weak absorption peak, and is more suitable for high content of phlorizin; MSC belongs to scattering correction, which can correct scattering interference and improve signal-to-noise ratio, and is more conducive to the detection of low content of 3-hydroxyphlorizin which is easily disturbed by background. This is consistent with the research results of DAIY et al. in the detection of salvia miltiorrhiza components, that is, MSC is more suitable for the detection of low content of tanshinone I (0.42 mg / g).
[0093] In addition, the CARS algorithm and the SPA algorithm were compared in this example. The results showed that the CARS algorithm outperformed the SPA algorithm in both origin discrimination and content prediction. In origin discrimination, the SPA algorithm and the CARS algorithm extracted 3.55% and 18.2% of the characteristic wavelengths of the full wavelength, respectively. Although the SPA algorithm had higher dimension reduction efficiency, the test set accuracy (92.9%) was lower than that of the CARS algorithm (100%). According to the algorithm principle, the SPA algorithm is based on the orthogonality constraint of the vector, which can strictly control the collinearity of variables to avoid redundancy; the CARS algorithm allows a certain redundancy to cover more spectral information through Monte Carlo sampling and adaptive weighting. Therefore, the SPA algorithm has fewer characteristic wavelengths, while the CARS algorithm covers more spectral information, resulting in higher accuracy. In content prediction, although the proportion of wavelengths selected by the two methods was less than 10%, the prediction performance of the CARS algorithm was still better than that of the SPA algorithm, indicating that the CARS algorithm can better balance the model complexity and the retention of effective information.
[0094] In this example, based on hyperspectral imaging technology and chemometrics methods, the origin discrimination and effective component content prediction models of Lithocarpus polystachyus were established. Through the comparison and analysis of different pretreatment methods, characteristic wavelength screening methods and machine learning models, the optimal model was established. The best origin classification model for origin traceability was the SD-CARS-SVM model, with a prediction accuracy of 100%; the best content prediction models of phlorizin and 3-hydroxyphlorizin were the SD-CARS-PLSR model and the MSC-CARS-PLSR model, respectively, with prediction set R 2 2 values of 0.93 and 0.81, respectively, and RPD values of 3.78 and 2.27, respectively. These results indicate that, as a rapid, high-throughput and environmentally friendly method, hyperspectral imaging technology has great potential in the quality detection of Lithocarpus polystachyus.
[0095] Lithocarpus geniculatus is rich in bioactive dihydrochalcone glycosides such as phloridzin and 3-hydroxyphloridzin, but the content of which varies significantly with geographical origin. In this example, hyperspectral imaging technology combined with chemometrics was used to rapidly identify the geographical origin of Lithocarpus geniculatus and accurately predict the content of active ingredients. By collecting Lithocarpus geniculatus samples from four producing areas in Jiangxi, Guizhou, Hunan and Yunnan, the original spectral data was obtained, and the SD pretreatment method, MSC pretreatment method and SNV pretreatment method were used for pretreatment. The origin classification used PLS-DA model, SVM model and KNN-Class model; the chemical component content prediction used PLSR model and KNN-Reg model. In addition, the feature wavelengths were screened by SPA algorithm and CARS algorithm. The best origin classification model was SD-CARS-SVM model, and the prediction accuracy was 100%. In chemical component prediction, the best content prediction model of phloridzin was SD-CARS-PLSR model, the prediction set R 2 was 0.93, and RPD was 3.78; the best content prediction model of 3-hydroxyphloridzin was MSC-CARS-PLSR model, the prediction set R 2 was 0.81, and RPD was 2.27. This example verifies the reliability of hyperspectral imaging combined with chemometrics in the rapid detection of the geographical origin of Lithocarpus geniculatus and its chemical components, and provides a theoretical basis for its quality evaluation and intelligent detection.
[0096] Based on the same inventive concept, the examples of the present application also provide a Lithocarpus geniculatus origin discrimination and chemical component content detection system for implementing the above-mentioned Lithocarpus geniculatus origin discrimination and chemical component content detection method. The problem-solving implementation scheme provided by the system is similar to the implementation scheme described in the above method, so the specific limitations in the following Lithocarpus geniculatus origin discrimination and chemical component content detection system examples can be referred to the limitations of the Lithocarpus geniculatus origin discrimination and chemical component content detection method in the above text, which will not be repeated here.
[0097] In an exemplary embodiment, a Lithocarpus geniculatus origin discrimination and chemical component content detection system is provided, which mainly includes the following functional modules:
[0098] A data acquisition module is used to acquire the origin information, hyperspectral data and content information of each chemical component of the Lithocarpus geniculatus sample.
[0099] An origin classification model construction and training module is used to construct an origin classification model, and train the origin classification model using the hyperspectral data and origin information of the Lithocarpus geniculatus sample to obtain a trained origin classification model.
[0100] The content prediction model construction and training module is configured to construct a content prediction model, train the content prediction model by using the hyperspectral data of the L. subcostata sample and the content information of each chemical component, and obtain a trained content prediction model.
[0101] The origin discrimination and content prediction module is configured to obtain hyperspectral data of a to-be-tested L. subcostata sample, input the hyperspectral data of the to-be-tested L. subcostata sample into the trained origin classification model to obtain an origin discrimination result of the to-be-tested L. subcostata sample, and input the hyperspectral data of the to-be-tested L. subcostata sample into the trained content prediction model to obtain a chemical component content prediction result of the to-be-tested L. subcostata sample.
[0102] In an exemplary embodiment, a computer device is provided, which can be a server or a terminal, and an internal structure diagram thereof can be as shown in Figure 11 The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is configured to store origin information of L. subcostata samples, hyperspectral data and content information of each chemical component, and hyperspectral data of a to-be-tested L. subcostata sample. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is executed by the processor to implement a L. subcostata origin discrimination and chemical component content detection method.
[0103] Those skilled in the art can understand that Figure 11 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. A specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0104] In an exemplary embodiment, a computer device is provided, which includes a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the steps in each of the method embodiments described above.
[0105] In an example embodiment, a computer readable storage medium is provided, storing a computer program which, when executed by a processor, implements the steps of any of the above method embodiments.
[0106] In an example embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the steps of any of the above method embodiments.
[0107] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned method embodiments. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0108] The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0109] Any technical features in the above embodiments can be combined, and for the sake of brevity, not all possible combinations are described above, however, it should be understood that the application encompasses all possible combinations of the technical features unless such a combination is not technically possible.
[0110] The principles and implementation manners of the present application are described herein by using specific examples, and the above embodiments are only used to help understand the method of the present application and its core idea; meanwhile, according to the idea of the present application, the specific implementation manners and application scopes will be changed by those skilled in the art. In conclusion, the content of the present specification should not be understood as a limitation of the present application.
Claims
1. A method for identifying the origin of Ligusticum chuanxiong and detecting the content of chemical components, characterized in that: The method for identifying the origin of the leaves and branches of the Ligusticum chuanxiong and detecting the content of chemical components comprises: Obtain the origin information, hyperspectral data and content information of each chemical component of the samples of Ligusticum chuanxiong; Constructing an origin classification model, and training the origin classification model using the hyperspectral data and the origin information of the Ligusticum chuanxiong sample to obtain a trained origin classification model; Constructing a content prediction model, and training the content prediction model using the hyperspectral data of the Ligusticum chuanxiong sample and the content information of each chemical component to obtain a trained content prediction model; Acquire the hyperspectral data of the tested Ligusticum chuanxiong sample, and input the hyperspectral data of the tested Ligusticum chuanxiong sample into the trained origin classification model to obtain the origin discrimination result of the tested Ligusticum chuanxiong sample; at the same time, input the hyperspectral data of the tested Ligusticum chuanxiong sample into the trained content prediction model to obtain the chemical component content prediction result of the tested Ligusticum chuanxiong sample.
2. The method for identifying the origin of Ligusticum chuanxiong and detecting the content of chemical components of Ligusticum chuanxiong according to claim 1, wherein: Obtain the origin information, hyperspectral data and content information of each chemical component of the Ligusticum chuanxiong sample, including: Obtaining a number of samples of Ligusticum chuanxiong from different origins, and determining the origin information corresponding to each of the samples of Ligusticum chuanxiong; respectively collecting hyperspectral images of each of the Ligusticum chuanxiong samples to obtain original spectral data; performing data correction and black-white correction on the original hyperspectral data to obtain corrected spectral data; Extracting regions of interest from the corrected spectral data, and determining the average relative reflectance within each region of interest as hyperspectral data of the Ligusticum chuanxiong sample; The chemical component content of each of the Ligusticum chuanxiong samples was determined respectively to obtain the content information of each chemical component of each of the Ligusticum chuanxiong samples.
3. The method for identifying the origin of Ligusticum chuanxiong and detecting the content of chemical components of Ligusticum chuanxiong according to claim 2, wherein: The chemical components include phlorizin and 3-hydroxyphlorizin.
4. The method for identifying the origin of Ligusticum chuanxiong and detecting the content of chemical components of Ligusticum chuanxiong according to claim 3, wherein: After the step of obtaining the origin information, hyperspectral data and content information of each chemical component of the sample of Ligusticum chuanxiong, the method for identifying the origin of Ligusticum chuanxiong and detecting the content of chemical components further includes: The hyperspectral data of the Ligusticum chuanxiong sample are preprocessed using an SD preprocessing method, an MSC preprocessing method or an SNV preprocessing method to obtain preprocessed hyperspectral data; the preprocessed hyperspectral data are used as hyperspectral data for training the origin classification model and the content prediction model.
5. The method for identifying the origin of Ligusticum chuanxiong and detecting the content of chemical components of Ligusticum chuanxiong according to claim 4, wherein: The origin classification model is the SD-CARS-SVM model; The SD-CARS-SVM model refers to an SVM model based on the SD preprocessing method and the CARS algorithm; wherein the SD preprocessing method is used to perform SD processing on the hyperspectral data of the Ligusticum chuanxiong sample to obtain preprocessed hyperspectral data; the CARS algorithm is used to perform characteristic wavelength screening based on the preprocessed hyperspectral data and the origin information; the SVM model is used to classify the Ligusticum chuanxiong sample by origin and output the origin discrimination result.
6. The method for identifying the origin of Ligusticum chuanxiong and detecting the content of chemical components of Ligusticum chuanxiong according to claim 4, characterized in that: The content prediction models include SD-CARS-PLSR model and MSC-CARS-PLSR model; The SD-CARS-PLSR model refers to a PLSR model based on an SD preprocessing method and a CARS algorithm; wherein the SD preprocessing method is used to perform SD processing on the hyperspectral data of the Ligusticum chuanxiong sample to obtain preprocessed hyperspectral data; the CARS algorithm is used to perform characteristic wavelength screening based on the preprocessed hyperspectral data and the content information of each chemical component; the PLSR model is used to predict the content of phlorizin in the Ligusticum chuanxiong sample and output a phlorizin content prediction result; The MSC-CARS-PLSR model refers to a PLSR model based on the MSC preprocessing method and the CARS algorithm; wherein the MSC preprocessing method is used to perform MSC processing on the hyperspectral data of the Ligusticum chuanxiong sample to obtain preprocessed hyperspectral data; the CARS algorithm is used to perform characteristic wavelength screening based on the preprocessed hyperspectral data and the content information of each chemical component; the PLSR model is used to predict the content of 3-hydroxyphlorizin in the Ligusticum chuanxiong sample and output the 3-hydroxyphlorizin content prediction result.
7. A system for identifying the origin of Ligusticum chuanxiong and detecting the content of chemical components, characterized in that: The system for identifying the origin of the leaves and branches of the Ligusticum chuanxiong and detecting the content of chemical components comprises: A data acquisition module is used to obtain the origin information, hyperspectral data and content information of each chemical component of the samples of Ligusticum chuanxiong; An origin classification model construction and training module is used to construct an origin classification model and train the origin classification model using the hyperspectral data and the origin information of the Ligusticum chuanxiong sample to obtain a trained origin classification model; A content prediction model construction and training module is used to construct a content prediction model and train the content prediction model using the hyperspectral data of the Ligusticum chuanxiong sample and the content information of each chemical component to obtain a trained content prediction model; The origin discrimination and content prediction module is used to obtain the hyperspectral data of the Ligusticum chuanxiong sample to be tested, and input the hyperspectral data of the Ligusticum chuanxiong sample to be tested into the trained origin classification model to obtain the origin discrimination result of the Ligusticum chuanxiong sample to be tested; at the same time, the hyperspectral data of the Ligusticum chuanxiong sample to be tested is input into the trained content prediction model to obtain the chemical component content prediction result of the Ligusticum chuanxiong sample to be tested.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the method for identifying the origin of the Ligusticum chuanxiong and detecting the chemical component content of any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying the origin of the Ligusticum chuanxiong and detecting the content of chemical components of any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for identifying the origin of the Ligusticum chuanxiong and detecting the content of chemical components of any one of claims 1 to 6 is implemented.