Method for rapidly determining content of one or more heavy metals in gastrodia elata tubers through laser-induced breakdown spectroscopy

By combining laser-induced breakdown spectroscopy with chemometrics and characteristic variable screening, the problems of long detection time and low accuracy in the detection of heavy metals in Gastrodia elata were solved, realizing rapid and accurate detection of heavy metals in Gastrodia elata tubers and improving detection speed and accuracy.

CN121933497APending Publication Date: 2026-04-28JINAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610146050.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-02
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies for heavy metal detection in Gastrodia elata suffer from problems such as long detection time, detection lag, and low accuracy of characteristic spectral lines and full spectrum modeling during LIBS spectroscopy, making it difficult to achieve rapid and accurate detection of heavy metal content.

Method used

A predictive model for heavy metal content in Gastrodia elata tubers was established by combining laser-induced breakdown spectroscopy with chemometrics, modeling with a one-dimensional convolutional neural network (1D-CNN), and using a fusion strategy of wavelength interval selection and wavelength point selection to screen variables.

Benefits of technology

This technology enables rapid and accurate detection of heavy metals in Gastrodia elata tubers, improving detection speed and accuracy, and providing important references for quality evaluation and food safety monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121933497A_ABST
    Figure CN121933497A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of traditional Chinese medicine detection, and particularly relates to a method for rapidly determining the content of one or more heavy metals in gastrodia elata tubers through laser-induced breakdown spectroscopy. The method for rapidly determining the content of one or more heavy metals in gastrodia elata tubers through laser-induced breakdown spectroscopy comprises the following steps: S1, pretreating a gastrodia elata sample, and preparing a tablet as a to-be-detected sample; s2, establishment of initial spectral data: detecting the tablet prepared in the step S1 by adopting a laser-induced breakdown spectroscopy instrument, and collecting the spectral data; s3, measuring heavy metals in the gastrodia elata sample; s4, preprocessing the initial spectral data, and establishing a gastrodia elata heavy metal element content prediction model according to the heavy metal content data of the gastrodia elata sample in combination with a chemometrics method. The detection method provided by the invention can accurately and efficiently realize rapid quantitative detection of the heavy metal elements in the gastrodia elata sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for quality testing of Gastrodia elata tubers, specifically a method for rapidly determining the content of one or more heavy metals in Gastrodia elata tubers using laser-induced breakdown spectroscopy. Background Technology

[0002] Gastrodia elata (Gastrodia elata Blume, G. elata) is an edible product widely cultivated in Asia. Its current uses have transcended its traditional medicinal applications, gradually becoming a widely used functional food ingredient with multiple health benefits worldwide. The China Food and Drug Administration's designation of Gastrodia elata as a functional food ingredient aligns with contemporary dietary trends and emphasizes the natural and healthy nutrients in it. In particular, the use of Gastrodia elata for neuroprotection and cognitive enhancement further highlights its advantages as a functional food, catering to the evolving needs of health-conscious consumers globally.

[0003] In recent years, industrial pollution has exposed soil to the risk of contamination, potentially causing heavy metals such as copper (Cu), cadmium (Cd), mercury (Hg), and lead (Pb) in Gastrodia elata to become toxic as their concentration increases. This not only harms plant growth but also accumulates within the plant over time, adversely affecting human health. Furthermore, studies have shown that Gastrodia elata possesses a strong adsorption capacity for these four heavy metals. Therefore, it is essential to establish a simple, rapid, and accurate model for detecting heavy metal content to ensure the quality and safety of Gastrodia elata for consumption.

[0004] Currently, heavy metal detection primarily employs methods such as atomic absorption spectroscopy (AAS), atomic fluorescence spectroscopy (AFS), inductively coupled plasma mass spectrometry (ICP-MS), inductively coupled plasma atomic emission spectrometry (ICP-AES), X-ray fluorescence spectrometry (XRF), and atomic transfer reaction fluorescence spectrometry (ATR-FTIR). However, while these methods yield accurate results, they also have numerous drawbacks, including complex pretreatment, cumbersome operation, high destructiveness, time-consuming and labor-intensive processes, the need for specialized personnel, expensive chemical reagents, and environmental pollution, failing to meet the requirements for simple, non-destructive, and rapid detection. Laser-induced breakdown spectroscopy (LIBS), as an atomic emission spectrometry technique, offers advantages such as rapid detection, no pollution, minimal sample preparation requirements, and multi-element analysis capabilities, making it widely used for elemental monitoring in biomedicine, environment, and food fields. LIBS, in particular, has recently attracted attention as a rapid detection method for toxic metals in agriculture, crucial for ensuring food safety and monitoring environmental pollution. Examples include the use of LIBS to determine heavy metals in Sargassum, the rapid determination of three heavy metals in Fritillaria thunbergii, and the rapid determination of potentially toxic metals in soil. However, no research has yet been reported on the application of LIBS technology in the study of Gastrodia elata.

[0005] Although LIBS uses high-power pulsed laser beams to generate thermal plasma containing atoms and ions, which then emits characteristic emission lines upon cooling, enabling real-time and rapid analysis of multiple elements in samples, its quantitative capabilities require a combination of chemometrics and band selection methods. The choice of these methods is crucial because an element often corresponds to multiple characteristic spectral lines, and traditional univariate analysis is insufficient for precise quantification. Furthermore, since Gastrodia elata is a natural product derived from its edible underground parts, matrix effects may interfere with elemental content determination. LIBS full-spectrum analysis, however, contains a wealth of information, including potential matrix effects or other redundant data.

[0006] Therefore, in order to further improve the accuracy of the heavy metal content in Gastrodia elata tubers, the selection of effective variables during modeling is the key to quality detection using LIBS spectroscopy. Summary of the Invention

[0007] The purpose of this invention is to overcome the problems of long detection time, detection lag, and low accuracy of characteristic spectral lines and full spectrum modeling in traditional heavy metal detection techniques for Gastrodia elata tubers, and to provide a method for rapid determination of one or more heavy metals in Gastrodia elata tubers by laser-induced breakdown spectroscopy, so as to improve detection speed and accuracy.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: a method for rapid determination of the content of one or more heavy metals in Gastrodia elata tubers using laser-induced breakdown spectroscopy, comprising the following steps:

[0009] S1. Pre-treat the Gastrodia elata sample and compress it into tablets as samples to be tested;

[0010] S2. Establishment of initial spectral data: The pellet obtained in step S1 is tested using a laser-induced breakdown spectrometer, and spectral data is collected.

[0011] S3. Determination of the content of one or more heavy metals in Gastrodia elata tubers;

[0012] S4. Preprocess the initial spectral data and combine it with chemometric methods. At the same time, based on the heavy metal content data of the Gastrodia elata sample, establish a predictive model for the heavy metal content of Gastrodia elata.

[0013] Preferably, the chemometric method described in step S4 includes at least one of partial least squares regression (PLSR), k-nearest neighbors (KNN), support vector regression (SVR), and one-dimensional convolutional neural network (1D-CNN) modeling algorithms.

[0014] More preferably, the chemometric method described in step S4 is a one-dimensional convolutional neural network modeling algorithm.

[0015] The 1D-CNN model based on the combination of wavelength spacing selection and wavelength point selection in this invention can greatly simplify the number of variables and has higher accuracy and generalization ability.

[0016] Preferably, step S4, after preprocessing the initial spectral data, also includes a variable filtering operation.

[0017] More preferably, the variable selection is performed by combining the backward interval partial least squares algorithm (bi-PLS) with at least one of the competitive adaptive reweighted sampling (CARS), genetic algorithm (GA), and random forest (RF) algorithm.

[0018] This invention employs four specific feature variable selection methods, belonging to a two-step fusion strategy combining wavelength interval selection (bi-PLS) and wavelength point selection (CARS, RF, GA). The latter method further optimizes the variables selected by the former. Its advantage lies in the fact that the wavelength interval selection algorithm can eliminate a large amount of redundant information while achieving optimal feature interval selection, effectively narrowing the variable space. Based on this, the wavelength point selection algorithm is then applied to further filter out spectral features highly correlated with the target label, determining the optimal variable combination and improving the predictive performance of the LIBS quantitative model.

[0019] This invention provides a method for rapid determination of one or more heavy metal contents in Gastrodia elata tubers using laser-induced breakdown spectroscopy. Only by employing the four processing methods described above can a satisfactory detection and analysis result be achieved. Because using bi-PLS, CARS, RF, or GA alone results in a large number of variables, LIBS spectroscopy can contain redundant information, leading to poor model results. However, the combined application significantly improves the robustness and predictive performance of the model. This invention utilizes the advantages of wavelength interval selection and wavelength point selection, combining the continuity characteristics of LIBS spectroscopy for the first time, and applying a fusion strategy of wavelength interval selection and wavelength point selection to the detection of heavy metal contents in Gastrodia elata. This enables accurate and efficient rapid quantitative detection of one or more heavy metal elements in Gastrodia elata tubers.

[0020] Preferably, the preprocessing in step S4 includes the following steps: using wavelet transform (WT) to denoise the laser-induced breakdown spectrum, then using baseline correction (BC) to eliminate baseline drift or baseline offset, so that the baseline of the signal becomes flat or stable, and finally using normalization to enhance the stability of the spectral signal.

[0021] Preferably, the preparation method of the tablet in step S1 includes the following steps: crushing the Gastrodia elata tuber and sieving it (60 mesh sieve), weighing the sample powder from it and placing it under a tablet press to prepare a tablet with a diameter of 10~15mm and a thickness of 3~5mm.

[0022] More preferably, the diameter of the compressed tablet is 13 mm and the thickness is 5 mm.

[0023] Preferably, the method for measuring heavy metals in the Gastrodia elata tuber in step S3 is as follows: the heavy metal content in the Gastrodia elata sample is determined by inductively coupled plasma mass spectrometry.

[0024] More preferably, the heavy metal elements measured in the Gastrodia elata sample are Cu, Cd, Hg and Pb.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] (1) This invention is the first to apply LIBS spectroscopy combined with chemometrics to the determination of one or more heavy metals (Cu, Cd, Hg, Pb) in Gastrodia elata, which can achieve accurate, efficient and rapid detection, and provides an important reference for the quality evaluation of Gastrodia elata and the safety monitoring of complex systems such as food.

[0027] (2) Based on LIBS spectroscopy combined with chemometrics, this invention employs a fusion strategy of specific wavelength interval selection and wavelength point selection for variable screening. This greatly simplifies the number of variables and significantly improves the accuracy of rapid detection of heavy metals in Gastrodia elata samples. The resulting prediction model also has higher adaptability and generalization ability. In future research, accurate and efficient quantitative models of heavy metal elements can also be embedded into LIBS instruments to replace traditional complex chemical analysis methods for detecting heavy metal content in Gastrodia elata. Attached Figure Description

[0028] Figure 1 This is a statistical histogram showing the content of four heavy metals in the Gastrodia elata sample of this invention.

[0029] Figure 2 This is a schematic diagram comparing the original (black), wavelet transform (red), wavelet transform, and baseline-corrected (blue) LIBS spectra of the Gastrodia elata sample of this invention.

[0030] Figure 3 This is a LIBS spectrum curve of Gastrodia elata after wavelet transform and baseline correction according to the present invention.

[0031] Figure 4 The correlation coefficient diagram of reference values ​​and predicted values ​​of heavy metal elements in Gastrodia elata in Example 1 of the present invention is shown (a: Cu, b: Pb, c: Hg, d: Pb).

[0032] Figure 5 This is a diagram showing the position of Cu element retained in the LIBS spectrum under different variable selection methods in Embodiment 2 of the present invention.

[0033] Figure 6 This is a diagram showing the position of Cd element in the LIBS spectrum under different variable selection methods in Embodiment 2 of the present invention.

[0034] Figure 7 This is a diagram showing the position of Hg element in LIBS spectra under different variable selection methods in Embodiment 2 of the present invention.

[0035] Figure 8 This is a diagram showing the position of Pb element retained in the LIBS spectrum under different variable selection methods in Embodiment 2 of the present invention.

[0036] Figure 9 The correlation coefficient diagram of reference values ​​and predicted values ​​of heavy metal elements in Gastrodia elata in Example 2 of the present invention is shown (a: Cu, b: Pb, c: Hg, d: Pb).

[0037] Figure 10 The correlation coefficient diagram of reference values ​​and predicted values ​​of heavy metal elements in Gastrodia elata in Comparative Example 1 of the present invention is shown (a: Cu, b: Pb, c: Hg, d: Pb).

[0038] Figure 11 The correlation coefficient diagram of reference values ​​and predicted values ​​of heavy metal elements in Gastrodia elata in Comparative Example 2 of the present invention is shown (a: Cu, b: Pb, c: Hg, d: Pb). Detailed Implementation

[0039] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0040] Unless otherwise specified, the experimental methods used in the examples and comparative examples are conventional methods, and the materials and reagents used are commercially available unless otherwise specified.

[0041] The experimental samples and methods used in the embodiments and comparative examples of this invention are as follows:

[0042] I. Experimental Samples

[0043] To obtain experimental samples, *Gastrodia elata* was artificially contaminated with Cu, Cd, Hg, and Pb. The experiment was conducted using soil cultivation in plastic pots (20cm × 30cm), with each pot containing 10kg of soil. The metals were mixed according to the designed concentration gradient and prepared as metal salts (…). , , , When added evenly to a flowerpot, the concentration gradient of each metal is as follows: (50, 100, 200, 300, 400, 500 mg / kg), (1, 5, 10, 15, 20, 25 mg / kg), (1, 5, 10, 15, 20, 25 mg / kg) and (50, 100, 200, 300, 400, 500 mg / kg), with soil without added heavy metals as a control, to simulate soils contaminated with different concentrations of heavy metals. Gastrodia elata was planted until it flowered and was harvested.

[0044] After collecting the Gastrodia elata samples, the surface soil was removed, and the samples were rinsed with distilled water to remove surface adhering substances (residual heavy metals, salts, etc.). Acidic or chelating reagents were not used to avoid altering the endogenous heavy metal content of the samples. The samples were then dried in a 45℃ oven to constant weight (the difference between two weighings 30 minutes apart was considered to be <0.1%). Each sample was pulverized and passed through a 60-mesh sieve. 2g of the powder was weighed and pressed into a tablet press at 30MPa for 5 minutes. A total of 280 tablets with a diameter of 13mm and a thickness of approximately 5mm were prepared for LIBS data acquisition.

[0045] II. Experimental Methods

[0046] 2.1 Acquisition of LIBS spectral data

[0047] The Gastrodia elata samples were placed in a LIBS instrument (MX2500+, Ocean Optics Co., Ltd.) using a solid-state laser—an Nd:YAG laser (Quantel, Big Sky Laser Ultra50)—with a wavelength of 1064 nm. The repetition rate was 10 Hz, the pulse width was 5-7 ns, the energy was 50 mJ, the beam diameter was 3 mm, and the beam quality M²=3.3. Spectra in the range of 198.71 nm–629.08 nm were acquired, with a spectral resolution of 0.1 nm. To reduce experimental errors and the influence of uneven elemental distribution, this invention acquired 6 LIBS spectra from each of the two sides of each compressed sample, and calculated the average spectrum from both sides to obtain the final LIBS spectrum of the sample, forming a LIBS spectral dataset.

[0048] 2.2 Measurement of heavy metals in Gastrodia elata samples

[0049] After LIBS data acquisition, each compressed Gastrodia elata sample was pulverized, and 0.3g of Gastrodia elata powder was accurately weighed out and placed in a microwave digestion vessel. 65% of the powder was then added. Digest 5 ml of the solution. Filter the digest and dilute with water to a final volume of 10 ml to obtain the sample solution. Then, according to the determination method of the People's Republic of China National Standard (GB5009.268-2016), the content of four heavy metals in Gastrodia elata was determined by inductively coupled plasma mass spectrometry (ICP-MS). The content statistics are as follows: Figure 1 As shown. By Figure 1The statistical results show that the content distribution trends of the calibration set and the prediction set are similar. This distribution characteristic is beneficial for establishing and evaluating the stability and generalization ability of the content calibration model. The results are shown in Table 1, indicating that the concentration distribution ranges of the four heavy metal elements are all relatively wide.

[0050] Table 1. Statistical table of heavy metal content in Gastrodia elata samples (mg / kg)

[0051] 2.3 Partitioning and Spectral Preprocessing of the Sample LIBS Spectral Dataset

[0052] This invention employs the KS algorithm to divide the LIBS spectral dataset into a training set and a test set at a 3:1 ratio, which are used respectively for training and calibrating the optimal calibration model and predicting its performance. The specific details are as follows: First, the two samples with the greatest Euclidean distance from the spectral dataset are selected as the initial training set samples. Then, the Euclidean distances between the remaining samples and the selected samples in the training set are calculated. According to the "maximize minimum distance" criterion, the sample with the greatest Euclidean distance from the current training set is added to the training set. This process is iterated until the number of training set samples reaches a specified number. The training set and the test set consist of 398 and 133 samples, respectively.

[0053] Typically, during the acquisition of LIBS spectra, the physical properties of the sample itself or systematic errors may cause undesirable phenomena such as background noise, baseline drift, or baseline shift in the spectrum. Therefore, this invention employs spectral preprocessing methods such as normalization, wavelet transform (WT), and baseline correction (BC) to reduce system noise and enhance the characteristic signals of LIBS spectra.

[0054] Figure 2 The image shows the average LIBS spectrum of Gastrodia elata in the 198-630 nm range, and the results after WT and BC preprocessing. The LIBS spectrum shows that the lower the heavy metal content in Gastrodia elata, the more the characteristic sensitive spectral lines are interfered with by other high-content matrices. This invention uses the National Institute of Standards and Technology (NIST) atomic spectral database to determine the characteristic spectral lines of Cu, Cd, Hg, and Pb, such as... Figure 3 As shown.

[0055] 2.4 Analysis of Multivariate Data

[0056] 2.4.1 Selection of Chemometric Methods (Calibration Model)

[0057] Partial Least Squares Regression (PLSR) is a classic chemometric method that builds a model by finding the optimal linear relationship between the input variable (LIBS spectrum) and the output variable (heavy metal content in Gastrodia elata). The number of principal components in PLSR determines the amount of key information retained in the model. By selecting an appropriate number of principal components, sufficient information can be retained while reducing model complexity. This ensures that the model can accurately fit the training data and generalize well to unseen test data.

[0058] K-Nearest Neighbors (KNN) is an instance-based supervised learning method that performs regression prediction by calculating the distance between the input variable (LIBS spectrum) and the output variable (heavy metal content in Gastrodia elata). The KNN algorithm finds the K nearest training samples to the sample to be predicted and makes predictions based on the labels or values ​​of these neighbors. The key parameter K represents the number of nearest neighbors, which determines the model's complexity and the stability of the prediction results. By adjusting the value of K, the number of neighbor samples can be controlled, thus affecting the model's generalization ability and accuracy.

[0059] Support Vector Regression (SVR) is a non-linear regression method that builds a model by mapping input variables to a high-dimensional feature space and finding the best-fit hyperplane within that space. The kernel function type is a key parameter in SVR, with the Radial Basis Function (RBF) kernel being a commonly used choice. This kernel primarily consists of two parameters: the kernel width (gamma, g) and the penalty parameter (c). The g value controls the range of the Gaussian distribution, while the c value balances the model's goodness of fit and complexity. Adjusting these two parameters controls the model's flexibility and smoothness, thereby avoiding overfitting and improving its generalization ability.

[0060] One-Dimensional Convolutional Neural Network (1D-CNN) is a deep learning algorithm that extracts features from an input variable (LIBS spectrum) by stacking multiple convolutional and pooling layers, and performs regression analysis through fully connected layers. The model's structure and complexity are determined by parameters such as kernel size, the number of convolutional and pooling layers, and the number of nodes in the fully connected layers. By adjusting these key parameters, 1D-CNN can automatically learn features from spectral data and perform regression analysis. Furthermore, it can be optimized by increasing network depth and adding regularization to improve model performance. Therefore, 1D-CNN is suitable for many spectral data analysis tasks and exhibits good generalization ability and robustness.

[0061] 2.4.2 Variable Selection Methods

[0062] bi-PLS is an algorithm for feature band selection, designed to choose the most relevant bands from full-spectrum data. Based on inverse partial least squares regression, it improves the model's predictive performance and interpretability by progressively eliminating irrelevant bands to obtain an optimal feature subset. The algorithm's steps include initialization, inverse selection, and model evaluation. First, all bands are included in the initial feature subset. Then, by calculating the contribution of each band to the model's predictive performance, the bands with the smallest contribution are selected and eliminated, and this selection process is iteratively repeated until a preset stopping criterion is met. In each iteration, after eliminating one band, the model is retrained and its performance metrics on the test set are evaluated. By comparing the performance metrics under different subsets, the best-performing subset is selected as the final feature subset. This algorithm reduces data redundancy and noise by progressively eliminating irrelevant bands, thereby improving the model's accuracy and interpretability.

[0063] CARS is an algorithm for feature band selection. First, it is adaptive, selecting bands with high discriminative power based on data characteristics, thus improving the accuracy and stability of feature selection. Second, by introducing a competition mechanism, the algorithm assigns higher weights to bands with high discriminative power, reducing redundancy and noise, and improving the quality of the feature subset. Furthermore, the algorithm uses an iterative approach for feature selection, quickly converging to the optimal feature subset, improving efficiency and consequently enhancing the model's predictive ability.

[0064] Genetic Algorithm (GA) is an optimization algorithm that simulates the natural evolutionary process. It possesses global search capabilities and can find relatively optimal solutions in large-scale search spaces. By simulating operations such as natural selection, crossover, and mutation, GA maintains population diversity and selects individuals with higher fitness in each generation for evolution, thus gradually approaching the optimal solution. Secondly, the algorithm has parallel computing capabilities, processing multiple individuals simultaneously and accelerating the search speed. Furthermore, it exhibits strong robustness, is insensitive to initial conditions and constraints, and can handle complex optimization problems.

[0065] Random Forest (RF) is an optimization algorithm based on frog behavior. Similar to Global Algorithm (GA), it possesses global search capabilities, enabling it to find the global optimum within the search space. By simulating the jumping behavior of a frog in its environment, the algorithm randomly selects a target position in each jump and updates its current position by comparing the fitness of the target position, thus gradually approaching the optimal solution. The algorithm is also adaptive, adjusting parameters according to the characteristics of the problem and the search progress, improving search efficiency. Furthermore, its strong robustness allows it to handle complex optimization problems.

[0066] 2.4.3 Evaluation of Prediction Model

[0067] This invention uses the coefficient of determination of the training set ( ), Root Mean Square Error (RMSEC) of the training set, Coefficient of Determination of the test set ( The optimal calibration model is selected based on the root mean square error (RMSEP) of the test set. The coefficient of determination determines the model's accuracy; the closer its value is to 1, the higher the accuracy of the prediction model. The RMSEP value determines the model's prediction error; the closer its value is to 0, the smaller the prediction error. The formulas for the root mean square error and the coefficient of determination are as follows:

[0068]

[0069] in, For the true value of the sample, For the predicted value of the sample, This is the average of the true values.

[0070] Example 1

[0071] This embodiment establishes a rapid detection model for the content of four heavy metal elements based on the LIBS full spectrum. Using the full spectrum (198.71nm-629.08nm) as the input variable, multivariate analysis of the four heavy metal elements is performed using PLSR, KNN, SVR, and 1D-CNN modeling algorithms. The results are shown in Table 2. It is found that among the full spectrum models for the four heavy metal elements, the 1D-CNN model performs best. The prediction models for Cu, Cd, Hg, and Pb are... The values ​​were 0.9901, 0.9852, 0.9914, and 0.9567, respectively, with RMSEP values ​​of 58.0969 mg / kg, 8.7678 mg / kg, 3.4851 mg / kg, and 8.2193 mg / kg, respectively. The PLSR and SVR models performed second best, with the prediction models... All values ​​are greater than 0.90, with the KNN model performing the worst.

[0072] Because LIBS spectra contain information about numerous elements, compared to the four heavy metal elements in this invention, current spectra contain a large amount of redundant information, which affects the performance of PLSR, KNN, and SVR models. 1D-CNN, as a deep learning algorithm, has the advantage of automatically extracting spectral features. In particular, it can gradually extract features at different levels through the stacking of multiple convolutional and pooling layers, thereby capturing feature information from the entire spectrum and presenting more information about the four target elements. Therefore, the 1D-CNN model has better adaptability and generalization ability compared to other models, improving the accuracy of quantitative analysis.

[0073] A correlation coefficient plot of reference values ​​and predicted values ​​was generated for the optimal full-spectrum multi-quantity model of the content of four heavy metal elements, as shown in the figure. Figure 4 As shown, the scatter plots of the best full spectrum for each element are relatively concentrated, with only a few points showing a scattered phenomenon.

[0074] Table 2. Multivariate analysis of four heavy metal elements based on LIBS full spectrum.

[0075] Example 2

[0076] LIBS full-spectral data typically contains a large number of bands, and redundant bands may exist, increasing the complexity of the analysis and potentially leading to overfitting. By selecting characteristic bands relevant to the target variable through feature band filtering, model performance can be improved while simplifying the model, and key information in the spectral data can be inferred, resulting in more accurate predictions.

[0077] This embodiment combines wavelength interval selection and wavelength point selection to apply LIBS spectroscopy for rapid detection of heavy metal content in Gastrodia elata. Specifically, it uses reverse partial least squares (biPLS) to effectively screen the entire LIBS spectrum across different regions. Figure 5-8As shown in the gray area, 27 highly correlated characteristic intervals for Cu in Gastrodia elata were identified, with 4682 variables; 29 highly correlated characteristic intervals for Cd, with 3544 variables; 21 highly correlated characteristic intervals for Hg, with 5268 variables; and 20 highly correlated characteristic intervals for Pb, with 3162 variables. Further effective variables were selected from the simplified LIBS spectral data using CARS, RF, and GA. Figures 5-8 Furthermore, PLSR, KNN, SVR, and 1D-CNN were combined to establish relevant multivariate models, and the best rapid detection model for heavy metal content was compared and selected. The results are shown in Table 3.

[0078] Table 3. Multivariate analysis of four heavy metal elements based on characteristic band screening.

[0079] The experimental data in Table 3 show that among the multivariate models of the four heavy metal elements, the number of variables after band selection is much smaller than that of the full spectrum. The 1D-CNN model after biPLS-CARS, biPLS-RF, and biPLS-GA processing performed the best, followed by the SVR model and the PLSR model, and the KNN model performed the worst.

[0080] See Figure 9 In the multivariate model of Cu, for biPLS, 4682 variables were selected, and the optimal model was biPLS-1D-CNN, which predicts the optimal performance. The value was 0.9887, and the RMSEP was 62.0744 mg / kg. For biPLS-CARS, after screening 90 variables within the biPLS interval, the optimal model was biPLS-CARS-1D-CNN, and the prediction model's... The value was 0.9791, and the RMSEP was 84.3500 mg / kg. For biPLS-RF, after screening 199 variables within the biPLS interval, the optimal model was biPLS-RF-1D-CNN, and the prediction model's... The value was 0.9882, and the RMSEP was 63.4154 mg / kg. For biPLS-GA, after screening 425 variables within the biPLS interval, the optimal model was biPLS-GA-1D-CNN, and the prediction model's... The value was 0.9892, and the RMSEP was 60.70 mg / kg. Based on the above comparison, it can be concluded that among the multivariate models for Cu, the biPLS-GA-1D-CNN yielded the best results. The concentration was 0.9892, and the RMSEP was 60.6997 mg / kg. (See the scatter plot for the predicted concentration.) Figure 9 a, with the best model for the full spectrum ( Compared with the full spectrum (0.9901, RMSEP 58.0969 mg / kg), although the prediction model's results were slightly worse than the full spectrum, it did not affect the model's prediction. Moreover, it greatly reduced the number of spectral variables, using only 425 variables to replace the full spectrum, proving that the model after variable screening by biPLS-GA has better adaptability and generalization ability.

[0081] In the multivariate model of Cd elements, for biPLS, 3544 variables were selected, and the optimal model was biPLS-1D-CNN, which predicts the model's... The value was 0.9841, and the RMSEP was 9.0748 mg / kg. For biPLS-CARS, 264 variables were screened within the biPLS interval, and the optimal model was biPLS-CARS-1D-CNN, with a prediction model of [missing value]. The value was 0.9870, and the RMSEP was 8.2195 mg / kg. For biPLS-RF, after screening 322 variables within the biPLS interval, the optimal model was biPLS-RF-1D-CNN, and the prediction model's... The value was 0.9827, and the RMSEP was 9.4777 mg / kg. For biPLS-GA, after screening 680 variables within the biPLS interval, the optimal model was biPLS-GA-1D-CNN, and the prediction model's... The value was 0.9847, and the RMSEP was 8.9024 mg / kg. Based on the above comparison, it can be concluded that among the multivariate models for Cd elements, biPLS-CARS-1D-CNN yielded the best results. The concentration was 0.9870, and the RMSEP was 8.2195 mg / kg. (See the scatter plot for the predicted concentration.) Figure 9 b, not only are the model results superior to the best full-spectrum model ( The RMSEP was 8.7678 mg / kg (0.9852), and the number of variables was greatly reduced, with only 264 variables replacing the full spectrum.

[0082] In the multivariate model of Hg elements, for biPLS, 5268 variables were selected, and the optimal model was biPLS-1D-CNN, which predicts the model's... The value was 0.9887, and the RMSEP was 3.9870 mg / kg. For biPLS-CARS, 422 variables were screened within the biPLS interval, and the optimal model was biPLS-CARS-1D-CNN, with a prediction model of [missing information]. The value was 0.9925, and the RMSEP was 3.2450 mg / kg. For biPLS-RF, after screening 370 variables within the biPLS interval, the optimal model was biPLS-RF-1D-CNN, and the prediction model's... The value was 0.9933, and the RMSEP was 3.0640 mg / kg. For biPLS-GA, after screening 718 variables within the biPLS interval, the optimal model was biPLS-GA-1D-CNN, and the prediction model's... The value was 0.9933, and the RMSEP was 3.0777 mg / kg. Based on the above comparison, it can be concluded that among the multivariate models for Hg, biPLS-RF-1D-CNN yielded the best results. The concentration was 0.9933, and the RMSEP was 3.0640 mg / kg. (See the scatter plot for the predicted concentration.) Figure 9 c, not only are the model results better than the best full-spectrum model ( The RMSEP was 0.9914 and 3.4851 mg / kg, and the number of variables was also greatly reduced, with only 370 variables replacing the full spectrum.

[0083] In the multivariate model with Pb elements, for biPLS, 3162 variables were selected, and the optimal model was biPLS-1D-CNN, which predicts the model's... The value was 0.9468, and the RMSEP was 9.1154 mg / kg. For biPLS-CARS, after screening 201 variables within the biPLS interval, the optimal model was biPLS-CARS-1D-CNN, and the prediction model's... The value was 0.9228, and the RMSEP was 10.9782 mg / kg. For biPLS-RF, after screening 377 variables within the biPLS interval, the optimal model was biPLS-RF-1D-CNN, and the prediction model's... The value was 0.9471, and the RMSEP was 9.0837 mg / kg. For biPLS-GA, 549 variables were screened within the biPLS interval, and the optimal model was biPLS-GA-1D-CNN, with a prediction model of [missing information]. The value was 0.9366, and the RMSEP was 9.9459 mg / kg. Based on the above comparison, it can be concluded that among the multivariate models for Pb elements, biPLS-RF-1D-CNN yielded the best results. The concentration was 0.9471, and the RMSEP was 9.0837 mg / kg. (See the scatter plot for the predicted content.) Figure 9 d, with the optimal model for the full spectrum ( Compared with the full spectrum (0.9567, RMSEP 8.2193 mg / kg), although the prediction model's results were slightly worse than the full spectrum, it did not affect the model's prediction. Moreover, it greatly reduced the number of spectral variables, replacing the full spectrum with only 377 variables, proving that the model after variable screening by biPLS-RF has better adaptability and generalization ability.

[0084] Comparative Example 1

[0085] Based on the LIBS spectra obtained from the dataset and the characteristic spectral lines of Gastrodia elata found in the NIST database and related literature, the following strong characteristic spectral lines (relative intensity greater than 600) were selected as follows: Cu I 324.71 nm, Cu I 327.39 nm, Cd I 228.62 nm, Cd II 257.44 nm, Hg II 219.07 nm, Hg II 260.39 nm, Pb I 280.12 nm, Pb I 405.85 nm.

[0086] The intensity of the characteristic spectral lines of the four heavy metal elements was fitted to reference values ​​to establish a univariate model of the content of the four heavy metal elements in Gastrodia elata. The results are shown in Table 4. It was found that among the Cu elements, CuI at 327.39 nm yielded the best result. The value was 0.5654, and the RMSEP was 418.9263 mg / kg; among Cd elements, the best result was for Cd I at 228.62 nm. The value was 0.7306, and the RMSEP was 37.3773 mg / kg; among Hg elements, the best result was for Hg II at 219.07 nm. The value was 0.7551, and the RMSEP was 18.5762 mg / kg; among Pb elements, the best result was for Pb I at 405.85 nm. The value was 0.4082, and the RMSEP was 36.7052 mg / kg.

[0087] Table 4 Univariate analysis of four heavy metal elements based on characteristic spectral lines

[0088] A correlation coefficient graph between reference values ​​and predicted values ​​was plotted for the optimal univariate model of the content of four heavy metal elements, such as... Figure 10As shown, the reference values ​​and predicted values ​​are significantly divergent, indicating that the prediction ability of the established optimal univariate model is poor.

[0089] Analysis of the above results revealed that the univariate model for the characteristic spectral lines of the four elements did not have ideal predictive ability for each element. Combined with the characteristics of atomic spectral analysis, this may be because Gastrodia elata grows in soil and contains a large amount of information about other elements, such as iron, magnesium, sodium, calcium, potassium, etc., which can also cause significant interference to the four heavy metal elements, resulting in a matrix effect.

[0090] Comparative Example 2

[0091] The selected characteristic spectral lines from Comparative Example 1 were combined as input variables. Specifically, Cu I 324.71 nm and Cu I 327.39 nm were used as the characteristic spectral line combination for Cu, Cd I 228.62 nm and Cd II 257.44 nm as the characteristic spectral line combination for Cd, Hg II 219.07 nm and Hg II 260.39 nm as the characteristic spectral line combination for Hg, and Pb I 280.12 nm and Pb I 405.85 nm as the characteristic spectral line combination for Pb. Multivariate analysis of the four heavy metal elements in Gastrodia elata was conducted using PLSR, KNN, and SVR modeling methods to establish a rapid detection model for heavy metal elements in Gastrodia elata based on characteristic spectral line combinations.

[0092] The results are shown in Table 5. Among the multivariate models of Cu characteristic spectral line combinations, the SVR model performed best, and the prediction model... The coefficient of performance was 0.8124, and the RMSEP was 252.8925 mg / kg, which is better than the univariate model based on CuI at 327.39 nm. Among the multivariate models of Cd characteristic spectral line combinations, both the KNN and SVR models performed well. The values ​​were 0.8928 and 0.9265, respectively, with RMSEPs of 23.5782 mg / kg and 19.5244 mg / kg, respectively, which were superior to the univariate model based on CdI at 228.62 nm. Among the multivariate models based on the combination of characteristic spectral lines of Hg, the SVR model performed best, with the highest prediction accuracy. The value was 0.8259, and the RMSEP was 15.6613 mg / kg, which was better than the univariate model based on Hg II at 219.07 nm. Among the multivariate models of Pb characteristic spectral line combinations, all three models performed poorly, with the PLSR model performing relatively better and predicting the model's... The result was 0.5040, and the RMSEP was 1.4198 mg / kg, which was also better than the univariate model based on Pb I at 405.85 nm. The reason why the multivariate model combining the characteristic spectral lines of Pb did not perform well may be that the selected characteristic spectral lines (Pb 280.12 nm and Pb 405.85 nm) are greatly affected by other elements in the matrix.

[0093] Table 5. Multivariate analysis of four heavy metal elements based on characteristic spectral line combinations.

[0094] A correlation coefficient plot was generated between reference and predicted values ​​for the multivariate model of the optimal characteristic spectral lines of the four heavy metal elements. Figure 11 As shown, the scatter plot of the best characteristic spectral lines of each element is more concentrated than the scatter plot of the best univariate in Comparative Example 1, but the result is still not ideal.

[0095] The experimental results above show that the results of the four elements of the univariate model of the characteristic spectral lines in Comparative Example 1 and the multivariate model of the combination of characteristic spectral lines in Comparative Example 2 are all worse than the full-spectrum multivariate models of Examples 1 and 2 of this invention.

[0096] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A method for rapid determination of one or more heavy metal contents in Gastrodia elata tubers using laser-induced breakdown spectroscopy, characterized in that, Includes the following steps: S1. Pre-treat the Gastrodia elata sample and compress it into tablets as samples to be tested; S2. Establishment of initial spectral data: The pellet obtained in step S1 is tested using a laser-induced breakdown spectrometer, and spectral data is collected. S3. Measurement of heavy metals in Gastrodia elata samples; S4. Preprocess the initial spectral data and combine it with chemometric methods. At the same time, based on the heavy metal content data of the Gastrodia elata sample, establish a predictive model for the heavy metal content of Gastrodia elata.

2. The detection method as described in claim 1, characterized in that, The chemometrics method described in step S4 includes at least one of partial least squares regression, K-nearest neighbor algorithm, support vector regression, and one-dimensional convolutional neural network modeling algorithm.

3. The detection method as described in claim 2, characterized in that, The chemometric method described in step S4 is a one-dimensional convolutional neural network modeling algorithm.

4. The detection method as described in claim 1, characterized in that, After preprocessing the initial spectral data as described in step S4, variable screening of the characteristic bands of the spectrum is also required.

5. The detection method as described in claim 4, characterized in that, The method for selecting variables for spectral feature bands is as follows: variable selection is performed by combining at least one of the following methods: backward-interval partial least squares algorithm with competitive adaptive reweighted sampling, genetic algorithm, and random forest algorithm.

6. The detection method as described in claim 1, characterized in that, The preprocessing described in step S4 includes the following steps: using wavelet transform to denoise the laser-induced breakdown spectrum, then using baseline elimination to remove baseline drift or baseline shift, making the baseline of the signal flat or stable, and finally using normalization to enhance the stability of the spectral signal.

7. The detection method as described in claim 1, characterized in that, The preparation method of the tablets described in step S1 includes the following steps: crushing and sieving the Gastrodia elata sample, and preparing the sample powder into tablets with a diameter of 10~15mm and a thickness of 3~5mm.

8. The detection method as described in claim 1, characterized in that, The method for measuring heavy metals in the Gastrodia elata sample in step S3 is as follows: the heavy metal content in the Gastrodia elata sample is determined by inductively coupled plasma mass spectrometry.