Online LIBS calibration method based on variable importance selection
The online LIBS calibration method based on variable importance selection, utilizing orthogonal partial least squares and principal component analysis, solves the problem of spectral signal instability caused by focal length variation, achieving efficient and accurate quantitative analysis of online LIBS equipment, suitable for real-time detection in industrial settings.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING INST OF TECH
- Filing Date
- 2023-11-30
- Publication Date
- 2026-05-19
AI Technical Summary
Existing online LIBS equipment suffers from unstable spectral signal intensity when the focal length changes, resulting in cumbersome and time-consuming calibration, making it difficult to achieve real-time quantitative analysis of materials across conveyor belts.
An online LIBS calibration method based on variable importance selection is adopted. Through orthogonal partial least squares discriminant analysis and principal component analysis, insensitive characteristic peaks are selected for calibration, and a quantitative regression model is established to achieve effective identification and quantitative analysis under focal length changes.
It simplifies the calibration process, improves the stability of spectral signals and the accuracy of quantitative models, and enhances work efficiency in industrial settings.
Smart Images

Figure CN117607124B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of online laser-induced breakdown spectroscopy, and more specifically to an online LIBS calibration method based on variable importance selection. Background Technology
[0002] The key feature of industrial applications of online laser-induced breakdown spectroscopy (LIBS) is the rapid and accurate quantitative analysis of materials transported across conveyor belts in real time. Traditional industrial applications of LIBS mainly rely on sampling and simple sample preparation of the conveyor belt material, followed by LIBS quantitative analysis. Although the analysis time is significantly reduced compared to laboratory testing, it is not strictly speaking online or real-time.
[0003] Currently, many industrial applications require real-time monitoring of materials on conveyor belts to provide quantitative data to guide industrial production and achieve intelligent control of factory capacity and efficiency. Generally, LIBS equipment calibration requires collecting spectral data from a certain number of standard samples with known industrial parameters at a fixed focal length. Specifically, this involves crushing and grinding block samples into powder with a particle diameter of less than 0.2 μm, and then pressing the powder sample into a regularly shaped, flat disc. Since typical LIBS equipment uses an off-axis optical path collection system—meaning the sample excitation and collection optical paths are independent and at an angle—this optical system is extremely sensitive to focal length changes. Even small changes in focal length can lead to a significant decrease in spectral intensity, or even complete loss of light signal. A LIBS system built using a coaxial zoom optical system can effectively overcome these difficulties. Because the light emission and collection paths are coaxial, changes in the focal length of the laser emission path are equivalent to changes in the focal length of the light signal collection path. The focal length can be automatically adjusted based on material fluctuations, thus enabling effective spectral acquisition of materials across conveyor belts. However, coaxial systems also present calibration challenges. Because the numerical aperture of the entire optical system changes during real-time focal length adjustment, this leads to variations in light collection efficiency. At long focal lengths, the numerical aperture is smaller, resulting in lower light collection efficiency and weaker spectral signals. Conversely, at short focal lengths, the numerical aperture increases, leading to relatively higher light collection efficiency and stronger spectral signals. Generally, online LIBS equipment has a relatively long working distance, resulting in a greater depth of focus in the optical system, meaning that the sample can be penetrated within a certain distance near the focal point. For iron ore, the depth of focus of a coaxial optical system with a working distance of 700 mm is approximately ±15 mm. Equipment calibration can be performed repeatedly at different working distances, using a specific calibration model at a given working distance. However, this calibration method is cumbersome, time-consuming, and labor-intensive, requiring the collection of spectral data from dozens of different working distances for a single calibration, making it impractical in real-world applications.
[0004] Therefore, how to provide an online LIBS calibration method based on variable importance selection is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides an online LIBS calibration method based on variable importance selection, which can calibrate the equipment at the initial working distance and achieve effective identification and quantitative analysis of samples at other zoom positions. The calibration method is simple and feasible, and the quantitative model has the characteristics of high accuracy and high robustness, which has great potential value for the widespread application of online LIBS equipment in industrial fields.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] Online LIBS calibration methods based on variable importance selection include:
[0008] Determine the initial working distance L0 and the working distance adjustment range [L1, L2] of the online LIBS system. n ];
[0009] Within the focal length adjustment range, select n focal lengths, n≥5, and collect no less than k spectra of the same sample at each focal length point, k≥100. After each collection of k spectra, form a batch of spectral data. The number of batches of spectral data for each focal length point shall be no less than m batches, m≥3.
[0010] Supervised classification of spectral data collected at n focal points is performed based on orthogonal partial least squares discriminant analysis.
[0011] Perform variable projection importance analysis on the classification model, mark the feature peaks with a variable projection importance greater than 1 in the classification model, and mark the positions of the feature peaks as the index matrix IDX. c ;
[0012] Select no fewer than t samples with different industrial parameters, adjust the online LIBS system to work at the initial working distance L0, and collect spectral data of t samples in sequence. The number of spectra collected for each sample is no less than k, and the spectral matrix of each sample is obtained.
[0013] The spectral variable matrix X of t samples and the industrial parameter regression matrix Y are input into a partial least squares regression algorithm based on principal component analysis to establish a quantitative regression model. The importance of variables in the quantitative regression model is calculated based on the spectral matrix, and characteristic peaks with a projected importance greater than 1 are marked. The positions of these characteristic peaks are then labeled as the index matrix IDX. r ;
[0014] Calculate the index matrix IDX c With the index matrix IDX rIntersection of IDX c∩r And calculate the final spectral index matrix;
[0015] The spectral intensities in the spectrum are calculated based on the final spectral index matrix, and a quantitative regression model is re-established based on the PCA-PLS algorithm.
[0016] Preferably, the formula for calculating the importance of the projection of the feature peak variable in the classification model is as follows:
[0017]
[0018] Among them, VIP c K represents the importance of the projection of the feature peak variables in the classification model. p a, A p These represent the total number of variables, the total number of principal components, and the OPLS-DA dimension, respectively. a The OPLS-DA weighting coefficients on the a-th principal component, SSX comp,a SSX is the sum of squares of the explanatory power of the a-th principal component on the spectral variable matrix X. cum It is the cumulative sum of squares of the interpretations of all principal components on the spectral variable matrix X, SSY comp,a SSY is the sum of squares of the explanatory power of the a-th principal component on the industrial parameter regressor matrix Y. cum It is the cumulative sum of squares of the explanations of all principal components on the industrial parameter regressor matrix Y;
[0019] Index Matrix IDX c The calculation formula is:
[0020]
[0021] in, It refers to the position of the variable with a projection importance greater than 1 in the spectral variable matrix X in the classification model.
[0022] Preferably, the formula for calculating the spectral matrix is:
[0023]
[0024] Among them, S t Represents the spectral matrix, matrix element I kJ This represents the signal strength of the J-th pixel on the detector during the k-th measurement of the t-th sample.
[0025] Preferably, the formula for calculating the importance of the projected peak variables in the quantitative regression model is as follows:
[0026]
[0027] Among them, VIP rThis indicates the importance of the projected peak variables in a quantitative regression model, where K is the total number of variables and w a The coefficients of the quantitative regression model for the a-th principal component are SSY. comp,a SSY is the sum of squares of the explanatory power of the a-th principal component on the industrial parameter regressor matrix Y. cum It is the cumulative sum of squares of the explanations of all principal components on the industrial parameter regressor matrix Y;
[0028] Index Matrix IDX r The calculation formula is:
[0029]
[0030] in, This indicates the position of the variable with a projection importance greater than 1 in the spectral variable matrix X in the quantitative regression model.
[0031] Preferably, the final spectral index matrix is:
[0032]
[0033] Where p is the variable position index value used in the final spectral index matrix.
[0034] Preferably, it further includes:
[0035] By changing the working distance L, selecting n focal lengths within the focal length adjustment range, randomly selecting a group of samples for testing, and predicting the corresponding industrial parameters based on a quantitative regression model, the accuracy of the model is verified.
[0036] The present invention has the following technical effects:
[0037] (1) The method of the present invention selects characteristic peaks in the spectrum that are not sensitive to zooming of the online LIBS system through variable importance analysis, and uses these characteristic peaks to calibrate the device at the initial position. The calibrated model can be applied to other focal positions.
[0038] (2) The method of the present invention significantly reduces the calibration time of zoom LIBS optical systems and improves work efficiency in industrial settings. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0040] Figure 1The flowchart of the online LIBS calibration method based on variable importance selection provided by the present invention is shown.
[0041] Figure 2 The image provided by this invention shows the classification results of the same sample obtained at 10 working distances.
[0042] Figure 3 This is a schematic diagram illustrating the distribution of important variables in the classification and regression models provided by this invention.
[0043] Figure 4 This invention provides a predictive model for the total iron content in iron ore, which is established at the initial position using important variables.
[0044] Figure 5 A comparison chart showing the results of predicting total iron content at different working distances using models that utilize important variables and those that do not, provided by this invention. Detailed Implementation
[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0046] This invention discloses an online LIBS calibration method based on variable importance selection, such as... Figure 1 As shown, it includes:
[0047] Step 1: Determine the initial working distance L0 of the online LIBS system, and the adjustment range of the working distance L is [L1, L2]. n Among them, the online LIBS system generally adopts a coaxial zoom system, in which the receiving and transmitting optical paths are coaxial.
[0048] Step 2: Select at least 5 focal lengths within the focal length adjustment range of the coaxial zoom LIBS system, namely L1, L2, L3, ... L n (n≥5), preferably n=10; collect no less than k (k≥100) spectra of the same sample at each focal point, preferably k=100; form a batch of spectral data after collecting k spectra, and the batch of spectral data for each focal point shall be no less than m (m≥3), preferably m=3.
[0049] Step 3: Use orthogonal partial least squares discriminant analysis (OPLS-DA) to perform supervised classification of the spectra collected at n focal points, which can be divided into n (n=5) classes.
[0050] Step 4: Perform variable projection importance (VIP) analysis on the classification model described in step 3 according to equation (1). Mark the characteristic peaks in the spectrum with a variable projection importance greater than 1 (VIP>1), and the index matrix of the characteristic peak positions is given by equation (2).
[0051]
[0052]
[0053] Among them, VIP c K represents the importance of the projection of the feature peak variables in the classification model. p a, A p These represent the total number of variables, the total number of principal components, and the OPLS-DA dimension, respectively. a The OPLS-DA weighting coefficients on the a-th principal component, SSX comp,a SSX is the sum of squares of the explanatory power of the a-th principal component on the spectral variable matrix X. cum It is the cumulative sum of squares of the interpretations of all principal components on the spectral variable matrix X, SSY comp,a SSY is the sum of squares of the explanatory power of the a-th principal component on the industrial parameter regressor matrix Y. cum It is the cumulative sum of squares of the explanations of the regressor matrix Y of the industrial parameters by all principal components.
[0054] Step 5: Select no less than t (preferably t = 50) samples with different industrial parameters, adjust the LIBS system to work at the initial working distance L0, and collect the spectral data of t samples in sequence. The number of spectra collected for each sample is no less than k (preferably k = 100), and the spectral matrix of each sample is shown in Equation (3).
[0055]
[0056] Among them, matrix element I kJ This represents the signal strength of the J-th pixel on the detector during the k-th measurement of the t-th sample.
[0057] Step 6: Input the spectral variable matrix X of t samples and the industrial parameter regression matrix Y into the partial least squares regression (PCA-PLS) algorithm based on principal component analysis to establish a quantitative regression model. Calculate the importance of variables in the quantitative regression model according to equation (4) and mark the characteristic peaks (VIPs) in the quantitative model whose projected importance of variables is greater than 1. r >1), mark the positions of the characteristic peaks to form an index matrix IDX as shown in (5). r .
[0058]
[0059]
[0060] Among them, VIP r This indicates the importance of the projected peak variables in a quantitative regression model, where K is the total number of variables and w a The coefficients of the quantitative regression model for the a-th principal component are SSY. comp,a SSY is the sum of squares of the explanatory power of the a-th principal component on the industrial parameter regressor matrix Y. cum It is the cumulative sum of squares of the explanations of all principal components on the industrial parameter regressor matrix Y;
[0061] Step 7: Calculate the index matrix IDX c With IDX r Intersection of IDX c∩r Calculate the final spectral index matrix: Where p is the variable position index value used in the final spectral index matrix.
[0062] Step 8: Utilize the index matrix IDX s The corresponding spectral intensities in the spectrum are indexed, and the quantitative regression model is rebuilt using the PCA-PLS algorithm.
[0063] In this embodiment, based on the above embodiments, it further includes:
[0064] Step 9: Change the working distance L to L1, L2, L3, ... L 10 A group of samples was randomly selected for testing, and the established quantitative regression model was used to predict the corresponding industrial parameters to verify the accuracy of the model.
[0065] The following examples illustrate how to calibrate an online LIBS system based on variable importance and achieve quantitative analysis of total iron content in iron ore.
[0066] Step 1: This example verifies the effectiveness of the method described in this invention by comparing the spectral signal fluctuation after preprocessing using the method of this invention with that after preprocessing using traditional methods. The samples used in this example are 50 iron ore powder samples with a particle size <200 μm. 10 g of powder from each sample was pressed into a disc with a diameter of 30 mm and a thickness of 4 mm using a pressure of 30 tons. The LIBS device uses a 200 mJ pulsed laser (Nd:YAG@1064 nm), and the spectrometer detector is a 4-channel × 2048 pixel linear CCD array.
[0067] Step 2: Determine the initial working distance L0 = 860mm for the online LIBS. The working distance L can be adjusted from 780mm to 960mm, i.e., L∈[780mm,960mm].
[0068] Step 3: Within the working distance adjustment range of the online LIBS system, select a focal length at 20mm intervals, for a total of 10 focal lengths, namely L1 = 780mm, L2 = 800mm, L3 = 820mm, L4 = 840mm…L 10 =960mm. Spectra of sample No. 47 were collected 100 times at the focal point of each working distance, and this was repeated 3 times to obtain the final spectral data for each focal point in 3 batches.
[0069] Step 4: Supervised classification of the spectra acquired at 10 focal points was performed using Orthogonal Partial Least Squares Discriminant Analysis (OPLS-DA), resulting in 10 classes. The classification results are represented by principal component score plots, as shown below. Figure 2 As shown, it is clear that the three batches of spectra at each distance have clustered into the same category.
[0070] Step 5: Perform variable projection importance (VIP) analysis on the classification model described in Step 4 according to equation (1). Label the feature peaks (VIPs) in the classification model where the variable projection importance is greater than 1. c >1), mark the positions of the characteristic peaks to form an index matrix IDX c .
[0071]
[0072]
[0073] In this example, IDX c The CCP contains 461 key characteristic peaks.
[0074] Step 6: Perform preliminary calibration of the equipment using 50 iron ore samples with different industrial parameters. Adjust the LIBS optical system to work at the initial working distance L0, and collect spectral data of 50 samples in sequence. 100 spectra are collected for each sample, and the spectral matrix of each sample is shown in Equation (3).
[0075]
[0076] Where matrix element I kJ This represents the signal strength of the J-th pixel on the detector during the k-th measurement of the t-th sample.
[0077] Step 7: Input the spectral variable matrix X of t samples and the industrial parameter regression matrix Y into the partial least squares regression (PCA-PLS) algorithm based on principal component analysis to establish a quantitative regression model. Calculate the importance of variables in the quantitative regression model according to equation (4) and mark the characteristic peaks (VIPs) in the quantitative regression model whose projected importance of variables is greater than 1. r>1), mark the positions of the characteristic peaks to form an index matrix IDX as shown in (5). r .
[0078]
[0079]
[0080] In this example, IDX r The index matrix contains a total of 227 important characteristic peaks.
[0081] Step 8: Calculate the index matrix IDX c With IDX r Intersection of IDX c∩r The final spectral index matrix In this example, IDX c∩r The number of important feature peak intersections contained in IDX is 13. s A total of 214 important feature peaks are included, and the distribution of feature peaks involved in classification and regression is as follows: Figure 3 As shown.
[0082] Step 9: Utilize the index matrix IDX s The corresponding spectral intensities in the spectrum are indexed, and the quantitative regression model is rebuilt using the PCA-PLS algorithm, such as... Figure 4 As shown.
[0083] Step 10: Change the working distance L to L1, L2, L3, ... L n (n≥5), a random sample group is selected for testing, and the established quantitative regression model is used to predict the corresponding industrial parameters to verify the model's accuracy. The prediction performance of the model recalibrated using this method and the initial calibration model on the same batch of samples at different working distances is compared. Figure 5 As shown, the root mean square error of prediction (RMSEP) and the average relative error (ARE) are calculated using conventional methods.
[0084] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0085] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An online LIBS calibration method based on variable importance selection, characterized in that, include: Determine the initial working distance of the online LIBS system and working distance adjustment range ; Select n focal lengths within the focal length adjustment range. At each focal point, at least k spectra of the same sample are collected, k≥100. Each batch of spectral data is formed after collecting k spectra. The number of batches of spectral data for each focal point is at least m, m≥3. Supervised classification of spectral data collected at n focal points is performed based on orthogonal partial least squares discriminant analysis. Perform variable projection importance analysis on the classification model, mark the feature peaks with a variable projection importance greater than 1 in the classification model, and mark the positions of the feature peaks as an index matrix. ; Select at least t samples with different industrial parameters and adjust the online LIBS system to operate at the initial working distance. Spectral data of t samples are collected sequentially, with at least k spectra collected for each sample, to obtain the spectral matrix of each sample; The spectral variable matrix X of t samples and the industrial parameter regression matrix Y are input into a partial least squares regression algorithm based on principal component analysis to establish a quantitative regression model. The importance of variables in the quantitative regression model is calculated based on the spectral matrix, and characteristic peaks with a projected importance greater than 1 are marked. The positions of these characteristic peaks are then used as an index matrix. ; Calculate the index matrix With index matrix intersection And calculate the final spectral index matrix; The corresponding spectral intensities in the spectrum are calculated based on the final spectral index matrix, and a quantitative regression model is re-established based on the PCA-PLS algorithm. The formula for calculating the importance of the projected peak variables in a classification model is as follows: ; in, Indicates the importance of the projection of the feature peak variables in the classification model. , , These represent the total number of variables, the total number of principal components, and the OPLS-DA dimension, respectively. It is in the OPLS-DA weighting coefficients on principal components It is the first The sum of squares of the explanatory power of the principal components on the spectral variable matrix X. It is the cumulative sum of squares of the interpretations of all principal components on the spectral variable matrix X. It is the first The sum of squares of the principal components' explanatory power for the industrial parameter regressor matrix Y. It is the cumulative sum of squares of the explanations of all principal components on the industrial parameter regressor matrix Y; Index Matrix The calculation formula is: ; in, It refers to the position of the variable with a projection importance > 1 in the spectral variable matrix X in the classification model; The formula for calculating the importance of the projected peak variables in a quantitative regression model is as follows: ; in, This indicates the importance of the projected peak variables in a quantitative regression model, where K is the total number of variables. It is the first The coefficients of the quantitative regression model of principal components, It is the first The sum of squares of the principal components' explanatory power for the industrial parameter regressor matrix Y. It is the cumulative sum of squares of the explanations of all principal components on the industrial parameter regressor matrix Y; Index Matrix The calculation formula is: ; in, This indicates the position of the variable with a projection importance greater than 1 in the spectral variable matrix X in the quantitative regression model.
2. The online LIBS calibration method based on variable importance selection according to claim 1, characterized in that, The formula for calculating the spectral matrix is: ; in, Represents the spectral matrix, matrix elements This represents the signal strength of the J-th pixel on the detector during the k-th measurement of the t-th sample.
3. The online LIBS calibration method based on variable importance selection according to claim 1, characterized in that, The final spectral index matrix is: ; Where p is the variable position index value used in the final spectral index matrix.
4. The online LIBS calibration method based on variable importance selection according to claim 1, characterized in that, Also includes: By changing the working distance L, selecting n focal lengths within the focal length adjustment range, randomly selecting a group of samples for testing, and predicting the corresponding industrial parameters based on a quantitative regression model, the accuracy of the model is verified.