Libs multi-distance hybrid spectrum classification method based on deep cnn and sample weight optimization
By assigning different weights to spectral samples at different detection distances in a deep CNN model, the problem of low classification accuracy in LIBS multi-distance mixed spectral classification is solved, achieving more efficient and accurate spectral classification, which is suitable for special scenarios such as Mars exploration.
Patent Information
- Application Number
- CN202510266824.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-03-07
AI Technical Summary
In existing LIBS multi-distance hybrid spectral classification methods, the differences in spectral characteristics caused by different detection distances are not fully considered during model training, resulting in low classification accuracy, especially in Mars exploration where there is insufficient sample size.
During the training of a deep CNN model, different sample weights are designed based on the absolute distance value and the relative distance difference of the spectral samples to optimize the weights of the spectral samples in the training set and improve the classification performance of the model.
It significantly improves the classification accuracy of LIBS multi-distance mixed spectra, saves computing resources and time, and is applicable to both identification and quantitative analysis tasks, making it more widely applicable.
Smart Images

Figure CN119992219B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of laser spectrum analysis, in particular to a LIBS multi-distance mixed spectrum classification method based on deep CNN and sample weight optimization. The method gives different weights to the LIBS spectrum samples collected at different detection distances, thereby optimizing the training process of the deep CNN model, and can realize accurate and efficient LIBS multi-distance mixed spectrum classification. BACKGROUND
[0002] Laser-induced breakdown spectroscopy (LIBS) is a chemical composition detection technology based on laser-induced plasma radiation, which has the advantages of fast response, sample micro-damage, multi-element simultaneous analysis, remote detection, etc., and is widely used in environmental monitoring, industrial detection, planetary exploration and other fields. In particular, the advantage of remote detection makes LIBS technology play an important role in planetary exploration. At present, LIBS technology has been successfully applied to Mars exploration missions three times. The "Curiosity" and "Perseverance" of NASA in the United States, and the "Zhurong" of China, all carry scientific payloads equipped with LIBS systems, namely the ChemCam, SuperCam and MarSCoDe. Based on LIBS spectrum data, people can use chemometrics models to identify and classify the soil, rock and other substances on the surface of Mars and quantitatively detect their chemical composition.
[0003] The performance of the chemometrics model depends on the extraction and learning ability of the model to the LIBS spectrum morphological features. For LIBS spectrum, a key factor affecting the spectrum morphological features is the detection distance. Even if the same device detects the spectrum of the same sample, the LIBS spectrum morphology will be significantly different if the detection distance is different, and the spectrum morphology difference caused by different detection distances is much larger than that caused by pulse-to-pulse fluctuations, as detailed in reference [1]. The reason why LIBS detection has a significant distance effect is that a series of factors such as laser focal spot diameter, intensity spatial distribution of the spot, sampling geometry, degree of absorption / scattering of laser, degree of absorption / scattering of the environment medium to the plasma signal, plasma temperature and electron number density will change with the change of detection distance, thereby affecting the final LIBS spectrum morphology.
[0004] In laboratory LIBS detection, it is relatively easy to keep the detection distance stable. However, in Mars in-situ detection, it is difficult to collect a large number of LIBS spectra at the same distance because the Mars rover moves frequently. For chemometrics models, especially those based on machine learning or deep learning, the number of LIBS spectral data samples at the same distance is usually insufficient to support effective model training. To address this challenge, one solution is to mix spectra collected at multiple distances to train the model, rather than using only a limited number of spectra collected at a specific distance. Although the number of mixed multi-distance spectral samples can be significantly increased, the spectral form differences caused by different detection distances may confuse the model and prevent it from extracting and learning the true features. Generally, when using a multi-distance mixed spectral dataset, the chemometrics model analysis results are better or worse than those of a single-distance dataset, depending on the similarity of spectral data at different distances and the learning ability of the model itself.
[0005] To improve the analysis results of chemometrics models, one of the most common method systems is distance correction. By performing certain data preprocessing on LIBS multi-distance mixed spectra, the form differences between spectra collected at different detection distances are reduced. As early as a decade ago, the ChemCam team designed a spectral response and distance correction function to correct the distance when processing LIBS multi-distance mixed spectra[2]. The parameters required to calculate the above distance correction function include photon spectral radiance, pixel size, stereo angle, unit integration time, and amplification conversion gain. Later, the ChemCam team proposed a distance correction scheme based on a distance calibration curve[3]. The core of this scheme is to find the appropriate characteristic radiation spectral line of the element to be analyzed, and use the spectral line intensity of the same characteristic radiation spectral line at different distances to construct a distance calibration curve, and then correct the multi-distance mixed spectra according to the distance calibration curve.
[0006] In addition to the above distance correction-based method system, there is another method system that does not perform distance correction, but directly processes LIBS multi-distance mixed spectral data by constructing a deep learning chemometrics model with strong learning ability. Our team has previously proposed a LIBS multi-distance mixed spectral classification method based on a deep convolutional neural network (CNN) algorithm model[4]. Without distance correction, the classification accuracy of the deep CNN model can be significantly higher than that of other conventional algorithm models such as support vector machines and back propagation neural networks.
[0007] The disadvantages of the above three existing methods are as follows:
[0008] For the method based on the distance correction function in reference [2], the parameters that change with the detection distance need to be considered comprehensively, including the photon spectral radiance, the pixel size, the stereo field of view, the unit integration time and the amplification conversion gain, etc. If any of them is missed, it will affect the effect of the distance correction function. The calculation process of each of the above parameters is not simple, so the calculation process of the entire distance correction function is very complex and tedious.
[0009] For the method based on the distance calibration curve in reference [3], the element to be analyzed needs to have a sufficient number of characteristic spectral lines to construct the distance calibration curve. In fact, the process of finding suitable spectral lines itself is relatively complex, and not all elements can find a sufficient number of suitable spectral lines. In addition, this method is suitable for quantitative analysis of chemical composition, and existing data results show that it can only achieve good results on single-variable regression tasks, and there is no relevant data to prove its effect on spectral recognition and classification tasks.
[0010] For the method based on the deep CNN model in reference [4], although it has the advantage of not needing distance correction, in the training process of the deep CNN model, the conventional default method of equal weight of all training set spectral samples is adopted, without fully considering the differences in spectral characteristics caused by different detection distances. On the one hand, the spectral signal-to-noise ratio of near-distance detection is relatively high, and it should be paid more attention by the model, and the spectral signal-to-noise ratio of far-distance detection is relatively low, and it should be paid less attention. On the other hand, the relative positions of the training set spectral samples and the test set spectral samples will also affect the classification effect. The training set spectral samples corresponding to the detection distance close to the test set should be paid more attention by the model, and the training set spectral samples corresponding to the detection distance far from the actual test set should be paid less attention. Therefore, when training the deep CNN model, if the training set spectral samples of different detection distances are simply given equal weight, it is not conducive to improving the classification accuracy of the model.
[0011] References
[0012] [1] Jie Feng, et al. Study to reduce laser-induced breakdown spectroscopy measurement uncertainty using plasma characteristic parameters. Spectrochimica Acta Part B: Atomic Spectroscopy 65 (2010) 549-556. [2] Jie Feng, et al. Study to reduce laser-induced breakdown spectroscopy measurement uncertainty using plasma characteristic parameters. Spectrochimica Acta Part B: Atomic Spectroscopy 65 (2010) 549-556.
[0013] [2] R. C. Wiens, et al. Pre-flight calibration and initial data processing for the ChemCam laser-induced breakdown spectroscopy instrument on the Mars Science Laboratory rover. Spectrochimica Acta Part B: Atomic Spectroscopy 82 (2013) 1-27.
[0014] [3] A. Mezzacappa, et al. Application of distance correction to ChemCam laser-induced breakdown spectroscopy measurements. Spectrochimica Acta Part B: Atomic Spectroscopy 120 (2016) 19-29.
[0015] [4] Fan Yang, et al. Laser-induced breakdown spectroscopy combined with a convolutional neural network: A promising methodology for geochemical sample identification in Tianwen-1 Mars mission. Spectrochimica Acta Part B: Atomic Spectroscopy 192 (2022) 106417. SUMMARY
[0016] In view of the above background and deficiencies of the prior art, the present application proposes a LIBS multi-distance mixed spectrum classification method based on deep CNN and sample weight optimization, which is suitable for analyzing LIBS spectrum mixed data collected at multiple different distances. The core innovation of the present application is to assign different sample weights to spectrum samples at different distances during CNN training. The present application can significantly improve the classification accuracy of LIBS multi-distance mixed spectrum, and provides important technical support and practical solutions for special LIBS application scenarios such as deep space exploration and field exploration.
[0017] The technical scheme of the present application is as follows:
[0018] The application discloses a LIBS multi-distance mixed spectrum classification method based on a deep CNN and sample weight optimization.
[0019] The overall flow of the technical solution can be divided into nine steps, as shown in the accompanying drawings of the specification. Figure 1 The specific description is as follows:
[0020] S1, preliminary preparation: prepare the detection samples of laser-induced breakdown spectroscopy (LIBS), check and save them in sample bags after no error is found, and record the name and category information of each sample. After the sample preparation is completed, a category total table covering all samples (i.e. N samples) is made, that is, all samples (i.e. N samples) belonging to the category are listed and integrated, and the categories of all training set and test set samples do not exceed the range of the total table.
[0021] After the category total table is made, the category vector of each sample is determined in the form of one-hot encoding. One-hot encoding converts each category into a binary vector, in which only one element is 1 and the rest are 0. Here, the application method of one-hot encoding is as follows: there are L different categories in the category total table, then the category vector C of one-hot encoding of each sample is a 1xL matrix, which contains 1 number 1 and (L-1) number 0. If sample i belongs to the jth category, then its category vector C i is
[0022] C i = [0, 0, … 1, 0, 0…] (1)
[0023] The jth element in formula (1) is 1.
[0024] S2, using the LIBS spectrum detection device, collect the LIBS spectrum of the detection sample under P different detection distances. When the spectrum is collected by the LIBS spectrum detection device, the detection distance is defined as the straight line distance between the laser external light outlet of the detection device and the geometric center of the surface of the detection sample, and the detection distance is changed by changing the position of the detection device or the detection sample. The spectrum is collected under P different detection distances, and it is ensured that the experimental conditions remain unchanged except the detection distance. The LIBS spectrum collected by the experiment is defined as the original spectrum data set.
[0025] S3, preprocessing the LIBS spectrum in the original spectrum dataset, usually the preprocessing steps include dark background removal, background baseline removal, wavelength calibration, invalid pixel screening and channel splicing. Among them, the dark background removal operation refers to obtaining the effective spectrum by subtracting the dark background spectrum from the original LIBS spectrum, wherein the dark background spectrum refers to the spectrum of the spectrometer response when there is no laser excitation; the background baseline removal operation refers to removing the continuous baseline in the denoised spectrum by using the asymmetric least squares baseline correction method; the wavelength calibration refers to converting the spectrometer pixel number into wavelength value by multivariate quadratic fitting; the invalid pixel screening refers to removing the pixel response value of each wavelength band of the LIBS spectrum beyond the wavelength range; the channel splicing refers to splicing the LIBS spectrum of multiple wavelength bands after invalid pixel screening into a whole according to the wavelength order; the LIBS spectrum after preprocessing is defined as the LIBS multi-distance mixed spectrum dataset.
[0026] S4, according to P different detection distances, the LIBS multi-distance mixed spectrum dataset in S3 is divided into P LIBS spectrum datasets: the spectrum dataset collected at the first distance is defined as d1; the spectrum dataset collected at the second distance is defined as d2; and so on, the spectrum dataset collected at the last distance, i.e. the Pth distance, is defined as d P . Each detection distance corresponds to a training set-test set division scheme of the LIBS multi-distance mixed spectrum dataset: the division scheme corresponding to the first distance is defined as Dataset1, indicating that the spectrum dataset d1 is used as the test set, and the remaining spectrum datasets d2 to d P are used as the training set; the division scheme corresponding to the second distance is defined as Dataset2, indicating that the spectrum dataset d2 is used as the test set, and the remaining spectrum datasets d1, d3 to d P are used as the training set; and so on, the division scheme corresponding to the Pth distance is defined as DatasetP, indicating that the spectrum dataset d P is used as the test set, and the remaining spectrum datasets d1 to d P-1 are used as the training set.
[0027] For each spectrum dataset division scheme, it is ensured that the test set and the training set have no intersection in the dimensions of detection distance and detection sample: taking the dataset division scheme DatasetK as an example, the spectrum dataset d K is used as the test set, and the remaining spectrum datasets d1…d K-1 , d K+1 …d PThe spectral data sets of other distances are used as the training set, and the "leave-one-out" strategy is used for testing. When testing the first sample, all spectral samples in the spectral data set of other distances except the first sample are used as the training set. Similarly, when testing the second sample, all spectral samples in the spectral data set of other distances except the second sample are used as the training set. In this way, when testing the Nth sample, all spectral samples in the spectral data set of other distances except the Nth sample are used as the training set.
[0028] S5, a deep CNN model is constructed, and the structure is designed as follows: the first layer is a batch normalization layer; the second layer, the fourth layer, the sixth layer, the seventh layer and the ninth layer are convolution layers, and the activation function is a linear rectifier function ReLU; the third layer, the fifth layer and the eighth layer are pooling layers, and the pooling method is a maximum pooling method; the tenth layer is a flattening layer; the eleventh layer is a fully connected layer, and the activation function is ReLU; the twelfth layer is a random inactivation layer; and the thirteenth layer is a fully connected layer, and the activation function is a sigmoid function. By the above method, the initial deep CNN model is obtained.
[0029] S6, a spectral sample weight optimization scheme of the training set is designed, the weight of each spectral sample is calculated according to the detection distance of each spectral sample, and the optimized spectral sample weight of the training set is input into the model in the training process of the deep CNN model.
[0030] The spectral sample weight of the training set is designed to be composed of an absolute distance weight w K1 and a relative distance weight w K2 .
[0031] The absolute distance weight w K of the spectral sample in the spectral data set d K1 is designed as
[0032]
[0033] In formula (2), r K represents the detection distance corresponding to the spectral data set d K .
[0034] When the spectral data set d Q is used as the test set, the relative distance weight w K of the spectral sample in the spectral data set d K2 is designed as
[0035]
[0036] In formula (3), r K represents the detection distance corresponding to the spectral data set d K , and r QThe distance total weight w Q The corresponding detection distance.
[0037] Taking both of the above into account, the distance total weight w K of the spectral samples in the spectral dataset d K is
[0038]
[0039] In formula (4), r K represents the distance total weight w K of the spectral samples in the spectral dataset d Q , and r Q represents the corresponding detection distance.
[0040] The distance total weight w K is normalized to be a dimensionless value between 0 and 1, and the normalized distance total weight w K of the spectral samples in the spectral dataset d K,norm is
[0041]
[0042] In formula (5), K = 1, 2, … Q-1, Q+1, … P.
[0043] In the training process of the deep CNN model, the calculated w K,norm is input into the model as the training set spectral sample weight of the spectral dataset d K .
[0044] S7, train the deep CNN model on the spectral dataset division schemes Dataset1 to DatasetP and perform model classification performance testing. The training of the deep CNN model adopts a batch training mode, the training iteration optimizer adopts the adaptive moment estimation Adam algorithm, and the loss function is the classification cross-entropy. For the training process, the input is the LIBS spectral sample of the training set sample, the weight corresponding to each training set spectral sample, and the category vector real label corresponding to each training set spectral sample, and the output is the calculated value of the category vector of each training set spectral sample. For the testing process, the input is the LIBS spectral sample of the test set sample, and the output is the calculated value of the category vector of the test set spectral sample.
[0045] S8, based on the number of correctly classified spectral samples index (i.e. classification accuracy evaluation index), evaluate the execution effect of the S7 step (i.e. evaluate the model classification performance) and optimize the related hyperparameters of the deep CNN model. The number of correctly classified spectral samples Ncorr is used as the classification accuracy evaluation index of the deep CNN model. Taking the data division scheme DatasetK as an example, the classification accuracy evaluation index of the deep CNN model is calculated as follows: KDuring the process of testing all the spectral samples, the Ncorr value is equal to the number of spectral samples classified correctly in the test, and the greater the Ncorr value, the better the classification performance of the deep CNN model.
[0046] According to the model classification performance, the related hyperparameters of the deep CNN model are optimized, including the batch training sample number batchsize, the learning rate initial value lr, and the iteration number epochs, in the process of optimizing the hyperparameters, the trial value range is set for each hyperparameter, in the specified value range of each hyperparameter, various hyperparameter combination schemes are tried for deep CNN model training, and the Ncorr value that can be obtained by each scheme in the test is calculated, until all the hyperparameter combination schemes are traversed, finally the hyperparameter combination scheme that can make the Ncorr value maximum is selected as the final scheme, thereby completing the optimization of the related hyperparameters.
[0047] S9, after the deep CNN model is constructed, unknown spectral samples can be classified.
[0048] The working principle of the present application is as follows:
[0049] For the training process of the deep CNN model, the default mode is usually to give equal weight to all training set spectral samples. The core innovation of the present application is to give different sample weights to spectral samples of different distances in the training process of the deep CNN model. The design idea of the training set spectral sample weight optimization scheme is as follows: on the one hand, the spectral signal-to-noise ratio of close-range detection is relatively high, and should be paid more attention by the model, and the spectral signal-to-noise ratio of long-range detection is relatively low, and should be paid less attention. On the other hand, the relative distance between the training set spectral sample and the test set spectral sample also affects the classification effect, and the training set spectral sample corresponding to the detection distance close to the test set should be paid more attention by the model, and the training set spectral sample corresponding to the detection distance greatly different from the actual test set should be paid less attention. If the training set spectral samples of different detection distances are simply given equal weight, and the deep CNN model gives equal attention to all training set spectral samples, it is obviously not conducive to improving the classification accuracy of the model in the actual test. Therefore, through the training set spectral sample weight optimization scheme proposed by the present application, the deep CNN model can extract and learn the truly important core feature information, thereby significantly improving the classification performance of the model in the actual test.
[0050] Advantages
[0051] Compared with the prior art, the advantages of the present application are:
[0052] Compared with the method based on the distance correction function, the present application does not need to explore how each experimental condition parameter changes with the detection distance, and does not need to carry out the calculation of the distance correction function, so the calculation time and resources can be significantly saved.
[0053] Compared with the method based on the distance calibration curve, the present application does not need to find and select suitable characteristic radiation spectral lines for each element to be analyzed, and does not need to design the construction scheme of the distance calibration curve, so the calculation time and resources can be significantly saved. In addition, although the method of the present application is designed for identification and classification tasks, its principle is also applicable to quantitative regression tasks, so it has wider applicability.
[0054] Compared with the existing distance correction-free method based on the deep CNN model, the present application fully considers the spectral characteristic differences caused by different detection distances, including the spectral signal-to-noise ratio difference and the relative distance difference between the training set spectral samples and the test set spectral samples, and optimizes the weight of the training set spectral samples on this basis, so as to fully exert the feature extraction and learning ability of the deep CNN model, so that the classification performance of the model is further improved on the basis of maintaining the advantages of distance correction-free.
[0055] In summary, in the training process of the deep CNN model, the present application gives different sample weights to the spectral samples at different detection distances, so that the weights of the spectral samples at different distances are optimized, thereby improving the classification effect of the deep CNN model. The present application has the advantages of no distance correction, efficient training and high accuracy, and can effectively classify the LIBS multi-distance mixed spectrum, and has important value for the field of laser spectrum analysis technology. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 The figure is a schematic diagram of the technical solution of the present application.
[0057] Figure 2 The figure is a distribution diagram of the training set spectral sample weight in each set of spectral data set division scheme with the detection distance.
[0058] Figure 3 The figure is a comparison diagram of the classification accuracy of the deep CNN model before and after the optimization of the training set spectral sample weight. DETAILED DESCRIPTION
[0059] The application process of the method described in the content part will be illustrated below in combination with a specific experimental case:
[0060] A LIBS multi-distance mixed spectrum classification method based on deep CNN and sample weight optimization, in the deep CNN training process, different distance spectrum samples are given different sample weights, that is, according to the absolute distance value of the training set spectrum sample and the distance difference value between the training set spectrum sample and the test set spectrum sample, each spectrum sample weight is specially designed, so that the weight of the spectrum sample at different distances is optimized, thereby improving the classification effect of the deep CNN model. The specific steps are as follows:
[0061] S1, preliminary preparation: prepare the detection sample of laser-induced breakdown spectroscopy (LIBS), record the name and category information of each sample, make a category table covering all N samples, and determine the category vector of each sample;
[0062] In this example, the LIBS experimental detection samples have a total of 37, that is, N=37, which are all national standard substances, and are marked as No. 1 to No. 37. The No. 1 to No. 37 standard sample substances are respectively: 1) clay 2) soft clay 3) carbonate rock 4) kaolin 5) basalt 6) pegmatite 7) dolomite 8) andesite 9) granite gneiss 10) siliceous sandstone 11) shale I type 12) quartz sandstone 13) argillaceous limestone 14) polymetallic ore 15) stream sediment I type 16) stream sediment II type 17) floodplain sediment 18) yellow red soil 19) laterite 20) saline-alkali soil I type 21) saline-alkali soil II type 22) gray calcareous soil 23) beach sediment 24) granite 25) shale II type 26) nickel ore 27) polymetallic lean ore 28) copper-rich ore 29) lead-zinc-rich ore 30) lead ore I type 31) lead ore II type 32) molybdenum ore 33) stream sediment III type 34) stream sediment IV type 35) stream sediment V type 36) stream sediment VI type 37) stream sediment VII type. The 37 detection samples can be divided into six categories, that is, L=6, which are rock I type, rock II type, soil I type, soil II type, sediment and ore. These categories are numbered 1-6, and a category table is made. When making the category table of the sample, list the categories to which each of the 37 samples belongs, then integrate, the categories of all training set and test set samples do not exceed the range of this table, when determining the category vector of each sample, the form of one-hot encoding is adopted, there are 6 different categories in the category table. Table 1 lists the category names and the numbers of the detection samples contained in each category.
[0063] Table 1
[0064] Class Probe sample number 1 Rock Type I 5,6,8,9,11,12,24,27,36 2 Rock Type II 3,7,13 3 Soil Type I 4,10,19 4 Soil Type II 1,2,18,20,21 5 Sediment 15,16,17,22,23,25,26,33,34,35,37 6 Ore 14,28,29,30,31,32
[0065] The category of each detection sample can be represented by a 1x6 vector, containing 1 number 1 and 5 number 0s. Taking sample No. 1 as an example, its category vector C1 can be represented as
[0066] C1 = [0, 0, 0, 1, 0, 0]
[0067] C1 represents the meaning that the No. 1 sample belongs to the category 4, i.e. soil type II.
[0068] S2, collecting LIBS spectral data of each sample at P different distances by using the LIBS spectral detection device. When the LIBS spectral detection device is used for spectral collection, the detection distance is defined as the straight-line distance between the laser external light outlet of the detection device and the geometric center of the surface of the detected sample, and the detection distance is changed by changing the position of the detection device or the detected sample; the spectrum is collected at P different detection distances, and it is ensured that the remaining experimental conditions remain unchanged except for the detection distance.
[0069] In this example, the detection device is a ground backup of the Mars surface composition detector MarSCoDe load, the laser single pulse energy is about 9 mJ, and the laser wavelength is 1064 nm. Eight different detection distances are provided, i.e. P = 8, which are 2.0 m, 2.3 m, 2.5 m, 3.0 m, 3.5 m, 4.0 m, 4.5 m and 5.0 m. At each detection distance, 60 effective LIBS spectra (not including dark background spectra) are collected for each sample, a total of 17760 LIBS spectra, which constitute the original spectral data set.
[0070] S3, preprocessing the LIBS spectra in the original spectral data set, including dark background removal, background baseline removal, wavelength calibration, invalid pixel exclusion and channel splicing, and the LIBS spectra after preprocessing are defined as the LIBS multi-distance mixed spectral data set. The dark background removal operation refers to obtaining the effective spectrum by subtracting the dark background spectrum from the original LIBS spectrum, wherein the dark background spectrum refers to the spectrum responded by the spectrometer when there is no laser excitation; the background baseline removal operation refers to removing the continuous baseline in the denoised spectrum by using the asymmetric least squares baseline correction method; the wavelength calibration refers to converting the spectrometer pixel number into wavelength value by using the multivariate quadratic fitting method; the invalid pixel exclusion refers to removing the pixel response value of each wavelength band of the LIBS spectrum beyond the wavelength range; and the channel splicing refers to splicing the LIBS spectra of multiple wavelength bands after invalid pixel exclusion into a whole according to the wavelength order.
[0071] In this example, the LIBS spectral detection device includes three spectral channels, each channel has 1800 pixels, and a total of 5400 pixels are included; in the invalid pixel exclusion step, 298, 299 and 289 invalid pixel data points are excluded from the three channels respectively; after invalid pixel exclusion and channel splicing, each LIBS spectrum includes 4514 data points, which can be represented as a 4514x1 matrix.
[0072] S4, according to 8 different detection distances, the LIBS multi-distance mixed spectral data set is divided into 8 LIBS spectral data sets under the same detection distance, wherein the spectral data set collected at 2.0 m distance is defined as d1; the spectral data set collected at 2.3 m distance is defined as d2, the spectral data set collected at 2.5 m distance is defined as d3, the spectral data set collected at 3.0 m distance is defined as d4, the spectral data set collected at 3.5 m distance is defined as d5, the spectral data set collected at 4.0 m distance is defined as d6, the spectral data set collected at 4.5 m distance is defined as d7, and the spectral data set collected at 5.0 m distance is defined as d8.
[0073] Each detection distance corresponds to a LIBS multi-distance mixed spectral data set training set-test set division scheme: wherein 2.0 m corresponds to the division scheme defined as Dataset1, indicating that the spectral data set d1 is used as the test set, and the remaining spectral data sets d2 to d8 are used as the training set; 2.3 m corresponds to the division scheme defined as Dataset2, indicating that the spectral data set d2 is used as the test set, and the remaining spectral data sets d1, d3 to d8 are used as the training set; and so on, 5.0 m corresponds to the division scheme defined as Dataset8, indicating that the spectral data set d8 is used as the test set, and the remaining spectral data sets d1 to d7 are used as the training set;
[0074] For each spectral data set division scheme, the test set and the training set have no intersection in the dimensions of detection distance and detection sample. Taking the data set division scheme Dataset1 as an example, the spectral data set d1 is used as the test set, and the spectral data sets of other distances d2 to d8 are used as the training set. When testing, the "leave-one-out" strategy is adopted, and all 37 samples are tested one by one; when testing the first sample, all spectral samples in the spectral data sets of other distances except the first sample are used as the training set; similarly, when testing the second sample, all spectral samples in the spectral data sets of other distances except the second sample are used as the training set; and so on, when testing the 37th sample, all spectral samples in the spectral data sets of other distances except the 37th sample are used as the training set. During the testing of each sample, the training set is 15120 spectral data, and the test set is 60 spectral data.
[0075] S5, a deep convolutional neural network (CNN) model is constructed, in this example, the structure of the deep CNN model is as follows: the first layer is a batch normalization layer; the second layer, the fourth layer, the sixth layer, the seventh layer, and the ninth layer are convolutional layers, and the activation function is a rectified linear unit (ReLU); the third layer, the fifth layer, and the eighth layer are pooling layers, and the pooling method is a max pooling method; the tenth layer is a flattening layer; the eleventh layer is a fully connected layer, and the activation function is ReLU; the twelfth layer is a random inactivation layer; the thirteenth layer is a fully connected layer, and the activation function is a sigmoid function. S6, a training set spectrum sample weight optimization scheme is designed, the weight of each spectrum sample is calculated according to the detection distance of each spectrum sample, and the optimized training set spectrum sample weight is input into the model during the training process of the deep CNN model.
[0076] In this example, programming is performed based on the Python language, and a deep learning Keras framework is used to construct a deep CNN model. In the training process, the weight of the training set spectrum sample is input into the deep CNN model through a sample_weight parameter. The training set spectrum sample weight is designed to consist of an absolute distance weight w K1 and a relative distance weight w K2 The following introduces the weight calculation method taking the spectrum dataset division scheme Dataset1 as an example. In the scheme Dataset1, the spectrum dataset d1 is used as a test set, and the spectrum datasets d2-d8 are used as training sets.
[0077] The absolute distance weight w 21 of the spectrum sample in the spectrum dataset d2 is
[0078]
[0079] The relative distance weight w 22 is
[0080]
[0081] The total weight w2 of the distance is
[0082] w2 = w 21 + w 22 = 0.44m -1 + 3.33m -1 = 3.77m -1
[0083] wherein r1 represents the detection distance corresponding to the spectrum dataset d1, and r2 represents the detection distance corresponding to the spectrum dataset d2;
[0084] Similarly, the total weight of each training set spectrum dataset can be calculated: w3 is 2.40m -1 , w4 is 1.33m -1w5 is 0.95 m -1 w6 is 0.75 m -1 w7 is 0.62 m -1 w8 is 0.53 m -1 The total weight of the above distances is normalized, wherein the normalized total weight of the distance w 2,norm
[0085]
[0086] Similarly, it can be calculated that w 3,norm is 0.23, w 4,norm is 0.13, w 5,norm is 0.09, w 6,norm is 0.07, w 7,norm is 0.06, and w 8,norm is 0.05.
[0087] By analogy, the sample weights of the spectral data sets for the training set in the schemes Dataset2 to Dataset8 can be calculated. Table 2 lists the training set spectral sample weights of the spectral data set division schemes at different detection distances. The description is accompanied by Figure 2 The training set spectral sample weights of the spectral data set division schemes at different detection distances are more intuitively shown (note: for each spectral data set division scheme, the detection distance corresponding to the test set does not exist in the training set, so the column body area corresponding to the distance is marked with a red cross in the figure).
[0088] Table 2
[0089] Dataset 2.0m 2.3m 2.5m 3.0m 3.5m 4.0m 4.5m 5.0m Dataset 1 / 0.36 0.23 0.13 0.09 0.07 0.06 0.05 Dataset 2 0.27 / 0.38 0.12 0.08 0.06 0.05 0.04 Dataset 3 0.18 0.39 / 0.17 0.09 0.07 0.05 0.04 Dataset 4 0.14 0.17 0.22 / 0.21 0.11 0.08 0.06 Dataset 5 0.11 0.12 0.13 0.22 / 0.21 0.12 0.08 Dataset 6 0.10 0.10 0.11 0.13 0.23 / 0.22 0.12 Dataset 7 0.10 0.09 0.10 0.11 0.14 0.24 / 0.23 Dataset 8 0.11 0.10 0.10 0.11 0.12 0.16 0.29 /
[0090] S7, the deep CNN model is trained on the spectral data set division schemes Dataset1 to Dataset8, and the model classification performance test is performed. In this example, the training of the deep CNN model adopts the batch training mode, the training iteration optimizer adopts the adaptive moment estimation Adam algorithm, and the loss function is set as the categorical cross entropy CategoricalCrossentropy. For the training process, the input is the LIBS spectral sample of the training set sample, the weight corresponding to each training set spectral sample, and the category vector real label corresponding to each training set spectral sample, and the output is the calculated value of the category vector of each training set spectral sample; for the test process, the input is the LIBS spectral sample of the test set sample, and the output is the calculated value of the category vector of the test set spectral sample.
[0091] S8, using the number of correctly classified spectral samples Ncorr as an evaluation index of the classification accuracy of the deep CNN model, evaluating the training and prediction effects of the deep CNN model (i.e., evaluating the classification performance of the model), and optimizing the related hyperparameters of the deep CNN model. In this example, taking the data division scheme Dataset1 as an example, in the process of testing all spectral samples in the spectral dataset d1, the value of Ncorr is equal to the number of correctly classified spectral samples in the test, and the larger the value of Ncorr, the better the classification performance of the deep CNN model. According to the classification performance of the model, the related hyperparameters of the deep CNN model are optimized, including the number of batch training samples batchsize, the initial value of the learning rate lr, and the number of iterations epochs. In the process of optimizing the hyperparameters, a trial value range is set for each hyperparameter. In the specified value range of each hyperparameter, various combinations of hyperparameters are tried to train the deep CNN model, and the value of Ncorr that can be achieved by each scheme in the test is calculated, until all combinations of hyperparameters are traversed, and the combination of hyperparameters that can make the value of Ncorr the largest is finally selected as the final scheme. In this example, the final combination of hyperparameters is batchsize = 512, lr = 2e -4 , and epochs = 601.
[0092] S9, after the deep CNN model is constructed, unknown spectral samples can be classified. In this example, in order to reflect the improvement of the classification performance of the deep CNN model by the training set spectral sample weight optimization, experimental groups (after weight optimization) and control groups (before weight optimization) are set according to whether the training set spectral sample weight optimization is performed. For each spectral dataset division scheme, the values of Ncorr of the two groups are calculated, and the results are shown in the accompanying drawings. Figure 3
[0093] As can be seen from the accompanying drawings, Figure 3 under the eight data set division schemes, the number of correctly classified spectral samples Ncorr of the deep CNN model after the training set spectral sample weight optimization is higher than that of the model without sample weight optimization, which can indicate that the training set spectral sample weight optimization can improve the classification performance of the deep CNN model. In summary, the present application has the advantages of no distance correction, high training efficiency, and high accuracy, and can effectively classify LIBS multi-distance mixed spectra, which has important value in the field of laser spectrum analysis technology.
[0094] The above specific embodiments are only an explanation of the present application, and are not a limitation of the present application. Those skilled in the art can make modifications to the embodiments without creative contribution after reading the specification, as long as the modifications are within the scope of the claims of the present application.
Claims
1. A LIBS multi-distance mixed spectrum classification method based on deep CNN and sample weight optimization, characterized in that, Different sample weights are given to spectral samples of different distances in the deep CNN training process, i.e., the spectral sample weights are specially designed according to the absolute distance values of the training set spectral samples and the distance difference values between the training set spectral samples and the test set spectral samples, so that the weights of spectral samples of different distances are optimized, thereby improving the classification effect of the deep CNN model; According to P different detection distances, the LIBS multi-distance mixed spectral data set is divided into P LIBS spectral data sets; In the deep CNN model training process, the optimized training set spectral sample weight is input into the model; the design idea of the training set spectral sample weight optimization scheme is as follows: on the one hand, the spectral signal-to-noise ratio of the near distance detection is relatively high, and should be paid more attention by the model, and the spectral signal-to-noise ratio of the far distance detection is relatively low, and should be paid less attention; on the other hand, the relative distance between the training set spectral samples and the test set spectral samples also affects the classification effect, the training set spectral samples corresponding to the detection distance close to the test set should be paid more attention by the model, and the training set spectral samples corresponding to the detection distance far from the actual test set should be paid less attention; The training set spectral sample weight is designed as the absolute distance weight w K1 and the relative distance weight w K2 Two parts; the spectral data set d K The absolute distance weight w K1 of the spectral sample in the middle is designed as where r K represents the spectral dataset d K corresponding to the detection distance; when the spectral dataset d Q is used as the test set, the spectral dataset d K the relative distance weight w K2 is designed to where r Q represents the distance of the spectral sample in the test set d Q corresponding to the detection distance; the distance total weight w K of the spectral sample in the spectral data set d K is The distance total weight w K is normalized to be a dimensionless value between 0 and 1, the spectral dataset d K The normalized distance total weight w K,norm is where K = 1, 2, Q-1, Q+1, P; in the training process of the deep CNN model, the calculated w K,norm is the training set spectrum sample weight of the spectrum data set d K .
2. The LIBS multi-distance hybrid spectral classification method based on deep CNN and sample weight optimization of claim 1, wherein, The following steps are included: S1, preliminary preparation: preparing a laser-induced breakdown spectroscopy (LIBS) detection sample, recording the name and category information of each sample, making a category table covering all N samples, and determining the category vector of each sample; S2, using a LIBS spectral detection device, collecting LIBS spectra of the detection sample at P different detection distances, and defining the collected LIBS spectra as an original spectral data set; S3, preprocessing the LIBS spectra in the original spectral data set, the preprocessing steps including dark background removal, background baseline removal, wavelength calibration, invalid pixel exclusion and channel splicing, and the preprocessed LIBS spectra are defined as a LIBS multi-distance mixed spectral data set; S4, the spectral dataset collected at the first distance is defined as d1; the spectral dataset collected at the second distance is defined as d2; and so on, the spectral dataset collected at the last distance, i.e., the Pth distance, is defined as dP. P Each detection distance corresponds to a LIBS multi-distance mixed spectral dataset training set-test set division scheme: the first distance corresponds to the division scheme defined as Dataset1, indicating that the spectral dataset d1 is used as the test set, and the remaining spectral datasets d2 to dP are used as the training set; the second distance corresponds to the division scheme defined as Dataset2, indicating that the spectral dataset d2 is used as the test set, and the remaining spectral datasets d1, d3 to dP are used as the training set; and so on. P The Pth distance corresponds to the division scheme defined as DatasetP, indicating that the spectral dataset dP is used as the test set, and the remaining spectral datasets d1 to dP-1 are used as the training set. P The Pth distance corresponds to the division scheme defined as DatasetP, indicating that the spectral dataset dP is used as the test set, and the remaining spectral datasets d1 to dP-1 are used as the training set. P The Pth distance corresponds to the division scheme defined as DatasetP, indicating that the spectral dataset dP is used as the test set, and the remaining spectral datasets d1 to dP-1 are used as the training set. P-1 The Pth distance corresponds to the division scheme defined as DatasetP, indicating that the spectral dataset dP is used as the test set, and the remaining spectral datasets d1 to dP-1 are used as the training set. In step S4, for each spectral dataset division scheme, it is ensured that the test set and the training set have no intersection in both dimensions of detection distance and detected sample: dataset division scheme DatasetK, spectral dataset d K As test set, d1...d K-1 , d K+1 ...d P These other-distance spectral datasets are used as training set, and a "leave-one-out" strategy is adopted for testing; when testing the 1st sample, all spectral samples in the other-distance spectral datasets except the 1st sample are used as training set; similarly, when testing the 2nd sample, all spectral samples in the other-distance spectral datasets except the 2nd sample are used as training set; and so on, when testing the Nth sample, all spectral samples in the other-distance spectral datasets except the Nth sample are used as training set; S5, constructing a deep convolutional neural network (CNN) model, the deep CNN model structure is designed as follows: the first layer is a batch normalization layer; the second layer, the fourth layer, the sixth layer, the seventh layer and the ninth layer are convolutional layers, and the activation function is a linear rectifier function (ReLU); the third layer, the fifth layer and the eighth layer are pooling layers, and the pooling method is maximum pooling; the tenth layer is a flat layer; the eleventh layer is a fully connected layer, and the activation function is ReLU; the twelfth layer is a random inactivation layer; the thirteenth layer is a fully connected layer, and the activation function is a sigmoid function; S6, designing a training set spectral sample weight optimization scheme, calculating the weight of each spectral sample according to the detection distance of each spectral sample; S7, training the deep CNN model on the spectral data set division schemes Dataset1 to DatasetP, and testing the classification performance of the model; S8, evaluating the classification performance of the model according to the classification accuracy evaluation index, and optimizing the related hyperparameters of the deep CNN model; In step S8, the number of correctly classified spectral samples Ncorr is used as an evaluation index of the classification accuracy of the deep CNN model; the data division scheme DatasetK is used to test all spectral samples in the spectral data set d K In the process of testing all spectral samples in the spectral data set d , the value of Ncorr is equal to the number of correctly classified spectral samples in the test, and the larger the value of Ncorr, the better the classification performance of the deep CNN model; according to the classification performance of the model, the related hyperparameters of the deep CNN model are optimized, including the number of batch training samples batchsize, the initial value of the learning rate lr, and the number of iterations epochs. In the process of optimizing the hyperparameters, a trial value range is set for each hyperparameter. In the specified value range of each hyperparameter, various combinations of hyperparameters are tried for deep CNN model training, and the value of Ncorr that can be achieved by each scheme in the test is calculated. Until all combinations of hyperparameters are traversed, the combination of hyperparameters that can maximize the value of Ncorr is finally selected as the final scheme, thereby completing the optimization of the related hyperparameters. S9, completing the construction of the deep CNN model.
3. The LIBS multi-distance hybrid spectral classification method based on deep CNN and sample weight optimization according to claim 2, characterized in that in In step S1, when making the total table of categories of the samples, list the categories to which the N samples respectively belong, then integrate, the categories of all the samples in the training set and the test set do not exceed the range of the total table; when determining the category vector of each sample, adopt the form of one-hot encoding, there are L different categories in the total table of categories, then the category vector C of each sample is a 1xL matrix, containing 1 number 1 and (L-1) number 0, if the sample i belongs to the jth category, then its category vector C i is where the jth element is 1.
4. The LIBS multi-distance hybrid spectral classification method based on deep CNN and sample weight optimization of claim 2, characterized in that in In step S2, the detection distance is defined as the straight-line distance between the laser external light outlet of the LIBS spectrum detection device and the geometric center of the surface of the detection sample. The detection distance is changed by changing the position of the detection device or the detection sample. Spectra are collected at P different detection distances, and the remaining experimental conditions are kept unchanged except for the detection distance.
5. The LIBS multi-distance hybrid spectral classification method based on deep CNN and sample weight optimization according to claim 2, characterized in that In step S3, the dark background removal operation refers to subtracting the dark background spectrum from the original LIBS spectrum to obtain the effective spectrum, wherein the dark background spectrum refers to the spectrum in response to the spectrometer without laser excitation. The background baseline removal operation refers to removing the continuous baseline in the denoised spectrum by using the asymmetric least squares baseline correction method. The wavelength calibration refers to converting the pixel number of the spectrometer into a wavelength value by using a multivariate quadratic fitting method. The invalid pixel screening refers to removing the pixel response value of each wavelength band of the LIBS spectrum that exceeds the wavelength range. The channel splicing refers to splicing the LIBS spectrum of multiple wavelength bands after the invalid pixel screening into a whole according to the wavelength order.
6. The LIBS multi-distance hybrid spectral classification method based on deep CNN and sample weight optimization of claim 2, characterized in that In step S7, the training of the deep CNN model adopts a batch training mode, the training iteration optimizer adopts the Adam algorithm, and the loss function is the classification cross-entropy. For the training process, the input is the LIBS spectrum sample of the training set sample, the weight corresponding to each training set spectrum sample, and the category vector real label corresponding to each training set spectrum sample. The output is the calculated value of the category vector of each training set spectrum sample. For the test process, the input is the LIBS spectrum sample of the test set sample, and the output is the calculated value of the category vector of the test set spectrum sample.
7. The LIBS multi-distance hybrid spectral classification method based on deep CNN and sample weight optimization of claim 2, characterized in that: In step S9, after the deep CNN model is constructed, the unknown spectrum sample is classified.
Citation Information
Patent Citations
Hyperspectral image noise label detection method based on super-pixel weight density
CN110046639A