Stable learning-based laser-induced breakdown spectroscopy chlorine element analysis method and system
By employing a laser-induced breakdown spectroscopy method based on stable learning, the accuracy problem of chlorine element analysis in different matrices and chloride substances was solved, enabling precise determination of chlorides in different chemical forms and improving the stability and interpretability of the model.
Patent Information
- Application Number
- CN202211279614.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-19
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2042-10-19
AI Technical Summary
Existing technologies cannot effectively analyze the chlorine content in different matrices and substances containing different chlorides. Conventional methods and machine learning models are affected by cation emission spectra and cannot be applied to the analysis of unknown chlorides.
A laser-induced breakdown spectroscopy method based on stable learning was adopted. Through sample preparation, spectral acquisition, preprocessing, regularization, feature selection, and training of a machine learning regression model based on the stable learning paradigm, an analytical model applicable to different matrices and chlorides was established.
It improves the accuracy and stability of chlorine elemental analysis, enabling precise determination of chlorides in different or mixed chemical forms, reducing the model's dependence on confounding variables, and enhancing the model's interpretability and stability.
Smart Images

Figure CN115579075B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of chlorine content determination technology in solid substances, specifically to a method and system for chlorine analysis based on laser-induced breakdown spectroscopy using stable learning, and more particularly to a method and system for determining the chlorine content of substances using laser-induced breakdown spectroscopy (LIBS) based on stable learning data processing. Background Technology
[0002] Non-metallic elements such as chlorine have high excitation energies and can combine with various cations in substances to form different compounds.
[0003] Currently, laser-induced breakdown spectroscopy is commonly used to determine the chlorine content in substances, and the resulting spectrum is dominated by cation emission lines. Conventional chemometrics or machine learning-based regression models therefore exhibit characteristics specific to a given cation and cannot be applied to the analysis of other cationic chlorides, single chlorides of unknown chlorides, or mixed chlorides.
[0004] Therefore, there is a need to invent a method and system that can be used with different matrices and substances containing different chlorides. Summary of the Invention
[0005] To address the shortcomings of existing technologies, the purpose of this invention is to provide a method and system for laser-induced breakdown spectroscopy chlorine element analysis based on stable learning.
[0006] The present invention provides a laser-induced breakdown spectroscopy method for chlorine element analysis based on stable learning, comprising an initial stage, a model training stage, and a model testing stage; the initial stage includes: S1.1, sample preparation: preparing multiple samples containing chlorine compounds, the chlorine content of the multiple samples exhibiting a gradient distribution, which serves as a label for the samples; randomly distinguishing training samples and test samples from the obtained samples, establishing a training sample set and a test sample set, the training sample set including m1 training samples, and the test sample set including m2 test samples; S1.2, spectral acquisition: performing multiple repeated measurements on each sample using laser-induced breakdown spectroscopy, acquiring the original repeated spectra of all samples; S1.3 Spectral Preprocessing: All original repeating spectra of all samples are preprocessed to obtain preprocessed repeating spectra for all samples, forming the training sample preprocessed repeating spectrum set and the test sample preprocessed repeating spectrum set. S1.4 Spectral Regularization: Regularization processing is performed on the training sample preprocessed repeating spectrum set and the test sample preprocessed repeating spectrum set for a given spectral channel. The training sample preprocessed repeating spectrum set is processed first, and the obtained regularization parameters are passed to the test sample preprocessed repeating spectrum set in a one-to-one correspondence with the spectral channels. The test sample preprocessed repeating spectrum set is then regularized according to these parameters to obtain the training sample regularization parameters. The model training phase includes: S2.1, spectral feature selection: For the regularized preprocessed repetitive spectral set of training samples, the importance of each spectral channel is calculated. Using a correlation data analysis algorithm, the correlation between all spectral intensities of a given channel and the corresponding chlorine content of the sample label is calculated. Based on the correlation between each channel and the chlorine content, the spectral channels are arranged from high to low correlation. The n spectral channels following the channel with the highest correlation in the chlorine emission correlation channels are selected as feature spectral channels; S2.2, calculation of sample weight vectors after spectral feature correlation: In the training sample regularized preprocessed repeating spectrum set, which retains only the characteristic spectral channels, the average regularized preprocessed repeating spectrum of each sample is calculated to obtain the average regularized preprocessed spectrum of each sample. The average regularized preprocessed spectrum is then regularized twice over the training sample set, introducing an m1-dimensional sample weight vector. The n averaged quadratic regularized preprocessed spectral features of each training sample are uniformly weighted, and the correlation of the sample weighted spectral features is calculated. An evaluation function is introduced to characterize the correlation. A global optimization algorithm is used to iteratively reduce the sample weight vector until the evaluation function value is reduced to the minimum or the number of iterations reaches a preset value, thus obtaining the optimized sample weight vector; S2.3. Training the machine learning regression model using a stable learning paradigm with feature-decorrelation sample weight vectors: Using the regularized preprocessed repeated spectra of the training samples as input variables, along with the chlorine content of the sample labels, supervised learning is performed under a stable learning paradigm. The loss function of a single input spectrum during training is weighted by the sample weight vector, and then all single-spectrum loss functions are weighted and averaged to obtain the model loss function. The model parameters are iteratively applied, and the gradient of the weighted average loss function is lowered until the expected value is reached. Model training terminates, and the model calibration performance parameters are output. The model testing phase includes: S3.1, Model Testing: Using the regularized preprocessed repeated spectra of the test samples as input variables, the model trained in S2.3 predicts the corresponding chlorine content and compares it with the chlorine content of the sample labels to calculate the model prediction performance parameters. For unknown substances, under the same sample preparation method, experimental method, spectral preprocessing, and spectral regularization conditions, the model is used to predict the chlorine content of the unknown substance. The performance of the results is estimated by the aforementioned model prediction performance parameters.
[0007] Preferably, step S1.1 includes the following steps: S1.1.1 The sample is prepared by cross-combining and physical mixing different matrix powders and different chlorine-containing compound powders within and between classes to prepare a certain amount of sample powder, which is then pressed into cakes. By changing the ratio of matrix powder and chlorine compound powder, a gradient distribution of chlorine content in multiple samples is achieved; S1.1.2 The prepared sample is divided into training sample and test sample at a typical ratio of 2:1 (m1:m2=2:1).
[0008] Preferably, step S1.2 includes the following steps: S1.2.1 Using a conventional LIBS device, optimize its parameters and record the chlorine emission line Cl I 837.59nm line in the spectrum; S1.2.2 Perform repeated measurements on a given sample in the form of a laser ablation pit array covering the sample surface. At the same time, the spectral range should be large enough, including the range from 230nm to 880nm. Continuous and repeated laser pulses are applied to each ablation pit, and the resulting spectra are hardware-accumulated to form a repeating spectrum.
[0009] Preferably, step S1.3 includes the following steps: S1.3.1, effective spectral segment extraction: based on the sensitivity range of the detector in the application device and specific application requirements, the effective portion of the original spectrum is selected for extraction; S1.3.2, spectral baseline removal; S1.3.3, spectral normalization, such as normalization of the total spectral intensity; S1.3.4, spectral averaging: using a moving average method to generate preprocessed repeat spectra of the sample, with the number of preprocessed repeat spectra for each sample maintained at 10. 1 -102 Magnitude.
[0010] Preferably, for step 1.4, firstly, the maximum and minimum light intensity values of the given spectral channel are extracted, and then the maximum and minimum light intensity values are used to linearly transform the light intensity values of the given spectral channel to the [0,1] interval. Similarly, the maximum and minimum light intensity values of the same channel are passed to the test sample set for regularization operation of the test sample preprocessed spectral set.
[0011] Preferably, step 2.1 includes the following steps: S2.1.1, spectral feature selection is performed on the regularized preprocessed repeated spectrum set of training samples. Based on the correlation between the spectral intensity in a given spectral channel and the chlorine content of the corresponding sample, the covariance is used to calculate the correlation. The corresponding Pearson correlation coefficient is used to arrange the spectral channels from most correlated to least correlated. S2.1.2, taking the channel with the highest correlation among the chlorine emission correlation channels as the benchmark, all channels with higher correlation are deleted, and the remaining n spectral channels are retained as spectral features. The choice of n matches the number of training samples and the number of preprocessed repeated spectra for each sample. n is generally 10. 2 Magnitude.
[0012] Preferably, for step 2.2, the average spectrum is calculated using a secondary regularization method for the training sample set: a global optimization algorithm is used to iteratively calculate the sample weight vector, and the Hilbert-Schmidt independence criterion is used to determine the nonlinear correlation between two sets of data. The global optimization algorithm includes a genetic algorithm and a particle swarm optimization algorithm. The algorithm can continuously optimize the specific value of the sample weight vector to continuously reduce the value of the evaluation function. When the number of algorithm iterations exceeds the set number or the value of the evaluation function remains unchanged after multiple iterations, the algorithm stops.
[0013] Preferably, for step 2.3, a stable learning paradigm machine learning regression model is trained using feature-decorrelated sample weight vectors: the loss function guiding the iterative optimization of the model is calculated by weighting the feature-decorrelated sample weight vectors, such as the mean squared error loss function as follows:
[0014]
[0015] Where w i The weights for the i-th sample in the optimized sample weight vector, Label its chlorine content, y ij Let l be the predicted chlorine content of the model for its j-th regularized preprocessed repeating spectrum, and l be the number of regularized preprocessed repeating spectra for each sample.
[0016] Preferably, step S3.1 includes the following steps: S3.1.1, the unknown sample is prepared according to the preparation methods of the training sample and the test sample; S3.1.2, the unknown sample is subjected to spectral acquisition according to the experimental methods of the training sample and the test sample; S3.1.3, the original repeated spectrum of the unknown sample is preprocessed according to the preprocessing method of the original repeated spectrum of the training sample and the test sample; S3.1.4, the preprocessed repeated spectrum of the unknown sample is regularized according to the regularization method of the preprocessed repeated spectrum of the test sample; S3.1.5, the regularized preprocessed repeated spectrum of the unknown sample is input into the model, and its corresponding chlorine content is output. The accuracy and precision of the chlorine content are estimated by the model prediction performance parameters in S3.1.
[0017] The present invention provides a laser-induced breakdown spectroscopy chlorine element analysis system based on stable learning, comprising: a first module for sample preparation: preparing multiple chlorine-containing compound samples, wherein the chlorine content of the multiple samples exhibits a gradient distribution and serves as a label for the samples; randomly distinguishing training samples and test samples from the obtained samples to establish a training sample set and a test sample set, wherein the training sample set includes m1 training samples and the test sample set includes m2 test samples; a second module for spectral acquisition: performing multiple repeated measurements on each sample using laser-induced breakdown spectroscopy to acquire the original repeated spectra of all samples; and a third module for spectral preprocessing: preprocessing the... The original repeating spectra of all samples are preprocessed to obtain preprocessed repeating spectra for all samples, forming the training sample preprocessed repeating spectrum set and the test sample preprocessed repeating spectrum set. The fourth module is used for spectral regularization: regularization processing is performed on the training sample preprocessed repeating spectrum set and the test sample preprocessed repeating spectrum set for a given spectral channel. The training sample preprocessed repeating spectrum set is first processed, and the obtained regularization parameters are passed to the test sample preprocessed repeating spectrum set in a one-to-one correspondence with the spectral channels. The test sample preprocessed repeating spectrum set is then regularized according to these parameters to obtain the training sample preprocessed repeating spectra. The system consists of six modules: Module 1: Multiple spectral sets and test sample regularized preprocessed repeated spectral sets; Module 5: Spectral feature selection: For the training sample regularized preprocessed repeated spectral sets, the importance of each spectral channel is calculated. Using a correlation data analysis algorithm, the correlation between all spectral intensities of a given channel and the corresponding sample's labeled chlorine content is calculated. Based on the correlation between each channel and chlorine content, the spectral channels are arranged from high to low correlation. The n spectral channels following the channel with the highest correlation in the chlorine emission correlation channels are selected as feature spectral channels; Module 6: Sample weight vector calculation for removing spectral feature correlation: After retaining only the feature... In the training sample regularized preprocessing repetitive spectrum set of the spectral channel, the average regularized preprocessed spectrum of each sample is calculated to obtain the average regularized preprocessed spectrum of each sample. The average regularized preprocessed spectrum is then regularized twice over the training sample set. An m1-dimensional sample weight vector is introduced, and the n average quadratic regularized preprocessed spectral features of each training sample are uniformly weighted. The correlation of the sample weighted spectral features is calculated, and an evaluation function is introduced to characterize the correlation. A global optimization algorithm is used to iteratively reduce the sample weight vector until the evaluation function value is reduced to the minimum or the number of iterations reaches the preset value, thus obtaining the optimized sample weight vector.Module 7: Used for training a stable learning paradigm machine learning regression model using feature-decorrelated sample weight vectors. It uses a regularized preprocessed repeating spectral set of training samples as input variables, along with the chlorine content of the sample label. Supervised learning is performed under a stable learning paradigm. The sample weight vector is used to weight the loss function of a single input spectrum during training, and then all single-spectral loss functions are weighted and averaged to obtain the model loss function. The model parameters are iterated repeatedly, and the gradient descent of the weighted average loss function is performed until the expected value is reached. Model training terminates, and the model calibration performance parameters are output. Module 8: Used for model testing. It uses a regularized preprocessed repeating spectral set of test samples as input variables. The model trained in S2.3 predicts the corresponding chlorine content, compares it with the chlorine content of the sample label, and calculates the model prediction performance parameters. For unknown substances, under the same sample preparation method, experimental method, spectral preprocessing, and spectral regularization conditions, the model predicts the chlorine content of the unknown substance. The performance of the results is estimated using the above model prediction performance parameters.
[0018] Compared with the prior art, the present invention has the following beneficial effects:
[0019] 1. This invention, by selecting spectral features based on the chlorine emission line, removes the confusing spectral features extracted by conventional feature selection algorithms that deviate from the essence of the chlorine content determination process in the spectral inversion process, while retaining the key chlorine emission line-related features. It can be applied to the analysis of chlorine in different matrices and substances containing different chlorides, which helps to improve the accuracy and stability of the model for chlorine analysis.
[0020] 2. This invention employs a global optimization algorithm to calculate and optimize the sample weights for decorrelation of spectral features. These sample weights are then used in the training of conventional machine learning regression models to perform a weighted average of the loss function for each sample, eliminating the correlation between features and further reducing the influence of confounding variables. This weakens the dependence of conventional machine learning models on confounding variables and helps improve the interpretability and stability of the model.
[0021] 3. This invention tests the model by setting different types of test sample sets, including test sample sets whose characteristics can be fully represented by the training samples and test sample sets whose characteristics can exceed the statistical distribution range of the training samples, and obtains model prediction performance evaluation parameters. It shows a significant improvement compared with conventional machine learning algorithms and realizes the accurate determination of chlorides or mixed chemical forms of chlorides in samples. Attached Figure Description
[0022] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0023] Figure 1 This invention mainly embodies the flowchart of the laser-induced breakdown spectroscopy chlorine element analysis method based on stable learning;
[0024] Figure 2 This invention primarily demonstrates the performance of machine learning models tested using two different test sample sets.
[0025] Figure 3 The main feature of this invention is the performance demonstration diagram of testing a non-optimized stable learning model using two different test sample sets.
[0026] Figure 4 The main feature of this invention is the performance demonstration diagram showing the testing of the optimized stable learning model using two different test sample sets. Detailed Implementation
[0027] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0028] Example 1
[0029] like Figure 1 , Figure 2 , Figure 3 as well as Figure 4 As shown, the laser-induced breakdown spectroscopy chlorine element analysis method based on stable learning provided by the present invention includes an initial stage, a model training stage, and a model testing stage.
[0030] The initial phase includes:
[0031] S1.1 Sample preparation: Prepare multiple samples containing chlorine compounds. The chlorine content of the multiple samples shows a gradient distribution and is used as a label for the samples. Randomly distinguish training samples and test samples from the obtained samples to establish a training sample set and a test sample set. The training sample set includes m1 training samples and the test sample set includes m2 test samples.
[0032] Specifically, regarding step S1.1:
[0033] S1.1.1 The samples are prepared by cross-combining and physical mixing different matrix powders and different chlorine-containing compound powders within and between classes to prepare a certain amount of sample powders. Then, they are pressed into cakes. By changing the ratio of matrix powder and chlorine compound powder, the chlorine content gradient distribution of multiple samples is achieved.
[0034] S1.1.2. Divide the prepared samples into training samples and test samples at a typical ratio of 2:1 (m1:m2=2:1).
[0035] S1.2 Spectral Acquisition: Laser-induced breakdown spectroscopy was used to perform multiple repeated measurements on each sample, and the original repeated spectra of all samples were acquired.
[0036] Specifically, regarding step S1.2:
[0037] S1.2.1 Using a conventional LIBS instrument, optimize its parameters to record the chlorine emission line Cl I 837.59nm in the spectrum. At the same time, the spectral range should be large enough, including the range from 230nm to 880nm.
[0038] S1.2.2. Repeated measurements are performed on a given sample by covering the sample surface with an array of laser ablation pits. At the same time, each ablation pit is continuously and repeatedly struck by laser pulses. The resulting spectra are accumulated in hardware to form a repeating spectrum.
[0039] S1.3 Spectral preprocessing: Preprocess all original repeat spectra of all samples to obtain preprocessed repeat spectra of all samples, forming the training sample preprocessed repeat spectra set and the test sample preprocessed repeat spectra set.
[0040] Specifically, regarding step S1.3:
[0041] S1.3.1 Effective spectral segment extraction: Based on the sensitivity range of the detector of the application equipment and the specific application requirements, the effective part of the original spectrum is selected for extraction.
[0042] S1.3.2, Spectral baseline removal.
[0043] S1.3.3 Spectral normalization, such as normalization of total spectral intensity.
[0044] S1.3.4 Spectral averaging: Pre-treated repeat spectra of the samples are generated using the moving average method, with the number of pre-treated repeat spectra for each sample maintained at 10. 1 -10 2 Magnitude.
[0045] S1.4 Spectral Regularization: Regularization processing is performed on the preprocessed repeated spectral sets of training samples and the preprocessed repeated spectral sets of test samples for a given spectral channel. The preprocessed repeated spectral sets of training samples are processed first, and the obtained regularization parameters are passed to the preprocessed repeated spectral sets of test samples in a one-to-one correspondence with the spectral channels. The preprocessed repeated spectral sets of test samples are regularized according to these parameters to obtain the regularized preprocessed repeated spectral sets of training samples and the regularized preprocessed repeated spectral sets of test samples, respectively.
[0046] Specifically, regarding step S1.4:
[0047] First, extract the maximum and minimum light intensity values of a given spectral channel. Then, use the maximum and minimum light intensity values to linearly transform the light intensity values of the given spectral channel to the [0,1] interval. Similarly, pass the maximum and minimum light intensity values of the same channel to the test sample set for regularization operation of the test sample preprocessed spectral set.
[0048] The model training phase includes:
[0049] S2.1 Spectral Feature Selection: For the regularized preprocessed repeated spectral set of training samples, the importance of each spectral channel is calculated. Using a correlation data analysis algorithm, the correlation between all spectral intensities of a given channel and the corresponding sample's labeled chlorine content is calculated. Based on the correlation between each channel and the chlorine content, the spectral channels are arranged from high correlation to low correlation. The n spectral channels following the channel with the highest correlation in the chlorine emission correlation channels are selected as feature spectral channels.
[0050] Specifically, regarding step S2.1:
[0051] S2.1.1. Spectral feature selection is performed on the regularized preprocessed repeated spectral set of training samples. Based on the correlation between the spectral intensity in a given spectral channel and the chlorine content of the corresponding sample, the covariance is used to calculate the correlation. The corresponding Pearson correlation coefficient is used to arrange the spectral channels from the most correlated to the least correlated.
[0052] S2.1.2. Using the channel with the highest correlation among the chlorine emission correlation channels as a benchmark, all channels with higher correlation are deleted, and the remaining n spectral channels are retained as spectral features. The choice of n is matched with the number of training samples and the number of preprocessed repeated spectra for each sample. n is generally around 10. 2 Magnitude.
[0053] S2.2 Calculation of Sample Weight Vector for De-spectral Feature Correlation: In the regularized preprocessed repeated spectrum set of training samples that retains only the feature spectral channels, the average regularized preprocessed repeated spectrum of each sample is calculated to obtain the average regularized preprocessed spectrum of each sample. The average regularized preprocessed spectrum is then regularized twice for the training sample set, introducing an m1-dimensional sample weight vector. The n averaged quadratic regularized preprocessed spectral features of each training sample are uniformly weighted, and the correlation of the sample weighted spectral features is calculated. An evaluation function is introduced to characterize the correlation. A global optimization algorithm is used to iteratively reduce the sample weight vector until the evaluation function value is reduced to the minimum or the number of iterations reaches the preset value, thus obtaining the optimized sample weight vector.
[0054] Specifically, regarding step S2.2:
[0055] The average spectrum is calculated using a secondary regularization method for the training sample set: a global optimization algorithm is used to iteratively calculate the sample weight vector. The Hilbert-Schmidt independence criterion is used to determine the nonlinear correlation between two sets of data. The global optimization algorithm includes genetic algorithm and particle swarm optimization algorithm. The algorithm can continuously optimize the specific value of the sample weight vector to continuously reduce the value of the evaluation function. The algorithm stops when the number of iterations exceeds the set number or the value of the evaluation function remains unchanged after multiple iterations.
[0056] S2.3. Using feature-decorrelation sample weight vectors for stable learning paradigm machine learning regression model training: Using the regularized preprocessed repeated spectrum set of training samples as input variables, along with its sample label chlorine content, supervised learning is performed under the stable learning paradigm. The loss function of a single input spectrum during training is weighted on a sample-by-sample basis using the sample weight vector. Then, the loss functions of all single spectra are weighted and averaged to obtain the model loss function. The model parameters are iterated cyclically, and the gradient descent of the weighted average loss function is performed until the expected value is reached. The model training terminates, and the model calibration performance parameters are output.
[0057] Specifically, regarding step S2.3:
[0058] Training a stable learning paradigm machine learning regression model using feature-decorrelated sample weight vectors: The loss function guiding the iterative optimization of the model is calculated by weighting the feature-decorrelated sample weight vectors, such as the mean squared error loss function as follows:
[0059]
[0060] Where w i The weights for the i-th sample in the optimized sample weight vector, Label its chlorine content, y ijLet l be the predicted chlorine content of the model for its j-th regularized preprocessed repeating spectrum, and l be the number of regularized preprocessed repeating spectra for each sample.
[0061] The model testing phase includes:
[0062] S3.1 Model Testing: Using the regularized preprocessed repeated spectral set of the test sample as the input variable, the model trained in S2.3 is used to predict the corresponding chlorine content, which is then compared with the chlorine content of the sample label to calculate the model prediction performance parameters.
[0063] Under the same conditions of sample preparation, experimentation, spectral preprocessing, and spectral regularization, the chlorine content of the unknown substance is predicted using a model, and the performance of the results is estimated by the performance parameters predicted by the model.
[0064] Specifically, for step S3.1: S3.1.1, the unknown sample is prepared according to the preparation methods of training samples and test samples;
[0065] S3.1.2 Unknown samples shall be subjected to spectral acquisition in accordance with the experimental methods used for training samples and test samples;
[0066] S3.1.3 The original repeated spectra of unknown samples are preprocessed according to the original repeated spectra preprocessing methods for training samples and test samples;
[0067] S3.1.4. The repeated spectra of the pre-processed unknown samples are regularized according to the regularization method of the repeated spectra of the pre-processed test samples.
[0068] S3.1.5. Input the regularized preprocessed repeated spectrum of the unknown sample into the model and output its corresponding chlorine content. The accuracy and precision of the chlorine content are estimated by the model prediction performance parameters in S3.1.
[0069] The present invention provides a laser-induced breakdown spectroscopy chlorine element analysis system based on stable learning, employing the aforementioned laser-induced breakdown spectroscopy chlorine element analysis method based on stable learning, comprising:
[0070] Module 1: Sample Preparation: Prepare multiple samples containing chlorine compounds. The chlorine content of the multiple samples exhibits a gradient distribution and serves as a label for the samples. Randomly distinguish between training samples and test samples from the obtained samples to establish a training sample set and a test sample set. The training sample set includes m1 training samples, and the test sample set includes m2 test samples.
[0071] The second module is for spectral acquisition: laser-induced breakdown spectroscopy is used to perform multiple repeated measurements on each sample, and the original repeated spectra of all samples are acquired.
[0072] The third module is used for spectral preprocessing: all original repeat spectra of all samples are preprocessed to obtain preprocessed repeat spectra of all samples, forming the preprocessed repeat spectra set of training samples and the preprocessed repeat spectra set of test samples.
[0073] Module 4: Spectral Regularization: Regularization is performed on the preprocessed repeatable spectra of the training samples and the preprocessed repeatable spectra of the test samples for a given spectral channel. The preprocessed repeatable spectra of the training samples are processed first, and the obtained regularization parameters are passed to the preprocessed repeatable spectra of the test samples in a one-to-one correspondence with the spectral channels. The preprocessed repeatable spectra of the test samples are then regularized according to these parameters to obtain the regularized preprocessed repeatable spectra of the training samples and the regularized preprocessed repeatable spectra of the test samples, respectively.
[0074] Module 5: Spectral Feature Selection: For the regularized preprocessed repeated spectral set of training samples, the importance of each spectral channel is calculated. Using a correlation data analysis algorithm, the correlation between all spectral intensities of a given channel and the corresponding sample's labeled chlorine content is calculated. Based on the correlation between each channel and the chlorine content, the spectral channels are arranged from high correlation to low correlation. The n spectral channels following the channel with the highest correlation in the chlorine emission correlation channels are selected as the feature spectral channels.
[0075] Module 6: Calculation of Sample Weight Vectors for De-spectral Feature Correlation: In the regularized preprocessed repeated spectrum set of training samples that retains only the feature spectral channels, the average of the regularized preprocessed repeated spectra of each sample is calculated to obtain the average regularized preprocessed spectrum of each sample. The average regularized preprocessed spectrum is then regularized twice over the training sample set, introducing an m1-dimensional sample weight vector. The n averaged quadratic regularized preprocessed spectral features of each training sample are uniformly weighted, and the correlation of the weighted spectral features of the samples is calculated. An evaluation function is introduced to characterize the correlation. A global optimization algorithm is used to iteratively reduce the sample weight vector until the evaluation function value is reduced to the minimum or the number of iterations reaches a preset value, thus obtaining the optimized sample weight vector.
[0076] Module 7: Training a stable learning paradigm machine learning regression model using feature-decorrelated sample weight vectors. It uses a regularized preprocessed set of repeated spectra of training samples as input variables, along with the chlorine content of the sample labels. Supervised learning is performed under a stable learning paradigm. The loss function of a single input spectrum during training is weighted by the sample weight vectors, and then all single-spectrum loss functions are weighted and averaged to obtain the model loss function. The model parameters are iterated iteratively, and the gradient of the weighted average loss function is used for gradient descent until the expected value is reached. Model training terminates, and the model calibration performance parameters are output.
[0077] Module 8: Model Testing: Using the regularized preprocessed repeated spectral set of the test sample as input variables, the model trained in S2.3 predicts the corresponding chlorine content, compares it with the chlorine content on the sample label, and calculates the model prediction performance parameters. For unknown substances, under the same sample preparation method, experimental method, spectral preprocessing, and spectral regularization conditions, the model predicts the chlorine content of the unknown substance, and the performance of the results is estimated using the aforementioned model prediction performance parameters.
[0078] Example 2
[0079] The present invention provides a laser-induced breakdown spectral analysis method and system for chlorine elemental analysis based on stable learning, comprising the following steps: an initialization stage, a model training stage, and a model testing stage.
[0080] The initial phase includes:
[0081] S1.1 Sample preparation: Using different matrix powders and different chlorine-containing compound powders, a certain number of sample powders are prepared by cross-combination and physical mixing of the two types of powders within and between the two types. Then, they are pressed into cake-shaped samples. The chlorine content of the samples shows a reasonable gradient distribution and is used as the label of the samples. The obtained samples are randomly distinguished into training samples and test samples to establish a training sample set (containing m1 training samples) and a test sample set (containing m2 test samples).
[0082] Specifically, step S1.1 includes the following steps:
[0083] S1.1.1 The different matrix powders and different chloride powders used to prepare the samples are selected according to the application requirements. Typically, 3 to 4 different matrix powders and 3 to 4 different chloride powders can be selected. Different series of sample powders are generated by combining the matrix and chloride. Typically, more than a dozen sample series can be obtained. By changing the ratio of matrix and chloride, the chlorine content gradient of different series of samples can be achieved. The chlorine content gradient range is also determined according to the application requirements.
[0084] S1.1.2 Furthermore, the matrix powders can be mixed with each other, and the chloride powders can also be mixed with each other. The combination of mixed matrix powders and mixed chlorides produces a new series of samples and achieves a chlorine content gradient.
[0085] S1.1.3. The prepared samples are divided into training samples and test samples in a typical 2:1 ratio (m1:m2=2:1). In particular, two types of test samples can be distinguished. The characteristics of the first type can be fully represented by the training samples, while the characteristics of the second type can exceed the statistical distribution range of the training samples.
[0086] S1.2 Spectral Acquisition: Using laser-induced breakdown spectroscopy, under optimized experimental conditions, the original repeat spectra of all samples were acquired by performing several repeated measurements for each sample, with particular attention paid to the excitation and detection of the Cl I 837.59nm line.
[0087] Specifically, step S2 includes the following steps:
[0088] S1.2.1 Using a conventional LIBS instrument, optimize its parameters to record the Cl I 837.59nm line in the spectrum with a relatively good signal-to-noise ratio. At the same time, the spectral range should be large enough, such as from 230nm to 880nm.
[0089] S1.2.2 Repeated measurements are performed on a given pressed cake sample, using a laser ablation pit array to fully cover the sample surface. Simultaneously, each ablation pit is subjected to continuous, repeated laser pulses. The resulting spectra are hardware-accumulated to form a repeating spectrum. The cumulative number of single-pulse spectra for a single ablation pit can reach 10. 1 The order of magnitude, with the final number of repeated spectra potentially reaching 10. 2 Magnitude.
[0090] S1.3 Spectral preprocessing: Preprocess all original repeat spectra of all samples to obtain preprocessed repeat spectra of all samples.
[0091] Specifically, step S3 includes the following steps:
[0092] S1.3.1 Effective spectral segment extraction: Based on the sensitivity range of the detector of the application equipment and the specific application requirements, the effective part of the original spectrum is selected for extraction.
[0093] S1.3.2, Spectral baseline removal.
[0094] S1.3.3 Spectral normalization, such as normalization of total spectral intensity.
[0095] S1.3.4 Spectral averaging: If a moving average method is used to generate pretreated repeat spectra for a sample, the number of pretreated repeat spectra for each sample should be maintained at 10. 1 -10 2 Magnitude.
[0096] S1.4 Spectral Regularization: The training sample preprocessed repeatable spectral set and the test sample preprocessed repeatable spectral set are respectively subjected to regularization processing for a given spectral channel. The training sample preprocessed spectral set is first processed, and the obtained regularization parameters are passed to the test sample preprocessed spectral set in a one-to-one correspondence with the spectral channels. The test sample preprocessed spectral set is regularized according to these parameters to obtain the training sample and test sample regularized preprocessed repeatable spectral sets respectively. Specifically, for step S4, the regularization operation of the training sample preprocessed spectral set is performed for each given spectral channel. First, the intensity maxima and minima of the channel are extracted, and then they are used to linearly transform the intensity values of the channel to the [0,1] interval. Similarly, the intensity maxima and minima of the same channel are passed to the test sample set for regularization operation of the test sample preprocessed spectral set. When the model trained by this method is used for unknown samples, under the same experimental conditions, the matching equipment is used for spectral acquisition, and the same spectral preprocessing is performed. The preprocessed spectrum of the unknown sample is regularized using the regularization parameters of the training sample preprocessed spectral set (intensity maxima and minima of each given spectral channel).
[0097] Based on the above steps, the method includes two stages: model training and model testing.
[0098] The model training phase includes the following steps related to the regularization preprocessing of the training sample repeating spectral set.
[0099] S2.1 Spectral Feature Selection: For the regularized preprocessed repeated spectral set of training samples, the importance of each spectral channel is calculated. A correlation data analysis algorithm is used to calculate the correlation between all spectral intensities of a given channel and the corresponding sample's labeled chlorine content. Based on the correlation between each channel and the chlorine content, the spectral channels are arranged from high to low correlation. The n spectral channels following the channel with the highest correlation at the Cl I 837.59nm line are selected as the characteristic spectral channels. Specifically, step S5 includes the following steps:
[0100] S2.1.1. Spectral feature selection is performed on the regularized preprocessed repeated spectral set of training samples. Based on the correlation between the spectral intensity in a given spectral channel and the chlorine content of the corresponding sample, it can be calculated by methods such as covariance. The corresponding Pearson correlation coefficient can be used to arrange the spectral channels from the most correlated to the least correlated.
[0101] S2.1.2. Using the channel with the highest correlation in the Cl I 837.59nm line correlation channels as the benchmark, all channels with higher correlation are deleted, and the remaining n spectral channels are retained as spectral features. The choice of n is matched with the number of training samples and the number of preprocessed repeated spectra for each sample. n is generally around 10. 2 The magnitude can be optimized in specific applications.
[0102] S2.2 Calculation of Sample Weight Vector to Remove Spectral Feature Correlation: In the regularized preprocessed repeated spectrum set of training samples that retains only the characteristic spectral channels, the average of the regularized preprocessed repeated spectra of each sample is calculated to obtain the average regularized preprocessed spectrum of each sample. The average spectrum is then regularized twice over the training sample set, introducing an m1-dimensional sample weight vector. The n averaged twice-regularized preprocessed spectral features of each training sample are uniformly weighted, and the correlation of the sample weighted spectral features is calculated. An evaluation function is introduced to characterize the correlation. A global optimization algorithm is used to iteratively reduce the sample weight vector until the evaluation function value is reduced to the minimum or the number of iterations reaches the preset value, thus obtaining the optimized sample weight vector. Specifically, for step S6, the method for calculating the secondary regularization of the average spectrum for the training sample set is the same as in step S1.4. A global optimization algorithm is used to iteratively calculate the sample weight vector. The evaluation function used can be a function such as the Hilbert-Schmidt independence criterion used to determine the nonlinear correlation between two sets of data. The selection of the method also needs to consider the computation time cost, because the algorithm complexity of the sample weight vector optimization process is O(n!). The global optimization algorithm includes genetic algorithm and particle swarm optimization algorithm. The algorithm can continuously optimize the specific value of the sample weight vector to continuously reduce the value of the evaluation function. When the number of algorithm iterations exceeds the set number or the value of the evaluation function remains unchanged after multiple iterations, the algorithm stops. Generally, the number of iterations can be set to 100.
[0103] S2.3. Training a stable learning paradigm machine learning regression model using feature-decorrelated sample weight vectors: Using a regularized preprocessed set of repeated spectra of training samples as input variables, along with the chlorine content of the sample labels, supervised learning is performed under a stable learning paradigm. Sample weight vectors are used to weight the loss function of a single input spectrum during training, and then all single-spectrum loss functions are weighted and averaged to obtain the model loss function. The model parameters are iteratively applied, and the gradient descent of the weighted average loss function is performed until the expected value is reached. Model training terminates, and the model calibration performance parameters are output, including root mean square error (RMSEC) and relative calibration error (REC). Specifically, for step S2.3, the only difference between training a stable learning paradigm machine learning regression model using feature-decorrelated sample weight vectors and typical machine learning regression model training is that the loss function guiding the model iteration optimization is calculated using feature-decorrelated sample weight vectors. That is, introducing sample weights changes the unweighted average calculation of typical machine learning models to a weighted average calculation. The modified RMSEC loss function is as follows:
[0104]
[0105] Where w i The weights for the i-th sample in the optimized sample weight vector, Label its chlorine content, y ij Let l be the predicted chlorine content of the model for its j-th regularized preprocessed repeating spectrum, and l be the number of regularized preprocessed repeating spectra for each sample.
[0106] Model testing phase:
[0107] The model testing phase includes the following steps related to the regularization preprocessing of the repeated spectral set for the test samples:
[0108] S3.1 Model Testing: Using the regularized preprocessed repeated spectral set of the test sample as input variables, the model trained in S2.3 is used to predict the corresponding chlorine content. This prediction is compared with the chlorine content on the sample label, and the model's prediction performance parameters are calculated, including the root mean square error of prediction (RMSEP) and the relative error of prediction (REP). These parameters are used to evaluate the performance of the model when applied to unknown substances, under the same sample preparation method, experimental method, spectral preprocessing, and spectral regularization conditions. The performance of the predicted chlorine content of the unknown substance is estimated using the aforementioned model prediction performance parameters. Specifically, step S3.1 includes the following steps:
[0109] S3.1.1 In practical applications, unknown samples are prepared according to the method for training and testing samples in this method.
[0110] S3.1.2 In practical applications, the unknown sample is subjected to spectral acquisition according to the experimental methods for training and testing the sample using the model of this method.
[0111] S3.1.3 In practical applications, the original repeating spectra of unknown samples are preprocessed according to the model training and test sample original repeating spectra preprocessing method of this method.
[0112] S3.1.4 In practical applications, the pre-processed repeated spectra of unknown samples are regularized according to the regularization method of the pre-processed repeated spectra of test samples in this method.
[0113] S3.1.5 In practical applications, the regularized preprocessed repeated spectra of unknown samples are input into the model trained by this method, and the corresponding chlorine content is output. The accuracy and precision of this content can be estimated by the model prediction performance parameters described in S8.
[0114] Preferred example
[0115] Based on Example 1 or Example 2, the present invention provides a laser-induced breakdown spectroscopy method and system for chlorine element analysis based on stable learning, under laboratory conditions: Martian simulated soil, basalt, and plagioclase amphibole powder are used as sample matrix powders. Sodium chloride, potassium chloride, basic copper chloride, and polyvinyl chloride are used as chloride powders, and these are used as the original powders to prepare the sample powder.
[0116] For a given combination of matrix powder and chloride powder, five chlorine-containing sample powders were generated by varying the powder mixing ratio, with chlorine content gradients covering the range of 1 wt% to 10 wt%. A total of twelve different combinations of matrix powder and chloride powder were used, with five samples of different chlorine contents prepared for each combination, plus three blank matrix powder samples, resulting in a first batch of 63 samples.
[0117] Fifty samples were randomly selected from the model training sample set, and 13 samples were randomly selected from the first test sample set. The test sample set has a similar data distribution to the training sample set (the characteristics of the test sample set can be represented by the training sample set).
[0118] Then, different matrix powders were randomly mixed to produce mixed matrix powder, and different chloride powders were randomly mixed to produce mixed chloride powder. The mixed matrix powder and mixed chloride powder were then randomly mixed to form a second test sample set of six samples. The data distribution of this test sample set did not overlap with the training sample set. The prepared chlorine-containing powder samples were pressed into chlorine-containing cakes, and the chlorine content of the samples was recorded to create sample labels.
[0119] LIBS spectra of the samples were collected using a laboratory LIBS experimental setup. The optimized experimental parameters were: laser energy set to 140 mJ, axial collection, simulating the Martian atmospheric environment, and an echelle spectrometer and ICCD detector set to a delay of 1600 ns and a gate width of 10 μs. Each repetitive spectrum was the sum of ten plasma excitation signals. A 4x10 laser ablation pit matrix was created on each sample surface, with pits spaced 1 mm apart. Forty spectra were collected to form the original repetitive spectrum dataset.
[0120] Based on the spectral characteristics corresponding to the Cl I 837.6nm line, n=200 spectral characteristics were selected.
[0121] A sample weight vector is introduced and globally optimized, specifically using particle swarm optimization and the Hilbert-Schmidt independence criterion. Therefore, the evaluation function is based on the Hilbert-Schmidt independence criterion, and the nonlinear correlation evaluation metric that can be selected in the application is not fixed, ultimately depending on the specific application requirements and model performance.
[0122] The optimized sample weight vector is then introduced into the training of the machine learning model to perform a weighted average of the loss function for the samples.
[0123] The training samples are preprocessed with regularized repeated spectral sets as input variables, and the loss function is applied to a sample-weighted average machine learning model. Cross-validation is used to avoid overfitting and underfitting, ultimately training a stable learning model for predicting chlorine content. The machine learning model specifically utilizes a backpropagation neural network.
[0124] During the model testing process, the two test sample sets mentioned above were used to optimize the testing of the stability of the analytical performance and the generalization ability of the model when predicting the spectra obtained from different types of samples.
[0125] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0126] In the description of this application, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.
[0127] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.
Claims
1. A laser-induced breakdown spectroscopy method for chlorine elemental analysis based on stable learning, characterized in that, This includes the initial phase, the model training phase, and the model testing phase. The initial phase includes: S1.1 Sample preparation: Prepare multiple samples containing chlorine compounds. The chlorine content of the multiple samples shows a gradient distribution and is used as a label for the samples. Randomly distinguish training samples and test samples from the obtained samples to establish a training sample set and a test sample set. The training sample set includes m1 training samples and the test sample set includes m2 test samples. S1.2 Spectral Acquisition: Laser-induced breakdown spectroscopy was used to perform multiple repeated measurements on each sample, and the original repeated spectra of all samples were acquired. S1.3 Spectral preprocessing: Preprocess all original repeat spectra of all samples to obtain preprocessed repeat spectra of all samples, forming the training sample preprocessed repeat spectra set and the test sample preprocessed repeat spectra set. S1.4 Spectral Regularization: The training sample preprocessed repeated spectral set and the test sample preprocessed repeated spectral set are respectively subjected to regularization processing for a given spectral channel. The training sample preprocessed repeated spectral set is first processed, and the obtained regularization parameters are passed to the test sample preprocessed repeated spectral set in a one-to-one correspondence with the spectral channels. The test sample preprocessed repeated spectral set is regularized according to these parameters to obtain the training sample regularized preprocessed repeated spectral set and the test sample regularized preprocessed repeated spectral set, respectively. The model training phase includes: S2.1 Spectral Feature Selection: For the regularized preprocessed repeated spectral set of training samples, the importance of each spectral channel is calculated. Using a correlation data analysis algorithm, the correlation between all spectral intensities of a given channel and the corresponding chlorine content of the sample label is calculated. According to the correlation between each channel and the chlorine content, the spectral channels are arranged from high correlation to low correlation. The n spectral channels after the channel with the highest correlation in the chlorine emission correlation channels are selected as the feature spectral channels. S2.2 Calculation of Sample Weight Vector for De-spectral Feature Correlation: In the regularized preprocessed repeated spectrum set of training samples that retains only the feature spectral channels, the average regularized preprocessed repeated spectrum of each sample is calculated to obtain the average regularized preprocessed spectrum of each sample. The average regularized preprocessed spectrum is then regularized twice for the training sample set, introducing an m1-dimensional sample weight vector. The n averaged quadratic regularized preprocessed spectral features of each training sample are uniformly weighted, and the correlation of the sample weighted spectral features is calculated. An evaluation function is introduced to characterize the correlation. A global optimization algorithm is used to iteratively reduce the sample weight vector until the evaluation function value is reduced to the minimum or the number of iterations reaches the preset value, thus obtaining the optimized sample weight vector. S2.3 Training of a stable learning paradigm machine learning regression model using feature-decorrelation sample weight vectors: Using a regularized preprocessed repeated spectrum set of training samples as input variables, along with the chlorine content of the sample label, supervised learning is performed under a stable learning paradigm. The loss function of a single input spectrum during training is weighted by the sample weight vector, and then all single spectrum loss functions are weighted and averaged to obtain the model loss function. The model parameters are iterated cyclically, and the gradient of the weighted average loss function is gradient-decreased until the expected value is reached. The model training is terminated, and the model calibration performance parameters are output. The model testing phase includes: S3.1 Model Testing: Using the regularized preprocessed repeated spectral set of the test sample as the input variable, the model trained in S2.3 is used to predict the corresponding chlorine content, which is then compared with the chlorine content of the sample label to calculate the model prediction performance parameters. Under the same conditions of sample preparation, experimentation, spectral preprocessing, and spectral regularization, the chlorine content of the unknown substance is predicted using a model, and the performance of the results is estimated by the performance parameters predicted by the model.
2. The laser-induced breakdown spectroscopy method for chlorine element analysis based on stable learning as described in claim 1, characterized in that, For step S1.1, the following steps are included: S1.1.1 The samples are prepared by cross-combining and physical mixing different matrix powders and different chlorine-containing compound powders within and between classes to prepare a certain amount of sample powders. Then, they are pressed into cakes. By changing the ratio of matrix powder and chlorine compound powder, the chlorine content gradient distribution of multiple samples is achieved. S1.1.
2. Divide the prepared samples into training samples and test samples at a typical ratio of 2:1 (m1:m2=2:1).
3. The laser-induced breakdown spectroscopy method for chlorine element analysis based on stable learning as described in claim 1, characterized in that, Step S1.2 includes the following steps: S1.2.1 Using a conventional LIBS instrument, optimize its parameters and record the chlorine emission line Cl I 837.59nm in the spectrum. At the same time, the spectral range should be large enough, including the range from 230nm to 880nm. S1.2.
2. Repeated measurements are performed on a given sample by covering the sample surface with an array of laser ablation pits. At the same time, each ablation pit is continuously and repeatedly struck by laser pulses. The resulting spectra are accumulated in hardware to form a repeating spectrum.
4. The laser-induced breakdown spectroscopy method for chlorine element analysis based on stable learning as described in claim 1, characterized in that, Step S1.3 includes the following steps: S1.3.1 Effective spectral segment extraction: Based on the sensitivity range of the detector of the application equipment and the specific application requirements, the effective part of the original spectrum is selected for extraction. S1.3.2, Spectral baseline removal; S1.3.3 Spectral normalization, such as normalization of total spectral intensity; S1.3.4 Spectral averaging: Pre-treated repeat spectra of the samples are generated using the moving average method, with the number of pre-treated repeat spectra for each sample maintained at 10. 1 -10 2 Magnitude.
5. The laser-induced breakdown spectroscopy method for chlorine element analysis based on stable learning as described in claim 1, characterized in that, Regarding step 1.4, First, extract the maximum and minimum light intensity values of a given spectral channel. Then, use the maximum and minimum light intensity values to linearly transform the light intensity values of the given spectral channel to the [0,1] interval. Similarly, pass the maximum and minimum light intensity values of the same channel to the test sample set for regularization operation of the test sample preprocessed spectral set.
6. The laser-induced breakdown spectroscopy method for chlorine element analysis based on stable learning as described in claim 1, characterized in that, Step 2.1 includes the following steps: S2.1.
1. Spectral feature selection is performed on the regularized preprocessed repeated spectrum set of training samples. Based on the correlation between the spectral intensity in a given spectral channel and the chlorine content of the corresponding sample, the covariance is used to calculate the correlation. The corresponding Pearson correlation coefficient is used to arrange the spectral channels from the most correlated to the least correlated. S2.1.
2. Using the channel with the highest correlation among the chlorine emission correlation channels as a benchmark, all channels with higher correlation are deleted, and the remaining n spectral channels are retained as spectral features. The choice of n is matched with the number of training samples and the number of preprocessed repeated spectra for each sample. n is generally around 10. 2 Magnitude.
7. The laser-induced breakdown spectroscopy method for chlorine element analysis based on stable learning as described in claim 1, characterized in that, Regarding step 2.2, The average spectrum is calculated using a secondary regularization method for the training sample set: a global optimization algorithm is used to iteratively calculate the sample weight vector. The Hilbert-Schmidt independence criterion is used to determine the nonlinear correlation between two sets of data. The global optimization algorithm includes genetic algorithm and particle swarm optimization algorithm. The algorithm can continuously optimize the specific value of the sample weight vector to continuously reduce the value of the evaluation function. The algorithm stops when the number of iterations exceeds the set number or the value of the evaluation function remains unchanged after multiple iterations.
8. The laser-induced breakdown spectroscopy method for chlorine element analysis based on stable learning as described in claim 1, characterized in that, For step 2.3, a stable learning paradigm machine learning regression model is trained using feature-decorrelated sample weight vectors. The loss function guiding the iterative optimization of the model is calculated by weighting the feature-decorrelated sample weight vectors, such as the mean squared error loss function as follows: Where w i The weights for the i-th sample in the optimized sample weight vector, Label its chlorine content, y ij Let l be the predicted chlorine content of the model for its j-th regularized preprocessed repeating spectrum, and l be the number of regularized preprocessed repeating spectra for each sample.
9. The laser-induced breakdown spectroscopy method for chlorine element analysis based on stable learning as described in claim 1, characterized in that, Step S3.1 includes the following steps: S3.1.1 Unknown samples shall be prepared in accordance with the preparation methods for training samples and test samples; S3.1.2 Unknown samples shall be subjected to spectral acquisition in accordance with the experimental methods used for training samples and test samples; S3.1.3 The original repeated spectra of unknown samples are preprocessed according to the original repeated spectra preprocessing methods for training samples and test samples; S3.1.
4. The repeated spectra of the pre-processed unknown samples are regularized according to the regularization method of the repeated spectra of the pre-processed test samples. S3.1.
5. Input the regularized preprocessed repeated spectrum of the unknown sample into the model and output its corresponding chlorine content. The accuracy and precision of the chlorine content are estimated by the model prediction performance parameters in S3.
1.
10. A laser-induced breakdown spectroscopy chlorine element analysis system based on stable learning, characterized in that, The laser-induced breakdown spectroscopy method for chlorine element analysis based on stable learning as described in any one of claims 1-9 includes: Module 1: Sample Preparation: Prepare multiple samples containing chlorine compounds. The chlorine content of the multiple samples shows a gradient distribution and serves as a label for the samples. Randomly distinguish training samples and test samples from the obtained samples to establish a training sample set and a test sample set. The training sample set includes m1 training samples, and the test sample set includes m2 test samples. The second module is used for spectral acquisition: laser-induced breakdown spectroscopy is used to perform multiple repeated measurements on each sample, and the original repeated spectra of all samples are acquired. The third module is used for spectral preprocessing: all original repeat spectra of all samples are preprocessed to obtain preprocessed repeat spectra of all samples, forming a training sample preprocessed repeat spectra set and a test sample preprocessed repeat spectra set. Module 4: Spectral Regularization: Regularization is performed on the preprocessed repeating spectra of the training samples and the preprocessed repeating spectra of the test samples for a given spectral channel. The preprocessed repeating spectra of the training samples are processed first, and the obtained regularization parameters are passed to the preprocessed repeating spectra of the test samples in a one-to-one correspondence with the spectral channels. The preprocessed repeating spectra of the test samples are regularized according to these parameters to obtain the preprocessed repeating spectra of the training samples and the preprocessed repeating spectra of the test samples. Module 5: Spectral Feature Selection: For the regularized preprocessed repeated spectral set of training samples, the importance of each spectral channel is calculated. Using a correlation data analysis algorithm, the correlation between all spectral intensities of a given channel and the corresponding chlorine content of the sample label is calculated. The spectral channels are arranged from high correlation to low correlation according to the correlation between each channel and the chlorine content. The n spectral channels after the channel with the highest correlation in the chlorine emission correlation channels are selected as the feature spectral channels. Module 6: Calculation of sample weight vector for removing spectral feature correlation: In the training sample regularized preprocessed repeated spectrum set that retains only the feature spectral channels, the average of the regularized preprocessed repeated spectra of each sample is calculated to obtain the average regularized preprocessed spectrum of each sample. The average regularized preprocessed spectrum is then regularized twice for the training sample set, introducing an m1-dimensional sample weight vector. The n averaged quadratic regularized preprocessed spectral features of each training sample are uniformly weighted, and the correlation of the sample weighted spectral features is calculated. An evaluation function is introduced to characterize the correlation. A global optimization algorithm is used to iteratively reduce the sample weight vector until the evaluation function value is reduced to the minimum or the number of iterations reaches the preset value, thus obtaining the optimized sample weight vector. Module 7: Training a stable learning paradigm machine learning regression model using feature-decorrelated sample weight vectors: Using a regularized preprocessed set of repeated spectra of training samples as input variables, along with the chlorine content of the sample labels, supervised learning is performed under a stable learning paradigm. The loss function of a single input spectrum during training is weighted by the sample weight vector, and then all single spectrum loss functions are weighted and averaged to obtain the model loss function. The model parameters are iterated cyclically, and the gradient descent of the weighted average loss function is performed until the expected value is reached. The model training terminates, and the model calibration performance parameters are output. Module 8: Model Testing: Using the regularized preprocessed repeated spectral set of the test sample as input variable, the model trained in S2.3 is used to predict the corresponding chlorine content, which is then compared with the chlorine content of the sample label to calculate the model prediction performance parameters. Under the same conditions of sample preparation, experimentation, spectral preprocessing, and spectral regularization, the chlorine content of the unknown substance is predicted using a model, and the performance of the results is estimated by the performance parameters predicted by the model.