Method and system for migration and normalization between heterogeneous laser-induced breakdown spectra

The spectral migration and standardization model is constructed through multi-module neural networks and convolutional neural networks, which solves the problem of spectral data differences caused by morphological differences in LIBS equipment, and realizes the standardization of spectral data and the stability of the inversion model, which is suitable for industrial online detection.

CN115436343BActive Publication Date: 2025-08-05SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211119341.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-14
Publication Date
2025-08-05
Estimated Expiration
2042-09-14

AI Technical Summary

Technical Problem

The prior art cannot effectively solve the spectral data differences caused by morphological differences in LIBS equipment, resulting in errors in prediction results. The existing methods such as PDS and PLS algorithms are not suitable for LIBS spectral characteristics, and cannot be implemented or the calculation costs are too high.

Method used

Multi-module parallel multi-layer neural network and convolutional neural network are adopted, combining optimized sample preparation, spectral data acquisition and preprocessing to build a spectral migration and standardization model, and the migration and standardization of spectral data are achieved through segmented regularization and cross-validation optimization model.

Benefits of technology

The generalization of the inversion model when multiple devices are used together is improved, ensuring the stability of the same device under different operating specifications and aging conditions, realizing the standardization of spectral data, and meeting the calculation time requirements of industrial online detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115436343B_ABST
    Figure CN115436343B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for migration and standardization between heterogeneous laser-induced breakdown spectra. After migration and standardization, the secondary spectrum can be effectively predicted by the main spectrum inversion model to obtain high-performance analysis results. The implementation of this method and system is divided into a spectrum migration and standardization model training stage and a model testing stage, under the premise of reasonable preparation of samples and reasonable division of the prepared samples into a model training sample set and a model test sample set, including spectrum preprocessing, regularization, model training and testing. Through a block-based multi-layer neural network, a migration and standardization model between the spectra of the primary and secondary device forms is constructed, which greatly eliminates the differences between the spectra generated by different device forms, improves the generalization of the main spectrum inversion model when multiple devices are used in combination, and improves the stability of the same device under different operating specifications and aging, and the comparability of the same device when experiments are carried out under different conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of laser-induced breakdown spectroscopy (LIBS), and in particular to a method and system for migrating and standardizing heterogeneous laser-induced breakdown spectra, and more particularly to a method and system for migrating and standardizing spectra generated by different laser-induced breakdown spectroscopy device forms based on a neural network. Background Art

[0002] In the practical application of laser-induced breakdown spectroscopy (LIBS), a spectral data inversion model (primary inversion model) is established using spectra (primary spectral data) collected using a given LIBS instrument, a given set of instrument operating parameters, or a given set of experimental conditions (collectively referred to as a given instrument configuration (primary instrument configuration)). This model is then used to calculate and predict material properties (element content, material type, etc.) from spectral data (secondary spectral data) collected using different LIBS instrument configurations (secondary instrument configurations). Differences between spectral data caused by factors unrelated to the measured material properties can lead to errors in the prediction results. Therefore, methods and systems for spectral data migration and standardization are needed. Through migration operations, standardized secondary spectral data are generated, and the differences between the secondary spectral data and the primary spectrum unrelated to the measured material properties are reduced to within the allowable error range.

[0003] Other spectroscopic techniques (such as near-infrared spectroscopy, mass spectrometry, and Raman spectroscopy) also meet these requirements, leading to the development of spectral transfer and normalization methods. For example, piecewise direct standardization (PDS) is based on the partial least squares (PLS) algorithm. However, these methods cannot be directly applied to LIBS spectroscopy: 1) PDS and other methods are optimized for near-infrared spectroscopy, mass spectrometry, and Raman spectroscopy, but lack the specificity of LIBS spectra. 2) The PLS algorithm underlying PDS and other methods is a linear multivariate fit and is not suitable for the nonlinear spectral variations in LIBS spectra caused by the physical processes involved. 3) The algorithms relied on by PDS and other methods are inadequate for practical calculations due to the large wavelength range, numerous channels and features, and resulting large data volumes of LIBS spectra, making them impractical or prohibitively computationally expensive. Therefore, a new technical solution is needed to address these issues.

[0004] The basic idea of this invention is to introduce multi-module parallel multi-layer neural networks and convolutional neural networks as the basic algorithms for LIBS spectral migration and standardization. At the same time, based on the characteristics of LIBS spectra, optimized sample preparation, spectral data acquisition, preprocessing and result generation are performed, so that they can be integrated with multi-module parallel multi-layer neural networks or convolutional neural networks to provide an effective migration and standardization model for LIBS spectra. Summary of the Invention

[0005] In view of the defects in the prior art, the present invention aims to provide a method and system for migration and standardization between heterogeneous laser-induced breakdown spectra.

[0006] According to a method for migration and standardization between heterogeneous laser-induced breakdown spectra provided by the present invention, the method includes a model training phase and a model testing phase, wherein the model training phase includes the following steps:

[0007] Step S1: Prepare samples and divide the samples into a model training sample set and a model testing sample set;

[0008] Step S2: defining the primary device form and the secondary device form;

[0009] Step S3: Collect the original spectra of the model training samples, conduct experiments using the primary and secondary device configurations, and collect the original primary spectra and original secondary spectra of all samples in the training sample set respectively;

[0010] Step S4: spectral preprocessing is performed. The original primary and secondary spectra are subjected to data preprocessing. The preprocessing steps include baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra.

[0011] Step S5: Spectral segment regularization: Each preprocessed primary and secondary spectrum is divided into several spectral segments. All primary and secondary spectral intensities in each spectral segment are divided by the maximum value of the primary and secondary spectrum set in that segment, and the primary and secondary spectral intensity values of each spectral segment are converted to the interval [0, 10].

[0012] Step S6: Generate model training spectrum pairs, randomly pair different regularized primary and secondary spectra of each sample in the training sample set to obtain model training spectrum pairs;

[0013] Step S7: Divide the training and validation spectral data sets, dividing the model training spectrum pairs into the training spectral data set and the validation spectral data set according to the ratio;

[0014] Step S8: Spectral migration and standardization model training, according to the spectrum of step S5, the spectrum is segmented, and a model is built for each segment. The training spectrum data set and the verification spectrum data set are input into the network. By rotation, cross-validation is performed to gradually reduce the difference between the primary and secondary spectra of each training spectrum pair in the segment, and the model is optimized and trained to obtain the corresponding segmented spectrum migration and standardization model. The cross-validation result gives the segmented model calibration performance; each spectrum segment is operated in parallel or sequentially to construct a multi-module segmented model, which corresponds to the spectrum segment one by one; the collection of segmented models is the migration and standardization model of the entire spectrum; the segmented model calibration performance is averaged to obtain the spectrum migration and standardization model calibration performance;

[0015] The model testing phase includes the following steps:

[0016] Step 1: Collect the original spectra of the model test samples. Based on steps S1 and S2, use the primary and secondary device forms to conduct experiments, and collect the original primary spectra and original secondary spectra of all samples in the test sample set respectively;

[0017] Step 2: Spectral preprocessing: the original primary and secondary spectra are subjected to data preprocessing. The preprocessing steps include baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra.

[0018] Step 3: Spectral segment regularization: Each preprocessed sub-spectrum is divided into several spectral segments. All sub-spectral intensities in each spectral segment are divided by the maximum value of the sub-spectral set of that segment, and the sub-spectral intensity values of each spectral segment are converted to the interval [0, 10].

[0019] Step 4: Spectral migration and normalization, regularizing the test sub-spectrum. Input the spectral migration and normalization model to calculate the normalized regularized sub-spectrum of the test sample.

[0020] Step 5: Deregularization: perform an inverse operation on the normalized regularized sub-spectrum of the test sample according to the regularization parameters of the main spectrum of the training sample to generate the normalized sub-spectrum of the test sample;

[0021] Step 6: Evaluation of model prediction performance: the standardized secondary spectrum of the test sample is compared with the pre-processed primary spectrum of the same test sample to give the model prediction performance; it is used to characterize the performance of the standardized secondary spectrum of the unknown sample obtained by using the model to calculate the spectrum generated by the unknown sample through the same secondary device morphology and performing the same spectral preprocessing and regularization / de-regularization as the test sample.

[0022] Preferably, the number of samples prepared in step S1 is determined according to specific application requirements, in particular, the concentration range of the elements contained in the samples must be consistent with the concentration range of the elements of the substances to be detected in the application; the surface physical properties of the samples can represent the surface physical properties of the substances to be detected in the application, and the surface physical properties refer to particle size, roughness, light absorption rate, hardness, and density; the training sample set and the test sample set can be divided according to the sample number ratio, training sample set: test sample set = 7:3 or 8:2; the division of the training samples and the test samples is performed randomly, but it is necessary to ensure that the characteristics of the training sample set are inclusive and representative of the characteristics of the test sample set;

[0023] Inclusiveness: The distribution of chemical and physical properties of the sample matrix of the training sample set includes the chemical and physical properties of the sample matrix of the test sample set; the content distribution of the analyte in the samples of the training sample set includes the content distribution of the analyte in the samples of the test sample set;

[0024] Representativeness: The content of the element to be measured in the test sample set is evenly distributed within the content range of the element to be measured in the training sample set.

[0025] Preferably, the equipment form defined in step S2 includes the equipment manufacturer, batch model, equipment usage parameters and its aging degree, and equipment usage experimental conditions; the defined main equipment form refers to a specific equipment form; the defined secondary equipment form includes at least one specific equipment form different from the main equipment form.

[0026] Preferably, in step S3, the number of pre-processed primary and secondary spectra of each training sample is within 10 2 Magnitude.

[0027] Preferably, the baseline removal in step S4 is to perform baseline removal and noise reduction processing on each original primary and secondary spectrum, and the required operations include finding the spectral peak position, spectral peak restoration, baseline fitting after spectral peak removal, and baseline subtraction; the normalization is to perform normalization processing on all pixels of each original primary and secondary spectrum, including internal standard normalization, total spectral intensity normalization, and laser pulse energy normalization, but not limited to these normalization methods; the order of baseline removal and normalization is specifically optimized and selected according to the data characteristics; the averaging is to accumulate multiple original primary and secondary spectra, and divide the accumulated result by the number of accumulated sheets.

[0028] Preferably, the number of spectral channels of each segment in step S5 is between 10 2 Magnitude.

[0029] Preferably, each model training spectrum pair in step S6 comprises a regular primary spectrum and a regular secondary spectrum from the same sample, and their pairing is performed randomly; generally, the number of model training spectrum pairs is between 10 4 -105 Magnitude.

[0030] Preferably, in step S7, all the model training spectrum pairs are taken as a complete database and divided into a training spectrum data set and a verification spectrum data set, wherein the data volume ratio is training set: verification set = 7:3 or 8:2, but is not limited to these values, and the training set and verification spectrum data set are dynamically divided, rotated during the model training process, and cross-validation is performed; the division of the training spectrum data set and the verification spectrum data set is optimized according to the samples, that is, the training spectrum pairs of a part of the samples are all divided into the training spectrum data set, and the training spectrum pairs of the remaining samples are all divided into the verification spectrum data set.

[0031] Preferably, the block-based multi-layer neural network in step S8 is composed of multiple fully connected multi-layer back-propagation neural networks or convolutional neural networks; the adjustable parameters of the back-propagation neural network include the size of the block, the number of neural network layers and the type of activation function; the adjustable parameters of the convolutional neural network include the convolution kernel size, the number of convolution kernels, the number of convolution network layers and the specific type of convolution layer; the model training belongs to supervised training, and the input information unit is the model training spectrum pair; the regular main spectrum in each pair is the target data for training, and the regular secondary spectrum is the starting data. The model training process is the process of approximating the starting data to the target data, which is controlled by a residual function such as the root mean square error; the validation set data is input into the trained model. If the validation set data does not meet the algorithm evaluation criteria, the selected adjustable parameters are modified and the model is retrained; when the validation set data meets the algorithm evaluation criteria RMSE<10%, the spectral migration and standardization model is established.

[0032] The present invention also provides a migration and standardization system between heterogeneous laser-induced breakdown spectra, the system comprising a model training phase and a model testing phase, the model training phase comprising the following modules:

[0033] Module M1: Prepare samples and divide them into model training sample set and model testing sample set;

[0034] Module M2: defines the primary and secondary equipment forms;

[0035] Module M3: Collect the original spectra of the model training samples, conduct experiments using primary and secondary equipment forms, and collect the original primary spectra and original secondary spectra of all samples in the training sample set respectively;

[0036] Module M4: Spectral preprocessing: The original primary and secondary spectra are preprocessed. The preprocessing module includes baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra.

[0037] Module M5: Spectral segment regularization: Each preprocessed primary and secondary spectrum is divided into several spectral segments. All primary and secondary spectral intensities in each spectral segment are divided by the maximum value of the primary and secondary spectrum set in that segment, and the primary and secondary spectral intensity values of each spectral segment are converted to the interval [0, 10].

[0038] Module M6: Generate model training spectrum pairs, randomly pair different regularized primary and secondary spectra of each sample in the training sample set to obtain model training spectrum pairs;

[0039] Module M7: Divide the training and validation spectral datasets, dividing the model training spectral pairs into the training spectral dataset and the validation spectral dataset according to the ratio;

[0040] Module M8: Spectral migration and standardization model training. According to the spectrum of module M5, each spectrum segment is modeled. The training spectrum data set and the verification spectrum data set are input into the network. Through rotation and cross-validation, the difference between the primary and secondary spectra of each training spectrum pair in the segment is gradually reduced. The model is optimized and trained to obtain the corresponding segmented spectrum migration and standardization model. The cross-validation result gives the segmented model calibration performance. Each spectral segment is operated in parallel or sequentially to construct a multi-module segmented model, which corresponds to the spectral segment one by one. The collection of segmented models is the migration and standardization model of the entire spectrum. The segmented model calibration performance is averaged to obtain the spectral migration and standardization model calibration performance.

[0041] The model testing phase includes the following modules:

[0042] Module 1: Collect the original spectra of the model test samples. Based on modules M1 and S2, the experiment is carried out using the primary and secondary equipment forms to collect the original primary spectra and original secondary spectra of all samples in the test sample set respectively;

[0043] Module 2: Spectral preprocessing: The original primary and secondary spectra are preprocessed. The preprocessing module includes baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra.

[0044] Module 3: Spectral segment regularization: Each preprocessed sub-spectrum is divided into several spectral segments. All sub-spectral intensities in each spectral segment are divided by the maximum value of the sub-spectral set in that segment, and the sub-spectral intensity values of each spectral segment are converted to the interval [0, 10].

[0045] Module 4: Spectral migration and standardization, regularized test sub-spectra are input into the spectral migration and standardization model, and the normalized regularized sub-spectra of the test sample are calculated;

[0046] Module 5: Deregularization: Perform an inverse operation on the normalized regularized sub-spectrum of the test sample according to the regularization parameters of the training sample main spectrum to generate the normalized sub-spectrum of the test sample;

[0047] Module 6: Model prediction performance evaluation, comparing the standardized secondary spectrum of the test sample with the pre-processed primary spectrum of the same test sample to give the model prediction performance; used to characterize the performance of the standardized secondary spectrum of the unknown sample obtained by using the model to calculate the spectrum produced by the unknown sample through the same secondary device form and performing the same spectral pre-processing and regularization / de-regularization as the test sample.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. The present invention uses a block-based multi-layer neural network to construct a migration and normalization model between the primary and secondary device morphology LIBS spectra. The migration operation generates standardized secondary spectral data, and the differences between the secondary spectrum and the primary spectrum that are not related to material properties are reduced to within the allowable error range.

[0050] 2. The present invention improves the generalization of the main inversion model when multiple devices are used together;

[0051] 3. The stability of the same device of the present invention under different operating specifications and aging;

[0052] 4. Comparability of experiments conducted using different conditions with the present invention and the same equipment;

[0053] 5. After the model of the present invention is established, the calculation time of migration calibration meets the requirements of industrial online detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:

[0055] Figure 1 A flowchart for training, testing, and using the spectral migration and normalization model of the present invention;

[0056] Figure 2 Schematic diagram of the multi-layer neural network structure of the present invention;

[0057] Figure 3 Schematic diagram of the block-based convolutional neural network structure of the present invention;

[0058] Figure 4 Specific example diagrams of the pre-processed primary spectrum, pre-processed secondary spectrum, and standardized secondary spectrum of a given sample in the spectrum migration and standardization process of the present invention;

[0059] Figure 5A partial display diagram of a pre-processed primary spectrum, a pre-processed secondary spectrum, and a standardized secondary spectrum of a given sample during the spectrum migration and standardization process of the present invention;

[0060] Figure 6 This is an intensity correlation diagram obtained by the primary and secondary device forms for a certain spectral line in a certain spectrum of a given sample before and after migration and standardization of the present invention. DETAILED DESCRIPTION

[0061] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0062] Example 1:

[0063] According to a method for migration and standardization between heterogeneous laser-induced breakdown spectra provided by the present invention, the method includes a model training phase and a model testing phase. The model training phase includes the following steps:

[0064] Step S1: Prepare samples and divide them into a model training sample set and a model test sample set. The number of samples prepared depends on the specific application requirements, especially the concentration range of the elements contained in the samples must be consistent with the concentration range of the elements of the substances to be detected in the application. The surface physical properties of the samples can represent the surface physical properties of the substances to be detected in the application. The surface physical properties refer to particle size, roughness, light absorption rate, hardness, and density. The training sample set and the test sample set can be divided according to the sample number ratio, training sample set: test sample set = 7:3 or 8:2, but are not limited to these values. The division of training samples and test samples is performed randomly, but it must ensure that the characteristics of the training sample set are inclusive and representative of the characteristics of the test sample set.

[0065] Inclusiveness: The distribution of chemical and physical properties of the sample matrix of the training sample set includes the chemical and physical properties of the sample matrix of the test sample set; the content distribution of the analyte in the samples of the training sample set includes the content distribution of the analyte in the samples of the test sample set;

[0066] Representativeness: The content of the element to be measured in the test sample set is evenly distributed within the content range of the element to be measured in the training sample set.

[0067] Step S2: Define the main equipment form and the secondary equipment form; the defined equipment form includes the equipment manufacturer, batch model, equipment usage parameters and its aging degree, and equipment usage test conditions; the defined main equipment form refers to a specific equipment form; the defined secondary equipment form includes at least one specific equipment form different from the main equipment form.

[0068] Step S3: Collect the original spectra of the model training samples, use the primary and secondary equipment forms to conduct experiments, and collect the original primary spectra and original secondary spectra of all samples in the training sample set respectively; the number of pre-processed primary and secondary spectra of each training sample is 10 2 Magnitude.

[0069] Step S4: spectral preprocessing is performed, and the original primary and secondary spectra are subjected to data preprocessing. The preprocessing steps include baseline removal, normalization and averaging to obtain the preprocessed primary and secondary spectra. Baseline removal is to remove the baseline and reduce the noise of each original primary and secondary spectrum. The required operations include finding the peak position of the spectral line, spectral peak restoration, baseline fitting after spectral peak removal, and baseline subtraction; normalization is to normalize all pixels of each original primary and secondary spectrum, including internal standard normalization, total spectral intensity normalization, and laser pulse energy normalization, but not limited to these normalization methods; the order of baseline removal and normalization is optimized and selected according to the data characteristics; averaging is to accumulate multiple original primary and secondary spectra, and the accumulated result is divided by the number of accumulated sheets.

[0070] Step S5: Spectral segmentation regularization: each pre-processed primary and secondary spectrum is divided into several spectral segments. All primary and secondary spectral intensities in each spectral segment are divided by the maximum value of the primary and secondary spectrum set of the segment, and the primary and secondary spectral intensity values of each spectral segment are converted to the range of [0, 10]. The number of spectral channels in each segment is 10. 2 Magnitude.

[0071] Step S6: Generate model training spectrum pairs, randomly pair different regularized primary and secondary spectra of each sample in the training sample set to obtain model training spectrum pairs; each model training spectrum pair contains a regularized primary spectrum and a regularized secondary spectrum from the same sample, and their pairing is random; generally, the number of model training spectrum pairs is 10 4 -10 5 Magnitude.

[0072] Step S7: Divide the training and validation spectral data sets, and divide the model training spectral pairs into training spectral data sets and validation spectral data sets in proportion; treat all the model training spectral pairs as a complete database, and divide them into training spectral data sets and validation spectral data sets, wherein the data volume ratio is training set: validation set = 7:3 or 8:2, but not limited to these values, and the training set and validation spectral data sets are dynamically divided, rotated during the model training process, and cross-validation is performed; the division of the training spectral data set and the validation spectral data set is optimized according to the samples, that is, the training spectral pairs of a part of the samples are all divided into the training spectral data sets, and the training spectral pairs of the remaining samples are all divided into the validation spectral data sets.

[0073] Step S8: Spectral migration and standardization model training, according to the spectrum of step S5, the spectrum is segmented, and each segment is modeled. The training spectrum data set and the verification spectrum data set are input into the network. By rotation, cross-validation is performed, and the difference between the main and secondary spectra of each training spectrum pair in the segment is gradually reduced. The model is optimized and trained to obtain the corresponding segmented spectrum migration and standardization model. The cross-validation result gives the segmented model calibration performance; each spectrum segment is operated in parallel or sequentially to construct a multi-module segmented model, which corresponds to the spectrum segment one by one; the set of segmented models is the migration and standardization model of the entire spectrum; the segmented model calibration performance is averaged to obtain the spectrum migration and standardization model calibration performance; the block multi-layer neural network is composed of multiple fully connected multi-layer back-propagation neural networks or convolutional neural networks. The network is composed of back propagation neural networks; the adjustable parameters of the back propagation neural network include the block size, the number of neural network layers and the type of activation function; the adjustable parameters of the convolutional neural network include the convolution kernel size, the number of convolution kernels, the number of convolution network layers and the specific type of convolution layer; the model training is supervised training, and the input information unit is the model training spectrum pair; the regular main spectrum in each pair is the target data for training, and the regular secondary spectrum is the starting data. The model training process is the process of approximating the starting data to the target data, which is controlled by a residual function such as the root mean square error; the validation set data is input into the trained model. If the validation set data does not meet the algorithm evaluation criteria, the selected adjustable parameters are modified and the model is retrained; when the validation set data meets the algorithm evaluation criteria RMSE < 10%, the spectral migration and standardization model is established.

[0074] The model testing phase includes the following steps:

[0075] Step 1: Collect the original spectra of the model test samples. Based on steps S1 and S2, use the primary and secondary device forms to conduct experiments, and collect the original primary spectra and original secondary spectra of all samples in the test sample set respectively;

[0076] Step 2: Spectral preprocessing: the original primary and secondary spectra are subjected to data preprocessing. The preprocessing steps include baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra.

[0077] Step 3: Spectral segment regularization: Each preprocessed sub-spectrum is divided into several spectral segments. All sub-spectral intensities in each spectral segment are divided by the maximum value of the sub-spectral set of that segment, and the sub-spectral intensity values of each spectral segment are converted to the interval [0, 10].

[0078] Step 4: Spectral migration and normalization, regularizing the test sub-spectrum. Input the spectral migration and normalization model to calculate the normalized regularized sub-spectrum of the test sample.

[0079] Step 5: Deregularization: perform an inverse operation on the normalized regularized sub-spectrum of the test sample according to the regularization parameters of the main spectrum of the training sample to generate the normalized sub-spectrum of the test sample;

[0080] Step 6: Evaluation of model prediction performance: the standardized secondary spectrum of the test sample is compared with the pre-processed primary spectrum of the same test sample to give the model prediction performance; it is used to characterize the performance of the standardized secondary spectrum of the unknown sample obtained by using the model to calculate the spectrum generated by the unknown sample through the same secondary device morphology and performing the same spectral preprocessing and regularization / de-regularization as the test sample.

[0081] Example 2:

[0082] Example 2 is a preferred example of Example 1 and is used to illustrate the present invention in more detail.

[0083] The present invention also provides a migration and standardization system between heterogeneous laser-induced breakdown spectra, which includes a model training phase and a model testing phase. The model training phase includes the following modules:

[0084] Module M1: Prepare samples and divide them into model training sample set and model testing sample set;

[0085] Module M2: defines the primary and secondary equipment forms;

[0086] Module M3: Collect the original spectra of the model training samples, conduct experiments using primary and secondary equipment forms, and collect the original primary spectra and original secondary spectra of all samples in the training sample set respectively;

[0087] Module M4: Spectral preprocessing: The original primary and secondary spectra are preprocessed. The preprocessing module includes baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra.

[0088] Module M5: Spectral segment regularization: Each preprocessed primary and secondary spectrum is divided into several spectral segments. All primary and secondary spectral intensities in each spectral segment are divided by the maximum value of the primary and secondary spectrum set in that segment, and the primary and secondary spectral intensity values of each spectral segment are converted to the interval [0, 10].

[0089] Module M6: Generate model training spectrum pairs, randomly pair different regularized primary and secondary spectra of each sample in the training sample set to obtain model training spectrum pairs;

[0090] Module M7: Divide the training and validation spectral datasets, dividing the model training spectral pairs into the training spectral dataset and the validation spectral dataset according to the ratio;

[0091] Module M8: Spectral migration and standardization model training. According to the spectrum of module M5, each spectrum segment is modeled. The training spectrum data set and the verification spectrum data set are input into the network. Through rotation and cross-validation, the difference between the primary and secondary spectra of each training spectrum pair in the segment is gradually reduced. The model is optimized and trained to obtain the corresponding segmented spectrum migration and standardization model. The cross-validation result gives the segmented model calibration performance. Each spectral segment is operated in parallel or sequentially to construct a multi-module segmented model, which corresponds to the spectral segment one by one. The collection of segmented models is the migration and standardization model of the entire spectrum. The segmented model calibration performance is averaged to obtain the spectral migration and standardization model calibration performance.

[0092] The model testing phase includes the following modules:

[0093] Module 1: Collect the original spectra of the model test samples. Based on modules M1 and S2, the experiment is carried out using the primary and secondary equipment forms to collect the original primary spectra and original secondary spectra of all samples in the test sample set respectively;

[0094] Module 2: Spectral preprocessing: The original primary and secondary spectra are preprocessed. The preprocessing module includes baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra.

[0095] Module 3: Spectral segment regularization: Each preprocessed sub-spectrum is divided into several spectral segments. All sub-spectral intensities in each spectral segment are divided by the maximum value of the sub-spectral set in that segment, and the sub-spectral intensity values of each spectral segment are converted to the interval [0, 10].

[0096] Module 4: Spectral migration and standardization, regularized test sub-spectra are input into the spectral migration and standardization model, and the normalized regularized sub-spectra of the test sample are calculated;

[0097] Module 5: Deregularization: Perform an inverse operation on the normalized regularized sub-spectrum of the test sample according to the regularization parameters of the training sample main spectrum to generate the normalized sub-spectrum of the test sample;

[0098] Module 6: Model prediction performance evaluation, comparing the standardized secondary spectrum of the test sample with the pre-processed primary spectrum of the same test sample to give the model prediction performance; used to characterize the performance of the standardized secondary spectrum of the unknown sample obtained by using the model to calculate the spectrum produced by the unknown sample through the same secondary device form and performing the same spectral pre-processing and regularization / de-regularization as the test sample.

[0099] Example 3:

[0100] Example 3 is a preferred example of Example 1 and is used to illustrate the present invention in more detail.

[0101] The present invention provides a neural network-based method for migrating and standardizing laser-induced breakdown spectra (LIBS) generated by different device configurations. In the practical application of laser-induced breakdown spectroscopy (LIBS), it is necessary to use a given LIBS device, a given set of device operating parameters, or a given set of experimental conditions, collectively referred to as a given device configuration (primary device configuration), to establish a spectral data inversion model (primary inversion model) based on the spectrum (primary spectral data) collected. This model is then used to calculate and predict material properties (element content, material type, etc.) of spectral data (secondary spectral data) collected by other different LIBS device configurations (secondary device configurations). Differences between spectral data caused by factors unrelated to the measured material properties can lead to errors in the prediction results. Therefore, a method and system for spectral data migration and standardization are needed. Through migration calculations, standardized secondary spectral data is generated, and the differences between the secondary spectral data and the primary spectrum unrelated to the measured material properties are reduced to within an allowable error. This paper introduces multi-module parallel multi-layer neural networks and convolutional neural networks as the basic algorithms for LIBS spectral migration and standardization. Simultaneously, based on the LIBS spectral characteristics, it optimizes sample preparation, spectral data acquisition, preprocessing, and result generation, enabling integration with these multi-module parallel multi-layer neural networks or convolutional neural networks to provide an effective migration and standardization model for LIBS spectra. This paper effectively improves the generalizability of the main inversion model when multiple devices are used in conjunction, the stability of the same device under different operating specifications and aging, and the comparability of experiments performed using the same device under different conditions.

[0102] Sample preparation steps:

[0103] Step 1: Prepare samples: Divide the samples into a model training sample set and a model testing sample set.

[0104] Model training steps:

[0105] Step 2: Define: Define the primary and secondary device forms.

[0106] Step 3: Collect original spectra of model training samples: Use primary and secondary equipment forms to conduct experiments and collect the original primary spectra and original secondary spectra of all samples in the training sample set respectively.

[0107] Step 4: Spectral preprocessing: The original primary and secondary spectra are preprocessed. The preprocessing steps may include baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra.

[0108] Step 5: Spectral segment regularization. Each preprocessed primary and secondary spectrum is divided into several spectral segments. All primary and secondary spectral intensities in each spectral segment are divided by the maximum value of the primary and secondary spectral sets in that segment, and the primary and secondary spectral intensity values of each spectral segment are converted to the range of [0,10].

[0109] Step 6: Generate model training spectrum pairs: Randomly pair different regularized primary and secondary spectra of each sample in the training sample set to obtain model training spectrum pairs.

[0110] Step 7: Divide the training and validation spectral datasets: Divide the model training spectral pairs into a training spectral dataset and a validation spectral dataset in proportion.

[0111] Step eight, spectral migration and standardization model training, according to the spectrum of step five, the spectrum is segmented, and modeling is performed for each segment. The training spectrum data set and the verification spectrum data set are input into the network. By rotation, cross-validation is achieved, and the difference between the primary and secondary spectra of each training spectrum pair in the segment is gradually reduced to achieve model optimization and training, and the corresponding segmented spectrum migration and standardization model is obtained. The cross-validation result gives the segmented model calibration performance, such as the root mean square error (RMSE). Each spectral segment is operated in parallel or sequentially to construct a multi-module segmented model, which corresponds to the spectral segment one by one. The set of segmented models is the migration and standardization model of the entire spectrum. The segmented model calibration performance is averaged to obtain the spectral migration and standardization model calibration performance.

[0112] Model testing steps:

[0113] Step 9: Collection of original spectra of model test samples: Based on steps 1 and 2, experiments are conducted using primary and secondary equipment configurations to collect the original primary spectra and original secondary spectra of all samples in the test sample set.

[0114] Step 10: Spectral preprocessing: The original primary and secondary spectra are subjected to data preprocessing. The preprocessing steps may include baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra.

[0115] Step 11: Spectral segment regularization. Each preprocessed sub-spectrum is divided into several spectral segments. All sub-spectral intensities in each spectral segment are divided by the maximum value of the sub-spectral set of that segment, and the sub-spectral intensity value of each spectral segment is converted to the interval [0,10].

[0116] Step 12, spectral migration and standardization: Regularized test sub-spectra are input into the spectral migration and standardization model to calculate the normalized regularized sub-spectra of the test sample.

[0117] Step 13, deregularization: perform an inverse operation on the normalized regularized sub-spectrum of the test sample according to the regularization parameters of the main spectrum of the training sample, deregularize, and generate a normalized sub-spectrum of the test sample.

[0118] Step 14: Model Prediction Performance Evaluation: Compare the standardized secondary spectrum of the test sample with the preprocessed primary spectrum of the same test sample to provide the model prediction performance, such as mean squared error. This is used to characterize the performance of the model when applying the model to the spectrum of an unknown sample generated by passing it through the same secondary device configuration and undergoing the same spectral preprocessing and regularization / deregularization as the test sample.

[0119] In step 1, the number of samples prepared is determined by the specific application requirements, especially the concentration range of the elements contained in the samples must be consistent with the element concentration range of the substances to be detected in the application; at the same time, the surface physical properties of the samples can represent the surface physical properties of the substances to be detected in the application, and the surface physical properties refer to particle size, roughness, light absorption rate, hardness, density, etc.

[0120] In step 1, the division of the training sample set and the test sample set can be based on the sample quantity ratio, training sample set:test sample set=7:3 or 8:2, but is not limited to these values.

[0121] In step 1, the division of the training samples and the test samples is performed randomly, but the representativeness and inclusiveness of the characteristics of the training sample set to the characteristics of the test sample set must be guaranteed, especially the characteristic element content distribution of the training sample set must be well representative of the characteristic element content distribution of the test sample set. If there are more than two characteristic elements, the sample distribution is displayed on a two-dimensional image through a dimensionality reduction algorithm, such as principal component analysis, to divide the training sample set and the test sample set. If the characteristic element content of the sample is unknown, the sample can be divided into the above-mentioned, random, and mutually representative sample sets according to other sample characteristics.

[0122] In step 2, the defined equipment form includes the equipment manufacturer, batch model, equipment usage parameters and its aging degree, and equipment usage experimental conditions; the defined primary equipment form refers to a specific equipment form; the defined secondary equipment form includes at least one specific equipment form different from the primary equipment form.

[0123] In step 3, the number of original primary and secondary spectra collected for each training sample must be large enough to ensure a large enough number of pre-processed primary and secondary spectra, and then to ensure a large enough number of model training spectrum pairs, as well as a large enough training and validation spectrum data set, and ultimately to ensure the optimization training of the spectral migration and standardization model based on the neural network; generally speaking, the number of pre-processed primary and secondary spectra for each training sample should be around 10 2 Magnitude.

[0124] In step 4, the baseline removal is to remove the baseline and reduce the noise of each original primary and secondary spectrum, and the required operations include finding the spectral peak position, spectral peak restoration, baseline fitting after spectral peak removal, and baseline subtraction; the normalization is to normalize all pixels of each original primary and secondary spectrum, including internal standard normalization, total spectral intensity normalization, laser pulse energy normalization, but not limited to these normalization methods; the order of baseline removal and normalization can be optimized and selected according to the data characteristics; the averaging is to accumulate multiple original primary and secondary spectra, and divide the accumulated result by the number of accumulated sheets.

[0125] In step 5, the number of segments of the spectrum is determined by the number of spectral channels in each segment that can be reasonably processed by the computer used. Generally speaking, the number of spectral channels in each segment can be between 10 and 20. 2 Magnitude.

[0126] In step 6, each model training spectrum pair includes a regular primary spectrum and a regular secondary spectrum from the same sample, and their pairing is performed randomly. 4 -10 5 Magnitude.

[0127] In step seven, all the model training spectrum pairs are taken as a complete database and divided into a training spectrum data set and a validation spectrum data set, wherein the data volume ratio is training set: validation set = 7:3 or 8:2, but not limited to these values.

[0128] In step seven, the training set and validation spectral data set are dynamically divided and rotated during the model training process to achieve cross-validation; the optimal division of the training spectral data set and the validation spectral data set can be performed according to the samples, that is, all the training spectral pairs of a part of the samples are divided into the training spectral data set, and all the training spectral pairs of the remaining samples are divided into the validation spectral data set. Such division helps to enhance the generalization of the trained model.

[0129] In step eight, the block-based multi-layer neural network is composed of multiple fully connected multi-layer back-propagation neural networks, or convolutional neural networks; the adjustable parameters of the back-propagation neural network include the size of the block, the number of neural network layers, and the type of activation function; the adjustable parameters of the convolutional neural network include the size of the convolution kernel, the number of convolution kernels, the number of convolution network layers, and the specific type of convolution layer; the model training is supervised training, and the input information unit is the model training spectrum pair; the regular primary spectrum in each pair is the target data for training, and the regular secondary spectrum is the starting data.

[0130] In step 8, the model training process is the process of approximating the starting data to the target data, which is controlled by a residual function such as the root mean square error. The validation set data is input into the trained model. If the validation set data does not meet the algorithm evaluation criteria (is the root mean square error (RMSE) between the starting data and the target data less than 10% of the spectral average intensity), the selected adjustable parameters are modified and the model is retrained. If the validation set data meets the algorithm evaluation criteria of RMSE < 10%, the spectral migration and normalization model is established.

[0131] In step eight, the block-based multi-layer neural network is a neural network structure designed specifically for LIBS spectrum calibration migration. Applying this structure to the calibration migration of LIBS spectra is an important innovation of the present invention. Its network structure is as follows: the block-based multi-layer neural network is composed of multiple parallel, small multi-layer neural networks. Each small multi-layer neural network takes different spectral segments of the LIBS spectrum as input. The number of blocks determines the scale of each small network, which will affect the calculation time and the final result and is a parameter that needs to be determined. The small multi-layer neural network is composed of several layers of fully connected layers. The number of layers is a parameter that needs to be determined. The number of neurons in each layer is fixed and is the same as the number of channels of the spectrum connected to it. Preferably, the spectrum with a total number of 22,000 channels is divided into 440 spectral segments with 50 channels, and the corresponding segmented multi-layer neural network is constructed. Each segmented network is a fully connected neural network. Sigmoid activation function is used between the first and second layers and between the second and third layers; ReLU activation function is used after the third layer, and mean squared error is used as the loss function. The above 440 segmented multi-layer neural networks constitute a complete neural network, which takes the training and validation spectral datasets as input, uses five-fold cross-validation to train the model, and finally obtains the spectral migration and standardization model.

[0132] In step eight, the convolutional neural network is a machine learning algorithm. Applying this algorithm to the calibration migration of LIBS spectra is an important innovation of the present invention. Its network structure is as follows: the convolutional neural network calculates the spectrum through a convolution kernel with fixed parameters. The convolution kernel length is very small. Each time it is calculated, the corresponding spectrum will be shifted to the right and then calculated. When the convolution kernel is shifted to the right to the tail of the spectrum, its calculation will obtain a data with the same length as the spectrum length. Each convolution kernel will obtain a data as described above after calculation. Multiple convolution kernels calculated simultaneously are called a convolution layer. After performing multiple convolution layers, the final data will be used as the result of the calibration migration, that is, the calibration spectrum. Preferably, for a spectrum with a total channel number of 22,000, a corresponding scale input and output convolutional neural network is constructed. The network uses three convolutional layers. The first and second layers both use four convolution kernels of size 8, and the third layer is a one-dimensional local connection convolution layer with a single convolution kernel of size 8. Sigmoid activation functions are used between the first and second layers and between the second and third layers, and ReLU activation functions are used after the third layer. Mean squared error (MSE) is used as the loss function. The convolutional neural network forms a complete network, using training and validation spectral datasets as input. Five-fold cross-validation is used for model training, ultimately resulting in a spectral transfer and normalization model.

[0133] In step nine, the test set samples are independent of the training set samples; the spectrum acquisition method is the same as that in step three.

[0134] In step 10, the spectrum preprocessing method is the same as that in step 4.

[0135] In step 11, the spectral segment regularization method is the same as in step 5, but only applied to the sub-spectra.

[0136] In step 12, the normalized sub-spectrum of the test sample is calculated using the trained model, and the normalized normalized sub-spectrum of the test sample is output.

[0137] In step thirteen, the deregularization is performed using the main spectrum regularization parameters of the training sample.

[0138] In step 14, the model calibration performance, such as mean squared error (MSE), is used to characterize the performance of the normalized secondary spectrum of an unknown sample obtained by applying the model to the spectrum generated by the unknown sample using the same secondary device configuration and undergoing the same spectral preprocessing and regularization / deregularization as the test sample. This is an important performance indicator for practical model applications.

[0139] Those skilled in the art may understand this embodiment as a more specific description of Embodiment 1 and Embodiment 2.

[0140] Those skilled in the art will appreciate that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, and units for implementing various functions can also be considered as both software modules implementing the method and structures within the hardware component.

[0141] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.

Claims

1. A method for migration and standardization between heterogeneous laser-induced breakdown spectra, characterized in that; The method includes a model training phase and a model testing phase, wherein the model training phase includes the following steps: Step S1: Prepare samples and divide the samples into a model training sample set and a model testing sample set; Step S2: defining the primary device form and the secondary device form; Step S3: Collect the original spectra of the model training samples, conduct experiments using the primary and secondary device configurations, and collect the original primary spectra and original secondary spectra of all samples in the training sample set respectively; Step S4: spectral preprocessing is performed. The original primary and secondary spectra are subjected to data preprocessing. The preprocessing steps include baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra. Step S5: Spectral segment regularization: Each preprocessed primary and secondary spectrum is divided into several spectral segments. All primary and secondary spectral intensities in each spectral segment are divided by the maximum value of the primary and secondary spectrum set in that segment, and the primary and secondary spectral intensity values of each spectral segment are converted to the interval [0, 10]. Step S6: Generate model training spectrum pairs, randomly pair different regularized primary and secondary spectra of each sample in the training sample set to obtain model training spectrum pairs; Step S7: Divide the training and validation spectral data sets, dividing the model training spectrum pairs into the training spectral data set and the validation spectral data set according to the ratio; Step S8: Spectral migration and standardization model training, according to the spectrum of step S5, the spectrum is segmented, and a model is built for each segment. The training spectrum data set and the verification spectrum data set are input into the network. By rotation, cross-validation is performed to gradually reduce the difference between the primary and secondary spectra of each training spectrum pair in the segment, and the model is optimized and trained to obtain the corresponding segmented spectrum migration and standardization model. The cross-validation result gives the segmented model calibration performance; each spectrum segment is operated in parallel or sequentially to construct a multi-module segmented model, which corresponds to the spectrum segment one by one; the collection of segmented models is the migration and standardization model of the entire spectrum; the segmented model calibration performance is averaged to obtain the spectrum migration and standardization model calibration performance; The model testing phase includes the following steps: Step 1: Collect the original spectra of the model test samples. Based on steps S1 and S2, use the primary and secondary device forms to conduct experiments, and collect the original primary spectra and original secondary spectra of all samples in the test sample set respectively; Step 2: Spectral preprocessing: the original primary and secondary spectra are subjected to data preprocessing. The preprocessing steps include baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra. Step 3: Spectral segment regularization: Each preprocessed sub-spectrum is divided into several spectral segments. All sub-spectral intensities in each spectral segment are divided by the maximum value of the sub-spectral set of that segment, and the sub-spectral intensity values of each spectral segment are converted to the interval [0, 10]. Step 4: Spectral migration and normalization, regularizing the test sub-spectrum. Input the spectral migration and normalization model to calculate the normalized regularized sub-spectrum of the test sample. Step 5: Deregularization: perform an inverse operation on the normalized regularized sub-spectrum of the test sample according to the regularization parameters of the main spectrum of the training sample to generate the normalized sub-spectrum of the test sample; Step 6: Model prediction performance evaluation: The standardized secondary spectrum of the test sample is compared with the pre-processed primary spectrum of the same test sample to provide the model prediction performance. This is used to characterize the performance of the standardized secondary spectrum of the unknown sample obtained by applying the model to the spectrum generated by the unknown sample through the same secondary device configuration and undergoing the same spectral pre-processing and regularization / de-regularization as the test sample. The equipment form defined in step S2 includes the equipment manufacturer, batch model, equipment usage parameters and its aging degree, and equipment usage experimental conditions; the defined main equipment form refers to a specific equipment form; the defined secondary equipment form includes at least one specific equipment form different from the main equipment form.

2. The migration and standardization method between heterogeneous laser-induced breakdown spectra according to claim 1, characterized in that: The number of samples prepared in step S1 is determined based on specific application requirements, and the concentration range of the elements contained in the samples must be consistent with the concentration range of the elements of the substances to be detected in the application. The surface physical properties of the samples can represent the surface physical properties of the substances to be detected in the application, and the surface physical properties refer to particle size, roughness, light absorption rate, hardness, and density. The training sample set and the test sample set are divided according to the sample number ratio, training sample set: test sample set = 7:3 or 8:

2. The division of the training samples and the test samples is performed randomly, but it is necessary to ensure that the characteristics of the training sample set are inclusive and representative of the characteristics of the test sample set. Inclusiveness: The distribution of chemical and physical properties of the sample matrix of the training sample set includes the chemical and physical properties of the sample matrix of the test sample set; the content distribution of the analyte in the samples of the training sample set includes the content distribution of the analyte in the samples of the test sample set; Representativeness: The content of the element to be measured in the test sample set is evenly distributed within the content range of the element to be measured in the training sample set.

3. The migration and standardization method between heterogeneous laser-induced breakdown spectra according to claim 1, characterized in that: In step S3, the number of pre-processed primary and secondary spectra of each training sample is 10 2 Magnitude.

4. The method for migration and standardization between heterogeneous laser-induced breakdown spectra according to claim 1, characterized in that: The baseline removal in step S4 is to perform baseline removal and noise reduction processing on each original primary and secondary spectrum, and the required operations include finding the peak position of the spectrum line, restoring the spectrum peak, baseline fitting after the spectrum peak is removed, and baseline subtraction; The normalization is a normalization process performed on all pixels of each original primary and secondary spectrum, including internal standard normalization, total spectrum intensity normalization, and laser pulse energy normalization, but is not limited to these normalization methods; The order of baseline removal and normalization is optimized according to the data characteristics; Averaging is to accumulate multiple original primary and secondary spectra and divide the accumulated result by the number of accumulated images.

5. The migration and standardization method between heterogeneous laser-induced breakdown spectra according to claim 1, characterized in that: The number of spectral channels in each segment in step S5 is 10 2 Magnitude.

6. The method for migration and standardization between heterogeneous laser-induced breakdown spectra according to claim 1, characterized in that: In step S6, each model training spectrum pair includes a regular primary spectrum and a regular secondary spectrum from the same sample, and their pairing is performed randomly; the number of model training spectrum pairs is 10 4 -10 5 Magnitude.

7. The method for migration and standardization between heterogeneous laser-induced breakdown spectra according to claim 1, characterized in that: In step S7, all model training spectrum pairs are treated as a complete database, which is divided into a training spectrum data set and a validation spectrum data set, wherein the data volume ratio is training set: validation set = 7:3 or 8:2, and the training set and validation spectrum data set are dynamically divided and rotated during the model training process for cross-validation; The optimal division of the training spectral dataset and the validation spectral dataset is performed based on samples, that is, all the training spectral pairs of a part of the samples are divided into the training spectral dataset, and all the training spectral pairs of the remaining samples are divided into the validation spectral dataset.

8. The method for migration and standardization between heterogeneous laser-induced breakdown spectra according to claim 1, characterized in that: The block-based multi-layer neural network in step S8 is composed of multiple fully connected multi-layer back-propagation neural networks or convolutional neural networks; the adjustable parameters of the back-propagation neural network include the block size, the number of neural network layers and the type of activation function; the adjustable parameters of the convolutional neural network include the convolution kernel size, the number of convolution kernels, the number of convolution network layers and the specific type of convolution layer; the model training is supervised training, and the input information unit is the model training spectrum pair; the regular primary spectrum in each pair is the target data for training, and the regular secondary spectrum is the starting data. The model training process is the process of approximating the starting data to the target data, which is controlled by the residual function; the validation set data is input into the trained model. If the validation set data does not meet the algorithm evaluation criteria, the selected adjustable parameters are modified and the model is retrained; when the validation set data meets the algorithm evaluation criteria RMSE <10%, the spectral migration and standardization model is established.

9. A system for migration and normalization between heterogeneous laser-induced breakdown spectra, characterized in that; The system includes a model training phase and a model testing phase. The model training phase includes the following modules: Module M1: Prepare samples and divide them into model training sample set and model testing sample set; Module M2: defines the primary and secondary equipment forms; Module M3: Collect the original spectra of the model training samples, conduct experiments using primary and secondary equipment forms, and collect the original primary spectra and original secondary spectra of all samples in the training sample set respectively; Module M4: Spectral preprocessing: The original primary and secondary spectra are preprocessed. The preprocessing module includes baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra. Module M5: Spectral segment regularization: Each preprocessed primary and secondary spectrum is divided into several spectral segments. All primary and secondary spectral intensities in each spectral segment are divided by the maximum value of the primary and secondary spectrum set in that segment, and the primary and secondary spectral intensity values of each spectral segment are converted to the interval [0, 10]. Module M6: Generate model training spectrum pairs, randomly pair different regularized primary and secondary spectra of each sample in the training sample set to obtain model training spectrum pairs; Module M7: Divide the training and validation spectral datasets, dividing the model training spectral pairs into the training spectral dataset and the validation spectral dataset according to the ratio; Module M8: Spectral migration and standardization model training. According to the spectrum of module M5, each spectrum segment is modeled. The training spectrum data set and the verification spectrum data set are input into the network. Through rotation and cross-validation, the difference between the primary and secondary spectra of each training spectrum pair in the segment is gradually reduced. The model is optimized and trained to obtain the corresponding segmented spectrum migration and standardization model. The cross-validation result gives the segmented model calibration performance. Each spectral segment is operated in parallel or sequentially to construct a multi-module segmented model, which corresponds to the spectral segment one by one. The collection of segmented models is the migration and standardization model of the entire spectrum. The segmented model calibration performance is averaged to obtain the spectral migration and standardization model calibration performance. The model testing phase includes the following modules: Module 1: Collect the original spectra of the model test samples. Based on modules M1 and S2, the experiment is carried out using the primary and secondary equipment forms to collect the original primary spectra and original secondary spectra of all samples in the test sample set respectively; Module 2: Spectral preprocessing: The original primary and secondary spectra are preprocessed. The preprocessing module includes baseline removal, normalization, and averaging to obtain preprocessed primary and secondary spectra. Module 3: Spectral segment regularization: Each preprocessed sub-spectrum is divided into several spectral segments. All sub-spectral intensities in each spectral segment are divided by the maximum value of the sub-spectral set in that segment, and the sub-spectral intensity values of each spectral segment are converted to the interval [0, 10]. Module 4: Spectral migration and standardization, regularized test sub-spectra are input into the spectral migration and standardization model, and the normalized regularized sub-spectra of the test sample are calculated; Module 5: Deregularization: Perform an inverse operation on the normalized regularized sub-spectrum of the test sample according to the regularization parameters of the training sample main spectrum to generate the normalized sub-spectrum of the test sample; Module 6: Model prediction performance evaluation: comparison of the normalized secondary spectrum of a test sample with the pre-processed primary spectrum of the same test sample to present the model prediction performance. This module is used to characterize the performance of the normalized secondary spectrum of an unknown sample obtained by applying the model to the spectrum generated by an unknown sample passing through the same secondary device configuration and undergoing the same spectral pre-processing and regularization / de-regularization as the test sample. The equipment form defined in the module M2 includes the equipment manufacturer, batch model, equipment usage parameters and its aging degree, and equipment use experimental conditions; the main equipment form defined refers to a specific equipment form; The defined secondary device form factors include at least one specific device form factor that is different from the primary device form factor.

Citation Information

Patent Citations

  • Drug detector standardization method based on dual-tree complex wavelet algorithm

    CN105784672A

  • Near-infrared spectrum multi-target calibration migration method based on affine transformation

    CN112414966A