Multi-element quantitative analysis method and system for laser-induced breakdown spectroscopy
By adopting a multi-element quantitative analysis model based on deep learning in LIBS technology, the problem of insufficient accuracy and robustness in quantitative analysis of low-content elements is solved, and efficient processing of complex spectral data and the accuracy of multi-element quantitative analysis is improved.
Patent Information
- Application Number
- CN202510199100.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The existing laser induced breakdown spectroscopy (LIBS) technology has problems with insufficient accuracy and robustness in quantitative analysis of low-content elements, especially when dealing with high-dimensional, complex nonlinear spectral data, which is difficult to effectively analyze in traditional methods.
A multi-element quantitative analysis model based on deep learning, including multi-level neural network structure and multi-level feature extraction module, is adopted to extract and fuse multi-scale features through attention fusion units and feature splicing sub-units, and optimize the model through adaptive weighted mean square error loss function.
It significantly improves the detection accuracy of low-content elements, enhances the processing ability of complex spectral data, improves the accuracy and robustness of multi-element quantitative analysis, and is suitable for a variety of spectral analysis tasks.
Smart Images

Figure CN119691407B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of laser induced breakdown spectroscopy (LIBS) analysis, and in particular to a multi-element quantitative analysis method and system for laser induced breakdown spectroscopy. Background Art
[0002] Laser Induced Breakdown Spectroscopy (LIBS) is an efficient elemental analysis technology that is non-destructive, fast, and accurate. It is widely used in material analysis, environmental monitoring, food safety, and other fields. LIBS technology uses high-energy laser pulses to excite the sample surface, causing it to break down and generate plasma. The spectral characteristics of the plasma emission are then analyzed to obtain the composition information of the elements in the sample. Although LIBS technology has many advantages in elemental analysis, its application in the quantitative analysis of low-content elements still faces many challenges.
[0003] First, LIBS spectral data are usually high-dimensional and complex, containing a large number of spectral peaks and noise components. In practical applications, the spectral signals of low-content elements are relatively weak and easily submerged by noise, resulting in reduced accuracy of quantitative analysis of low-content elements. Secondly, the processing of LIBS spectral data involves many factors, such as laser power, sample surface state, environmental conditions, etc., which further increase the complexity of spectral data. Therefore, traditional analysis methods (such as quantitative analysis methods based on peak area and peak height) are often difficult to effectively process high-dimensional, nonlinear, and noisy spectral data, resulting in poor application results in low-content element analysis.
[0004] At present, machine learning and deep learning methods have been applied to the analysis of LIBS spectral data in order to improve the accuracy and robustness of the analysis. Traditional neural network (NN), support vector machine (SVM), random forest (RF), linear regression (LR) and XGBoost methods can solve the analysis problem of LIBS spectral data to a certain extent. However, these methods still have some limitations in practical applications. For example, when facing complex spectral data, traditional neural networks are prone to gradient vanishing or exploding problems, and the prediction accuracy for low-content elements is low; support vector machine (SVM) has high computational complexity when processing large-scale data and is more sensitive to noise; although random forest (RF) can resist overfitting well, its model has poor interpretability and consumes a lot of computing resources; linear regression (LR) assumes that there is a linear relationship between data and cannot effectively handle complex nonlinear features; although XGBoost has strong prediction ability, it has great challenges in hyperparameter adjustment and computational overhead.
[0005] In order to solve the above problems, deep learning technology (especially convolutional neural networks and deep neural networks) has gradually become a research hotspot in the field of LIBS spectral analysis in recent years. Deep learning methods can automatically extract high-level features from data, reduce the reliance on manual feature design, and achieve efficient model optimization through end-to-end training. However, the existing deep learning models still have some shortcomings, especially when dealing with quantitative analysis tasks of low-content elements, their prediction accuracy and generalization ability still need to be further improved.
[0006] Therefore, this technical field is in urgent need of a new optimization method that can propose an efficient and accurate multi-element quantitative analysis method based on deep learning technology and combined with the characteristics of LIBS spectral data, especially in the analysis of low-content elements, which can maintain high accuracy and robustness. Summary of the invention
[0007] The invention provides a multi-element quantitative analysis method and system of laser induced breakdown spectroscopy, which are used to solve the technical problem of low accuracy of the multi-element quantitative analysis method of laser induced breakdown spectroscopy.
[0008] In order to solve the above technical problems, the technical solution proposed by the present invention is:
[0009] A multi-element quantitative analysis method of laser induced breakdown spectroscopy comprises the following steps:
[0010] A multi-element quantitative analysis model of laser induced breakdown spectroscopy is constructed, wherein the multi-element quantitative analysis model comprises an input layer, a feature extraction layer and an output layer connected in sequence, wherein the feature extraction layer comprises at least one multi-level feature extraction module, wherein the multi-level feature extraction module comprises an attention fusion unit, wherein the attention fusion unit comprises a plurality of parallel feature extraction branches and a feature splicing sub-unit, wherein the plurality of feature extraction branches are respectively used to perform convolution operations of different sizes and dimensions to extract feature maps of different scales, wherein the plurality of feature extraction branches are sorted according to the convolution depth, wherein each feature extraction branch is jump-connected to other feature extraction branches with adjacent sorting positions and a greater convolution depth than the branch; wherein the feature splicing sub-unit is used to splice the feature maps output by the plurality of feature extraction branches;
[0011] The spectral characteristic data of the laser induced breakdown spectrum are obtained from historical data to construct a training data set, and the multi-element quantitative analysis model is trained with the training data set, and the element content of the target laser induced breakdown spectrum is predicted with the multi-element quantitative analysis model.
[0012] Preferably, the attention fusion unit includes a small-scale feature extraction branch, a medium-scale feature extraction branch, and a high-scale feature extraction branch;
[0013] The small-scale feature extraction branch includes a first convolution, a second convolution, and a feature fusion block, the output end of the first convolution is connected to the input end of the second convolution and the feature fusion block respectively, and the output end of the second convolution is connected to the input end of the feature fusion block;
[0014] The mesoscale feature extraction branch includes a third convolution, a fourth convolution, and a first feature splicing block, the output end of the third convolution is connected to the input end of the fourth convolution and the first feature splicing block respectively, the output end of the fourth convolution is connected to the input end of the first feature splicing block; the output end of the second convolution is also jump-connected to the input end of the first feature splicing block;
[0015] The high-scale feature extraction branch includes a fifth convolution, a sixth convolution, and a second feature splicing block. The output end of the fifth convolution is respectively connected to the input ends of the sixth convolution and the second feature splicing block, and the output end of the fifth convolution is connected to the input end of the second feature splicing block; the output end of the sixth convolution is also jump-connected to the input end of the second feature splicing block.
[0016] Preferably, the convolution scale of the first convolution is larger than the convolution scale of the second convolution, the convolution scale of the third convolution is larger than the convolution scale of the fourth convolution, and the convolution scale of the fifth convolution is larger than the convolution scale of the sixth convolution; the convolution scale of the first convolution is smaller than the convolution scale of the third convolution, and the convolution scale of the third convolution is smaller than the convolution scale of the fifth convolution; the convolution scale of the second convolution is smaller than the convolution scale of the fourth convolution, and the convolution scale of the fourth convolution is smaller than the convolution scale of the sixth convolution;
[0017] and / or
[0018] The attention fusion unit also includes multiple feature refining blocks, which correspond one-to-one to multiple feature extraction branches. The output end of each feature extraction branch is connected to the input end of the feature splicing sub-unit through its corresponding feature refining block.
[0019] Preferably, the feature splicing subunit includes a third feature splicing block and a dimension reduction block, the output end of each feature extraction branch is connected to the input end of the third feature splicing block through its corresponding feature refinement block, and the output block of the third feature splicing block is connected to the input end of the dimension reduction block;
[0020] and / or
[0021] The attention fusion unit includes a dimensionality reduction subunit, and the output end of the dimensionality reduction subunit is respectively connected to the input ends of the small-scale feature extraction branch, the medium-scale feature extraction branch, and the high-scale feature extraction branch.
[0022] Preferably, the multi-level feature extraction module further includes: a downsampling unit and a feature enhancement unit, and the downsampling unit, the attention fusion unit and the feature enhancement unit are connected in series in sequence.
[0023] Preferably, the multi-level feature extraction module includes a plurality of multi-level feature extraction modules, and the plurality of multi-level feature extraction modules are sequentially connected in series to form a feature extraction layer;
[0024] and / or
[0025] The output layer includes a global feature integration module and a fully connected module, and the global feature integration module and the fully connected module are connected in series in sequence.
[0026] Preferably, the loss function of the multi-element quantitative analysis model satisfies:
[0027] ;
[0028] in, is the loss function value; is the total number of samples; is the number of predicted elements; For sample The true value of The content of each element; For sample The predicted value of The content of each element; The weights are adaptive and dynamically adjusted for low-content elements and large error values.
[0029] Preferably, the adaptive weight satisfies:
[0030] ;
[0031] in, is an inverse weight term, which is used to increase the weight of low-content elements. is a smoothing constant to prevent the denominator from being zero; is the error weight term, and the weight is dynamically increased as the prediction error increases; It is a weight adjustment parameter used to balance the influence of the inverse weight term and the error weight term.
[0032] Preferably, spectral characteristic data of the laser induced breakdown spectroscopy are obtained from historical data to construct a training data set, including:
[0033] Constructing a spectral feature matrix based on the acquired spectral feature data, obtaining the actual element content corresponding to the spectral feature data and constructing an element content matrix corresponding to the spectral feature matrix;
[0034] The spectral feature matrix is normalized, and a smoothing filter algorithm is used to remove noise from low-content elements in the element content matrix. The smoothing filter algorithm includes one or a combination of the following:
[0035] Gaussian smoothing algorithm, threshold truncation algorithm, weighted average smoothing algorithm;
[0036] The element contents in the denoised element content matrix are used as labels of the corresponding spectral features in the spectral feature matrix.
[0037] A computer system comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the computer program.
[0038] The present invention has the following beneficial effects:
[0039] 1. The present invention adopts a multi-element quantitative analysis model based on deep learning. The multi-element quantitative analysis model deeply mines the potential laws in the spectral data through a multi-level neural network structure, optimizes the modeling ability of the interaction between elements, can effectively reduce the noise interference in the laser-induced spectrum, and improve the detection accuracy of low-content elements, which is equivalent to the traditional method. The present invention greatly improves the accuracy of prediction in the quantitative analysis of low-concentration elements. In addition, the present invention realizes the efficient extraction and fusion of multi-scale features through a multi-branch parallel structure and a feature fusion mechanism, significantly improves the network's ability to express complex features, and reduces the computational complexity.
[0040] 2. In the preferred embodiment, the present invention can efficiently extract multi-scale information from spectral data through a multi-level feature extraction module (downsampling, AFBlock, multi-scale fusion) and a global feature integration layer, and is suitable for a variety of spectral analysis tasks. The above network structure design is modular and highly scalable, and the network depth or number of channels can be adjusted according to actual needs to adapt to different application scenarios.
[0041] In addition to the above-described purposes, features and advantages, the present invention has other purposes, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0043] Figure 1 A structural block diagram of DeepLIBSNet provided by an embodiment of the present invention;
[0044] Figure 2 AFBlock structure diagram provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0045] The embodiments of the present invention are described in detail below with reference to the accompanying drawings, but the present invention can be implemented in many different ways as defined and covered by the claims.
[0046] With the widespread application of laser-induced breakdown spectroscopy (LIBS) technology, how to improve the accuracy and stability of LIBS technology in multi-element quantitative analysis, especially in the analysis of low-content elements, has become an important issue in current research. Although the existing LIBS analysis method can quickly measure the elements in the sample, the traditional analysis method has poor accuracy and robustness in the detection of low-content elements due to the large background noise and weak signal intensity of the laser-induced spectral signal in the measurement of low-content elements. In addition, although traditional machine learning models, such as support vector machines (SVM), random forests (RF), linear regression (LR), XGBoost, etc., can model and analyze LIBS spectral data to a certain extent, they still have certain limitations when dealing with high-dimensional data, complex nonlinear relationships, and sample imbalance. It is difficult to fully explore the potential laws in the spectral data, resulting in limited accuracy and reliability of the analysis results.
[0047] Therefore, the technical problem of the present invention is how to improve the accuracy, stability and robustness of multi-element quantitative analysis in laser induced spectroscopy analysis through deep learning optimization methods, especially the problem of high-precision detection of low-content elements.
[0048] In order to solve the above technical problems, the present invention provides a multi-element quantitative analysis method of laser induced breakdown spectroscopy, the method comprising:
[0049] 1. Construction of a multi-element quantitative analysis model for laser-induced breakdown spectroscopy (hereinafter referred to as DeepLIBSNet)
[0050] In order to improve the accuracy, stability and robustness of multi-element quantitative analysis in laser induced spectroscopy, this paper proposes a spectroscopy analysis network structure based on deep learning, called DeepLIBSNet. Through a specific modular design, the network gradually extracts the multi-scale features of spectral data and finally completes the classification or regression task.
[0051] Among them, Figure 1 As shown, the multi-element quantitative analysis model (DeepLIBSNet) includes an input layer, a feature extraction layer, and an output layer connected in sequence.
[0052] 1. Input layer
[0053] Specifically, the input layer is used to extract local correlation features in spectral data and obtain preliminary spatial information; preferably, Conv , step length The convolutional layer;
[0054] 2. Feature extraction layer
[0055] Specifically, the feature extraction layer includes at least one multi-level feature extraction module; the multi-level feature extraction module may include one or more, and the network depth is determined by specific task requirements;
[0056] In a preferred embodiment, the feature extraction layer includes a plurality of multi-level feature extraction modules, and the plurality of multi-level feature extraction modules are connected in series in sequence;
[0057] Specifically, the multi-level feature extraction module includes a downsampling unit, an attention fusion unit, and a feature enhancement unit; the downsampling unit, the attention fusion unit, and the feature enhancement unit are connected in series in sequence;
[0058] 11. Downsampling unit
[0059] The downsampling unit is used to downsample the output features of the previous layer, shorten the spectrum length, extract more global features, and reduce the computational complexity;
[0060] In a preferred embodiment, the downsampling unit is preferably Conv , step length The convolutional block of
[0061] 12. Attention Fusion Block (AFBlock)
[0062] The attention fusion unit, through a multi-branch parallel structure and feature fusion mechanism, achieves efficient extraction and fusion of multi-scale features, significantly improving the network's ability to express complex features while reducing computational complexity.
[0063] Specifically, the attention fusion unit includes multiple parallel feature extraction branches and a feature splicing subunit, the multiple feature extraction branches are respectively used to perform convolution operations of different sizes to extract feature maps of different scales, the multiple feature extraction branches are sorted according to the convolution depth, and each feature extraction branch is jump-connected to other feature extraction branches with adjacent sorting positions and greater convolution depth than it; the feature splicing subunit is used to splice the feature maps output by the multiple feature extraction branches;
[0064] In a preferred embodiment, Figure 2As shown, the attention fusion unit includes a dimensionality reduction subunit, a small-scale feature extraction branch, a medium-scale feature extraction branch, a high-scale feature extraction branch, and a feature splicing subunit;
[0065] The dimension reduction subunit is used to perform a dimension reduction operation on the input feature map to reduce the computational complexity while retaining key information. In this embodiment, the downsampling unit is preferably Convolutional layers;
[0066] The small-scale feature extraction branch includes a first convolution, a second convolution, and a feature fusion block. The output end of the first convolution is connected to the input end of the second convolution and the feature fusion block respectively, and the output end of the second convolution is connected to the input end of the feature fusion block; the small-scale feature extraction branch is used to capture small-scale features, and fuses the results of the first convolution and the second convolution through an element-by-element addition operation (ADD) to enhance the local feature expression capability.
[0067] The mesoscale feature extraction branch includes a third convolution, a fourth convolution, and a first feature splicing block, the output end of the third convolution is connected to the input end of the fourth convolution and the first feature splicing block respectively, the output end of the fourth convolution is connected to the input end of the first feature splicing block; the output end of the second convolution is also jump-connected to the input end of the first feature splicing block;
[0068] The medium-scale feature extraction branch is used to extract medium-scale features, and fuse the third convolution, the fourth convolution results and the jump connection results through a concatenation operation (Concat) to retain the richness of multi-scale features;
[0069] The high-scale feature extraction branch includes a fifth convolution, a sixth convolution, and a second feature splicing block. The output end of the fifth convolution is respectively connected to the input ends of the sixth convolution and the second feature splicing block, and the output end of the fifth convolution is connected to the input end of the second feature splicing block; the output end of the sixth convolution is also jump-connected to the input end of the second feature splicing block.
[0070] The high-scale feature extraction branch is used to capture large-scale features; the convolution results are also fused through a concatenation operation (Concat).
[0071] In a preferred embodiment, the convolution scale of the first convolution is larger than the convolution scale of the second convolution, the convolution scale of the third convolution is larger than the convolution scale of the fourth convolution, and the convolution scale of the fifth convolution is larger than the convolution scale of the sixth convolution; the convolution scale of the first convolution is smaller than the convolution scale of the third convolution, and the convolution scale of the third convolution is smaller than the convolution scale of the fifth convolution; the convolution scale of the second convolution is smaller than the convolution scale of the fourth convolution, and the convolution scale of the fourth convolution is smaller than the convolution scale of the sixth convolution;
[0072] In a preferred embodiment, the first convolution is preferably Convolution, the second convolution is preferably Convolution, the third convolution is preferably Convolution, the fourth convolution is preferably Convolution, the fifth convolution is preferably Convolution, the sixth convolution is preferably Convolution; the feature fusion block uses element-by-element addition operation (ADD) to fuse the results of two convolutions; the first feature splicing block and the second feature splicing block both use splicing operation (Concat) to fuse the convolution results;
[0073] In a preferred embodiment, the attention fusion unit further includes a plurality of feature refining blocks, the plurality of feature refining blocks correspond one to one to the plurality of feature extraction branches, and the output end of each feature extraction branch is connected to the input end of the feature splicing subunit through its corresponding feature refining block;
[0074] Specifically, it includes three feature refining blocks, which are divided into a first feature refining block, a second feature refining block, and a third feature refining block; the input end of the first feature refining block is connected to the output end of the feature fusion block, the input end of the second feature refining block is connected to the output end of the first feature splicing block; the input end of the third feature refining block is connected to the output end of the second feature splicing block;
[0075] The feature refining block is used to refine the input feature map, further compress redundant information and enhance the feature expression capability; in this embodiment, the feature refining block is preferably Convolutional layers;
[0076] The feature splicing subunit includes a third feature splicing block and a dimension reduction block, the output end of each feature extraction branch is connected to the input end of the third feature splicing block through its corresponding feature refinement block, and the output block of the third feature splicing block is connected to the input end of the dimension reduction block;
[0077] The third feature splicing block is used to integrate the output results of the small-scale feature extraction branch, the medium-scale feature extraction branch, and the high-scale feature extraction branch to form a unified multi-scale feature representation;
[0078] In a preferred embodiment, the third feature concatenation block uses a concatenation operation (Concat) to integrate the above branch output results.
[0079] The dimensionality reduction is fast through The design of the convolutional layer improves the feature representation capability while effectively controlling the computational complexity of the module, making it suitable for scenarios with limited resources.
[0080] The AFBlock unit has the following advantages:
[0081] 1. Multi-scale feature extraction capability: Through multi-branch structure design, convolution kernels of different sizes are used (such as , , etc.), significantly improving the ability to capture multi-scale features.
[0082] 2. Efficient feature fusion mechanism: Through element-by-element addition and channel concatenation operations, it can efficiently integrate features from different branches while maintaining the integrity of the features.
[0083] 3. The amount of calculation is controllable: through the dimension reduction layer and The design of the convolutional layer improves the feature representation capability while effectively controlling the computational complexity of the module, making it suitable for scenarios with limited resources.
[0084] 13. Feature Enhancement Unit
[0085] The output of the AFBlock unit is locally enhanced. In this embodiment, the feature enhancement unit preferably adopts The convolution kernel is set to 1.
[0086] (III) Output layer
[0087] The output layer includes a global feature integration module and a fully connected module, and the global feature integration module and the fully connected module are connected in series in sequence.
[0088] The global feature integration module compresses the channel features of each position through the convolution kernel, while retaining the most important information to generate the final global feature representation. In this embodiment, the global feature integration module is preferably Conv , step length convolution.
[0089] The fully connected module is used to output a probability distribution of the number of categories, which is applicable to spectral classification problems. The fully connected module is preferably a Dense fully connected layer.
[0090] (IV) Workflow:
[0091] Step 1: Input the spectral data to be analyzed into the network. The shape of the input data is ,in is the spectrum length, The number of input channels is usually 1, indicating a single-channel spectral signal.
[0092] Step 2: Use The convolution kernel with a step size of 1 performs the first step of feature extraction on the input spectral data, extracts the local correlation features in the spectral data, obtains preliminary spatial information, and outputs the original length with a shape of The feature map of is the number of intermediate channels.
[0093] Step 3: Repeat the following steps as needed:
[0094] Step 31: Use , the convolution kernel with a step size of 2 downsamples the output features of the previous layer, shortening the spectrum length to 1 / 2 of the original, extracting more global features while reducing the computational complexity; the shape of the feature extracted in this step is ,in is the level of the current module;
[0095] Step 32: On the downsampled features, use the AFBlock unit to perform multi-scale feature extraction and fusion, including:
[0096] Dimensionality reduction steps:
[0097] Assume the input feature map size is ,in , and Respectively represent the height, width and number of channels of the feature map; Convolution is used to reduce the dimension, and the output feature map size becomes ,in Much smaller than , to reduce computational complexity while retaining key information.
[0098] Multi-level feature extraction steps:
[0099] The reduced feature map is divided into three parallel branches, each of which performs different convolution operations to extract multi-scale features:
[0100] First branch: pass through Convolution and Convolution is used to capture small-scale features. Subsequently, the two convolution results are fused through element-by-element addition (ADD) to enhance the local feature expression capability.
[0101] Second branch: pass through Convolution and Convolution is used to extract medium-scale features. The two convolution results are fused through concatenation to preserve the richness of multi-scale features.
[0102] The third branch: through Convolution and Convolution is used to capture large-scale features. The convolution results are also fused through concatenation.
[0103] Feature refinement steps:
[0104] The fusion results of each branch are separately The convolutional layer performs feature refinement, further compresses redundant information and enhances feature expression capabilities.
[0105] Specifically: The output size of the concatenation or addition operation of each branch is ,go through After convolution, the output size remains .
[0106] Global feature fusion steps:
[0107] The refined features of the above three branches are integrated through concatenation operation (Concat) to form a unified multi-scale feature representation with a size of Then, through a Convolutional layer, which reduces the fused features to the final output size , to adapt to the input requirements of subsequent networks.
[0108] Step 33: Adoption The convolution kernel with a step size of 1 is used to enhance the local features of the output of the AFBlock module, refine the multi-scale features of the output, and improve the feature expression ability of the network. The final output feature shape is feature map.
[0109] Step 4: Input the final output of the multi-stage module into Convolutional layer, used for global feature integration.
[0110] The convolution kernel is used to compress the channel features at each position while retaining the most important information to generate the final global feature representation.
[0111] Step 5: Input the global features into the fully connected layer to generate the final prediction results: the output of the fully connected layer is the probability distribution of the number of categories, which is suitable for spectral classification problems; the output of the fully connected layer is a continuous value, which is used for tasks such as material composition prediction.
[0112] 2. Constructing a training data set and using the training data set to train the multi-element quantitative analysis model
[0113] 1.1 Loss Function Construction
[0114] The present invention provides an adaptive weighted mean square error loss function, which introduces a dynamic weight adjustment mechanism based on element content and prediction error, so that low-content elements obtain higher weights in the loss calculation, thereby significantly improving the prediction accuracy of low-content elements while maintaining the overall regression performance.
[0115] The specific calculation formula of the adaptive weighted mean square error loss function designed by the present invention is as follows:
[0116] ;
[0117] in:
[0118] : loss function value;
[0119] : total number of samples;
[0120] : The number of predicted elements;
[0121] :sample The true value of The content of each element;
[0122] :sample The predicted value of The content of each element;
[0123] : Adaptive weights, dynamically adjusted for low-content elements and large error values.
[0124] Adaptive Weight The specific calculation formula is:
[0125] ;
[0126] in, is an inverse weight term, which is used to increase the weight of low-content elements. is a smoothing constant to prevent the denominator from being zero; is the error weight term, and the weight is dynamically increased as the prediction error increases; It is a weight adjustment parameter used to balance the influence of the inverse weight term and the error weight term.
[0127] The steps to use the training function are as follows:
[0128] 1. Initialize model parameters and input training data, including true values and predicted values .
[0129] 2. For each sample and elements , calculate the adaptive weight ,in:
[0130] When the element content is low, Increase, improve the weight of low-content elements;
[0131] When the prediction error is large, Increase, dynamically increase the weight of difficult-to-predict samples.
[0132] 3. Based on adaptive weight , calculate the weighted mean square error .
[0133] 4. Update model parameters using gradient descent or other optimization methods.
[0134] 5. Repeat steps 2-4 until the loss function converges or meets the preset conditions.
[0135] This training function has the following advantages:
[0136] 1. Dynamicity: The weights are adjusted dynamically with the element content and prediction error, avoiding the limitations of the fixed weight method.
[0137] 2. Low-content optimization: Assign higher weights to low-content elements to significantly improve their prediction accuracy.
[0138] 3. Overall performance is taken into consideration: while optimizing low-content elements, the prediction accuracy of high-content elements is ensured.
[0139] (II) Constructing a training sample set
[0140] (1) Sample collection:
[0141] Obtain LIBS experimental data, including the spectral characteristics of each sample (input data) and the target element content (true label).
[0142] Input data: spectral feature matrix ,in Indicates Spectral characteristics of the samples.
[0143] Output data: element content matrix ,in Indicates Samples The actual content of the element.
[0144] (2) Data preprocessing
[0145] Data preprocessing is an important part of DeepLIBSNet model training. Its purpose is to process the raw data into a format suitable for model training, while improving the model's adaptability to data distribution and its ability to predict low-content elements. Data preprocessing includes the following specific steps:
[0146] a) Spectral feature normalization
[0147] Purpose: Since the numerical ranges of different spectral features may vary greatly, direct input may lead to numerical instability during model training and even affect the convergence of the model. Normalization can unify the numerical range of spectral features into a fixed interval, thereby improving the stability of training and the prediction ability of the model.
[0148] Method: For each spectral feature For normalization, commonly used methods include:
[0149] Minimum-maximum normalization: Scale the data to the [0,1] interval. The formula is as follows:
[0150]
[0151] Standardization: Adjust the data to zero mean and unit variance. The formula is as follows:
[0152]
[0153] in, is the mean of the spectral features, is the standard deviation.
[0154] Implementation method: Normalize the spectral feature matrix of each sample column by column to ensure that the normalized data distribution is suitable for the model input.
[0155] b) Low-content element label smoothing
[0156] Purpose: The measured values of low-content elements are usually noisy, which makes it difficult for the model to accurately fit the true value of low-content elements during training. Smoothing can reduce the impact of noise and improve the prediction accuracy of the model for low-content elements.
[0157] Method: The following smoothing techniques are used to process the true values of low-content elements , i.e. training labels:
[0158] Gaussian smoothing: Apply a Gaussian kernel function to the labeled data to reduce the impact of outliers on training.
[0159]
[0160] in, is the smoothing parameter, Indicates the nearby sample index.
[0161] Threshold cutoff: Set the noise filtering threshold for low-content elements and set the noise part smaller than the noise filtering threshold to zero.
[0162] Weighted average smoothing: Take the weighted average of the label values of adjacent data points and give higher weights to the center point.
[0163] Implementation method: For the target element content of each sample, smoothing is performed element by element, and a smoothed label matrix is generated.
[0164] (3) Data partitioning
[0165] The purpose of data partitioning is to allocate the original data set into training sets, validation sets, and test sets according to different task requirements to ensure the scientificity and fairness of model training, performance evaluation, and final testing. The partitioning includes the following specific steps:
[0166] 1. Division ratio setting
[0167] The entire data set is divided into training set, validation set and test set in the ratio of 8:1:1.
[0168] Training set: accounts for 80% of the data set and is used for learning model parameters.
[0169] Validation set: 10% of the dataset, used for tuning model hyperparameters and detecting overfitting.
[0170] Test set: 10% of the dataset, used to evaluate the performance of the final model.
[0171] Importance: Reasonable ratio division can ensure that there is enough data for learning during training, and also provide independent test samples for model performance evaluation.
[0172] 2. Division method
[0173] Random partitioning: After randomly shuffling the data set samples, they are allocated to the training set, validation set, and test set according to the partition ratio. Random partitioning can avoid deviations caused by uneven data distribution.
[0174] Stratification: If the content distribution of certain elements in the sample is obviously unbalanced (for example, there are fewer samples of low-content elements), a stratification method is required. That is:
[0175] The samples are divided into several categories according to the content range of the target elements;
[0176] Data is extracted proportionally from each category and distributed into training sets, validation sets, and test sets to ensure that the category distribution in each set is consistent with the original dataset.
[0177] 3. Data Validation
[0178] Make sure that the divided data sets do not overlap with each other, that is, there is no duplication of samples between the training set, validation set, and test set.
[0179] Check the sample distribution in each set to verify whether it meets the requirements of stratified sampling.
[0180] Through the above data preprocessing and data partitioning steps, the training efficiency and generalization ability of the DeepLIBSNet model can be effectively improved, laying a solid foundation for subsequent model training and performance evaluation.
[0181] (4) Model initialization steps
[0182] 1. Build the DeepLIBSNet model structure
[0183] In the model initialization stage, we first need to build the DeepLIBSNet model structure to ensure that the network's hierarchical relationship and functional modules are correctly connected. The specific steps are as follows:
[0184] 1.1 Input layer:
[0185] The function of the input layer is to receive spectral feature data. The dimension of the input data is ,in represents the number of samples, Represents the number of spectral features.
[0186] The input layer directly inputs the spectral features to the next convolution layer.
[0187] 1.2. Convolutional layer:
[0188] The first convolutional layer uses Convolution operation, stride ,The role of this layer is to extract local spectral features and generate feature maps.
[0189] The next several times It is superimposed with the AFBlock module to capture deeper spectral features.
[0190] The role of is to refine feature extraction and provide higher resolution input data for subsequent modules.
[0191] The AFBlock module further enhances the feature recognition capability of low-content elements through an adaptive feature extraction mechanism.
[0192] 1.3. Fully connected layer:
[0193] After the convolutional layer, The layer further compresses the feature map and connects the fully connected layer to integrate all the extracted deep features.
[0194] 1.4. Output layer:
[0195] The output layer is a fully connected layer with the same output dimension as the number of target elements. Each value in the output represents the predicted concentration of each element.
[0196] 2. Initialize network parameters
[0197] Network parameter initialization plays a key role in the training stability and convergence speed of the model. The parameter initialization of this model includes the following:
[0198] 2.1 Convolution kernel parameter initialization:
[0199] The convolution kernel parameters of all convolutional layers are initialized to randomly distributed values, using a better initialization method, such as Xavier initialization or He initialization, to avoid the gradient vanishing or exploding problem of deep networks.
[0200] 2.2 Weight adjustment parameter initialization:
[0201] Weight adjustment parameters involved in the loss function and the smoothing constant Initialize to default values. For example:
[0202] : Used to control the weight of low-content elements in the loss function.
[0203] : Used to avoid zero denominator in numerical calculations.
[0204] 2.3 Setting the Optimizer
[0205] The choice of optimizer directly affects the training efficiency and performance of the model. This model uses the Adam optimizer, which combines the advantages of momentum and adaptive learning rate. The specific configuration is as follows:
[0206] Optimizer selection: Adam optimizer.
[0207] Initial learning rate: set to 0.001 to ensure stability at the beginning of training.
[0208] Optimizer parameter adjustment: The learning rate can be adjusted dynamically according to the training results, and the learning rate decay strategy stepDecay is adopted.
[0209] 2.4 Configuring the Loss Function
[0210] The loss function is used to measure the error between the model prediction value and the true value. This model uses the adaptive weighted mean square error loss function (AdaptiveWeightedMSELoss), which is designed to improve the prediction accuracy of low-content elements. The formula of the loss function is as follows:
[0211] ;
[0212] in, is the loss function value; is the total number of samples; is the number of predicted elements; For sample The true value of The content of each element; For sample The predicted value of The content of each element; The weights are adaptive and dynamically adjusted for low-content elements and large error values.
[0213] Among them, the adaptive weight for:
[0214] ;
[0215] in, is an inverse weight term, which is used to increase the weight of low-content elements. is a smoothing constant to prevent the denominator from being zero; is the error weight term, and the weight is dynamically increased as the prediction error increases; It is a weight adjustment parameter used to balance the influence of the inverse weight term and the error weight term.
[0216] By setting the above loss function, the model can pay more attention to the prediction errors of low-content elements, thereby improving the prediction accuracy of these elements.
[0217] (5) Model training
[0218] The following is a detailed description of the training steps of the DeepLIBSNet model:
[0219] 1 Training data input
[0220] The spectral features of the preprocessed training set samples Input into the DeepLIBSNet model, and after multi-layer feature extraction and nonlinear transformation of the network, the corresponding prediction value is generated. .
[0221] The dimension of the input data is ,in represents the number of training samples, Represents the dimension of the spectral features.
[0222] The dimensions of the output prediction value are ,in is the number of target element types.
[0223] 2. Loss calculation
[0224] The adaptive weighted mean square error loss function (AdaptiveWeightedMSELoss) is used to calculate the loss value of the training set.
[0225] 3. Back Propagation
[0226] Based on loss value , execute the backpropagation algorithm, calculate the gradients of each layer of the network, and update the model parameters.
[0227] Parameter updates use an optimizer (such as the Adam optimizer) to iteratively adjust the gradient and momentum information.
[0228] After each update, the network's weight parameters and bias parameters are gradually optimized in the direction of reducing the loss value.
[0229] 4. Learning Rate Adjustment
[0230] In order to improve training efficiency and avoid the model falling into local optimality in the later stage of training, the learning rate is adjusted dynamically.
[0231] Use a learning rate decay strategy, for example:
[0232] Cosine Annealing: Periodically reduce the learning rate during training to improve model performance.
[0233] Exponential Decay: The learning rate decreases exponentially, and the formula is:
[0234] ;
[0235] in is the initial learning rate, is the attenuation factor, is the current training round number.
[0236] 5. Verify Model Performance
[0237] After each training cycle, use the validation set and the true label Evaluate model performance:
[0238] Input validation set samples and generate predicted values .
[0239] Calculate validation set loss , monitor the changing trend of validation loss.
[0240] If the validation loss starts to increase (early stopping condition), it may be a sign that the model is overfitting.
[0241] 6. Repeat the training
[0242] Repeat the above steps and continue to iterate the training model until one of the following termination conditions is met:
[0243] Condition 1: Validation loss no longer decreases significantly:
[0244] If the loss value of the validation set does not decrease significantly over a number of consecutive epochs (e.g., 5 epochs), training is stopped to prevent overfitting.
[0245] Condition 2: Reach the preset maximum number of training rounds:
[0246] Set the maximum number of training rounds, for example 100 rounds, and terminate the training after reaching this number of rounds.
[0247] Through the above training steps, it is ensured that the DeepLIBSNet model is fully learned on the training set while maintaining the generalization ability on the validation set, providing a reliable basis for the final evaluation of the test set.
[0248] (6) Model testing
[0249] 1 Input test data
[0250] The test set samples Input the trained DeepLIBSNet model, and generate the corresponding prediction value through the network forward propagation. .
[0251] The dimensions of the test data are consistent with the training data, ensuring that the model can process it seamlessly.
[0252] Prediction value output dimension and number of target element types Be consistent.
[0253] 2 Performance Evaluation
[0254] The following performance indicators are used to conduct a comprehensive performance evaluation of the prediction results of the test set:
[0255] Overall performance indicators:
[0256] Mean Squared Error (MSE):
[0257] ;
[0258] A measure of the overall deviation between the predicted values and the true values.
[0259] Mean Absolute Error (MAE):
[0260] ;
[0261] Reflects the average absolute magnitude of the model prediction error.
[0262] Low-level element assessment:
[0263] For low-content elements, MSE and MAE were calculated separately to analyze the prediction performance of the model in the low concentration range.
[0264] Statistics of accuracy, deviation range and limit of detection (LOD) of low-content elements.
[0265] Verify the actual effect of the adaptive weighted mean square error loss function on the optimization of low-content elements.
[0266] Through overall and local performance evaluation, the model is ensured to have good generalization ability and significantly improve the prediction accuracy of low-content elements.
[0267] Below is an example table showing the comparison results of the DeepLIBSNet model with traditional neural networks, support vector machines (SVM), random forests (RF) and other methods in spectral analysis tasks. It is assumed that the evaluation indicators include mean square error (MSE), mean absolute error (MAE) and accuracy for low concentration elements (ALCE).
[0268] Performance comparison of DeepLIBSNet model with other models
[0269]
[0270] As can be seen from the above, the method of the present invention has the following advantages:
[0271] Improve the detection accuracy of low-content elements: By adopting the DeepLIBSNet model based on deep learning, the noise interference in the laser-induced spectrum can be effectively reduced, and the detection accuracy of low-content elements can be improved, especially in the quantitative analysis of low-concentration elements, which shows a significant improvement over traditional methods.
[0272] Optimizing the nonlinear characteristics of spectral data: Compared with traditional machine learning methods, DeepLIBSNet can better handle the complex nonlinear relationships in spectral data. Through a multi-level neural network structure, it deeply mines the potential laws in spectral data and optimizes the modeling ability of interactions between elements, thereby improving the accuracy and reliability of multi-element quantitative analysis.
[0273] Enhanced robustness of the analysis process: This invention combines the powerful feature extraction and learning capabilities of deep neural networks to cope with problems such as sample imbalance and data missing, enhances the robustness of the model in practical applications, and adapts to the analysis needs under different samples and experimental conditions.
[0274] Automation and efficiency: After adopting the deep learning model, the analysis process does not require human intervention and can automatically extract features from large amounts of spectral data, reducing the complexity of manual feature extraction and improving the efficiency and accuracy of data processing.
[0275] Wide applicability: The method of the present invention is applicable to various types of laser induced spectroscopy (LIBS) data and can be flexibly applied in a variety of industrial and scientific research fields, and is particularly suitable for complex element analysis in the fields of environmental monitoring, material science, metallurgical engineering, etc.
[0276] In summary, the present invention has successfully solved the problems of insufficient accuracy in low-content element detection and insufficient processing capabilities of traditional machine learning methods for complex spectral data by combining deep learning optimization technology. It has significant technical advantages and broad application prospects.
[0277] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-element quantitative analysis method of laser induced breakdown spectroscopy, characterized in that: The following steps are involved: Constructing a multi-element quantitative analysis model of laser induced breakdown spectroscopy, the multi-element quantitative analysis model comprises an input layer, a feature extraction layer and an output layer connected in sequence, the feature extraction layer comprises at least one multi-level feature extraction module, the multi-level feature extraction module comprises an attention fusion unit, the attention fusion unit comprises a plurality of parallel feature extraction branches and a feature splicing sub-unit, the plurality of feature extraction branches are respectively used to perform convolution operations of different sizes and dimensions to extract feature maps of different scales, the plurality of feature extraction branches are sorted according to the convolution depth, and each feature extraction branch is jump-connected to other feature extraction branches with adjacent sorting positions and a greater convolution depth than the branch; The feature splicing subunit is used to splice the feature graphs output by multiple feature extraction branches; Acquire spectral characteristic data of the laser induced breakdown spectrum from historical data to construct a training data set, use the training data set to train the multi-element quantitative analysis model, and use the multi-element quantitative analysis model to predict the element content of the target laser induced breakdown spectrum; The attention fusion unit includes a small-scale feature extraction branch, a medium-scale feature extraction branch, and a high-scale feature extraction branch; The small-scale feature extraction branch includes a first convolution, a second convolution, and a feature fusion block, the output end of the first convolution is connected to the input end of the second convolution and the feature fusion block respectively, and the output end of the second convolution is connected to the input end of the feature fusion block; The mesoscale feature extraction branch includes a third convolution, a fourth convolution, and a first feature splicing block, the output end of the third convolution is connected to the input end of the fourth convolution and the first feature splicing block respectively, the output end of the fourth convolution is connected to the input end of the first feature splicing block; the output end of the second convolution is also jump-connected to the input end of the first feature splicing block; The high-scale feature extraction branch includes a fifth convolution, a sixth convolution, and a second feature splicing block, the output end of the fifth convolution is connected to the input end of the sixth convolution and the second feature splicing block respectively, the output end of the fifth convolution is connected to the input end of the second feature splicing block; the output end of the sixth convolution is also jump-connected to the input end of the second feature splicing block; The convolution scale of the first convolution is larger than the convolution scale of the second convolution, the convolution scale of the third convolution is larger than the convolution scale of the fourth convolution, and the convolution scale of the fifth convolution is larger than the convolution scale of the sixth convolution; the convolution scale of the first convolution is smaller than the convolution scale of the third convolution, and the convolution scale of the third convolution is smaller than the convolution scale of the fifth convolution; the convolution scale of the second convolution is smaller than the convolution scale of the fourth convolution, and the convolution scale of the fourth convolution is smaller than the convolution scale of the sixth convolution; and / or The attention fusion unit also includes multiple feature refining blocks, which correspond one-to-one to multiple feature extraction branches. The output end of each feature extraction branch is connected to the input end of the feature splicing sub-unit through its corresponding feature refining block.
2. The multi-element quantitative analysis method of laser induced breakdown spectroscopy according to claim 1, characterized in that: The feature splicing subunit includes a third feature splicing block and a dimension reduction block, the output end of each feature extraction branch is connected to the input end of the third feature splicing block through its corresponding feature refinement block, and the output block of the third feature splicing block is connected to the input end of the dimension reduction block; and / or The attention fusion unit includes a dimensionality reduction subunit, and the output end of the dimensionality reduction subunit is respectively connected to the input ends of the small-scale feature extraction branch, the medium-scale feature extraction branch, and the high-scale feature extraction branch.
3. The multi-element quantitative analysis method of laser induced breakdown spectroscopy according to claim 1, characterized in that: The multi-level feature extraction module also includes: a downsampling unit and a feature enhancement unit, and the downsampling unit, the attention fusion unit, and the feature enhancement unit are connected in series in sequence.
4. The multi-element quantitative analysis method of laser induced breakdown spectroscopy according to claim 1, characterized in that: The multi-level feature extraction module includes a plurality of multi-level feature extraction modules, and the plurality of multi-level feature extraction modules are connected in series in sequence to form a feature extraction layer; and / or The output layer includes a global feature integration module and a fully connected module, and the global feature integration module and the fully connected module are connected in series in sequence.
5. The multi-element quantitative analysis method of laser induced breakdown spectroscopy according to claim 4, characterized in that: The loss function of the multi-element quantitative analysis model satisfies: ; in, is the loss function value; is the total number of samples; is the number of predicted elements; For sample The true value of The content of each element; For sample The predicted value of The content of each element; The weights are adaptive and dynamically adjusted for low-content elements and large error values.
6. The multi-element quantitative analysis method of laser induced breakdown spectroscopy according to claim 5, characterized in that: The adaptive weights satisfy: ; in, is an inverse weight term, which is used to increase the weight of low-content elements. is a smoothing constant to prevent the denominator from being zero; is the error weight term, and the weight is dynamically increased as the prediction error increases; It is a weight adjustment parameter used to balance the influence of the inverse weight term and the error weight term.
7. The multi-element quantitative analysis method of laser induced breakdown spectroscopy according to claim 1, characterized in that: The spectral feature data of laser-induced breakdown spectroscopy are obtained from historical data to construct a training dataset, including: Constructing a spectral feature matrix based on the acquired spectral feature data, obtaining the actual element content corresponding to the spectral feature data and constructing an element content matrix corresponding to the spectral feature matrix; The spectral feature matrix is normalized, and a smoothing filter algorithm is used to remove noise from low-content elements in the element content matrix. The smoothing filter algorithm includes one or a combination of the following: Gaussian smoothing algorithm, threshold truncation algorithm, weighted average smoothing algorithm; The element contents in the denoised element content matrix are used as labels of the corresponding spectral features in the spectral feature matrix.
8. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Non-vision-field target detection method and device and storage medium
CN113204010A
Image restoration method based on multi-feature fusion network
CN113362242A