Libs spectral grade prediction method based on multi-head attention

CN121958991BActive Publication Date: 2026-08-07CHANGSHA RES INST OF MINING & METALLURGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHANGSHA RES INST OF MINING & METALLURGY CO LTD
Filing Date
2026-04-02
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0007]本发明提供了基于多头注意力的LIBS光谱品位预测方法,以解决现有的计算模型难以基于LIBS光谱数据进行品位预测的问题

Benefits of technology

[0040] The LIBS spectral grade prediction method provided by this invention combines multi-scale convolution with a dual-layer attention mechanism (intra- and inter-spectral bands). The multi-scale convolution module enhances the ability to identify local peak shapes, weak peaks, and overlapping spectral peaks, enabling the model to extract more refined local structural information of the spectrum. The dual-layer attention mechanism (intra- and inter-spectral bands) can model fine-grained dependencies within a spectral band and capture long-distance correlations across spectral bands, thus significantly improving the ore grade prediction performance based on LIBS spectral data while taking into account both local details and global trends. Compared with traditional methods and conventional models, this invention shows significant improvements in prediction accuracy, fitting ability, and stability, and can maintain low prediction errors and small error fluctuations under complex spectral conditions, thereby meeting the needs of on-site online detection and process control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958991B_ABST
    Figure CN121958991B_ABST
Patent Text Reader

Abstract

The present application relates to the technical fields of spectral analysis and ore composition detection, and discloses a LIBS spectral grade prediction method based on multi-head attention. A multi-scale extraction module is constructed based on one-dimensional convolution with different convolution kernel sizes, a feature extraction module is constructed based on rotation position coding and multi-head attention, a spectral grade prediction model is constructed based on the multi-scale extraction module, the feature extraction module, residual connection and multi-layer perception mechanism; spectral data of an ore composition detection site is acquired, spectral features are obtained by preprocessing the spectral data and input into the spectral grade detection model, the multi-scale extraction module extracts local multi-scale features based on the spectral features, the feature extraction module obtains global features from the local multi-scale features, fusion features are obtained based on the global features, the local multi-scale features and the residual connection, and ore grade prediction results are obtained based on the fusion features and the multi-layer perception mechanism; the problem that existing calculation models are difficult to perform grade prediction based on spectral data is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spectral analysis and ore composition detection technology, and in particular to an ore grade prediction method based on laser-induced breakdown spectroscopy (LIBS) and its deep learning modeling technology. Background Technology

[0002] Ore grade testing is a crucial step in mine production, mineral processing control, and resource assessment. Traditional ore grade analysis primarily relies on laboratory testing methods such as wet chemical analysis and X-ray fluorescence spectroscopy (XRF). While these methods offer high accuracy, they generally suffer from long testing cycles, complex sample pretreatment, and the inability to achieve online real-time monitoring, making it difficult to meet the demands of modern mines for rapid testing and process control.

[0003] In recent years, laser-induced breakdown spectroscopy (LIBS) has gradually become an important technology for ore composition analysis due to its advantages such as no need for complex sample preparation, fast detection speed, and the ability to perform simultaneous multi-element analysis. LIBS technology uses high-energy laser pulses to excite samples to generate plasma and collects their emission spectra, thereby achieving qualitative and quantitative elemental analysis. However, affected by factors such as the complexity of ore composition, matrix effects, spectral noise, peak overlap, and laser energy fluctuations, LIBS spectral signals often have high dimensionality, strong noise, and nonlinear characteristics, making it difficult for traditional spectroscopic quantitative methods to obtain stable and reliable prediction results.

[0004] In spectral data modeling, traditional methods such as Partial Least Squares Regression (PLS) and Support Vector Regression (SVR) can handle high-dimensional spectral data to some extent, but their modeling capabilities are limited, making it difficult to fully capture complex nonlinear relationships and cross-spectral correlations in the spectrum. With the development of deep learning, models such as Convolutional Neural Networks (1D-CNN), Residual Networks (ResNet-1D), and Transformers are increasingly being applied to spectral analysis tasks. Convolutional models can extract local peak features, but are limited by a fixed receptive field, making it difficult to model long-distance dependencies; while Transformers have global modeling capabilities, they lack fine-grained representation of local peaks and are prone to insufficient feature extraction in noisy spectra.

[0005] Furthermore, existing deep learning models generally lack a collaborative modeling mechanism for local spectral structure and global trends, making it difficult to simultaneously consider peak details, inter-segment correlations, and overall spectral morphology, resulting in significant room for improvement in prediction accuracy and stability. In ore LIBS spectra, local peak shapes, weak peaks, background variations, and inter-segment coupling relationships have a significant impact on grade prediction, but existing methods often fail to effectively integrate this multi-scale, multi-structure information.

[0006] Therefore, there is an urgent need for a spectral modeling method that can simultaneously capture local spectral details and global dependencies, possesses strong robustness and high prediction accuracy, in order to improve the application effect of LIBS in rapid detection and online monitoring of ore grade. Summary of the Invention

[0007] This invention provides a LIBS spectral grade prediction method based on multi-head attention to solve the problem that existing computational models are difficult to use for grade prediction based on LIBS spectral data.

[0008] To achieve the above objectives, the present invention employs the following technical solution:

[0009] This invention provides a LIBS spectral grade prediction method based on multi-head attention, comprising the following steps:

[0010] Step 1: Construct a multi-scale extraction module for extracting local features at different scales based on one-dimensional convolution with different kernel sizes; construct a first spectral feature extraction unit and a second spectral feature extraction unit based on rotational position encoding and multi-head attention; construct a feature extraction module based on the spectral feature extraction unit and the global spectral feature extraction unit; and construct a spectral grade prediction model based on the multi-scale extraction module, the feature extraction module, residual connections, and a multilayer perceptron.

[0011] Step 2: Acquire LIBS spectral data from the ore composition detection site, preprocess the LIBS spectral data to obtain spectral features, input the spectral features into the spectral grade detection model, the multi-scale extraction module extracts local multi-scale features based on the spectral features, the feature extraction module sequentially performs intra-spectral and inter-spectral modeling on the local multi-scale features to obtain global features, and obtain fused features based on the global features, local multi-scale features, and residual connections, and obtain ore grade prediction results based on the fused features and the multilayer perceptron.

[0012] Based on the above design, a spectral grade prediction model was designed, which effectively addresses the core industrial pain points of LIBS spectral detection of ores in industrial sites, such as strong noise interference, complex overlapping spectral lines, incomplete extraction of single-scale features, inability to accurately model intra-spectral line correlation and long-distance dependence across spectral bands, rigid local-global feature fusion, poor generalization to industrial conditions, and low grade prediction accuracy.

[0013] Furthermore, the network structure of the multi-scale extraction module includes: constructing a multi-channel parallel structure based on one-dimensional convolution with different kernel sizes;

[0014] In the multi-channel parallel structure, each convolutional kernel branch corresponds to a local feature at a single scale, and each convolutional kernel branch operates independently without interfering with each other.

[0015] The multi-scale extraction module obtains local features at multiple different scales based on one-dimensional convolution with different kernel sizes, and concatenates the multiple local features along the channel dimension to form local multi-scale features.

[0016] Furthermore, the multi-scale feature extraction module combines the router probability distribution with all feature vectors in the input spectral features to select convolution kernel branches that are adapted to the industrial spectral characteristics, obtains local features at multiple different scales, and concatenates multiple local features along the channel dimension into local multi-scale features.

[0017] Through the above design, multi-scale one-dimensional convolution parallel design is used to extract local peak shapes and weak peaks in the corresponding spectral features of LIBS spectral data. By using convolution kernels of different scales, narrow peak and wide peak information are captured simultaneously and details within the spectral band are preserved. While maintaining computational efficiency, the ability to express the diversity of spectral peak shapes is significantly enhanced.

[0018] Furthermore, in the first spectral feature extraction unit and the second spectral feature extraction unit, the query matrix of the multi-head attention is obtained based on the expert convolution set and the router probability distribution, the key matrix is ​​obtained based on multiple depthwise separable convolutions and residual connections, and the value matrix is ​​obtained based on depthwise separable convolutions and linear projections.

[0019] The expert convolution set includes convolution kernels that are... The vertical convolution and the convolution kernel are The horizontal convolution and the convolution kernel are The standard convolution.

[0020] Furthermore, in both the first and second spectral feature extraction units, the query matrix and key matrix of the multi-head attention are configured with rotational position encoding.

[0021] The above design effectively reduces noise interference in the industrial environment and enhances the effective spectral characteristics related to grade.

[0022] Furthermore, the feature extraction module is constructed by stacking the first spectral feature extraction unit and the second spectral feature extraction unit;

[0023] The first spectral feature extraction unit models the spectral line dependence within the spectral band based on multi-head attention and rotational position encoding to obtain local structural features;

[0024] The second spectral feature extraction unit models the long-distance cross-spectral dependence of local structural features based on multi-head attention and rotational position encoding to obtain global features.

[0025] Through the above design, a two-layer attention architecture within and between spectral bands was designed, and rotational position coding was introduced into the attention calculation to enhance the continuity of spectral position information and the ability to model relative positions. This not only retains the ability to model global dependencies, but also strengthens the perception of local order within spectral bands, making up for the shortcomings of a single attention structure in fine-grained spectral expression.

[0026] Furthermore, the method of obtaining fusion features based on global features, local multi-scale features, and residual connections includes obtaining fusion features based on global features, local multi-scale features, residual connections, and weighted fusion.

[0027] The fusion weights of the weighted fusion are adaptively generated based on global features and local multi-scale features.

[0028] Furthermore, the spectral grade prediction model also includes an adaptive fusion module, which adaptively generates fusion weights based on global features and local multi-scale features, and then fuses the global features and local multi-scale features element by element based on the fusion weights.

[0029] The adaptive fusion module includes a first input branch, a second input branch, and an output branch;

[0030] Both the first input branch and the second input branch are equipped with a lightweight multilayer perceptron and a layer normalization mechanism for generating fusion weights;

[0031] The output branch concatenates and fuses the features output from the first input branch and the second input branch to obtain fused features.

[0032] Furthermore, the objective function of the spectral grade prediction model is constructed based on the basic regression loss, parameter regularization loss, attention entropy regularization loss, spectral smoothing consistency loss, and local-global consistency loss.

[0033] The basic regression loss is constructed based on the deviation between the predicted grade of the spectral grade prediction model and the actual grade measured in the industrial field, combined with the mean square error.

[0034] The parameter regularization loss is constructed based on all learnable parameters of the spectral grade prediction model combined with L2 regularization loss;

[0035] The attention entropy regularization loss is constructed based on the attention weight distribution entropy value of each spectral sampling point position in the multi-head attention mechanism;

[0036] The spectral smoothing consistency loss is constructed based on the eigenvector differences of the fused features between adjacent wavelength positions;

[0037] The local-global consistency loss is constructed based on the difference between local multi-scale features and global features combined with the L2 norm.

[0038] Based on the above design, we propose to jointly optimize the basic regression loss with attention entropy regularization loss, parameter regularization loss, spectral smoothing consistency loss and local-global consistency loss, in order to constrain the attention distribution, maintain the consistency of spectral prediction and suppress the influence of noise.

[0039] Beneficial effects:

[0040] The LIBS spectral grade prediction method provided by this invention combines multi-scale convolution with a dual-layer attention mechanism (intra- and inter-spectral bands). The multi-scale convolution module enhances the ability to identify local peak shapes, weak peaks, and overlapping spectral peaks, enabling the model to extract more refined local structural information of the spectrum. The dual-layer attention mechanism (intra- and inter-spectral bands) can model fine-grained dependencies within a spectral band and capture long-distance correlations across spectral bands, thus significantly improving the ore grade prediction performance based on LIBS spectral data while taking into account both local details and global trends. Compared with traditional methods and conventional models, this invention shows significant improvements in prediction accuracy, fitting ability, and stability, and can maintain low prediction errors and small error fluctuations under complex spectral conditions, thereby meeting the needs of on-site online detection and process control.

[0041] Secondly, the adaptive fusion module effectively integrates local and global features through dynamic weight allocation, which improves the model's generalization ability and robustness under different samples and noise conditions.

[0042] This invention also offers advantages in interpretability and reliability. The attention weights in the attention mechanism and the fusion weights in the adaptive fusion module can be used to indicate the key spectral bands and important features that the model focuses on, thereby providing a basis for the analysis of the physical meaning of spectral features, diagnosis of abnormal samples, and assessment of model reliability. Compared to traditional methods that require frequent manual calibration, this invention reduces reliance on complex preprocessing and human experience, and improves automation and field applicability.

[0043] In summary, this invention has significant effects in improving the accuracy of grade prediction, enhancing model stability and robustness, and improving the interpretability of results. It has good application value and industrialization prospects. Attached Figure Description

[0044] Figure 1 This is a schematic diagram of the network structure of the spectral grade prediction model according to a preferred embodiment of the present invention;

[0045] Wherein, AdaptiveFusion represents the adaptive fusion module; RoPE represents rotational position encoding;

[0046] Figure 2 This is a schematic diagram of the network structure of the adaptive fusion module in a preferred embodiment of the present invention;

[0047] Here, Hadamard represents element-wise multiplication. Detailed Implementation

[0048] The technical solution of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms "an" or "a" and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms "connected" or "linked" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up," "down," "left," "right," etc., are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship also changes accordingly.

[0050] This application provides a LIBS spectral grade prediction method based on multi-head attention, including the following steps:

[0051] Step 1: Construct a multi-scale extraction module for extracting local features at different scales based on one-dimensional convolution with different kernel sizes; construct a first spectral feature extraction unit and a second spectral feature extraction unit based on rotational position encoding and multi-head attention; construct a feature extraction module based on the spectral feature extraction unit and the global spectral feature extraction unit; and construct a spectral grade prediction model based on the multi-scale extraction module, the feature extraction module, residual connections, and a multilayer perceptron.

[0052] Please see Figure 1 The network structure of the multi-scale extraction module includes: a multi-channel parallel structure constructed based on one-dimensional convolution with different kernel sizes;

[0053] In a multi-channel parallel structure, each convolutional kernel branch corresponds to a local feature at a single scale. Each convolutional kernel branch operates independently without interfering with each other. In this embodiment, the multi-channel parallel structure uses a one-dimensional convolution with a kernel of 3 to process small-scale local features, a one-dimensional convolution with a kernel of 5 to process medium-scale local features, and a one-dimensional convolution with a kernel of 9 to process large-scale local features.

[0054] The multi-scale extraction module obtains local features at multiple different scales based on one-dimensional convolution with different kernel sizes, and concatenates multiple local features along the channel dimension to form local multi-scale features.

[0055] The multi-scale feature extraction module combines the router probability distribution with all feature vectors in the input spectral features to select convolutional kernel branches that are suitable for industrial spectral characteristics, obtain local features at multiple different scales, and concatenates multiple local features along the channel dimension into local multi-scale features.

[0056] The design of the multi-scale feature extraction module enables the extraction of information from multiple receptive fields when facing multiple components in the LIBS spectrum, such as sharp spectral peaks, broad peak structures, and background noise. This allows the model to obtain a more comprehensive representation of local feature structures and effectively distinguish different types of spectral line structures, such as narrow peaks and broad peaks, local abrupt changes and smooth trends. This improves the accuracy of subsequent attention mechanisms in modeling spectral line relationships. At the same time, the lightweight one-dimensional convolution also reduces computational complexity, making it more suitable for online monitoring and embedded deployment in industrial settings.

[0057] The feature extraction module includes a first spectral feature extraction unit and a second spectral feature extraction unit. The network structure of the first spectral feature extraction unit and the second spectral feature extraction unit is the same, which is a combination of multi-head attention and rotational position encoding. The rotational position encoding is set in the query matrix and key matrix of the multi-head attention. However, in actual industrial problem applications, the first spectral feature extraction unit is mainly aimed at the spectral line dependence relationship within a local spectral band, capturing the structural correlation and physical coupling features within the local range, while the second spectral feature extraction unit is mainly aimed at the long-distance dependence relationship across spectral bands in the entire spectral range, capturing the physical correlation, energy coupling mode and cooperative change features between different spectral bands.

[0058] Specifically, in the first spectral feature extraction unit and the second spectral feature extraction unit, the query matrix of multi-head attention is obtained based on the expert convolution set and the router probability distribution, the key matrix is ​​obtained based on multiple depthwise separable convolutions and residual connections, and the value matrix is ​​obtained based on depthwise separable convolutions and linear projections.

[0059] The expert convolution set includes convolution kernels as The vertical convolution and the convolution kernel are The horizontal convolution and the convolution kernel are Standard convolution;

[0060] In both the first and second spectral feature extraction units, the query matrix and key matrix of the multi-head attention are set with rotational position encoding;

[0061] The feature extraction module is constructed by stacking the first spectral feature extraction unit and the second spectral feature extraction unit;

[0062] The first spectral feature extraction unit models the spectral line dependence within the spectral band based on multi-head attention and rotational position encoding to obtain local structural features;

[0063] The second spectral feature extraction unit models the long-distance cross-spectral dependencies of local structural features based on multi-head attention and rotational position encoding to obtain global features;

[0064] Based on the different industrial purposes of the first and second spectral feature extraction units, the meaning of the rotational position encoding also differs. In the first spectral feature extraction unit, the rotational position encoding enables the attention mechanism to naturally express the relative wavelength position information of the spectrum, thus allowing the model to not only focus on the numerical characteristics of spectral lines but also understand the relative positional relationship between spectral lines, i.e., the positional relationship of spectral line A relative to spectral line B. In the second spectral feature extraction unit, the rotational position encoding enables the attention mechanism to explicitly express the relative wavelength structure of the spectrum, and to better identify the synergistic changes between spectral lines of different elements, the energy distribution trend across spectral bands, the coupling relationship between distant spectral lines, and the global changes caused by instrument response or plasma state.

[0065] Please see Figure 2 In this embodiment, the spectral grade prediction model also includes an adaptive fusion module to handle the fusion of global features and local multi-scale features. In other embodiments, this can also be handled by residual connection.

[0066] For the adaptive fusion module, fusion weights are adaptively generated based on global features and local multi-scale features, and then the global features and local multi-scale features are fused element by element based on the fusion weights.

[0067] The specific network structure of the adaptive fusion module includes a first input branch, a second input branch, and an output branch. The first input branch is connected to the second spectral feature extraction unit, and the second input branch is connected to the multi-scale feature extraction module.

[0068] Both the first and second input branches are equipped with a lightweight multilayer perceptron and a layer normalization mechanism for generating fusion weights.

[0069] The output branch concatenates and fuses the features output from the first input branch and the second input branch to obtain the fused features;

[0070] The fusion mechanism used in the adaptive fusion module improves the model's generalization ability with its adaptive weights. It also achieves complementary fusion of local and global information for the fused features, eliminating the problem that a single path is difficult to express the spectral structure. The fused features also enhance the robustness to noise and local abnormal data, enabling more robust feature representation.

[0071] Finally, after the feature is spliced ​​and fused by the adaptive fusion module, the fused features are obtained, and then the multilayer perceptron is used to predict the fused features to obtain the final prediction result.

[0072] The network structure of the spectral grade prediction model consists of a multi-scale feature extraction module, a first spectral feature extraction unit, a second spectral feature extraction unit, an adaptive fusion module, and a multilayer perceptron.

[0073] The objective function of the spectral grade prediction model is constructed based on the basic regression loss, parameter regularization loss, attention entropy regularization loss, spectral smoothing consistency loss, and local-global consistency loss, and is expressed by the following formula:

[0074] ;

[0075] in, Represent the objective function; Indicates the basic regression loss; This represents the loss due to parameter regularization. This represents the attention entropy regularization loss; This indicates the loss of spectral smoothing consistency; This represents the loss of local-to-global consistency. This represents the weighting coefficient of each loss term, used to balance the contributions of different losses;

[0076] The basic regression loss is used to measure the deviation between the model's predicted value and the actual grade measured in the industrial field. The basic regression loss is constructed by combining the deviation between the predicted grade of the spectral grade prediction model and the actual grade measured in the industrial field with the mean square error, and is expressed by the following formula:

[0077] ;

[0078] in, Indicates the number of training samples; Indicates the first The true grade value of each sample; This represents the grade value predicted by the spectral grade prediction model;

[0079] The parameter regularization loss is used to limit the excessive parameters of the spectral grade prediction model, suppress overfitting, and improve the generalization ability of the spectral grade prediction model under different spectral conditions. The parameter regularization loss is constructed based on all learnable parameters of the spectral grade prediction model combined with the L2 regularization loss, and is expressed by the following formula:

[0080] ;

[0081] in, This represents the learnable parameters (including all weights of convolutional layers, attention layers, fusion layers, MLPs, etc.) in the spectral grade prediction model. It represents the unique index (serial number / identifier) ​​of each independent learnable parameter matrix / tensor in the spectral grade prediction model, and is the traversal variable for traversing all learnable weight parameters of the model;

[0082] Attention entropy regularization loss is used to avoid excessive concentration or dispersion of attention, improve the model's ability to maintain stable attention to key spectral segments, and enhance the model's interpretability.

[0083] The attention entropy regularization loss is constructed based on the attention weight distribution entropy values ​​at each spectral sampling point location in the multi-head attention mechanism, and is expressed by the following formula:

[0084] ;

[0085] in, This indicates the length of the spectral sequence, i.e., the number of sampling points for each spectrum; Indicates the first... Attention weights for each position.

[0086] Spectral smoothing consistency loss is used to ensure the continuity of fused features in the wavelength dimension, suppress local abrupt changes caused by noise, and improve the robustness of the model to spectral noise.

[0087] The spectral smoothing consistency loss is constructed based on the eigenvector difference of the fused features between adjacent wavelength positions, and is expressed by the following formula:

[0088] ;

[0089] in, Indicates the th element in the fused feature sequence Feature vectors corresponding to each wavelength position;

[0090] Local-global consistency loss is used to strengthen the synergistic relationship between local convolutional features and global attention features, thereby improving the stability of the AdaptiveFusion module;

[0091] The local-global consistency loss is constructed based on the difference between local multi-scale features and global features, combined with the L2 norm, and is expressed by the following formula:

[0092] ;

[0093] in, This represents the local multi-scale features extracted by the multi-scale extraction module; This represents the global features obtained by the feature extraction module;

[0094] After constructing the spectral grade prediction model, the training of the spectral grade prediction model is carried out in an end-to-end manner. The parameters of the spectral grade prediction model are jointly optimized through a unified objective function. The training process includes data loading, batch construction, forward propagation, loss calculation, back propagation, and parameter update.

[0095] First, the preprocessed LIBS spectral data is divided into training, validation, and test sets. During the training phase, a multi-threaded data loading mechanism is employed, inputting the spectral data into the spectral grade prediction model in batches to improve data reading efficiency and reduce I / O latency. Each batch of data contains a fixed number of spectral samples and their corresponding true grade labels, ensuring the stability and controllability of the training process.

[0096] During the forward propagation phase, spectral features are sequentially extracted through a multi-scale extraction module, a feature extraction module, and an adaptive fusion module. This process progressively extracts local multi-scale features, local structural features, and global features. The local multi-scale features and global features are then fused using the adaptive fusion module to obtain the fused features. Finally, these features are input into a multilayer perceptron to generate the grade prediction value for the current batch.

[0097] After obtaining the prediction results, the training process calculates the loss value according to the pre-designed objective function. The objective function consists of multiple components, including the basic regression loss, parameter regularization loss, attention entropy regularization loss, spectral smoothing consistency loss, and local-global consistency loss, which are used to simultaneously optimize prediction accuracy, model stability, spectral physical consistency, and feature fusion performance. After the loss value is calculated, the gradient of all learnable parameters in the network is calculated using the backpropagation algorithm, and the parameters are updated using the AdamW optimization algorithm.

[0098] To prevent overfitting and improve the generalization ability of the spectral grade prediction model during training, this embodiment employs a validation set monitoring mechanism. After every few training epochs, the model is evaluated on the validation set, and the learning rate is dynamically adjusted or an early stopping strategy is triggered based on changes in the validation loss. Furthermore, gradient clipping, Dropout, and weight decay strategies can be enabled during training to further enhance training stability.

[0099] The entire training process runs on an NVIDIA RTX 4090 GPU, enabling high throughput and fast convergence on large-scale spectral datasets. After training, the spectral grade prediction model undergoes a final evaluation on the test set to verify its predictive performance on unseen samples. The resulting spectral grade prediction model can be directly used for real-time spectral analysis and grade prediction tasks in industrial settings.

[0100] Step 2: Acquire LIBS spectral data from the ore composition detection site, preprocess the LIBS spectral data to obtain spectral features, input the spectral features into the spectral grade detection model, the multi-scale extraction module extracts local multi-scale features based on the spectral features, the feature extraction module sequentially performs intra-spectral and inter-spectral modeling on the local multi-scale features to obtain global features, and obtain fused features based on the global features, local multi-scale features, and residual connections, and obtain ore grade prediction results based on the fused features and multilayer perceptron.

[0101] LIBS spectral data in the 320–800 nm band were acquired using an Aurora 4000 three-channel spectrometer. The LIBS spectral data was then preprocessed, including wavelength alignment, intensity normalization, and noise suppression, to obtain spectral features. These spectral features were then input into a trained spectral grade prediction model to obtain the corresponding ore grade prediction results.

[0102] To verify the effectiveness of the LIBS spectral grade prediction method and model based on multi-head attention provided in this embodiment on LIBS spectral data, this embodiment selects several representative traditional and deep learning methods as baseline models for comparison, including PLS, SVR, 1D-CNN, ResNet-1D, and Transformer. All models are trained and tested under the same dataset partitioning, training strategy, and evaluation metrics to ensure the fairness and reproducibility of the comparison results. The evaluation metric used is the mean absolute error (MAE). ), mean square error ( ) and coefficient of determination ( (), used to measure the stability of the model under different spectral conditions.

[0103] Table 1 shows the prediction performance of different models on the test set. The results show that traditional methods such as PLS and SVR exhibit significant underfitting on high-dimensional spectral data, resulting in large prediction errors. Convolutional models such as 1D-CNN and ResNet-1D can capture local peak features, but they are insufficient in cross-spectral correlation modeling. The Transformer model has strong global modeling capabilities, but its ability to express fine-grained local peaks is limited, leading to room for improvement in performance in complex spectral regions. In contrast, the model of this invention... , and It achieved the best performance across all three metrics, indicating that it can more accurately capture the local structure and global dependencies of the spectrum.

[0104] Table 1: Performance comparison of this application with different models.

[0105]

[0106] As can be seen from the results in Table 1, the model of this invention... , and It achieved top performance across all three metrics, demonstrating significantly superior performance compared to traditional methods and existing deep learning models in spectral feature extraction, nonlinear relationship modeling, and overall fitting ability. Traditional methods such as PLS and SVR exhibit significant underfitting characteristics on high-dimensional spectral data, with generally high error levels. This is primarily because these methods struggle to effectively capture complex peak structures and cross-spectral correlation information in the spectrum. Convolutional models such as 1D-CNN and ResNet-1D can extract local peak features, thus showing a significant performance improvement over traditional methods. However, due to the limitations of the local receptive field of the convolutional kernel, their ability to model long-range dependencies is insufficient, resulting in limited overall fitting accuracy.

[0107] The Transformer model has advantages in global modeling, therefore... While exhibiting good performance on the metrics, its sensitivity to local peak shapes is insufficient, especially in regions with weak peaks, overlapping peaks, or strong background fluctuations, leading to inadequate representation of local features. This invention's model enhances its ability to capture local peak shapes through a multi-scale convolution module, effectively models local dependencies and global correlations of the spectrum through a two-layer attention mechanism within and between spectral bands, and achieves dynamic weighted integration of local and global features through an adaptive fusion module. This allows the model to establish a more reasonable representation of spectral features at different scales and with different structures. Therefore, this invention's model... and The results show that it is significantly better than Transformer in all metrics, indicating that it not only has a stronger fitting ability, but also has a clear advantage in fine-grained prediction accuracy.

[0108] Furthermore, considering the overall fitting ability of the model, the model of this invention... A value of 0.958 indicates that it can explain the vast majority of the relationship between spectral variations and grade, demonstrating strong generalization ability. In contrast, traditional methods... The scores are generally below 0.85, indicating that they are unable to capture the complex nonlinear structures in spectral data. Although convolutional models have improved, they still cannot reach the fitting level of the model in this invention. Although Transformer performs well in global modeling, its fitting ability is still slightly inferior to the model in this invention due to the lack of fine-grained expression of local peak shapes.

[0109] In summary, the model of this invention demonstrates significant advantages in prediction accuracy, stability, robustness, and adaptability to complex spectral structures, fully proving its effectiveness and engineering application value in LIBS spectral grade prediction tasks.

[0110] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A LIBS spectral grade prediction method based on multi-head attention, characterized in that, Includes the following steps: Step 1: Construct a multi-scale extraction module for extracting local features at different scales based on one-dimensional convolution with different kernel sizes; construct a first spectral feature extraction unit and a second spectral feature extraction unit based on rotational position encoding and multi-head attention; construct a feature extraction module based on the first spectral feature extraction unit and the second spectral feature extraction unit; and construct a spectral grade prediction model based on the multi-scale extraction module, the feature extraction module, residual connections, and a multilayer perceptron. In the first spectral feature extraction unit and the second spectral feature extraction unit, the query matrix of the multi-head attention is obtained based on the expert convolution set and the router probability distribution, the key matrix is ​​obtained based on multiple depthwise separable convolutions and residual connections, and the value matrix is ​​obtained based on depthwise separable convolutions and linear projections. The expert convolution set includes convolution kernels that are... The vertical convolution and the convolution kernel are The horizontal convolution and the convolution kernel are Standard convolution; In both the first spectral feature extraction unit and the second spectral feature extraction unit, the query matrix and key matrix of the multi-head attention are set with rotation position encoding; The feature extraction module is constructed by stacking a first spectral feature extraction unit and a second spectral feature extraction unit. The first spectral feature extraction unit models the spectral line dependence within the spectral band based on multi-head attention and rotational position encoding to obtain local structural features; The second spectral feature extraction unit models the long-distance cross-spectral dependence of local structural features based on multi-head attention and rotational position encoding to obtain global features; Step 2: Acquire LIBS spectral data from the ore composition detection site, preprocess the LIBS spectral data to obtain spectral features, input the spectral features into the spectral grade detection model, the multi-scale extraction module extracts local multi-scale features based on the spectral features, the feature extraction module sequentially performs intra-spectral and inter-spectral modeling on the local multi-scale features to obtain global features, and obtain fused features based on the global features, local multi-scale features, and residual connections, and obtain ore grade prediction results based on the fused features and the multilayer perceptron.

2. The LIBS spectral grade prediction method based on multi-head attention according to claim 1, characterized in that, The network structure of the multi-scale extraction module includes: a multi-channel parallel structure constructed based on one-dimensional convolution with different kernel sizes; In the multi-channel parallel structure, each convolutional kernel branch corresponds to a local feature at a single scale, and each convolutional kernel branch operates independently without interfering with each other. The multi-scale extraction module obtains local features at multiple different scales based on one-dimensional convolution with different kernel sizes, and concatenates the multiple local features along the channel dimension to form local multi-scale features.

3. The LIBS spectral grade prediction method based on multi-head attention according to claim 2, characterized in that, The multi-scale feature extraction module combines the router probability distribution with all feature vectors in the input spectral features to select convolution kernel branches that are suitable for industrial spectral characteristics, obtains local features at multiple different scales, and concatenates multiple local features along the channel dimension to form local multi-scale features.

4. The LIBS spectral grade prediction method based on multi-head attention according to claim 1, characterized in that, The method of obtaining fusion features based on global features, local multi-scale features, and residual connections includes obtaining fusion features based on global features, local multi-scale features, residual connections, and weighted fusion. The fusion weights of the weighted fusion are adaptively generated based on global features and local multi-scale features.

5. The LIBS spectral grade prediction method based on multi-head attention according to claim 4, characterized in that, The spectral grade prediction model also includes an adaptive fusion module, which adaptively generates fusion weights based on global features and local multi-scale features, and then fuses the global features and local multi-scale features element by element based on the fusion weights. The adaptive fusion module includes a first input branch, a second input branch, and an output branch; Both the first input branch and the second input branch are equipped with a lightweight multilayer perceptron and a layer normalization mechanism for generating fusion weights; The output branch concatenates and fuses the features output from the first input branch and the second input branch to obtain fused features.

6. The LIBS spectral grade prediction method based on multi-head attention according to any one of claims 1-5, characterized in that, The objective function of the spectral grade prediction model is constructed based on the basic regression loss, parameter regularization loss, attention entropy regularization loss, spectral smoothing consistency loss, and local-global consistency loss. The basic regression loss is constructed based on the deviation between the predicted grade of the spectral grade prediction model and the actual grade measured in the industrial field, combined with the mean square error. The parameter regularization loss is constructed based on all learnable parameters of the spectral grade prediction model combined with L2 regularization loss; The attention entropy regularization loss is constructed based on the attention weight distribution entropy value of each spectral sampling point position in the multi-head attention mechanism; The spectral smoothing consistency loss is constructed based on the eigenvector differences of the fused features between adjacent wavelength positions; The local-global consistency loss is constructed based on the difference between local multi-scale features and global features combined with the L2 norm.

Citation Information

Patent Citations

  • Multi-element quantitative analysis method and system based on laser-induced breakdown spectroscopy

    CN119691407A

  • Interpretable deep feature fusion network-based industrial intelligent predictive maintenance method

    WO2026021130A1