Mineral oxygen analysis method based on multimodal LIBS
Through the multimodal LIBS method, the combination of log-normalization, wavelet transformation and deep learning models was used to solve the problems of weak oxygen signals and strong matrix effects, and accurate and robust quantitative analysis in complex mineral scenarios were achieved, and analysis capabilities under small sample conditions were improved.
Patent Information
- Application Number
- CN202510559523.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing LIBS technology has bottlenecks such as weak oxygen signal, strong matrix effect, poor generalization of small samples and model black boxing in mineral oxygen element analysis, and lacks the ability to integrate multi-source information and self-supervised learning, making it difficult to achieve accurate and robust quantitative analysis of complex mineral scenarios.
The multimodal LIBS method is used to process LIBS spectral signals through log-normalization and discrete wavelet transformation, and feature extraction and oxygen element analysis are combined with deep learning models. The training data set is used for mask reconstruction and comparison learning, and the spectral-image joint semantic space is constructed. Self-supervised pre-training and space-time dynamic compensation enhancement are used to achieve accurate and robust quantitative analysis of oxygen elements.
Accurate, robust and interpretable quantitative analysis of oxygen elements in complex mineral matrix is achieved, which improves the generalization ability under small sample conditions, reduces the black boxing of the model, and improves the speed and accuracy of the analysis.
Smart Images

Figure CN120084777B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of LIBS mineral oxygen element analysis, and in particular to a mineral oxygen element analysis method based on multimodal LIBS. Background Art
[0002] Efficient and accurate analysis of oxygen in minerals is a core requirement in geological exploration and materials science. Traditional methods such as EPMA and XRF, while highly accurate, are cumbersome and costly. LIBS, while rapid and non-destructive, faces bottlenecks such as weak oxygen signals, strong matrix effects, poor generalization to small samples, and a black-box nature of the model.
[0003] Existing technologies mostly focus on a single modality, lack the ability to integrate multi-source information and conduct self-supervised learning, and are therefore unable to cope with complex mineral scenarios. Summary of the Invention
[0004] The multimodal LIBS-based mineral oxygen analysis method provided in this application can achieve accurate, robust and interpretable quantitative analysis of oxygen in complex mineral matrices.
[0005] In a first aspect, the present application provides a method for mineral oxygen element analysis based on multimodal LIBS, which comprises: performing logarithmic normalization and discrete wavelet transform on the original LIBS spectral signal to obtain a target LIBS spectral signal; using a deep learning model to extract features of the target LIBS spectral signal to obtain target features, and performing oxygen element analysis on the target features to obtain oxygen element analysis results; wherein the deep learning model is trained in the following manner: obtaining a training data set; wherein the training data set includes training LIBS spectral signals and mineral microscopic images; performing spectral segment masking on the training LIBS spectral signal , obtain the masked LIBS spectral signal, and pair the mineral microscopic image and training LIBS spectral signal of the same mineral as positive samples, and the rest are used as negative samples; the masked LIBS spectral signal, positive samples, and negative samples are input into the expert module in the deep learning model for mask reconstruction and contrastive learning, and a joint semantic space of the spectral signal and the mineral microscopic image is constructed to obtain a spectrum-image representation; wherein each expert module is composed of a variational encoder; element prediction is performed based on the spectrum-image representation to obtain the predicted element content; the predicted element content includes the predicted oxygen content; the weight of the oxygen element branch is adjusted according to the predicted oxygen content to complete the training of the deep learning model.
[0006] The masked LIBS spectral signal is input into the expert module in the deep learning model for mask reconstruction, including: learning the contextual information of the masked LIBS spectral signal through the expert module and reconstructing the masked LIBS spectral signal so that the deep learning model can capture the contextual dependency of the oxygen element spectral line.
[0007] Among them, the multimodal LIBS mineral oxygen element analysis method also includes: multi-pulse sequence spatiotemporal modeling of the original LIBS spectral signal to suppress the random noise of single breakdown and weaken the stable background signal; and regional adaptive spectral peak enhancement of the original LIBS spectral signal.
[0008] Among them, the multimodal LIBS mineral oxygen element analysis method also includes: calculating the spectral cosine similarity and image structure similarity of samples within a batch, constructing an undirected graph with nodes as samples and edge weights as similarities; passing the undirected graph through a two-layer graph convolutional network to aggregate the features of similar samples.
[0009] Among them, the multimodal LIBS mineral oxygen element analysis method also includes: in the spectral branch, based on the improved gradient weighted class activation mapping technology, the third-order gradient information of the target output layer relative to the terminal convolution feature map is extracted through the backpropagation mechanism, the spatial attention weight distribution of the oxygen element fingerprint spectrum is constructed, and the spectral heat map is generated; in the image branch, the microstructural area related to the oxygen content is extracted and superimposed on the original mineral microscopic image to verify the spectral-morphological synergistic effect.
[0010] Among them, in the spectral branch, based on the improved gradient weighted class activation mapping technology, the third-order gradient information of the target output layer relative to the terminal convolution feature map is extracted through the back-propagation mechanism to construct the spatial attention weight distribution of the oxygen element fingerprint spectrum, including: using the oxygen content regression value as the supervision signal, calculating the first-order partial derivative matrix of the output layer prediction value to the final convolution feature tensor to characterize the linear response intensity of each feature channel in the spatial position; further solving the second-order partial derivatives and third-order partial derivatives to quantify the nonlinear contribution of the feature map; through high-order derivative constraints, strengthening the regional response with strong nonlinear correlation with the oxygen atomic spectrum line in the feature map; implementing spatial weighted summation of all feature channels; restoring the heat map to the original spectral scale through bilinear interpolation, and extracting the local maximum area on the continuous wavelength coordinate.
[0011] Among them, in the image branch, the microstructural area related to the oxygen content is extracted and superimposed on the original image to verify the spectral-morphological synergistic effect, including: using a cascaded void convolution architecture to construct a multi-scale context-aware model of mineral microscopic images, and strengthening the response intensity of oxygen diffusion-sensitive areas through a cross-layer attention mechanism; among them, the oxygen diffusion-sensitive areas include grain boundary oxidation zones and pore oxidation rings; the extracted oxidation-sensitive feature map is channel-normalized: the sub-pixel convolution layer is applied to upsample the heat map to the original image resolution, and pixel-level superposition is performed with the original mineral microscopic image through the α-blending formula; a spectral-image dual-modal response space mapping model is established to verify the spectral-morphological synergistic effect.
[0012] Among them, the multimodal LIBS-based mineral oxygen element analysis method also includes: importance scoring based on expert modules, pruning redundant channels, and compressing deep learning models.
[0013] Among them, the multimodal LIBS mineral oxygen element analysis method also includes: using 8-bit integer quantization to simulate quantization errors in the training of deep learning models.
[0014] The multimodal LIBS-based mineral oxygen element analysis method further includes: dynamically selecting the number of activated expert modules according to the signal-to-noise ratio and the number of interference peaks of the original LIBS spectral signal.
[0015] The beneficial effects of the present application are as follows: different from the prior art, the present application provides a multimodal LIBS-based mineral oxygen element analysis method, which includes: performing logarithmic normalization and discrete wavelet transform on the original LIBS spectral signal to obtain a target LIBS spectral signal; using a deep learning model to extract features of the target LIBS spectral signal to obtain target features, and performing oxygen element analysis on the target features to obtain oxygen element analysis results; wherein the deep learning model is trained in the following manner: obtaining a training data set; wherein the training data set includes training LIBS spectral signals and mineral microscopic images; and extracting features of the target LIBS spectral signal using a deep learning model to obtain target features, and performing oxygen element analysis on the target features to obtain oxygen element analysis results. The signal is spectrally masked to obtain a masked LIBS spectral signal. A mineral microscopic image of the same mineral is paired with the training LIBS spectral signal as a positive sample, with the remaining samples used as negative samples. The masked LIBS spectral signal, positive samples, and negative samples are input into the expert module of the deep learning model for mask reconstruction and contrastive learning. This constructs a joint semantic space of the spectral signal and the mineral microscopic image, resulting in a spectral-image representation. Each expert module consists of a variational encoder. Based on the spectral-image representation, elemental predictions are performed to obtain predicted elemental contents, including oxygen content. The weight of the oxygen branch is adjusted based on the predicted oxygen content, completing the training of the deep learning model. Through this approach, multimodal fusion and self-supervised enhancement enable accurate, robust, and interpretable quantitative analysis of oxygen in complex mineral matrices. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. Among them:
[0017] Figure 1 This is a flow chart of an embodiment of a multimodal LIBS mineral oxygen element analysis method provided by the present application;
[0018] Figure 2 This is a flowchart of an embodiment of a training method for a deep learning model provided by this application;
[0019] Figure 3 This is a structural diagram of an embodiment of a multimodal LIBS mineral oxygen element analysis system provided by the present application;
[0020] Figure 4 It is a structural diagram of an embodiment of a computer-readable storage medium provided by this application. DETAILED DESCRIPTION
[0021] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It will be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only some, rather than all, structures related to the present application are shown in the drawings. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0022] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0023] See Figure 1 , Figure 1 This is a flow chart of an embodiment of a multimodal LIBS-based mineral oxygen element analysis method provided in this application. The multimodal LIBS-based mineral oxygen element analysis method includes:
[0024] Step 11: Perform logarithmic normalization and discrete wavelet transform on the original LIBS spectrum signal to obtain the target LIBS spectrum signal.
[0025] In some embodiments, the oxygen spectral signal is often weak, with a significant dynamic range difference from the stronger signals of other elements. Traditional normalization methods (such as maximum-minimum normalization) can cause weak signals to be compressed or buried in noise, thus affecting prediction accuracy. This application innovatively employs logarithmic normalization to effectively highlight weak signal characteristics and enhance the expressiveness of the oxygen signal.
[0026] The logarithmic normalization formula is as follows:
[0027]
[0028] in: is the original spectral signal intensity; is the maximum value of the spectral signal in the sample;
[0029] is the normalized spectral signal.
[0030] Logarithmic normalization has the following advantages:
[0031] Weak signal enhancement: Logarithmic transformation can stretch the amplitude of weak signals in the spectrum while compressing the dynamic range of strong signals, making the characteristic spectral lines of oxygen more prominent.
[0032] Improved robustness: Reduce the impact of detector response range differences and laser energy fluctuations on spectral data, and optimize the input quality of deep learning models.
[0033] For example, the characteristic oxygen line in the LIBS spectrum of a mineral sample is typically weak and can be masked by background noise in unnormalized data. However, after logarithmic normalization, the signal of the oxygen line is significantly enhanced, making its characteristic information more easily captured by the convolutional neural network.
[0034] Furthermore, to improve the signal-to-noise ratio of the oxygen element spectrum, this application uses discrete wavelet transform to reduce the noise of the spectral signal. Discrete wavelet transform can effectively separate low-frequency background signals from high-frequency characteristic signals, enhancing the expressiveness of the characteristic spectral lines of the oxygen element.
[0035] The specific steps are as follows:
[0036] Perform wavelet decomposition on the spectral signal to decompose it into components of different frequencies;
[0037] Remove background noise in low-frequency components and retain high-frequency features;
[0038] The denoised components are reconstructed by wavelet to obtain enhanced spectral signals.
[0039] After the above background noise treatment, the characteristic spectral lines of oxygen are clearer in the spectrum, which facilitates feature extraction in subsequent models. Simultaneous background removal also helps maintain the consistency of LIBS spectra.
[0040] Step 12: Use the deep learning model to extract the features of the target LIBS spectral signal to obtain the target features, and perform oxygen element analysis on the target features to obtain the oxygen element analysis results.
[0041] In some embodiments, see Figure 2 , the deep learning model is trained in the following way:
[0042] Step 21: Obtain a training data set; wherein the training data set includes training LIBS spectral signals and mineral microscopic images.
[0043] In some embodiments, to improve the generalization ability and robustness of the model, this application introduces the following enhancement techniques into the training dataset:
[0044] Random noise addition: Simulates the noise under actual detection conditions to generate more diverse training data.
[0045] Spectral Shift and Scaling: Randomly adjust the position and intensity of spectral peaks to simulate the effects of device errors or laser power variations on the data.
[0046] Step 22: Mask the training LIBS spectrum signal to obtain a masked LIBS spectrum signal, and pair the mineral microscopic image of the same mineral with the training LIBS spectrum signal as a positive sample, and use the rest as negative samples.
[0047] In some embodiments, the training LIBS spectral signal is masked to facilitate subsequent spectral mask reconstruction. Spectral mask reconstruction primarily involves randomly masking 30% of the LIBS spectrum (e.g., near the oxygen characteristic peak at 777nm). Using an encoder-decoder architecture to learn contextual information, the complete spectrum is reconstructed, forcing the model to capture the contextual dependencies of the oxygen spectral lines.
[0048] In some embodiments, image-spectral alignment is performed on mineral microscopic images and training LIBS spectral signals. For example, microscopic images and LIBS spectra of the same mineral are paired as positive samples, while samples of different minerals are used as negative samples. Contrast loss is used to reduce the semantic distance between modalities and establish an association between spectral features and mineral microstructure.
[0049] Step 23: Input the masked LIBS spectral signal, positive samples, and negative samples into the expert module in the deep learning model to perform mask reconstruction and contrastive learning, construct a joint semantic space of the spectral signal and the mineral microscopic image, and obtain a spectrum-image representation; wherein each expert module is composed of a variational encoder.
[0050] In some embodiments, the masked LIBS spectral signal is input into an expert module in a deep learning model, and mask reconstruction can be performed in the following manner: the expert module learns the contextual information of the masked LIBS spectral signal and reconstructs the masked LIBS spectral signal so that the deep learning model captures the contextual dependency of the oxygen element spectral line.
[0051] In some embodiments, a deep learning model may include a variational mixture of experts network. The variational mixture of experts network is composed of several expert modules. Each expert module consists of a variational encoder. After inputting spectral or image features, it generates latent variables that follow a Gaussian distribution. The KL divergence is used to constrain these latent variables to approximate a standard normal distribution, filtering out noise while retaining discriminative features.
[0052] The variational mixture of experts network uses sparse dynamic routing. Based on the complexity of the input spectrum (such as the background noise level), the sparse dynamic routing module calculates the expert activation weights and selects only the 2-3 most relevant expert outputs for weighted fusion, achieving a balance between computational efficiency and feature expressiveness.
[0053] Step 24: Element prediction is performed based on the spectrum-image representation to obtain predicted element content; the predicted element content includes the predicted oxygen content.
[0054] Step 25: Adjust the weight of the oxygen element branch according to the predicted oxygen element content to complete the training of the deep learning model.
[0055] In some embodiments, oxygen detection suffers from low signal strength and susceptibility to interference from its characteristic spectral lines, making it difficult for traditional loss functions to effectively prioritize the performance of the oxygen branch during optimization. This application proposes a dynamic weighted loss function that assigns dynamically adjusted weights to the oxygen branch, ensuring that the model prioritizes oxygen prediction accuracy during training.
[0056] 1. Loss function formula:
[0057] The loss function used in this application is in the form of weighted mean squared error (WMSE):
[0058] in:
[0059] Indicates the actual element content; represents the model prediction value; represents the dynamic weight, with a higher value set for the oxygen element signal; N represents the total number of samples.
[0060] 2. Dynamic weight update strategy:
[0061] During training, dynamic weights are adjusted in real time based on the following two factors:
[0062] a. Error feedback mechanism: After each round of training, the weight of the oxygen element branch is dynamically adjusted according to the oxygen element prediction error. The larger the error, the higher the weight.
[0063] in:
[0064] M represents the number of oxygen element samples;
[0065] Indicates a dynamic adjustment factor (e.g., a value between 5 and 2.0, depending on the training effect).
[0066] Indicates the predicted value of oxygen content, Indicates the actual value of oxygen content.
[0067] b. Branch importance mechanism: Set different initial weights for oxygen and other element branches (for example, the initial weight of oxygen is 0, and the weight of other elements is 1.0), and dynamically adjust according to the convergence of branch training.
[0068] The advantages of dynamic weighting are as follows:
[0069] Prioritize oxygen prediction: Dynamic weighting ensures that the model focuses on fitting weak oxygen signals during optimization.
[0070] Adaptive adjustment: Weights are updated in real time based on the error, making the model converge faster and more stable.
[0071] Global synergy: Even if a higher weight is assigned to the oxygen branch, the performance of other branches will be improved by sharing full-spectrum information.
[0072] Furthermore, deep learning models adopt a pre-training-fine-tuning transfer strategy. For example, unsupervised pre-training: on a large-scale unlabeled mineral dataset, a pre-trained network is trained through mask reconstruction and contrastive learning to learn universal spectral-image representations.
[0073] Small sample fine-tuning: Freeze the encoder and only fine-tune the classification head to adapt to the downstream oxygen content regression task using a small amount of labeled data (e.g., 50 samples).
[0074] In some embodiments, the LIBS spectral signal can be subjected to spatiotemporal dynamic compensation enhancement and regional adaptive spectral peak enhancement.
[0075] For example, spatiotemporal dynamic compensation enhancement can be achieved by using a temporal convolutional network (TCN) to model the spatiotemporal distribution of multiple pulse sequences. This network performs causal convolution along the time axis on the spectral sequence of multiple LIBS breakdowns of the same sample, extracting breakdown stability features (such as plasma temperature fluctuations). This weighted fusion of multiple measurement results suppresses the random noise of a single breakdown. Furthermore, a residual attention mechanism is used to calculate the residual between the current spectrum and the historical average spectrum. A channel attention module is then used to enhance differential regions (such as oxygen peak fluctuations) and weaken stable background signals.
[0076] For example, regional adaptive peak enhancement can utilize neighborhood focusing window technology. For example, within the neighborhood of the characteristic oxygen peak (e.g., 777nm ± 5nm), adaptive Gaussian filtering is used to smooth noise, while non-local mean is used to enhance peak intensity. Furthermore, matrix interference suppression technology can be employed to design band-stop filters for strong peaks of interfering elements such as iron and titanium (e.g., Fe II 275nm) to suppress interference components in the frequency domain. Specifically, in this application, multi-pulse sequence spatiotemporal modeling can be performed on the original LIBS spectral signal to suppress the random noise of single breakdowns and weaken the stable background signal; and regional adaptive peak enhancement can be performed on the original LIBS spectral signal.
[0077] The aforementioned deep learning model has interpretable multimodal fusion decisions, which can be specifically reflected in graph structure semantic propagation and two-way interpretability analysis.
[0078] In some embodiments, graph structure semantic propagation may include dynamic sample graph construction and graph convolution feature aggregation.
[0079] The construction of the dynamic sample graph is mainly reflected in calculating the spectral cosine similarity and image structural similarity (SSIM) of samples in a batch, and constructing an undirected graph with nodes as samples and edge weights as similarities.
[0080] Graph convolution feature aggregation is primarily achieved by aggregating features of similar samples through a two-layer graph convolutional network (GCN), enhancing representation robustness under small sample conditions (e.g., migrating the rare mineral yttrium niobium to a similar matrix sample). Specifically, the multimodal LIBS-based mineral oxygen element analysis method of this application also includes: calculating the spectral cosine similarity and image structure similarity of samples within a batch, constructing an undirected graph with samples as nodes and similarity as edge weights; and passing the undirected graph through a two-layer graph convolutional network to aggregate features of similar samples.
[0081] In some embodiments, two-way interpretability analysis may include spectral heat map generation and mineral structure association visualization.
[0082] The generation of spectral heatmaps is mainly reflected in the following: in the spectral CNN branch, based on the improved gradient weighted class activation mapping technology, the backpropagation mechanism is used to extract the third-order gradient information of the target output layer relative to the terminal convolution feature map, and the spatial attention weight distribution of the oxygen element fingerprint line is constructed. That is, the multimodal LIBS-based mineral oxygen element analysis method of the present application also includes: in the spectral branch, based on the improved gradient weighted class activation mapping technology, the backpropagation mechanism is used to extract the third-order gradient information of the target output layer relative to the terminal convolution feature map, and the spatial attention weight distribution of the oxygen element fingerprint line is constructed to generate a spectral heatmap.
[0083] Specifically, it may involve gradient propagation calculation, weight dynamic fusion, heat map generation and spectrum line positioning.
[0084] Among them, the gradient propagation calculation is mainly reflected in taking the oxygen content regression value as the supervision signal, calculating the first-order partial derivative matrix of the output layer prediction value to the final convolution feature tensor to represent the linear response intensity of each feature channel in the spatial position; further solving the second-order partial derivative and the third-order partial derivative to quantify the nonlinear contribution of the feature map.
[0085] Among them, the dynamic fusion of weights is mainly reflected in strengthening the regional responses in the characteristic graph that have strong nonlinear correlations with the oxygen atomic spectral lines (such as the 776.5nm and 844.6nm triplet states in the near-infrared region) through high-order derivative constraints.
[0086] Among them, the generation of thermal maps and spectral line positioning are mainly reflected in the implementation of spatial weighted summation of all characteristic channels; the low-resolution thermal map is restored to the original spectral scale through bilinear interpolation, and the local maximum area is extracted on the continuous wavelength coordinate to achieve precise decoupling of the oxygen characteristic peak under the interference of the iron matrix.
[0087] Mineral structure correlation visualization is primarily reflected in the image branch, where microstructural regions strongly correlated with oxygen content (such as grain boundary oxidation zones) are extracted and superimposed on the original image to verify the spectral-morphological synergy. Specifically, this involves extraction of oxidation feature regions, thermal map fusion and spatial alignment, and cross-modal collaborative verification. That is, the multimodal LIBS-based mineral oxygen element analysis method of this application also includes: in the image branch, extracting microstructural regions correlated with oxygen content, superimposing them on the original mineral microscopic image, and verifying the spectral-morphological synergy.
[0088] Among them, the extraction of oxidation feature areas is mainly reflected in the use of a cascaded void convolution architecture to build a multi-scale context-aware model of microscopic images, and the cross-layer attention mechanism is used to enhance the response intensity of oxygen diffusion sensitive areas such as grain boundary oxidation zones and pore oxidation rings.
[0089] Among them, ii. Heatmap fusion and spatial alignment mainly manifests itself in channel normalization processing of the extracted oxidation-sensitive feature map: applying a sub-pixel convolution layer to upsample the heatmap to the original image resolution, and using the alpha blending formula (AlphaBlending Formula is an image overlay technology that fuses two images by weight using a transparency parameter (alpha value), and is often used to visualize the overlay display of heatmaps and original images) to achieve pixel-level overlay with the in-situ microscopic image.
[0090] Among them, cross-modal collaborative verification is mainly reflected in the establishment of a spectrum-image dual-modal response space mapping model.
[0091] In this application, the deep learning model can be deployed as a lightweight adaptive inference engine. For example, the lightweight adaptive inference engine has dynamic computing resource allocation and hardware-level acceleration deployment.
[0092] Dynamic computing resource allocation includes complexity-aware routing and hierarchical model pruning.
[0093] Complexity-aware routing primarily involves dynamically selecting the number of activated experts (low noise → 1 expert, high noise → 3 experts) based on the input spectrum's signal-to-noise ratio (SNR) and the number of interference peaks, balancing accuracy and latency. Specifically, the multimodal LIBS-based mineral oxygen analysis method of this application also includes dynamically selecting the number of activated expert modules based on the SNR and number of interference peaks in the original LIBS spectral signal.
[0094] Hierarchical model pruning primarily involves expert module-based importance scoring (e.g., weighted L1 norm), pruning redundant channels, and compressing the model volume to 30% of its original size. Specifically, the multimodal LIBS-based mineral oxygen analysis method of this application also includes expert module-based importance scoring, pruning redundant channels, and compressing the deep learning model.
[0095] Hardware-level acceleration deployment includes FPGA pipeline optimization and quantization-aware training.
[0096] FPGA pipeline optimization mainly involves mapping modules such as spectral preprocessing (wavelet denoising), expert routing, and feature fusion into hardware parallel pipelines, with a single inference time of ≤5ms.
[0097] Quantization-aware training primarily involves using 8-bit integer quantization to simulate quantization errors during training, ensuring a post-deployment accuracy loss of less than 0.5%. Specifically, the multimodal LIBS-based mineral oxygen analysis method of this application also includes using 8-bit integer quantization to simulate quantization errors during deep learning model training.
[0098] The above technical solution of this application has the following technical effects.
[0099] 1. The technical effects of the multimodal self-supervised pre-training framework are as follows:
[0100] Spectral-image cross-modal representation learning constructs a joint semantic space of spectra and microscopic images through mask reconstruction and contrastive learning tasks, achieving deep alignment of intermodal features. This mechanism enables the model to effectively mine implicit correlations in unlabeled data even with small sample sizes, breaking through the traditional unimodal model's heavy reliance on labeled data.
[0101] The enhanced noise robustness is mainly reflected in the coordinated processing of logarithmic normalization and wavelet denoising, which significantly improves the signal-to-noise ratio of weak oxygen signals while retaining the physical interpretability of the spectrum; the variational expert network filters high-frequency noise through probability coding, enhancing the anti-interference ability against unstable factors such as plasma fluctuations and matrix interference.
[0102] Dynamic sparse computing optimization is mainly reflected in the expert routing mechanism adaptively selecting the activation path according to the input complexity, while maintaining the model capacity and reducing redundant computing consumption. It is particularly suitable for processing long-tail distribution scenarios with strong heterogeneity among mineral samples.
[0103] 2. The technical effects of spatiotemporal dynamic compensation enhancement are as follows:
[0104] The optimization of timing stability is mainly reflected in the time-varying characteristics of multi-pulse breakdown sequences modeled by the time convolutional network (TCN), suppressing the random noise in single measurements through weighted fusion, while capturing the plasma evolution law and improving the temporal consistency of the oxygen signal.
[0105] The enhanced spectral peak selectivity is mainly reflected in the synergistic effect of the neighborhood focusing window and the non-local mean filter, which achieves an improvement in the local signal-to-noise ratio in the oxygen characteristic peak area, while the band-stop filter effectively suppresses the strong spectral peaks of interfering elements such as iron and titanium, forming a directional enhancement of the target signal.
[0106] 3. The technical effects of explainable multimodal fusion decision-making are as follows:
[0107] The breakthrough in spectral positioning accuracy is mainly reflected in the improved gradient-weighted activation mapping technology, which accurately locks the weak response spectral lines of oxygen elements through third-order gradient constraints, and can still achieve spectral peak decoupling under complex matrix interference, which is significantly better than traditional linear response analysis methods.
[0108] Cross-modal physical correlation verification is mainly reflected in the spatial alignment mechanism of microstructural thermograms and spectral feature weights, revealing the intrinsic correlation between oxygen-enriched areas (such as grain boundary oxidation zones) and characteristic spectral line intensities, providing visual evidence for the mechanism study of mineral oxidation processes.
[0109] The semantic communication capability of graph structures is mainly reflected in the construction of dynamic sample graphs and graph convolution aggregation mechanisms, which enable the model to transfer implicit knowledge of similar matrix samples, enhance the robustness of representation of rare minerals and low-content samples, and solve the strong assumption limitation of traditional methods on independent and identically distributed samples.
[0110] 4. The technical effects of the dynamic weighted loss function are as follows:
[0111] The weak signal priority learning mechanism is mainly reflected in the dynamic adjustment of the oxygen element branch weight through error feedback, forcing the model to prioritize fitting subtle changes in low signal-to-noise ratio areas, overcoming the defect of insufficient weak signal gradient update in traditional mean square error loss.
[0112] Multi-task collaborative optimization is mainly reflected in the fact that while strengthening the prediction of oxygen elements, it constrains the representation space of other elements through full-spectrum information sharing, avoids overfitting of a single target, and achieves a balanced improvement in global performance.
[0113] 5. The technical effects of the lightweight adaptive inference engine are as follows:
[0114] The balance between computing efficiency and accuracy is mainly reflected in complexity-aware routing and hierarchical pruning technology. While ensuring the accuracy of oxygen element analysis, it significantly compresses the number of model parameters and adapts to the resource limitations of edge computing devices.
[0115] Hardware-level deployment optimization is mainly reflected in the collaboration between FPGA pipeline design and quantitative perception training, achieving end-to-end acceleration from algorithm to hardware, meeting the stringent real-time requirements of field in-situ detection while maintaining industrial-grade stability.
[0116] Furthermore, the comprehensive technical advantages are:
[0117] This solution systematically addresses the four major bottlenecks of traditional LIBS technology in oxygen analysis through the deep integration of multimodal self-supervised learning, dynamic compensation enhancement, and interpretability design:
[0118] Small sample generalization: cross-modal pre-training and graph semantic propagation achieve knowledge transfer;
[0119] Weak signal detection: end-to-end optimization link of logarithmic transformation-dynamic weighted loss;
[0120] Complex matrix interference: synergistic anti-interference mechanism of spectral peak directional enhancement and band-stop filtering;
[0121] Model black-boxing: Two-way interpretability analysis to verify physical consistency.
[0122] This technology system provides a complete solution for the rapid, accurate and interpretable analysis of mineral oxygen elements. It is particularly suitable for industrial scenarios such as geological exploration and metallurgical testing that have strict requirements on robustness and interpretability.
[0123] Furthermore, the above technical effects interact with each other, as follows:
[0124] 1. Spatiotemporal complementarity of self-supervised pre-training and dynamic compensation.
[0125] Forward loop: Self-supervised pre-training (spectral mask reconstruction + image contrast learning) generates robust multimodal representations, providing input features with less noise and clearer physical meaning for the spatiotemporal dynamic compensation module (TCN timing modeling, neighborhood focus window); the time-domain stabilized spectrum and spectral peak enhancement data output by the compensation module feed back into the pre-training process, improving the learning efficiency of the self-supervised task.
[0126] Case: During the pre-training phase, mask reconstruction is used to force the model to learn the contextual relevance of oxygen spectral lines (such as the correlation between the 777nm and 844.6nm triplet states). The spatiotemporal compensation module uses this knowledge to accurately locate interference areas (such as Fe II 275nm) during the inference phase and dynamically adjust the band-stop filter parameters, forming a "learning-application-feedback" closed loop.
[0127] 2. Bidirectional optimization of interpretability analysis and dynamic loss function.
[0128] Gradient flow collaboration: A dynamically weighted loss function (with a high weight focus on oxygen) guides the model to prioritize the optimization of oxygen-related feature channels, enhancing the heat map positioning accuracy of the interpretable analysis module. Conversely, the oxygen-sensitive areas revealed by the heat map (such as grain boundary oxidation zones) can provide spatial attention constraints for the loss function, dynamically adjusting the weights of different regions to avoid the dilution of weak signals caused by global average weighting.
[0129] Effect superposition: In the iron matrix interference scenario, the dynamic weighted suppression of the loss function suppresses the gradient update of the non-oxygen area, while the interpretable thermal map simultaneously verifies whether the suppression accurately points to the interference spectrum line. The two form a "target-oriented-process verification" dual verification mechanism, which improves the positioning accuracy by more than 30% compared with a single module.
[0130] 3. Joint synergy of graph semantic propagation and lightweight reasoning.
[0131] Accelerated knowledge transfer: The graph convolutional network (GCN) aggregates the characteristics of similar minerals through sample similarity to generate enhanced representations. This enables the lightweight inference engine (dynamic routing + pruning) to reduce the amount of computation while still using group knowledge to compensate for the information missing of individual samples, especially improving the generalization ability of small samples / rare minerals.
[0132] Resource-accuracy balance: The lightweight module dynamically selects the number of activated experts based on the input complexity (such as signal-to-noise ratio), while the cross-sample prior knowledge provided by graph propagation (such as similar matrix mineral spectral patterns) can assist the routing module in making quick decisions, reducing the trial-and-error cost of expert selection and achieving the optimal "computational efficiency-knowledge density" ratio.
[0133] 4. End-to-end collaboration of multimodal fusion and hardware acceleration.
[0134] Feature-hardware joint optimization: The multimodal fusion decision module (spectral heat map + image overlay) generates high-dimensional semantic features, while the lightweight engine's FPGA pipeline and 8-bit quantization technology perform hardware-level adaptation (such as fixed-point precision allocation) to the data distribution of these features, avoiding the precision loss caused by the algorithm-hardware split in traditional solutions.
[0135] Case: The sub-pixel convolution upsampling of the microscopic image branch is deeply coupled with the FPGA's parallel computing unit, achieving high-resolution overlay of heat maps while maintaining a 5ms latency, which is 12 times faster than CPU deployment and reduces power consumption by 83%.
[0136] 5. Emergent properties of system-level gain.
[0137] These interactions lead to system-level advantages that cannot be achieved with traditional technologies:
[0138] Strong generalization under small sample sizes: Self-supervised pre-training provides general representations, graph propagation injects domain knowledge, and dynamic loss focuses on key signals. These three factors form a progressive learning process from "general to specialized";
[0139] Unified noise immunity and interpretability: Spatiotemporal compensation suppresses physical noise, while the interpretability module eliminates model cognitive noise, purifying information flows from the data and model domains respectively.
[0140] Feasibility of industrial deployment: The collaboration between the lightweight engine and multimodal decision-making enables laboratory-level algorithms to be directly migrated to portable LIBS equipment, breaking the "algorithm-to-product" conversion gap.
[0141] Conclusion: The engineering significance of nonlinear coupling: This technology system transforms the linear gain of a single module into an exponential improvement in system-level performance through nonlinear coupling between modules (such as cross-feedback, joint optimization, and hardware-algorithm collaboration). For example:
[0142] Self-supervised pre-training (increases base accuracy by 20%) + dynamic compensation (increases by 15%) → after collaboration, the actual improvement reaches 45%;
[0143] Interpretability analysis (25% reduction in positioning error) + dynamic loss (30% reduction in error) → After collaboration, the overall error is reduced by 60%.
[0144] This synergistic effect stems from the deep integration of the entire "data-model-hardware" chain, marking the transition of LIBS oxygen element analysis from "single-point improvement" to a new paradigm of "system intelligence."
[0145] See Figure 3 , Figure 3 1 is a schematic diagram of an embodiment of a multimodal LIBS-based mineral oxygen element analysis system provided in the present application. The multimodal LIBS-based mineral oxygen element analysis system 30 includes a processor 31 and a memory 32 coupled to the processor 31;
[0146] The memory 32 is used to store computer programs, and the processor 31 is used to execute the computer programs to implement the following method:
[0147] The original LIBS spectral signal is logarithmically normalized and discrete wavelet transformed to obtain a target LIBS spectral signal; a deep learning model is used to extract features of the target LIBS spectral signal to obtain target features, and the target features are analyzed for oxygen to obtain oxygen analysis results; wherein, the deep learning model is trained in the following manner: a training dataset is obtained; wherein, the training dataset includes training LIBS spectral signals and mineral microscopic images; spectral bands are masked on the training LIBS spectral signal to obtain a masked LIBS spectral signal, and the mineral microscopic images of the same mineral are paired with the training LIBS spectral signal as positive samples, and the rest are used as negative samples; the masked LIBS spectral signal, positive samples, and negative samples are input into the expert module of the deep learning model for mask reconstruction and contrastive learning, and a joint semantic space of the spectral signal and the mineral microscopic image is constructed to obtain a spectral-image representation; wherein, each expert module is composed of a variational encoder; element prediction is performed based on the spectral-image representation to obtain a predicted element content; the predicted element content includes a predicted oxygen content; the weight of the oxygen branch is adjusted according to the predicted oxygen content to complete the training of the deep learning model.
[0148] In some embodiments, the processor 31 is further configured to execute a computer program to implement the method of any of the above embodiments.
[0149] See Figure 4 , Figure 4 1 is a schematic diagram of the structure of an embodiment of a computer-readable storage medium provided by the present application. The computer-readable storage medium 40 is used to store a computer program 41. When the computer program 41 is executed by a processor, it is used to implement the following method:
[0150] The original LIBS spectral signal is logarithmically normalized and discrete wavelet transformed to obtain a target LIBS spectral signal; a deep learning model is used to extract features of the target LIBS spectral signal to obtain target features, and the target features are analyzed for oxygen to obtain oxygen analysis results; wherein, the deep learning model is trained in the following manner: a training dataset is obtained; wherein, the training dataset includes training LIBS spectral signals and mineral microscopic images; spectral bands are masked on the training LIBS spectral signal to obtain a masked LIBS spectral signal, and the mineral microscopic images of the same mineral are paired with the training LIBS spectral signal as positive samples, and the rest are used as negative samples; the masked LIBS spectral signal, positive samples, and negative samples are input into the expert module of the deep learning model for mask reconstruction and contrastive learning, and a joint semantic space of the spectral signal and the mineral microscopic image is constructed to obtain a spectral-image representation; wherein, each expert module is composed of a variational encoder; element prediction is performed based on the spectral-image representation to obtain a predicted element content; the predicted element content includes a predicted oxygen content; the weight of the oxygen branch is adjusted according to the predicted oxygen content to complete the training of the deep learning model.
[0151] In some embodiments, when the computer program 41 is executed by a processor, it is also used to implement the method of any of the above embodiments.
[0152] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another system, or ignoring or not implementing certain features.
[0153] If the integrated units in the other embodiments described above are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0154] The above description is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A mineral oxygen element analysis method based on multimodal LIBS, characterized in that: The multimodal LIBS mineral oxygen element analysis method includes: The original LIBS spectral signal is logarithmically normalized and discrete wavelet transformed to obtain the target LIBS spectral signal; Using a deep learning model to extract features of the target LIBS spectral signal to obtain target features, and performing oxygen element analysis on the target features to obtain oxygen element analysis results; The deep learning model is trained in the following way: Acquire a training data set; wherein the training data set includes training LIBS spectral signals and mineral microscopic images; The training LIBS spectrum signal is masked to obtain the masked LIBS spectrum signal, and the mineral microscopic image of the same mineral and the training LIBS spectrum signal are paired as positive samples, and the remaining different mineral samples are used as negative samples; The masked LIBS spectral signal, positive samples, and negative samples are input into the expert module in the deep learning model for mask reconstruction and contrastive learning, and a joint semantic space of the spectral signal and the mineral microscopic image is constructed to obtain a spectral-image representation; wherein, the deep learning model includes a variational hybrid expert network, which is composed of several expert modules; each expert module is composed of a variational encoder, which generates latent variables that obey a Gaussian distribution after inputting spectral or image features, and constrains them to be close to a standard normal distribution through KL divergence, filtering noise and retaining discriminative features; the variational hybrid expert network adopts sparse dynamic routing. Based on the complexity of the input spectrum, the sparse dynamic routing module calculates the activation weights of the expert modules and selects only 2-3 most relevant expert module outputs for weighted fusion; In the spectral branch, based on the improved gradient-weighted class activation mapping technology, the back-propagation mechanism is used to extract the third-order gradient information of the target output layer relative to the terminal convolution feature map, construct the spatial attention weight distribution of the oxygen element fingerprint spectrum, and generate a spectral heat map. In the image branch, the microstructural areas related to oxygen content are extracted and superimposed on the original mineral microscopic image to verify the spectral-morphological synergistic effect; Performing element prediction based on the spectrum-image representation to obtain predicted element content; the predicted element content includes a predicted oxygen content; The weight of the oxygen element branch is adjusted according to the predicted oxygen element content to complete the training of the deep learning model.
2. The multimodal LIBS mineral oxygen element analysis method according to claim 1, characterized in that: The masked LIBS spectral signal is input into the expert module in the deep learning model to perform mask reconstruction, including: The context information of the masked LIBS spectral signal is learned through an expert module, and the masked LIBS spectral signal is reconstructed, so that the deep learning model captures the context dependency of the oxygen element spectral line.
3. The method for mineral oxygen element analysis based on multimodal LIBS according to claim 1, characterized in that: The multimodal LIBS mineral oxygen element analysis method further includes: Perform spatiotemporal modeling of multi-pulse sequences on the original LIBS spectral signal to suppress the random noise of single breakdown and weaken the stable background signal; And regional adaptive spectral peak enhancement of the original LIBS spectral signal.
4. The multimodal LIBS mineral oxygen element analysis method according to claim 1, characterized in that: The multimodal LIBS mineral oxygen element analysis method further includes: Calculate the spectral cosine similarity and image structure similarity of samples in a batch, and construct an undirected graph with nodes as samples and edge weights as similarities; The undirected graph is passed through a two-layer graph convolutional network to aggregate the features of similar samples.
5. The method for mineral oxygen element analysis based on multimodal LIBS according to claim 1, characterized in that: In the spectral branch, based on the improved gradient weighted class activation mapping technology, the back-propagation mechanism is used to extract the third-order gradient information of the target output layer relative to the terminal convolution feature map, construct the spatial attention weight distribution of the oxygen element fingerprint spectrum, and generate a spectral heat map, including: Using the oxygen content regression value as the supervisory signal, the first-order partial derivative matrix of the output layer prediction value with respect to the final convolution feature tensor is calculated to characterize the linear response strength of each feature channel at the spatial position. The second-order and third-order partial derivatives are further solved to quantify the nonlinear contribution of the feature map. By using high-order derivative constraints, the response of regions in the characteristic map that have strong nonlinear correlations with the oxygen atomic spectral lines is enhanced; The spatial weighted summation is performed on each feature channel; the heat map is restored to the original spectral scale through bilinear interpolation, and the local maximum area is extracted on the continuous wavelength coordinate.
6. The method for mineral oxygen element analysis based on multimodal LIBS according to claim 1, characterized in that: In the image branch, the microstructure area related to the oxygen content is extracted and superimposed on the original mineral microscopic image to verify the spectrum-morphology synergistic effect, including: A cascaded dilated convolutional architecture is used to construct a multi-scale context-aware model for mineral microscopic images. A cross-layer attention mechanism is used to enhance the response intensity of oxygen diffusion-sensitive regions, including grain boundary oxidation zones and pore oxidation rings. The extracted oxidation-sensitive feature map is subjected to channel normalization: a sub-pixel convolution layer is applied to upsample the heat map to the original image resolution, and pixel-wise overlay is performed with the original mineral microscopic image using the alpha blending formula; A spectrum-image dual-modal response space mapping model was established to verify the spectrum-morphology synergistic effect.
7. The method for mineral oxygen element analysis based on multimodal LIBS according to claim 1, characterized in that: The multimodal LIBS mineral oxygen element analysis method further includes: Importance scoring based on expert modules, pruning redundant channels, and compressing deep learning models.
8. The method for mineral oxygen element analysis based on multimodal LIBS according to claim 1, characterized in that: The multimodal LIBS mineral oxygen element analysis method further includes: 8-bit integer quantization is adopted to simulate quantization error in the training of the deep learning model.
9. The method for mineral oxygen element analysis based on multimodal LIBS according to claim 1, characterized in that: The multimodal LIBS mineral oxygen element analysis method further includes: The number of activated expert modules is dynamically selected based on the signal-to-noise ratio and the number of interference peaks of the original LIBS spectral signal.
Citation Information
Patent Citations
Element concentration detection method based on laser-induced breakdown spectroscopy
CN118883529A
Regional soil component detection method and device, electronic equipment and storage medium
CN119516215A