Multi-element quantitative analysis method for laser ablation inductively coupled plasma mass spectrometry
By constructing a signal extraction and classification model based on machine learning, the problems of standard reliance and human experience in LA-ICP-MS quantitative technology have been solved, realizing fully automated and high-precision multi-element quantitative analysis, which is applicable to different instrument platforms and expands the application of elemental quantitative imaging.
Patent Information
- Application Number
- CN202610055983.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-16
- Publication Date
- 2026-03-06
AI Technical Summary
Existing quantitative techniques for laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS) rely on external standard methods, internal standard methods, or signal summation normalization methods. These methods suffer from difficulties in obtaining standards, cumbersome procedures, poor universality, and high dependence on human experience, resulting in low efficiency and poor repeatability. They are also unable to effectively process multi-channel time-series mass spectrometry signals from LA-ICP-MS.
A machine learning-based approach was adopted to construct a one-dimensional convolutional neural network (1D-CNN) for signal extraction. A hybrid architecture of convolutional neural network and long short-term memory network (CNN-LSTM) was used for signal classification and quantitative prediction. Combined with element ratio normalization, PowerTransformer transformation and concentration grouping modeling, standard-free multi-element quantitative analysis was achieved.
It achieves full-process automation and standard-free quantification, improves analytical efficiency and accuracy, eliminates human error, is compatible with different instrument platforms, expands the application of elemental quantitative imaging, and reduces analytical costs.
Smart Images

Figure CN121612970A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of analytical instruments and artificial intelligence technology. Specifically, it relates to a machine learning model, quantitative analysis method, and system for laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS), which enables standard-free, automated, and multi-element quantitative analysis of solid samples. Background Technology
[0002] Laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS) has become an important tool for elemental quantification in fields such as geology, environment, materials, and biology due to its advantages, including no need for complex pretreatment, the ability to perform in-situ analysis of micro-areas, and simultaneous multi-element detection. However, its quantitative accuracy is limited by elemental fractionation effects and matrix effects, and currently, it mainly relies on external standard methods, internal standard methods, or signal summation normalization methods for correction. These methods have the following limitations: external standard methods require matrix-matched standard substances, which are difficult to obtain; internal standard methods require prior knowledge of the content of the internal standard element, which is cumbersome and cannot be implemented in applications such as elemental imaging; signal summation normalization methods are limited by the integrity of sample composition and have poor universality.
[0003] In recent years, machine learning (ML) has shown its potential in analytical chemistry to handle complex nonlinear relationships, but its application in LA-ICP-MS has lagged significantly. Existing research largely focuses on laser-induced breakdown spectroscopy (LIBS), while the multi-channel time-series mass spectrometry signals generated by LA-ICP-MS have complex data structures that differ greatly from LIBS spectral data, making direct transfer of existing algorithms impossible. LIBS data consists of single-excitation, spatially resolved spectra, while LA-ICP-MS data is a continuous ablation, time-resolved mass spectrometry sequence, with completely different data structure dimensions and physical meanings. Existing LIBS machine learning models cannot handle background drift, transition region identification, and temporal coupling effects between multiple elements in time-series signals. Furthermore, key steps in traditional quantitative procedures, such as signal extraction and background subtraction, heavily rely on human experience, resulting in low efficiency and poor repeatability. This invention addresses and overcomes this technological barrier by designing, for the first time, a machine learning model system (including signal extraction, classification, and quantification models) fully adapted to the time-resolved mass spectrometry data structure of LA-ICP-MS. This successfully introduces the advantages of machine learning into the LA-ICP-MS quantitative field, solving problems that LIBS algorithms cannot address. Therefore, developing an intelligent quantitative method that can automatically process complex LA-ICP-MS signals and fundamentally eliminate reliance on standard samples has become a pressing technical challenge in this field. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing LA-ICP-MS quantitative techniques and provide a fully automated, standard-free, high-precision multi-element quantitative analysis method and system based on a machine learning model. This method aims to eliminate the subjectivity of manual operation, overcome the dependence on internal / external standards, significantly improve analytical efficiency and the comparability of data between different instrument platforms, and expand the application of LA-ICP-MS in cutting-edge fields such as elemental quantitative imaging.
[0005] A multi-element quantitative analysis method for laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS) includes the following steps:
[0006] Acquire raw time-resolved mass spectrometry data of the sample using LA-ICP-MS; input the raw data into a trained signal extraction model, which is a model built on a one-dimensional convolutional neural network (1D-CNN) specifically designed to identify and separate background signals, sample signals and transition signals from LA-ICP-MS multi-channel time series, and output the effective sample signal range;
[0007] The valid sample signal is input into a trained signal classification model, which is a model built on a hybrid architecture of convolutional neural network and long short-term memory network (CNN-LSTM) specifically designed to classify LA-ICP-MS signal patterns to identify normal signals, no signals and inclusion signals, and outputs the classified normal sample signal.
[0008] The normal sample signal is input into a trained element quantitative prediction model. The element quantitative prediction model is a model built on a CNN-LSTM hybrid architecture and uses the target element signal and at least one other major element signal as combined inputs. It is specifically designed to directly predict the element content and outputs the predicted content of multiple elements in the sample.
[0009] Furthermore, before inputting the signal into the quantitative prediction model, the following preprocessing steps are included: normalizing the element ratios of the signal; and performing a PowerTransformer transformation on the normalized signal to correct data skew.
[0010] Further, in step S4, the training method of the quantitative prediction model includes: dividing the concentration range of the target element in the training data into multiple concentration intervals, and training an independent sub-model for each concentration interval.
[0011] Furthermore, in step S2, during the training of the signal extraction model, the input data undergoes dimensionality reduction and feature compression processing based on Pearson correlation coefficient and sliding window standard deviation, compressing the original multi-channel data into an input with a fixed number of channels.
[0012] Furthermore, in step S3, the input features of the signal classification model are fused with the original time series signal and its derived features. The derived features include at least one of the following: the logarithmic form of the signal, the sliding standard deviation, the amplitude of the fast Fourier transform (FFT), the total signal intensity ratio, the peak relative to the background intensity, and the peak stability.
[0013] Furthermore, the method does not require any internal standard elements or external standard substances for quantitative correction.
[0014] Furthermore, a multi-element quantitative analysis system for laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS) includes: a data processing module for receiving raw time-resolved mass spectrometry data from LA-ICP-MS; a model storage and execution module for storing the signal extraction model, signal classification model, and quantitative prediction model trained by the above method, and for performing corresponding calculations; and a result output module for outputting the processed effective signal range, signal classification results, and element quantitative prediction content.
[0015] Further, it includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor implements the steps of the above method when executing the program.
[0016] Furthermore, when the program is executed by the processor, it implements the steps of the above method.
[0017] This application addresses the following three core issues:
[0018] 1. Originality and Breakthrough: For the first time in the field of LA-ICP-MS, a complete, end-to-end machine learning quantitative analysis model was constructed, realizing the simultaneous standard-free quantification of 9 major elements and 30 trace elements, which is an important breakthrough in the development of this technology.
[0019] 2. Full-process automation and high objectivity: It realizes full-process automation from raw signal extraction and classification to final content prediction, completely replacing manual operation that relies on experience, eliminating human error, and ensuring the objectivity and consistency of data processing results.
[0020] 3. Fundamentally eliminates reliance on standard samples: The model learns and corrects relationships autonomously from the data through algorithms, without the need for any internal or external standards, which greatly reduces analysis costs and time, and makes it possible to accurately quantify samples without suitable standards or with unknown internal standard information (such as elemental imaging samples).
[0021] 4. High Precision and Strong Robustness: Through a series of targeted designs such as "multi-element input," "Si normalization," and "concentration grouping modeling," the model achieves high precision in both major and trace elements, meeting the needs of practical analysis (e.g., ...). Figure 7 (As shown), and it can adapt well to data from different laboratories and instruments, demonstrating excellent robustness and generalization ability.
[0022] 5. Broad application prospects: This method provides a direct and efficient solution for LA-ICP-MS elemental quantitative imaging, and is expected to promote the standardization and wider application of this technology in the study of micro-area elemental distribution. Attached Figure Description
[0023] This specification will further illustrate embodiments by way of exemplary models, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:
[0024] Figure 1 This is a schematic diagram of the signal extraction model (a) and the signal classification model (b) in this invention.
[0025] Figure 2 This is a schematic diagram of the structure of the element quantitative prediction model (CNN-LSTM) in this invention.
[0026] Figure 3 This is the overall flowchart of the end-to-end element content prediction algorithm of the present invention.
[0027] Figure 4 This is an evaluation chart showing the machine learning model's ability to extract and classify raw LA-ICP-MS signals.
[0028] Figure 5 This is a comparison chart showing the impact of different signal preprocessing strategies on the accuracy of MgO content prediction.
[0029] Figure 6 A comparative chart showing the impact of optimizing quantitative models using different input data dimensions (single element vs. multi-element combination) on prediction accuracy.
[0030] Figure 7 This is a comparison chart showing the prediction results of the model of this invention with the results of the traditional internal standard method on different laboratory data. Detailed Implementation
[0031] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. It should be understood that these exemplary embodiments are given merely to enable those skilled in the art to better understand and implement this specification, and are not intended to limit the scope of this specification in any way. Unless obvious from the linguistic context or otherwise, the same reference numerals in the figures represent the same structures or operations.
[0032] As introduced in the background section, existing algorithm research mainly focuses on laser-induced breakdown spectroscopy (LIBS). LIBS data is a single-excitation, spatially resolved spectrum, while LA-ICP-MS data is a continuous ablation, time-resolved mass spectrometry sequence. The data structure and physical meaning are completely different, therefore existing algorithms cannot be directly transferred. Machine learning models have opened up new avenues for application scenarios where accurate internal standard information is difficult to obtain. For example, in elemental imaging of geological samples using LA-ICP-MS, machine learning models completely solve the bottleneck problem of obtaining the content of internal standard elements, directly performing elemental quantification calculations on the raw mass spectrometry data, providing a more reliable and efficient solution for quantitative elemental imaging. This technology not only has the potential to promote the standardized application of LA-ICP-MS in geological imaging but also to extend to the analysis of environmental samples, biological tissues, and other substances, further promoting the development of microscopic elemental quantitative analysis towards automation and high precision.
[0033] This model automates LA-ICP-MS data processing, improving efficiency and eliminating data quality fluctuations caused by differences in researcher experience. It also enhances the reproducibility and comparability of quantitative data across different laboratories and instruments. The model eliminates the reliance on internal standard elements, external standard samples, and linear correction formulas found in traditional analyses. This reduces the need for repeated determinations of standard substances during analysis and eliminates the need for pre-obtaining internal standard elements from samples using other techniques, significantly improving analytical efficiency. In addition to traditional methods such as external standard correction, internal standard correction, and signal summation normalization, this machine learning-based quantitative model provides a new option for elemental quantitative analysis in LA-ICP-MS.
[0034] This patent proposes a machine learning model for elemental quantification in laser ablation plasma mass spectrometry (LAPS), addressing the elemental fractionation problem in LPS and enabling standard-free and standard-free multi-element quantification of 9 major elements and 30 trace elements in silicate samples. Simultaneously, the model automates data extraction and processing, reducing manual intervention. To this end, the algorithm consists of three deep learning models with different objectives: identifying and extracting background and sample signals from the original signal, classifying elemental signals, and quantifying elements. These are described below:
[0035] Background signal and sample signal identification and extraction: This model is built on a one-dimensional convolutional neural network (1D-CNN), such as... Figure 1 As shown in Figure a, its input is the raw time-resolved mass spectrometry signal with labels including "background signal," "sample signal," and "transition signal." The model captures the temporal features of the signal through convolutional layers and outputs the classification result for each time point, thereby achieving automatic and accurate separation of background and sample signals. To address the problem of overfitting caused by a large number of channels in the original data (e.g., more than 40 elements), the data is dimensionality-reduced before training based on Pearson correlation coefficient and sliding window standard deviation, compressing the multi-channel data into a smaller, more representative number of channels (e.g., five channels). A specific example is the construction of a (1D-CNN) for separating time-series background signals, shown in Figure c. The input of this model is a labeled one-dimensional array, namely the raw time-resolved mass spectrometry signal and its three labels (background signal, sample signal, and transition signal). The preprocessed model is then input into a one-dimensional convolutional neural network. This model mainly consists of convolutional layers, fully connected layers, and an output layer. The convolutional backbone network includes three convolutional layers. Since larger convolutional kernels help enhance the smoothness of the time series and the ability to capture longer temporal information, the kernel size is 41 for each layer, followed by a ReLU activation function. The output layer projects each time step to three output classes through a linear mapping, achieving time-series classification. Because the original data contains many channels and measures over 40 elements simultaneously, with strong correlations and redundancy among the elements, directly inputting it into a one-dimensional convolutional neural network model can easily lead to overfitting. Therefore, before modeling, dimensionality reduction and feature compression of the original data are performed by calculating the Pearson correlation coefficient and the sliding window standard deviation, ultimately obtaining five-channel input data. This significantly reduces dimensionality while ensuring information integrity, improving model stability and generalization ability. During training, normalization is used, and the loss function is cross-entropy loss. The optimizer is Adam to achieve smooth updates of the adaptive learning rate and gradient momentum.
[0036] Element signal classification: This model employs a hybrid architecture of CNN and LSTM, such as... Figure 1As shown in b, a feature extractor sensitive to the local morphology (such as spikes and plateaus) of the LA-ICP-MS signal was designed for the CNN part, while the LSTM part specifically models the signal's unique attenuation and fluctuation patterns. Based on geological interpretation, its input is the effective signal segment processed by the signal extraction model, labeled as "normal signal," "no signal," and "inclusion signal." The CNN part is responsible for extracting the local spatiotemporal features of the signal, while the LSTM part models the overall temporal dependencies of the signal. To improve model performance, the input features not only include the original signal values but also incorporate various manually designed derived features, such as logarithmic intensity, moving standard deviation, and FFT amplitude, enabling the model to more comprehensively understand the signal patterns. Specifically, a deep learning algorithm was designed to perform element signal classification tasks to replace the user's manual observation of the features of each sample signal. Figure 1 (b) This model employs a hybrid CNN and LSTM architecture to simultaneously capture local patterns and temporal dependencies. Its input is a labeled one-dimensional array, representing the original time-resolved mass spectrometry signal and its three labels (normal signal, no signal, and inclusion signal). The CNN part includes two convolutional layers, ReLU activation, and max pooling operations to extract local temporal features and progressively compress the temporal dimension. The LSTM part models the dynamic changes of the signal over time, using its final state as the input to the output layer. A linear mapping then projects each set of temporal signals to three output classes, achieving signal classification. To balance the original signal with prior information, enabling faster convergence and stronger generalization performance under limited experimental data, the input features integrate the original signal with domain knowledge: including seven complementary features such as the original value and its logarithmic form, moving standard deviation, fast Fourier transform amplitude, total signal intensity ratio, peak relative to background intensity, and peak stability. The model uses normalization processing, employs cross-entropy loss, and uses Adam as the optimizer.
[0037] Specifically, a hybrid deep learning model based on convolutional neural networks (CNN) and long short-term memory networks (LSTM) was constructed to extract features and predict element content in multi-channel LA-ICP-MS time series signals. Figure 2 Unlike traditional correction models, this model architecture captures both local mutation features and long-term temporal dependencies through convolutional feature extraction and temporal modeling. This model is the core of this invention and also employs a CNN-LSTM hybrid architecture, such as... Figure 2 As shown. Its innovation lies in:
[0038] Multi-element input strategy: For predicting a target element (e.g., Mg), the model input includes not only the signal of that element but also the signal of at least one other major element (e.g., Al, Fe, Na, K, Ca, Ti). This design allows the model to learn and correct for common interferences affecting all elements caused by factors such as instrument drift, thereby significantly improving prediction accuracy and model generalization ability (e.g., ...). Figure 6 (As shown). This design enables the model to internally learn and correct for common instrument state drift, which is key to achieving standard-free quantification. Furthermore, combined with preprocessing and training strategies specifically designed for LA-ICP-MS quantification, such as elemental normalization and concentration-grouped modeling, a complete solution is ultimately formed.
[0039] Key preprocessing steps:
[0040] a. Normalization: Ratio of all element signals to element signals, using the stability of the element ratio to offset some system fluctuations.
[0041] b. PowerTransformer: Transforms the normalized signal to correct its skewed distribution and make it more consistent with the numerical assumptions of the model.
[0042] c. Concentration Grouping Modeling: To address the issue of elemental concentrations spanning multiple orders of magnitude, the training data is divided into several intervals based on the concentration of the target element, and a dedicated sub-model is trained for each interval. This effectively solves the problem of decreased prediction accuracy in low-concentration regions (e.g., ...). Figure 5 (As shown).
[0043] Ensemble learning strategy: For each element or each concentration range, multiple network models are trained independently, and the outputs of multiple models are combined during prediction (e.g., by averaging or averaging after removing outliers) to enhance the stability and robustness of the prediction.
[0044] Data preprocessing strategy: After comparing various approaches, we applied the following process for each sample and each element channel before inputting it into the neural network:
[0045] (1) Background subtraction: The background signal is generated by instrument noise, gas phase fluctuations and pre-ablation readings. The background signal is estimated from the pre-ablation section and subtracted to obtain the net analysis signal.
[0046] (2) Element normalization.
[0047] (3) Peak removal and noise reduction: The signal after contrast value normalization is subjected to peak removal and smoothing to remove occasional peaks and random noise, improve signal quality, and achieve meaningful time feature extraction.
[0048] (4) Label Grouping: Since element concentrations span multiple orders of magnitude, training the model directly across the entire range often biases it towards high-concentration samples, leading to decreased accuracy in low-concentration regions. To address this issue, the concentration range of each element is divided into multiple intervals, and an independent model is trained for each interval. The number of intervals is determined by the order of magnitude covered by the training data. This strategy improves the model's adaptability across different concentration intervals.
[0049] Model Structure: The CNN-LSTM model consists of three main parts: a convolutional feature extraction module, a temporal modeling module, and a regression prediction module. The convolutional module contains two one-dimensional convolutional layers, each followed by a ReLU activation function and a max-pooling layer. The first convolutional layer focuses on capturing local temporal structures, such as short-period fluctuations and sudden spikes. The subsequent max-pooling operation reduces the temporal resolution while preserving significant structural information, effectively suppressing noise and reducing computational cost. The second convolutional layer further extracts higher-order local temporal patterns, enabling the model to learn the key features of complex signals. The convolutional features are then fed into an LSTM network, which models long-range temporal dependencies in the sequence. The hidden states at the final time step are extracted as a global representation for subsequent regression. The regression module consists of two fully connected layers with ReLU activation and dropout regularization to enhance non-linear expressiveness and reduce overfitting. The final output is a single continuous value representing the predicted concentration of the target element. For each element and each concentration group, the optimal set of hyperparameters selected by Optuna for this model is used to train the final model for that specific element and interval, thereby ensuring optimal architecture configuration across different concentration ranges.
[0050] The above describes the data preprocessing and training process of this invention. The core objective of this algorithm is to use the existing LA-ICP-MS raw dataset, label it, and input it into the model for training. In subsequent uses, only the original LA-ICP-MS test data is needed. Figure 3 This model can predict the elemental content of a sample directly without any internal or external standards, promoting the application of LA-ICP-MS in scenarios where it is difficult to obtain internal standard elemental information.
[0051] Specific Example: Multi-element quantitative analysis of silicate standard samples
[0052] Data preparation: Collect raw time-resolved mass spectrometry data of LA-ICP-MS for various silicate standard samples with known elemental contents (such as USGS glass standards BCR-2G, BHVO-2G, BIR-1G). The data should cover the complete process from background, signal rise to signal decay.
[0053] Model training:
[0054] Signal extraction model training: The original data is labeled (background, sample, transition) and then dimensionality reduced to train a three-layer convolutional 1D-CNN model.
[0055] Signal classification model training: The extracted sample signals are labeled (normal, no signal, inclusions), and seven derived features are calculated and input into the CNN-LSTM hybrid model for training.
[0056] Quantitative prediction model training: Known element concentrations are used as labels. For each target element, its signal is combined with the signal of a selected principal element to form the input. First, Si normalization and PowerTransformer transformation are performed. Then, the data are grouped according to the concentration distribution of the element, and a CNN-LSTM sub-model is trained for each group. The Adam optimizer and cross-entropy / mean squared error loss function are used for training.
[0057] Model application and validation:
[0058] 1. Obtain raw LA-ICP-MS data of the silicate sample to be tested.
[0059] 2. Input the raw data into the trained signal extraction model to obtain the net sample signal segment.
[0060] 3. Input the net sample signal into the signal classification model to identify and filter out "normal signals".
[0061] 4. For each element to be measured, its "normal signal" and the selected major element signal are subjected to the same Si normalization and PowerTransformer transformation, and the corresponding quantitative sub-model is selected for prediction based on its approximate concentration range.
[0062] 5. Finally, the predicted content values of multiple elements in the sample are obtained.
[0063] Specifically:
[0064] The machine learning model constructed in this invention enables the extraction and classification of raw time-resolved mass spectrometry signals. On the one hand, it eliminates the drawbacks of traditional manual selection of signal intervals, and on the other hand, it is a key step in realizing automated data processing. Figure 4 This demonstrates the effectiveness of the algorithm's reliability assessment. A typical LA-ICP-MS signal, such as... Figure 4 As shown in Figure a, the signal consists of a background signal, a sample signal, and an intermediate transition signal range. Using the signal extraction function, the background and sample signals are accurately distinguished, achieving a 99% accuracy rate in identifying them. Figure 4(b) Only a very small number of transition signals are misclassified because the signal transition is inherently continuous and gradual, resulting in unclear boundary regions. This ambiguity often leads to discrepancies even in manual annotation due to subjective interpretation.
[0065] Furthermore, the mass spectrometry signal patterns generated by laser ablation can be used to determine the elemental distribution within a sample. A normal elemental distribution produces a steadily decreasing elemental signal pattern. If the elemental content is too low, the resulting signal is close to the gas background signal. In geological samples, tiny mineral inclusions are common. When these inclusions are ablated during laser ablation, regular perturbation signals are generated; for example, the signals of certain trace elements may suddenly rise sharply before falling back to normal levels. Accurately identifying the type of elemental signal helps researchers determine the distribution of that element and provides important data support for interpreting the geological significance of elemental content. Utilizing the signal classification function of machine learning models, normal signals and no-signal conditions can be accurately distinguished, achieving an accuracy of 98% in identifying normal and no-signal data, sufficient to replace manual operation. Figure 4 c). A very small number of misclassifications were observed, primarily occurring when the signal intensity was close to the background level. Even experienced analysts are prone to misclassification in such cases. Therefore, elements classified as inclusion signals require manual verification.
[0066] When optimizing the elemental quantification function of the machine learning model, it was found that simply inputting elemental signals into the model could not obtain accurate prediction data. Two important factors strongly affected the performance of the model's data prediction. The first is signal preprocessing. Taking MgO content prediction as an example, the raw MgO signal and label were input into the model, and 10-fold cross-validation was performed, using the relative deviation between the predicted value and the label value as the evaluation criterion. As shown in Figure 6a, for data with MgO concentrations higher than 5.0%, the relative deviation fluctuated between -7% ± 25% (1SD), with only a few outliers. However, for data with MgO concentrations lower than 5.0%, the relative deviation was 87% ± 69% (1SD), indicating that the reliability of the prediction dropped sharply for low MgO concentration data. In traditional ICP-MS data processing, elemental ratios are more stable than individual elements because ratios can eliminate signal interference from ICP sources and changes in signal sensitivity over time. We chose Si as the normalization element because the training data mainly consisted of silicate rocks and silicate minerals.
[0067] Furthermore, PowerTransformer and Standardscale were used to correct skew and normalize. Results show that the preprocessing method combining PowerTransformer transformation with Si normalization improves the accuracy of MgO concentration prediction. For MgO concentrations above 1.0%, the relative bias is tightly clustered around 0.7% ± 12% (1SD), while for MgO concentrations between 0.1% and 1.0%, the relative bias remains relatively large, at 27% ± 36% (1SD) (Fig. 5c). This may be due to the wide range of MgO concentrations in the data, spanning three orders of magnitude from 0.1% to 52%. To address this issue, we divided the training data into three concentration ranges for independent parameter optimization. Using this strategy, the relative bias for low-concentration MgO reached 1.7% ± 10% (1SD) (Fig. 5d), while maintaining good data quality in the high-MgO region (relative bias = 1.7 ± 8.3%). Therefore, this invention proposes that the optimal method for preprocessing LA-ICP-MS raw data is "data grouping + PowerTransformer + Si normalization".
[0068] The second influencing factor is what data should be input when building a quantitative model for an element. For example, when building a model for MgO, models optimized by inputting only MgO data and label data often produce relatively discrete prediction results. Figure 6 a) We believe this is due to the decoupling of the linear relationship between signal and concentration caused by changes in instrument status during LA-ICP-MS analysis. This change in instrument status is not simply due to sensitivity drift within a single analysis in the same laboratory, but rather to variations in instrument status across different laboratories, operators, and testing times. If this impact of status variation cannot be addressed, the existing machine learning model can only be applied to element prediction of test data generated simultaneously with the training data, and cannot be extended to data generated under other instrument conditions, severely limiting the model's application.
[0069] When LA-ICP-MS simultaneously acquires multi-element content information, an inherent relationship exists between elemental signals. For example, elemental sensitivity increases or decreases with changes in ablation yield. During ionization in inductively coupled plasma, processes such as ionization temperature, ion transport, and ion loss simultaneously affect all elements. This inherent relationship is preserved in the elemental signals even under different experimental conditions. Therefore, we propose using multi-element data to construct quantitative models for single elements. For example, when constructing a quantitative model for MgO, signals from Mg, Na, and Fe are selected as inputs to the model to optimize its parameters. Figure 6As shown in b, the prediction accuracy for MgO is improved to 0.1% ± 6.0% (1SD). If Al, Fe, Na, K, Ca, and Ti, commonly used silicate minerals, are used, the prediction accuracy for MgO can be further improved to -0.1% ± 3.5% (1SD). Figure 6 c). Therefore, by using a quantitative model that combines principal elements with target elements for joint optimization, the prediction accuracy can be improved by 2-3 times, and the problem of a small number of discrete data points can be completely solved.
[0070] In this study, each element had its own quantitative model, resulting in a quantitative matrix of 39 models for 9 major elements and 30 trace elements. The relative deviations for the 9 major elements ranged from 1.8% to 7.7%, while the relative deviations for the trace elements ranged from 5.2% to 11.6%. This meets the quantitative requirements for elements in actual silicate geological samples.
[0071] The machine learning model developed in this study automatically processed LA-ICP-MS elemental quantitative data from three different laboratories. All three laboratories used silicate glass standards from the U.S. Geological Survey: BCR-2G, BHVO-2G, and BIR-1G. These standard data were processed using the traditional internal standard correction method, with silicon as the internal standard element. By comparison, the model's predicted values for major and trace elements deviated from the internal standard correction method by -2.1% ± 5.8% (1 SD) and 5.0% ± 11.7% (1 SD), respectively. The data show that the model's predicted values are within the conventional analytical accuracy range of LA-ICP-MS and are largely consistent with the data obtained using the internal standard correction method. Furthermore, the model can be applied to LA-ICP-MS data from different laboratories, demonstrating its good accuracy and stability. Figure 7 As shown, the results of using this method to predict standard samples are highly consistent with those of the traditional internal standard method, verifying the accuracy and reliability of this method.
[0072] In this embodiment, the data used for training is primarily silicate samples, therefore silicon is chosen for normalization. For other types of minerals, such as carbonate minerals, raw carbonate mineral data can be used for training, with Ca selected as the normalization element. In summary, the normalization element in the data preprocessing strategy can be selected based on the sample type.
[0073] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.
[0074] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.
[0075] For each patent, patent application, patent application publication, and other material such as articles, books, specifications, publications, and documents referenced in this specification, the entire contents of which are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this specification, as well as documents that limit the broadest scope of the claims in this specification (currently or subsequently appended to this specification). It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or terminology used in the supplementary materials to this specification and the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.
[0076] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.
Claims
1. A method for multi-element quantitative analysis of laser ablation inductively coupled plasma mass spectrometry, characterized by, The method comprises the following steps: obtaining LA-ICP-MS raw time-resolved mass spectrum data of a sample; inputting the raw data into a trained signal extraction model, which is a model constructed based on a one-dimensional convolutional neural network and is dedicated to identifying and separating background signals, sample signals and transition signals from LA-ICP-MS multi-channel time series, and outputting an effective sample signal interval; inputting the effective sample signal into a trained signal classification model, which is a model constructed based on a mixed architecture of a convolutional neural network and a long short-term memory network and is dedicated to classifying LA-ICP-MS signal patterns to identify normal signals, no signals and inclusion signals, and outputting classified normal sample signals; inputting the normal sample signals into a trained element quantitative prediction model, which is a model constructed based on the mixed architecture and taking target element signals and at least one other major element signal as combined input and is dedicated to directly predicting element content, and outputting predicted contents of multiple elements in the sample.
2. The method of claim 1, wherein, In step S4, before inputting the signal into the quantitative prediction model, the following preprocessing steps are further included: performing element ratio normalization processing on the signal; and performing PowerTransformer transformation on the normalized signal to correct data skewness.
3. The method according to claim 1 or 2, characterized in that, In step S4, the training method of the quantitative prediction model comprises: according to the concentration range of the target element in the training data, dividing it into multiple concentration intervals, and training an independent sub-model for each concentration interval.
4. The method of claim 1, wherein, In step S2, when the signal extraction model is trained, the input data is subjected to dimensionality reduction and feature compression processing based on Pearson correlation coefficient and sliding window standard deviation, so as to compress the original multi-channel data into input of fixed channel number.
5. The method of claim 1, wherein, In step S3, the input features of the signal classification model include the original time series signal and its derived features, and the derived features include at least one of the following: logarithmic form of the signal, sliding standard deviation, fast Fourier transform amplitude, signal total intensity ratio, peak relative background intensity and peak stability.
6. The method of claim 1, wherein, The method does not need to rely on internal standard elements or external standard materials for quantitative correction.
7. A multi-element quantitative analysis system for laser ablation inductively coupled plasma mass spectrometry, characterized by, The method comprises: a data processing module for receiving LA-ICP-MS raw time-resolved mass spectrum data; a model storage and execution module storing a signal extraction model, a signal classification model and a quantitative prediction model trained according to the method of any one of claims 1-6, and used for performing corresponding calculation; and a result output module for outputting processed effective signal interval, signal classification result and element quantitative prediction content.
8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the steps of the method of any one of claims 1-6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the steps of the method of any one of claims 1-6.