A method for constructing a chemical rapid LIBS classification and identification model, medium and system

By calibrating the spectrum and constructing a decoupled dual-flow conditional graph network for plasma parameters, the problem of unstable classification features in laser-induced breakdown spectroscopy under cross-instrument and cross-matrix conditions was solved, achieving higher recognition accuracy.

CN122310239BActive Publication Date: 2026-07-28INSPECTION & QUARANTINE TECH CENT SHANDONG ENTRY EXIT INSPECTION & QUARANTINE BUREAU
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSPECTION & QUARANTINE TECH CENT SHANDONG ENTRY EXIT INSPECTION & QUARANTINE BUREAU
Filing Date
2026-05-29
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

In existing technologies, laser-induced breakdown spectroscopy exhibits unstable classification characteristics under cross-instrument and cross-matrix conditions, leading to a significant decrease in classification accuracy.

Method used

By inserting a mercury-argon reference lamp to calibrate the spectrum, using a sparse autoencoder to compress the feature dimension, and combining the Boltzmann plot slope and Stark broadening to invert electron temperature and density, a plasma parameter decoupled two-flow conditional graph network is constructed. Combined with an adaptive plasma state-space iterative classification algorithm and Fredholm first-type integral equation regularized inversion spectral decomposition, real-time diagnosis and classification of element concentration vectors are achieved.

Benefits of technology

It achieves stability and accuracy of classification features under different instrument and substrate conditions, and improves the recognition capability under different conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122310239B_ABST
    Figure CN122310239B_ABST
Patent Text Reader

Abstract

The application provides a kind of management and control chemical fast LIBS classification identification model construction method, medium and system, belongs to management and control chemical technical field, the application is by collecting mercury argon reference lamp spectrum to carry out wavelength calibration, according to atomic spectrum database extraction management and control element characteristic transition wavelength window and compressed to low-dimensional feature space by sparse self-encoder, with hydrogen line starker broadening and Boltzmann graph slope method real-time inversion electron density and electron temperature and to compressed characteristic matrix is physically normalized, calculates plasma drift index to adaptively adjust plasma parameter decoupling double-flow condition graph network update strategy, the normalized characteristic matrix and physical diagnostic value are input into the artificial intelligence network to output element concentration vector and class prediction probability, solve the technical problems that plasma parameter fluctuation leads to laser-induced breakdown spectroscopy classification feature instability, classification accuracy significantly decreases under the condition of cross-instrument and cross-matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of controlled chemicals technology, and specifically relates to a method, medium and system for constructing a rapid LIBS classification and identification model for controlled chemicals. Background Technology

[0002] Laser-induced breakdown spectroscopy (LAS) is an in-situ detection technique that uses a high-power pulsed laser focused on the sample surface to induce plasma emission of characteristic spectra, thereby enabling qualitative and quantitative analysis of elemental composition. It has been widely applied in scenarios such as rapid on-site screening of controlled chemicals, port inspection of explosive precursors, and online monitoring of hazardous chemicals. In these applications, traditional methods typically use the original spectral line intensity or peak area as classification features, combining principal component analysis, linear discriminant analysis, or support vector machines (SVMs) to construct classification models.

[0003] However, in existing technologies, due to successive fluctuations in laser pulse energy, differences in the physical morphology of the sample matrix, and variations in ambient temperature and humidity, the electron temperature and electron density of the plasma drift significantly between different measurement batches, causing the absolute intensity of the spectral lines of the same element to change nonlinearly with the plasma state. This nonlinear change renders the classification features constructed based on the original spectral line intensities untransferable across instruments, matrices, and environments, and the classification boundaries learned by the model on the training set quickly become invalid in actual deployment scenarios. In other words, existing technologies suffer from the technical problem of unstable laser-induced breakdown spectral classification features due to fluctuations in plasma parameters, and a significant decrease in classification accuracy across instruments and matrices. Summary of the Invention

[0004] In view of this, the present invention provides a method, medium and system for constructing a rapid LIBS classification and identification model for controlled chemicals, which can solve the technical problems in the prior art where fluctuations in plasma parameters lead to unstable laser-induced breakdown spectral classification characteristics and a significant decrease in classification accuracy under cross-instrument and cross-matrix conditions.

[0005] The present invention is implemented as follows: The first aspect of the present invention provides a method for constructing a rapid LIBS classification and identification model for controlled chemicals, comprising the following steps:

[0006] Insert a mercury-argon reference lamp, acquire the reference lamp spectrum, update the pixel wavelength mapping polynomial, and output the original spectral dataset after wavelength calibration.

[0007] Based on the atomic spectral database, the characteristic transition wavelength windows of the controlled elements are extracted, and the feature dimension is compressed to within the target dimension by a sparse autoencoder, and the compressed feature matrix is ​​output.

[0008] The electron density is inverted using the hydrogen line Stark broadening, and the electron temperature is inverted using the Boltzmann plot slope. The compressed feature matrix is ​​then physically normalized, and the normalized feature matrix is ​​output.

[0009] Calculate the plasma drift index and adjust the update strategy of the plasma parameter decoupling dual-flow condition graph network according to the range to which the plasma drift index belongs;

[0010] The normalized feature matrix, electron temperature, and electron density are input into the plasma parameters and decoupled into the two-current conditional graph network. An adaptive plasma state-space iterative classification algorithm is executed synchronously, and the element concentration vector and category prediction probability are output.

[0011] Calculate the Mahalanobis distance between the element concentration vector and the standard concentration vector library. Based on whether the Mahalanobis distance exceeds the discrimination threshold, decide whether to enable the Fredholm first-type integral equation regularization inversion spectral decomposition classification algorithm and output the mixture composition classification results.

[0012] The three classification results are merged using a voting weighting fusion strategy to output the final classification label; if any two results are inconsistent, a dual-pulse laser-induced breakdown spectral mode is triggered, and the physical normalization to mixture component classification step is re-executed.

[0013] Specifically, the update of the pixel wavelength mapping polynomial involves acquiring the mercury-argon reference lamp spectrum before each laser trigger, fitting a third-order polynomial with the center pixel position of the known spectral line, writing the polynomial coefficients into the wavelength lookup table for the current measurement, and controlling the fitting error of the center pixel position within the fitting error threshold.

[0014] The sparse autoencoder employs a three-layer fully connected structure with input dimensions of 512, 256, and target dimensions. The activation function is a linear rectified function, the decoding layer is symmetrically expanded, and the training objective is the sum of reconstruction error and sparsity penalty.

[0015] Specifically, the Boltzmann plot slope inversion of electron temperature involves selecting several weak spectral lines within the same excited state multivariate and fitting a straight line with the excitation energy as the horizontal axis and the logarithm of the normalized intensity as the vertical axis. The slope is inversely proportional to the electron temperature. The selection criteria for weak spectral lines are that the signal-to-noise ratio is greater than the signal-to-noise ratio threshold and the peak intensity is lower than the intensity ratio threshold of the strongest spectral line of the same element.

[0016] Specifically, the hydrogen line Stark broadening inversion electron density is extracted... The full width at half maximum (FWHM) of the spectral line is subtracted from the instrument broadening to obtain the pure Stark broadening value. The electron density is then calculated by substituting this value into the Stark broadening coefficient lookup table. The instrument broadening is obtained by calibrating the measured FWHM of known narrow spectral lines from the mercury-argon reference lamp.

[0017] Specifically, the physical normalization involves modifying the intensity component of each spectral line in the compressed characteristic matrix by means of the electron temperature and electron density using the Boltzmann factor, thereby eliminating the influence of plasma parameter fluctuations on the absolute intensity of the spectral lines.

[0018] The plasma drift index is calculated by summing the normalized deviations of electron temperature and electron density relative to the weighted average of the indexes of the previous several frames. When the plasma drift index is greater than or equal to the high drift threshold, the adaptive plasma state space iterative classification algorithm is triggered to re-iterate the entire algorithm and the condition vector input is forcibly updated. When the plasma drift index is between the low drift threshold and the high drift threshold, only the linear layer weights corresponding to the condition vector are updated. When the plasma drift index is less than the low drift threshold, the cached spectral intensity ratio matrix is ​​used directly for classification.

[0019] The plasma parameter decoupled dual-flow conditional graph network consists of two streams: a first stream, the spectral graph construction stream, which treats each detected spectral line as a graph node, and a node feature vector consisting of center wavelength, peak intensity, full width at half maximum (FWHM), integral area, and asymmetry. The graph neural network backbone adopts a graph attention network structure, and the attention weights are calculated by concatenating electron temperature and electron density as conditional vectors and fusing them into the attention scoring function. The second stream is the global spectral flow, which inputs the normalized full spectrum after background subtraction into a one-dimensional convolutional neural network. The features of the two streams are fused through a cross-attention mechanism and then output as a category prediction probability by a classification head.

[0020] The adaptive plasma state-space iterative classification algorithm constructs a positive physical model by combining the Boltzmann population equation and the Saha ionization equilibrium equation. It synchronously inverts the electron temperature, electron density and relative concentration of each element through Newton iteration. The converged relative concentration of each element forms an element concentration vector, and the Mahalanobis distance is used to perform nearest neighbor classification in the element concentration vector space.

[0021] Specifically, the Fredholm first-class integral equation regularization inversion spectral decomposition classification algorithm constructs a kernel matrix using a pure component standard spectral library, iterates the sparse concentration vector using the alternating direction multiplier method with Tikhorov first-norm joint regularization, automatically selects the regularization parameter through generalized cross-validation, and inputs the non-zero support set of the sparse concentration vector and the corresponding amplitude into the decision tree classifier to output the mixture component classification result.

[0022] Specifically, the voting weight fusion strategy assigns weights based on the classification accuracy of each category on the validation set using the three-way algorithm. The category corresponding to the maximum value after summing the weighted probabilities of the three-way algorithm is taken as the final classification label. The weights are dynamically updated according to the validation results of each batch of new samples.

[0023] In the dual-pulse laser-induced breakdown spectral mode, the first pulse preheats the plasma to expand and reduce the effective optical depth, and the second pulse triggers acquisition within a delay time range, which is determined by experimental measurement of the self-absorption coefficient at different delay times.

[0024] The training dataset of the plasma parameter decoupled dual-flow conditional graph network collects controlled chemical samples covering three physical forms: powder, liquid, and solid bulk. Simultaneously, the electronic temperature and electronic density diagnostic values ​​are recorded as regression labels, and random Gaussian noise, random wavelength shift, and random baseline drift are applied to the original spectrum for data augmentation.

[0025] The discrimination threshold of the Mahalanobis distance is determined by statistically analyzing the Mahalanobis distance distributions of known pure substances and known mixtures, taking the critical value that best separates the two distributions, and then analyzing the receiver operating characteristic curve. When the number of training set samples is less than the dimension of the element concentration vector, ridge estimation is used to ensure that the covariance matrix is ​​invertible.

[0026] The fitting error threshold is 0.02. The target dimension is 100; the signal-to-noise ratio threshold is 10; the intensity ratio threshold is 30%; the high drift threshold is 0.15; the low drift threshold is 0.05; the graph attention network has 4 layers, with 8 attention heads per layer; the one-dimensional convolutional neural network has 5 layers, with a kernel size of 7; the alternating direction multiplier method has 30 to 50 iteration steps.

[0027] A second aspect of the present invention provides a computer-readable storage medium storing program instructions, which, when executed in a computer, are used to perform the above-described method for constructing a rapid LIBS classification and identification model for controlled chemicals.

[0028] A third aspect of the present invention provides a rapid LIBS classification and identification model construction system for controlled chemicals, comprising the aforementioned computer-readable storage medium, wherein the system is a computer, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

[0029] This invention constructs a multi-algorithm collaborative classification framework driven by a physical model, embedding real-time diagnostic results of plasma electron temperature and electron density into the entire process of feature extraction and model inference, thereby eliminating the interference of plasma parameter fluctuations on classification features from the root.

[0030] This invention uses the Boltzmann plot slope method and the hydrogen line Stark broadening method to invert electron temperature and electron density in real time, and performs physical normalization of the compressed feature matrix using the Boltzmann factor, transforming the original spectral line intensity into a physical quantity decoupled from the plasma state. A plasma parameter decoupled dual-flow conditional graph network embeds these physical quantities as conditional vectors into the attention scoring function, enabling adaptive redistribution of the intensity ratio between spectral lines of the same element under different matrix and instrument conditions at the feature layer. An adaptive plasma state-space iterative classification algorithm constructs a positive physical model by simultaneously applying the Boltzmann population equation and the Saha ionization equilibrium equation, replacing the classification features with element concentration vectors that are robust across instruments. In summary, this invention solves the technical problems mentioned in the background art, such as the instability of laser-induced breakdown spectral classification features caused by plasma parameter fluctuations and the significant decrease in classification accuracy under cross-instrument and cross-matrix conditions. Attached Figure Description

[0031] Figure 1 This is a flowchart of the method of the present invention.

[0032] Figure 2 The image shows frame-by-frame diagnostic curves of electron temperature and electron density for a typical batch.

[0033] Figure 3 Scatter plot of the spatial distribution characteristics of the final layer of the global spectrum of 7 types of controlled chemicals.

[0034] Figure 4 For single-pulse and double-pulse modes Curve showing the change of the spectral self-absorption coefficient with time delay. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below.

[0036] like Figure 1 The diagram shown is a flowchart of a method for constructing a rapid LIBS classification and identification model for controlled chemicals, provided by the first aspect of this invention. This method includes the following steps:

[0037] S01. Insert a mercury-argon reference lamp, acquire the reference lamp spectrum, update the pixel wavelength mapping polynomial, and output the original spectral dataset after wavelength calibration.

[0038] S02. Extract the characteristic transition wavelength windows of controlled elements based on the National Atomic Spectrum Database, compress the feature dimension to less than 100 dimensions using a sparse autoencoder, and output the compressed feature matrix.

[0039] S03. Invert the electron density using the hydrogen line Stark broadening and the electron temperature using the Boltzmann plot slope. Perform physical normalization on the compressed feature matrix and output the normalized feature matrix.

[0040] S04. Calculate the plasma drift index and adjust the update strategy of the plasma parameter decoupling dual-flow condition graph network according to the range to which the plasma drift index belongs.

[0041] S05. Input the normalized feature matrix, electron temperature and electron density into the plasma parameters and decouple the two-flow conditional graph network. Simultaneously execute the adaptive plasma state-space iterative classification algorithm and output the element concentration vector and the category prediction probability.

[0042] S06. Calculate the Mahalanobis distance between the element concentration vector and the standard concentration vector library. Based on whether the Mahalanobis distance exceeds the discrimination threshold, decide whether to enable the Fredholm first type integral equation regularization inversion spectral decomposition classification algorithm and output the mixture component classification results.

[0043] S07. Merge the three classification results using a voting weight fusion strategy and output the final classification label. If any two results are inconsistent, trigger the dual-pulse laser-induced breakdown spectral mode and re-execute S03 to S06.

[0044] The update method for the pixel wavelength mapping polynomial is as follows: before each laser trigger, the spectrum of the mercury-argon reference lamp is acquired, and a third-order polynomial is fitted with the center pixel position of the known spectral line. The polynomial coefficients are written into the wavelength lookup table for the current measurement. The fitting error of the center pixel position is controlled within 0.02 nanometers. The 0.02 nanometer control boundary is determined by repeatedly acquiring the mercury-argon reference lamp spectrum under different temperature gradients and vibration amplitudes, statistically analyzing the drift distribution, taking the 95th quantile of the distribution, and verifying it through iterative experiments.

[0045] The National Atomic Spectrum Database is the atomic spectrum database released by the Beijing Institute of Applied Physics and Computational Mathematics; when the database lacks data on corresponding controlled elements, the atomic spectrum database of the National Institute of Standards and Technology (NIST) in the United States is used as a supplementary reference.

[0046] The sparse autoencoder employs a three-layer fully connected structure with input dimensions of 512, 256, and 100, with a linear rectified function as the activation function. The decoding layer is symmetrically expanded, and the training objective is the sum of reconstruction error and sparsity penalty. The 100-dimensional compression objective is determined by enumerating 50 to 300 dimensions on controlled product datasets with different numbers of elements, and using leave-one-out cross-validation to determine the inflection point value of classification accuracy.

[0047] The Boltzmann plot slope method is implemented as follows: within the same excited-state multivariate, 3 to 5 weak spectral lines are selected, and a straight line is fitted with the excitation energy as the horizontal axis and the logarithm of the normalized intensity as the vertical axis. The slope is inversely proportional to the electron temperature, as expressed by the following formula: ;in This is a standard electronic temperature reference value. Boltzmann's constant, To fit the slope of the straight line, The weak spectral lines are determined by taking the average of 100 repeated measurements under stable laboratory conditions. The screening criteria for weak spectral lines are a signal-to-noise ratio greater than 10 and a peak intensity lower than 30% of the intensity of the strongest spectral line of the same element. The 30% threshold is determined by repeatedly testing standard samples of known concentrations to minimize the electron temperature inversion error.

[0048] The step of obtaining the electron density from the hydrogen line Stark broadening inversion is as follows: extracting... The full width at half maximum (FWHM) of the spectral line is subtracted from the instrumental broadening to obtain the pure Stark broadening value. The electron density is then calculated by substituting this value into a Stark broadening coefficient lookup table, as shown in the following formula: ;in This is the standard electron density reference value. For pure Stark stretch value, The reference value is widened to the standard Stark. and The value was determined by averaging 100 repeated measurements under stable laboratory conditions; the instrument broadening was obtained by calibrating the measured full width at half maximum (FWHM) of a known narrow spectral line from a mercury-argon reference lamp.

[0049] The physical normalization step involves: using electron temperature and electron density to correct each spectral line intensity component in the compressed feature matrix component by Boltzmann factor, eliminating the influence of plasma parameter fluctuations on the absolute intensity of the spectral lines, and outputting the normalized feature matrix.

[0050] The formula for calculating the plasma drift index is as follows: ;in and This is the exponentially weighted average of electron temperature and electron density from the first 20 frames. and The meaning is the same as above; when When this occurs, the adaptive plasma state-space iterative classification algorithm is triggered to undergo a full re-iteration, and the condition vector input of the decoupled two-stream conditional graph network is forcibly updated; when When, only the linear layer weights corresponding to the condition vectors in the decoupled two-flow conditional graph network of plasma parameters are updated, while the parameters of the remaining layers are kept frozen; when At that time, the current model parameters remain unchanged, and the cached spectral intensity ratio matrix is ​​used directly for classification; the thresholds of 0.05 and 0.15 are determined by enumerating threshold combinations and taking the combination with the smallest accuracy loss in a control experiment with artificially applied plasma perturbations of different amplitudes, with the decrease in classification accuracy as the objective function, and after 5-fold cross-validation.

[0051] The specific structure of the plasma parameter decoupling dual-flow conditional graph network is as follows: the first flow is the spectral graph construction flow, where each detected spectral line is considered a graph node, and the node feature vector consists of five components: center wavelength, peak intensity, full width at half maximum (FWHM), integral area, and asymmetry. Strong edges are connected between spectral lines of the same element, and the edge weights are calculated by normalizing the inverse of the corresponding energy level difference. Weak edges are connected between spectral lines of different elements, and the edge weights are calculated by statistically normalizing the co-occurrence frequencies in the training set. The graph neural network backbone adopts a graph attention network structure with four layers, each with eight attention heads. Attention weights are calculated... During computation, electron temperature and electron density are concatenated into a conditional vector, which is then fused to the attention scoring function through a linear layer. Node embeddings from each graph attention network layer are concatenated via cross-layer skip connections and fed into the graph readout layer, where global summation pooling is used to obtain the graph-level representation vector. Neuron activation values ​​from each layer of the graph attention network are distributed by channel to eight parallel attention heads after batch normalization. The key-value dimension of each attention head is the node embedding dimension divided by the number of heads. The second stream is a global spectral stream, which inputs the normalized full spectrum after background subtraction into a one-dimensional convolutional neural network with five convolutional layers and a kernel size of 7. A filter is used. The number of parameters increases from 32 to 256, capturing broadband structural features such as molecular band emission and continuous background residue. During inference, the convolutional kernels of each layer of the one-dimensional convolutional neural network are allocated using CUDA streams for asynchronous layer-by-layer computation, running in parallel with the message passing computation of the graph attention network. The two stream features are fused after the graph readout layer via a cross-attention mechanism, using the graph stream output as the query vector and the global spectral stream output as the key and value vectors. The fused feature vector is obtained by calculating the cross-attention weights and then weighted summing them. The attention matrix of the cross-attention layer is calculated in blocks according to the query sequence length, with the block size determined by the current... The available memory is dynamically determined; the fused feature vectors are processed through two fully connected classification heads to output the predicted class probabilities, while two independent regression heads output the predicted electron temperature and electron density values ​​respectively; the weight matrices of the fully connected layers of the regression and classification heads are stored in memory as half-precision floating-point numbers and converted to single-precision numbers for calculation during forward propagation; training uses a multi-task loss function, with the main task being the cross-entropy classification loss and the auxiliary tasks being the mean square error regression loss for electron temperature and electron density. The weights of the two auxiliary losses are each between 0.1 and 0.3, and the weight range is determined by grid search on the validation set;

[0052] The steps for establishing the training dataset for the plasma parameter decoupled dual-flow conditional graph network specifically include: collecting controlled chemical samples covering three physical forms: powder, liquid, and solid bulk material; collecting no less than 200 spectra for each physical form and each type of sample; simultaneously recording electron temperature and electron density diagnostic values ​​as regression labels; supplementing data for the corresponding physical form when the variance of classification accuracy under different physical forms exceeds an acceptable threshold; applying random Gaussian noise, random wavelength shift of 0.01 nm to 0.05 nm, and random baseline drift to the original spectra for data augmentation, with the shift range determined by the 90% confidence interval of the measured instrument drift distribution; and dividing the training set, validation set, and test set into an 8:1:1 ratio.

[0053] The specific steps for training the plasma parameter decoupled dual-flow conditional graph network include: using an adaptive moment estimation optimization algorithm with an initial learning rate of 0.001; multiplying the learning rate by 0.5 when the classification loss on the validation set does not decrease for 5 consecutive rounds; setting the batch size to 32; terminating training when the classification accuracy on the validation set does not improve for 10 consecutive rounds; and evaluating the final performance using a test set after training.

[0054] The technical effects of the plasma parameter decoupling dual-stream conditional graph network are as follows: by embedding electron temperature and electron density as conditional information into the attention scoring function of the graph attention network, the attention weights are adaptively redistributed according to changes in plasma state, so that the intensity ratio relationship between the same element spectral lines under different physical states and different matrix conditions is physically decoupled at the feature layer; dual-stream fusion makes the element-level local spectral line features and global spectral shape features complement each other; the cross-attention mechanism enables broadband emission structures in the global spectral shape that are difficult to be individually attributed to enhance the embedding representation of the corresponding element nodes; multi-task training forces the intermediate layers of the network to encode physically meaningful plasma state information, improving the transfer robustness of features under different instrument and measurement conditions; cross-layer skip connections retain shallow local spectral line structure information and avoid the loss of detailed features during deep propagation.

[0055] The principle and implementation of the adaptive plasma state-space iterative classification algorithm are as follows: The algorithm constructs a forward physical model between electron temperature, electron density, and observed spectral line intensity by simultaneously applying the Boltzmann population equation and the Saha ionization equilibrium equation. Then, it iterates and inverts to simultaneously estimate the electron temperature, electron density, and relative concentrations of each element. The first step is to give the initial electron temperature using the Boltzmann plot slope method. The second step is to substitute the initial electron temperature into the Saha ionization equilibrium equation, iteratively solve for the electron density, and establish the theoretical spectral line intensity ratio matrix for each element's valence state. The third step is to construct a least-squares linear system using the theoretical spectral line intensity ratio matrix, solve for the relative concentrations of each element, and calculate the residual vector of the difference between the observed and theoretical intensities. The fourth step is to update the electron temperature and electron density using the gradient of the residual vector according to the Newton iteration format, return to the second step, and repeat the iteration until the L2 norm of the residual is less than the convergence threshold. The convergence threshold is determined by iteratively testing standard samples of known concentrations until the residual no longer decreases. In the fifth step, the relative concentrations of each element after convergence are used to construct an element concentration vector. Nearest neighbor classification is performed in the element concentration vector space using Mahalanobis distance, and the predicted probability of each category is output. The algorithm replaces the original spectral line intensity with a physically meaningful element concentration vector, eliminating inter-class variations caused by plasma parameter fluctuations and making the classification features robust across instruments. The Saha ionization equilibrium equation correlates electron temperature with the number density of particles in the neutral and primary ionized states of each element. After simultaneously solving the Boltzmann population equation, a closed set of equations is formed regarding electron temperature, electron density, and the relative concentrations of each element. Newton iteration converges rapidly under the guidance of the Jacobian matrix of this closed set of equations, enabling simultaneous estimation of physical parameters and analysis of chemical composition, and providing self-consistent compensation for matrix effects and changes in plasma conditions.

[0056] The principle and implementation of the spectral graph Laplacian manifold regularized sparse discriminative projection algorithm are as follows: A 7-nearest neighbor graph is constructed using all training spectral samples, with edge weights calculated using spectral cosine similarity; an intra-class graph is constructed using edges between samples of the same class, and an inter-class graph is constructed using edges between samples of different classes, with the intra-class Laplacian matrix and inter-class Laplacian matrix calculated respectively; the solution of the projection matrix is ​​transformed into a generalized eigenvalue problem with graph manifold regularization terms and norm sparse constraints, solved by alternating direction multiplier splitting, iterating for 50 to 100 steps to converge; after projection, classification is performed using a support vector machine in a low-dimensional discriminative space, outputting the predicted class probability; the nearest neighbor number 7 is evaluated by assessing the ratio of intra-class divergence to inter-class divergence after discriminative projection under different nearest neighbor number values, and the ratio is taken. The minimum nearest neighbor number is determined by 5-fold cross-validation. The physical starting point of the algorithm is that laser-induced breakdown spectral data are distributed in a high-dimensional space on a low-dimensional manifold determined by physical laws. Graph Laplacian regularization explicitly models the structure of the low-dimensional manifold, ensuring that physically adjacent samples remain close after low-dimensional projection. Sparse constraints automatically select feature wavelengths carrying discriminative information, making the non-zero rows of the projection matrix correspond to the physically meaningful element feature transition wavelengths, significantly improving generalization ability under the condition of scarce labeled samples. Under the constraint of physical manifold, the algorithm ensures that the construction of classification boundaries does not depend on a large number of labeled samples, because manifold regularization utilizes the geometric information of unlabeled samples, and the transfer performance is guaranteed under new matrix morphology or new instrument conditions.

[0057] The discrimination threshold of the Mahalanobis distance is determined by statistically analyzing the Mahalanobis distance distributions of known pure substances and known mixtures, taking the critical value that best separates the two distributions, and then analyzing the receiver operating characteristic curve. The calculation of the Mahalanobis distance depends on the covariance matrix, which is obtained by statistically analyzing the element concentration vectors of each category of samples in the training set. When the number of samples in the training set is less than the dimension of the element concentration vector, ridge estimation is used to ensure that the covariance matrix is ​​invertible, and the ridge parameter is determined by leave-one-out cross-validation.

[0058] The implementation of the Fredholm first-type integral equation regularized inversion spectral decomposition and classification algorithm is as follows: a kernel matrix is ​​constructed using a standard spectral library of pure components, and the kernel matrix contains the nonlinear response of matrix effects; the sparse concentration vector is solved by 30 to 50 iterations using the Tikhorov first-norm joint regularization and the alternating direction multiplier method; the regularization parameter is automatically selected through generalized cross-validation; the non-zero support set and corresponding amplitude of the sparse concentration vector are used as the classification features of the mixture components, which are input into a decision tree classifier, and the classification results of the mixture components are output; the algorithm unifies physical decomposition and chemical classification within the same inversion framework, and the classification features are directly the relative content of each component, with clear physical meaning; the introduction of Tikhorov sparse regularization makes the inversion process of the ill-conditioned kernel matrix numerically stable, ensuring the reliability of mixture component identification.

[0059] The voting weight fusion strategy is as follows: the class prediction probability output by the plasma parameter decoupled dual-flow conditional graph network, the class prediction probability output by the adaptive plasma state space iterative classification algorithm, and the class prediction probability output by the spectral graph Laplace manifold regularized sparse discriminative projection algorithm are used as three inputs. The classification accuracy of each class on the validation set of the three algorithms is used as the basis for assigning weights. The class corresponding to the maximum value after summing the three weighted probabilities is taken as the final classification label. The weights are dynamically updated with the validation results of each batch of new samples. The update step size is the maximum step size value that makes the weighted fusion accuracy of the validation set monotonically converge. The maximum step size value is determined by enumerating step size candidate values ​​and observing convergence.

[0060] In the dual-pulse laser-induced breakdown spectral mode, the first pulse preheats the plasma to expand, reducing the effective optical depth and suppressing the self-absorption effect; the second pulse triggers acquisition within a delay time range; the delay time range is determined by experimentally measuring the self-absorption coefficient at different delay times, taking the delay time interval corresponding to the self-absorption coefficient dropping below 20% under single-pulse conditions; the 20% is determined by acquiring spectra of known high-concentration standard samples under single-pulse and dual-pulse conditions respectively, verifying the temperature diagnosis error using the Boltzmann plot slope method, and taking the proportion of the self-absorption coefficient corresponding to the error being below an acceptable level.

[0061] The spectral cosine similarity is the dot product of two spectral vectors divided by the product of their respective L2 norms. It reflects the degree of consistency of the directions of the two spectra in high-dimensional space, is independent of the absolute value of the intensity, and is robust to baseline drift and laser energy fluctuations.

[0062] The cross-layer skip connection involves concatenating the node embedding vectors output by each layer of the graph attention network along the feature dimension and then directly feeding them into subsequent layers, so that the shallow local spectral structure information bypasses the intermediate layers and is directly transmitted to the graph readout layer.

[0063] The global summation pooling involves summing the embedding vectors of all nodes in the graph according to their components to obtain a graph-level representation vector of fixed dimensions, which is independent of the number of nodes in the graph.

[0064] The first-norm sparsity constraint involves summing the absolute values ​​of each component of the projection matrix or concentration vector and limiting the summation value to no more than a preset upper limit, thereby forcing a large number of components in the solution vector to approach zero and achieving sparse selection.

[0065] The specific implementation of step S01 is as follows: Before each laser trigger, a mercury-argon reference lamp is inserted into the optical path, and its emission spectrum is collected. Using the center pixel position of a known spectral line of the mercury-argon reference lamp as the control point, a third-order polynomial is used to perform least-squares fitting of the mapping relationship between pixel coordinates and wavelength. The coefficients of the fitted polynomial are written into the wavelength lookup table for the current measurement, thereby completing the dynamic calibration from pixel to wavelength. The fitting error is controlled within 0.02. Within this range, the boundary was determined through iterative experiments after repeatedly acquiring the spectra of the mercury-argon reference lamp under different temperature gradients and vibration amplitudes, statistically analyzing the drift distribution, and taking the 95th percentile. Following the above calibration process, the original wavelength-calibrated spectral dataset is output, eliminating wavelength shifts caused by instrument thermal drift and mechanical vibration, providing an accurate wavelength reference for subsequent feature extraction.

[0066] The specific implementation of step S02 is as follows: Based on the atomic spectral database, primarily the atomic spectral database released by the Beijing Institute of Applied Physics and Computational Mathematics, supplemented by the atomic spectral database of the National Institute of Standards and Technology (NIST) when data is missing, characteristic transition wavelength windows of each controlled element are extracted to form an initial high-dimensional spectral feature vector. Subsequently, a sparse autoencoder is used to compress the feature dimension. The encoding layer adopts a three-layer fully connected structure with input dimensions of 512, 256, and 100, with a linear rectified function as the activation function. The decoding layer is symmetrically expanded, and the training objective is the sum of reconstruction error and sparsity penalty. The target compression dimension of 100 is determined by enumerating 50 to 300 dimensions on controlled product datasets with different numbers of elements, and using leave-one-out cross-validation to determine the inflection point value of classification accuracy. Through the above process, a compressed feature matrix of less than 100 dimensions is output, which significantly reduces subsequent computational overhead while retaining key spectral line discrimination information.

[0067] The specific implementation of step S03 is as follows: First, perform hydrogen line Stark broadening inversion to determine electron density. Extract... The pure Stark broadening value is obtained by subtracting the instrument broadening obtained from the measured half-width at half-maximum (FWHM) of the spectral lines calibrated using the known narrow spectral lines from the mercury-argon reference lamp. Substitute into the Stark broadening factor lookup table, and follow the formula. Calculate the electron density, where and The electron temperature was determined by averaging 100 repeated measurements under stable laboratory conditions. Subsequently, the electron temperature was retrieved using the Boltzmann plot slope method. Within the same excited-state multivariate, 3 to 5 weak spectral lines were selected, with screening criteria including a signal-to-noise ratio greater than 10 and peak intensity less than 30% of the strongest spectral line intensity of the same element. A straight line was fitted with the excitation energy as the horizontal axis and the logarithm of the normalized intensity as the vertical axis. The electron temperature was calculated based on the inverse relationship between the slope and the electron temperature. Standard electronic temperature reference value Similarly, it is determined by taking the average of 100 repeated measurements under stable conditions. and Then, the intensity component of each spectral line in the compressed characteristic matrix is ​​corrected component by component using the Boltzmann factor to eliminate the modulation of the absolute intensity of the spectral line by the fluctuation of plasma parameters, and the normalized characteristic matrix is ​​output.

[0068] The specific implementation of step S04 is as follows: the exponentially weighted average of the electron temperature and electron density of the previous 20 frames. and Based on the formula Calculating the plasma drift index .in accordance with The update strategy for the decoupled two-flow conditional graph network, which adjusts plasma parameters in three levels within the specified interval, is as follows: When this happens, a full re-iteration is triggered and the conditional vector input is forcibly updated; when When the condition vector is updated, only the linear layer weights are updated, while the parameters of the other layers remain frozen; when At this time, keeping the current model parameters unchanged, classification is performed directly using the cached spectral line intensity ratio matrix. The above thresholds were determined through 5-fold cross-validation by enumerating threshold combinations and using the decrease in classification accuracy as the objective function in a control experiment with artificially applied plasma perturbations of different amplitudes.

[0069] The specific implementation of step S05 is as follows: The plasma parameter decoupling dual-flow conditional graph network consists of a spectral graph construction flow and a global spectral shape flow in parallel. The spectral graph construction flow treats each detected spectral line as a graph node, and the node feature vector consists of five components: center wavelength, peak intensity, full width at half maximum (FWHM), integral area, and asymmetry. Strong edges are connected between spectral lines of the same element, and weak edges are connected between spectral lines of different elements. The graph attention network backbone has four layers, with eight attention heads in each layer. The attention scoring function incorporates a conditional vector concatenated from electron temperature and electron density. The node embeddings in each layer are concatenated through cross-layer skip connections and then global summation pooling to obtain the graph-level representation vector. The global spectral shape flow inputs the normalized full spectrum after background subtraction into a five-layer one-dimensional convolutional neural network with a kernel size of 7 and the number of filters increasing from 32 to 256. The two feature streams are fused through a cross-attention mechanism and output as category prediction probabilities by a classification head, and as predicted electron temperature and electron density values ​​by an independent regression head. Meanwhile, the adaptive plasma state-space iterative classification algorithm establishes a positive physical model by combining the Boltzmann population equation and the Saha ionization equilibrium equation. It then goes through four steps in sequence: initializing the Boltzmann diagram, solving for the electron density using the Saha equation, solving for the relative element concentration using least squares, and updating the parameters using Newton iteration. Once the residual L2 norm is lower than the convergence threshold, the relative element concentration is used to construct the element concentration vector, and the nearest neighbor classification using Mahalanobis distance is used to output the predicted probability of the category.

[0070] The specific implementation of step S06 is as follows: The covariance matrix of the element concentration vectors of each category of samples in the training set is statistically analyzed. When the number of samples is less than the vector dimension, ridge estimation is used to ensure invertibility. The Mahalanobis distance between the current element concentration vector and the center of each category in the standard concentration vector library is calculated. The discrimination threshold is determined by statistically analyzing the Mahalanobis distance distribution of known pure substances and known mixtures, taking the optimal critical value for distribution separation, and then analyzing the receiver operating characteristic curve. If the Mahalanobis distance exceeds the discrimination threshold, the Fredholm first-type integral equation regularized inversion spectral decomposition classification algorithm is activated: a kernel matrix containing matrix effect nonlinear response is constructed using the standard spectral library of pure components. Tikhorov first-norm joint regularization is applied, and the sparse concentration vector is solved through 30 to 50 iterations using the alternating direction multiplier method. The regularization parameter is automatically selected by generalized cross-validation. The non-zero support set and amplitude of the sparse concentration vector are input into the decision tree classifier, and the mixture component classification result is output.

[0071] The specific implementation of step S07 is as follows: The predicted class probabilities from the three outputs of the plasma parameter decoupled dual-flow conditional graph network, the adaptive plasma state-space iterative classification algorithm, and the Fredholm first-type integral equation regularized inversion spectral decomposition classification algorithm are used as inputs. Weights are assigned based on the classification accuracy of each algorithm on the validation set. The class corresponding to the maximum value after summing the weighted probabilities of the three paths is taken as the final classification label. The weights are dynamically updated according to the validation results of each batch of new samples, and the update step size is the maximum step size that makes the weighted fusion accuracy of the validation set monotonically converge. If the results of any two paths are inconsistent, a dual-pulse laser-induced breakdown spectral mode is triggered: the first pulse preheats the plasma to expand and reduce the effective optical depth to suppress the self-absorption effect; the second pulse triggers acquisition within the delay time range. The delay time range is determined by measuring the self-absorption coefficient at different delay times and taking the interval corresponding to the self-absorption coefficient dropping below 20% under single-pulse conditions. Then, steps S03 to S06 are re-executed.

[0072] It should be noted that the key technologies of this invention include three core technologies: a physical normalization mechanism, a plasma parameter decoupling dual-flow conditional graph network, and an adaptive plasma state-space iterative classification algorithm. These three technologies work closely together. The physical normalization mechanism corrects spectral line intensity component by component using Boltzmann factors, eliminating the modulation effect of plasma parameter fluctuations at the data level and decoupling the input features received by the subsequent classifier from the instrument state. The plasma parameter decoupling dual-flow conditional graph network embeds physical diagnostic values ​​into the attention scoring function, enabling the network to adaptively redistribute the correlation weights between spectral lines based on the current plasma state during inference. Dual-flow cross-attention fusion further compensates for the deficiency of single-flow structures in capturing broadband background features. The adaptive plasma state-space iterative classification algorithm replaces the classification features with element concentration vectors obtained by inverting a system of simultaneous physical equations, eliminating inter-class variations introduced by matrix effects and instrument differences at the physical mechanism level, and giving the classification features cross-instrument robustness. The three technologies apply physical constraints at the feature layer, network layer, and physical inversion layer, forming a multi-layered complementary decoupling system, enabling the overall classification framework to maintain a stable classification boundary even when plasma parameters drift over a wide range.

[0073] It should be noted that traditional spectral decomposition methods face serious ill-conditioned inversion problems when performing laser-induced breakdown spectral classification and identification of multi-component mixed controlled chemicals with strong spectral line overlap. The spectral line overlap occurs because the characteristic transition wavelengths of different elements overlap under the limited resolution of the spectrometer. This results in the observed spectral line intensities being a linear superposition or even nonlinear coupling of contributions from multiple components, making methods relying solely on spectral line peaks or areas for decoupling ineffective in separating the contributions of each component. The reason for this technical problem is that the plasma excitation temperatures of the components in the mixture are similar, leading to significant radiation from different components at adjacent wavelengths. The broadening effect of the instrument function further exacerbates the degree of spectral line overlap. Furthermore, the plasma matrix effect causes the proportional relationship between the spectral line intensities of each component to change nonlinearly with concentration, resulting in an extremely large condition number in the kernel matrix. This makes the inversion process extremely sensitive to measurement noise. The usual solutions to the aforementioned technical problems are linear methods such as least-squares spectral decomposition or non-negative matrix decomposition. However, these methods can produce numerically unstable solutions under ill-conditioned kernel matrix conditions, resulting in negative concentrations or large concentration oscillations in the inversion results. While introducing simple Tikhonov regularization can improve numerical stability, it lacks constraints on the sparse component structure of the mixture, making it difficult to distinguish between the contributions of real components and spurious peaks caused by noise. This invention effectively solves this technical problem. The Fredholm first-type integral equation regularized inversion spectral decomposition classification algorithm incorporates the nonlinear response of matrix effects into the kernel matrix construction process. It applies smoothness and sparsity constraints simultaneously using Tikhonov L1 regularization. The alternating direction multiplier method splits the originally coupled optimization problem into several subproblems that are solved alternately. Each subproblem has a closed-form solution, and the iterative process is numerically stable. The L1-norm sparsity constraint forces a large number of components in the obtained concentration vector to approach zero, automatically filtering out real mixed components and eliminating spurious peak interference. The regularization parameters are automatically selected by generalized cross-validation without the need for additional labeled data, avoiding subjective errors introduced by manual parameter tuning. The above mechanisms together ensure the numerical reliability and physical interpretability of the mixture component classification results under conditions of strong spectral line overlap and high noise.

[0074] A second aspect of the present invention provides a computer-readable storage medium storing program instructions, which, when executed in a computer, are used to perform the above-described method for constructing a rapid LIBS classification and identification model for controlled chemicals.

[0075] A third aspect of the present invention provides a rapid LIBS classification and identification model construction system for controlled chemicals, comprising the aforementioned computer-readable storage medium. The system can be any one of a computer, a server, or a microcontroller. The computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.

[0076] Specifically, the principle of this invention is:

[0077] The fundamental reason why this invention can solve the above-mentioned technical problems is that it deeply couples the plasma physics model with the machine learning classifier, so that the classification feature is transformed from the original spectral line intensity that is sensitive to the plasma state into a physical decoupling factor that is insensitive to the plasma state.

[0078] Specifically, the reason why the intensity of the original spectral lines is sensitive to plasma parameters is that the radiation intensity of the spectral lines is determined by the Boltzmann population distribution. A small change in electron temperature causes an exponential change in the excited-state particle population distribution, leading to a nonlinear drift in the spectral line intensity. This invention first compresses the high-dimensional original spectrum into a low-dimensional feature space using a sparse autoencoder. Then, it performs component-by-component physical normalization of each feature component using the real-time inverted electron temperature and electron density according to the Boltzmann factor. This eliminates the modulation effect of plasma parameter fluctuations on the spectral line intensity at the mechanistic level, ensuring that the normalized feature matrix retains only information related to elemental composition.

[0079] Building upon this foundation, the plasma parameter decoupling dual-flow conditional graph network inputs the normalized feature matrix into the spectral graph to construct a flow. This flow is then modeled using a graph attention network to model the physical relationships between spectral lines of different elements. Simultaneously, electron temperature and electron density are concatenated into a conditional vector and incorporated into the attention scoring function, allowing the attention weights to adaptively adjust with the plasma state. This further decouples the influence of matrix effects at the network level. The global spectral flow captures background structural features such as broadband molecular emission in parallel, and the two feature streams are fused through a cross-attention mechanism to achieve complementarity between local spectral line features and global spectral features.

[0080] The adaptive plasma state-space iterative classification algorithm takes a different approach. It constructs a forward physical model by simultaneously solving the Boltzmann population equation and the Saha ionization equilibrium equation. Newton's iterative method synchronously inverts electron temperature, electron density, and the relative concentrations of each element, replacing the classification features with element concentration vectors that have clear physical meaning, thereby eliminating inter-class variations caused by changes in plasma conditions. The Fredholm first-type integral equation regularized inversion spectral decomposition classification algorithm is activated when the Mahalanobis distance exceeds the discrimination threshold, using Tikhorov sparse regularization to ensure the numerical stability of the mixture composition inversion. The three classification results are fused through dynamic weighted voting; when inconsistent, a double-pulse mode is triggered to suppress self-absorption effects, further improving the physical reliability of the classification. The synergistic effect of these multi-level physical decoupling mechanisms ensures that the classification features remain stable under different instruments, substrates, and environmental conditions, fundamentally solving the problem of decreased classification accuracy caused by fluctuations in plasma parameters.

[0081] The following provides a specific embodiment 1 of the present invention, and the specific implementation of each step in this embodiment 1 is described in detail below.

[0082] The specific implementation of step S01 is as follows: Before each laser trigger, the spectrum of the mercury-argon reference lamp is acquired, the center pixel position of the known spectral line is identified, and a third-order polynomial is fitted with the pixel number as the independent variable and the known wavelength as the dependent variable. The polynomial coefficients are written into the wavelength lookup table for the current measurement to complete wavelength calibration and output the original spectral dataset. The fitting error of the center pixel position is controlled within 0.02 nm. This control boundary is determined by repeatedly acquiring the mercury-argon reference lamp spectrum under different temperature gradients and vibration amplitudes, statistically analyzing the drift distribution, taking the 95th quantile of the distribution, and verifying iterative experiments.

[0083] The specific implementation of step S02 is as follows: Based on the National Atomic Spectrum Database, characteristic transition wavelength windows of controlled elements are extracted, and spectral intensity components are truncated within the corresponding wavelength range to form a high-dimensional original feature vector. The sparse autoencoder's encoding layer adopts a three-layer fully connected structure with input dimensions of 512, 256, and 100 dimensions. The activation function is a linear rectified function, the decoding layer is symmetrically expanded, and the training objective is the sum of reconstruction error and sparsity penalty, expressed by the following formula:

[0084] ;

[0085] In the formula, For training loss, The number of training samples, For the first 1 input feature vector For the corresponding reconstructed vector, for The 2-norm, This represents the number of hidden layer neurons. The target sparsity is empirically set to 0.05. For the first The average activation value of each hidden layer neuron in the training batch. The sparsity penalty weight has an empirical value of 3. It is a natural logarithm function. The reconstruction error term is... Normalization and sparsity penalty terms are based on target sparsity. Compared with actual activation value The cross-entropy form. The compressed target of 100 dimensions is determined by enumerating 50 to 300 dimensions on control product datasets with different numbers of elements, and taking the inflection point value of classification accuracy using leave-one-out cross-validation.

[0086] The specific implementation of step S03 is as follows: Electron temperature is inverted using the Boltzmann plot slope method. Within the same excited-state multivariate, 3 to 5 weak spectral lines (signal-to-noise ratio greater than 10 and peak intensity less than 30% of the intensity of the strongest spectral line of the same element) are selected. A straight line is fitted with the excitation energy as the horizontal axis and the logarithm of the normalized intensity as the vertical axis. The electron temperature formula is expressed as follows:

[0087] ;

[0088] In the formula, The temperature of the electron to be measured. This is a standard electronic temperature reference value, determined by averaging 100 repeated measurements under stable laboratory conditions. Boltzmann's constant, The slope of the straight line fitted to the Boltzmann plot is given, with dimensions in units of excitation energy divided by the normalized logarithm of intensity. (Take the excitation energy unit as) hour), and With identical dimensions, both sides of the equation are dimensionless ratios. Electron density is inverted using the Stark broadening of the hydrogen line to extract... The Stark broadening value is obtained by subtracting the instrument broadening (calibrated by the measured half-width of a known narrow spectral line from the mercury-argon reference lamp) from the full width at half maximum (FWHM) of the spectral line. The electron density formula is expressed as follows:

[0089] ;

[0090] In the formula, The electron density to be measured, This is the standard electron density reference value. This is the pure Stark broadening value after deducting instrument broadening. The reference value is widened to the standard Stark. and Both values ​​were determined by averaging 100 repeated measurements under stable laboratory conditions, and both sides of the equation are dimensionless ratios. Physical normalization is used... and Each spectral line intensity component in the compressed feature matrix is ​​corrected component by component using the Boltzmann factor to eliminate the influence of plasma parameter fluctuations on the absolute intensity of the spectral line, and a normalized feature matrix is ​​output.

[0091] The specific implementation of step S04 is as follows: Calculate the plasma drift index, and the formula is expressed as follows:

[0092] ;

[0093] In the formula, The plasma drift index is a dimensionless quantity. The exponentially weighted average of the electronic temperatures of the first 20 frames. The exponentially weighted average of the electron density of the first 20 frames. and The meaning is the same as above, with the two terms on the right side of the equation respectively... and Normalization results in a dimensionless quantity as a whole. When this occurs, the adaptive plasma state-space iterative classification algorithm is triggered to undergo a full re-iteration, and the condition vector input of the decoupled two-stream conditional graph network is forcibly updated; when When, only the linear layer weights corresponding to the condition vector are updated, while the parameters of the remaining layers are frozen; when At this time, keeping the current model parameters unchanged, the cached spectral intensity ratio matrix is ​​used directly for classification. The thresholds of 0.05 and 0.15 were determined by enumerating threshold combinations and selecting the combination with the minimum accuracy loss in a control experiment with artificially applied plasma perturbations of different amplitudes, using the decrease in classification accuracy as the objective function, and then performing 5-fold cross-validation.

[0094] The specific implementation of step S05 is as follows: the first flow of the plasma parameter decoupling dual-flow conditional graph network is the spectral graph construction flow, and each detected spectral line is regarded as a graph node, and the node feature vector... From the center wavelength Peak intensity Half height and width Integral area Asymmetry It consists of 5 components, and the formula is expressed as follows:

[0095] ;

[0096] In the formula, , , , , They are nodes The reference normalized values ​​for the center wavelength, peak intensity, full width at half maximum (FWHM), integral area, and asymmetry of the corresponding spectral line are all determined by the statistical mean of the training set. For dimensionless node feature vectors, the superscript is... Indicates transpose. Strong edges are connected between spectral lines of the same element, with edge weights. Calculated by normalization of the inverse of the corresponding energy level difference; weak edges are connected between spectral lines of different elements, with edge weights. The co-occurrence frequency statistics in the training set are normalized. The graph attention network has 4 layers, with 8 attention heads in each layer. The attention weights are calculated based on... and The concatenation results in a conditional vector, expressed by the following formula:

[0097] ;

[0098] In the formula, The dimensionless conditional vector is fused to the attention scoring function through a linear layer. Layer The formula for the attention coefficient of a person's height is expressed as follows:

[0099] ;

[0100] In the formula, For the first Layer Node in the head For nodes Attention coefficient For the first The learnable attention parameter vector of the size, For the first Layer node feature linear transformation matrix, For the first Layer condition vector linear transformation matrix, For the first Layer nodes Embedded vector, For the first Layer nodes Embedded vector, For nodes The set of neighboring nodes, This represents a vector concatenation operation. The function is a linear rectified function with leakage. The node embeddings from each graph attention network layer are concatenated along the feature dimension via cross-layer skip connections and then fed into the graph readout layer. Global summation and pooling are then performed to obtain the graph-level representation vector. , The vector is dimensionless. The second stream is the global spectral stream, which takes the normalized full spectrum after background subtraction and inputs it into a 5-layer one-dimensional convolutional neural network. The kernel size is 7, and the number of filters increases from 32 to 256. The output is a global spectral representation vector. , It is a dimensionless vector. The two streams are fused through a cross-attention mechanism, in order to... For query vectors, The key vector and value vector are fused together with the feature vector. The formula is expressed as follows:

[0101] ;

[0102] In the formula, , , These are linear transformation matrices for the query, key, and value, respectively. The dimension of the key vector. For normalized exponential functions, This is a dimensionless fused feature vector. The output class prediction probability is obtained through a two-layer fully connected classification head. Simultaneously, each is output via an independent regression head. and The predicted value. The multi-task loss function formula is expressed as follows:

[0103] ;

[0104] In the formula, Total loss due to multitasking Cross-entropy classification loss, The mean square error loss is due to the electronic temperature regression. The mean square error loss for electron density regression. and To assist in loss weighting, the values ​​range from 0.1 to 0.3, determined through grid search on the validation set. All three loss terms are dimensionless after being normalized to the square of their respective reference values. The adaptive plasma state-space iterative classification algorithm is executed synchronously: First, the Boltzmann diagram slope method is used to give... Initial value; second step, Substitute into the Saha ionization equilibrium equation and solve iteratively. Establish the theoretical spectral line intensity ratio matrix for each element and each valence state. , It is a dimensionless matrix, whose elements are the ratios of the theoretical intensities of each spectral line; the third step is to use... Construct a least-squares linear system and solve for the relative concentration vectors of each element. And calculate the residual vector, as expressed in the following formula:

[0105] ;

[0106] In the formula, For the normalized residual vector, To observe the spectral line intensity vector, Let be the vector of relative concentrations of each element. for The L2 norm is used for normalization. It is a dimensionless vector; the fourth step is to use... The gradient is updated according to the Newton-Raphson iteration scheme. and Return to step two and repeat until... Less than the convergence threshold; Step 5, using the converged... Calculate the Mahalanobis distance in the element concentration vector space and output the class prediction probability. .

[0107] The specific implementation of step S06 is: calculating the element concentration vector. Compared with the mean vectors of each category in the standard concentration vector library The Mahalanobis distance between them is expressed by the following formula:

[0108] ;

[0109] In the formula, Let be the Mahalanobis distance, and be a dimensionless quantity. For the first The mean vector of the concentration vectors of each sample element in the training set. For the first The covariance matrix of each class is obtained by statistically analyzing the element concentration vectors of samples from each class in the training set. When the number of samples is less than the vector dimension, ridge estimation is used to ensure invertibility. The ridge parameters are determined by leave-one-out cross-validation. , All are relative concentrations (dimensionless). for The inverse matrix. Discrimination threshold. The optimal separation threshold between the two distributions was determined by statistical analysis of the Mahalanobis distance distributions of known pure substances and known mixtures, followed by analysis of the receiver operating characteristic (ROC) curve. At that time, the Fredholm first-type integral equation regularized inversion spectral decomposition classification algorithm is activated: a kernel matrix is ​​constructed using a pure component standard spectral library. The sparse concentration vector is solved by employing Tikhonov first-norm joint regularization and alternating direction multiplier method iteratively for 30 to 50 steps. The regularization problem is expressed by the following formula:

[0110] ;

[0111] In the formula, To obtain the sparse concentration vector, To optimize the variables (concentration vector). This is a kernel matrix constructed from a standard spectral library of pure components, with its columns representing the normalized standard spectra of each pure component. For the normalized data fitting term, Here are the Tikhorov 2-norm regularization parameters. Both are L1 norm sparse regularization parameters, and are automatically selected through generalized cross-validation. The reference concentration vector, determined by the mean of the concentration vectors of each pure component sample in the training set, is used to normalize the regularization term so that all terms are dimensionless. It is a norm. The non-zero support set and its corresponding amplitude are input into the decision tree classifier, which outputs the mixture component classification result. .

[0112] The specific implementation of step S07 is as follows: decouple the plasma parameters from the output of the two-flow condition diagram network. The output of the adaptive plasma state-space iterative classification algorithm The output of the Laplacian manifold regularized sparse discriminative projection algorithm for spectral plots The merging probability formula for the voting weighted merging strategy is as follows:

[0113] ;

[0114] In the formula, Predict probability vectors for the fused categories. , , These are the fusion weights corresponding to the three-way algorithm, assigned based on the classification accuracy of each category on the validation set, satisfying the following conditions. The weights are dynamically updated based on the validation results of each new batch of samples, and the update step size is the maximum step size that makes the weighted fusion accuracy of the validation set monotonically converge. The category corresponding to the maximum value is used as the final classification label. If the results of any two paths are inconsistent, the dual-pulse laser-induced breakdown spectroscopy mode is triggered. The first pulse preheats the plasma to expand, reduce the effective optical depth, and suppress the self-absorption effect. The second pulse triggers acquisition within the corresponding delay time interval. The delay time interval is determined by experimentally measuring the self-absorption coefficient under different delay times and taking the delay time interval corresponding to the self-absorption coefficient dropping to below 20% under the single-pulse condition. Then, S03 to S06 are executed again.

[0115] To better understand and implement this invention, a specific application scenario of the invention is provided below as an embodiment 2: To illustrate the technical solution of this invention, technicians have built a laser-induced breakdown spectroscopy detection platform, using a pulsed Nd:YAG laser with a laser wavelength of 1064 nm. The pulse width is 8 The single pulse energy is 80. The spectrometer has a resolution of 0.05. The coverage wavelength range is 200. Up to 900 The detector is an enhanced charge-coupled device array. Seven types of controlled chemicals were selected as test samples, covering three physical forms: powder, liquid, and solid bulk. No less than 200 spectra of each type of sample were collected under normal temperature and pressure, and the diagnostic values ​​of electron temperature and electron density were recorded simultaneously.

[0116] Technicians first performed wavelength calibration according to step S01. Before each laser trigger, a mercury-argon reference lamp was inserted, and its characteristic spectral lines were acquired. A third-order polynomial was fitted using the center pixel position of the known spectral lines, and the polynomial coefficients were written into the wavelength lookup table for that measurement. After calibration, the pixel-to-wavelength fitting residuals were all below 0.02. This effectively eliminates wavelength drift caused by changes in ambient temperature and mechanical vibration. Subsequently, step S02 is used to extract the characteristic transition wavelength windows of the controlled elements, and a sparse autoencoder is used to compress the high-dimensional spectral features to 100 dimensions. During the compression process, leave-one-out cross-validation is used to determine the inflection point of the target dimension, thus preserving the discrimination information at the key transition wavelengths.

[0117] Following step S03, the technicians... After subtracting instrument broadening from the full width at half maximum (FWHM) of the spectral lines, the pure Stark broadening value is obtained, and the electron density is inverted. The electron temperature is then inverted by selecting 3 to 5 weak spectral lines within the same excited-state multivariate using the Boltzmann plot slope method. In typical measurement batches, the electron density range is [range missing]. ~ The electronic temperature range is 8000–12000. ,like Figure 2 As shown, Figure 2 Frame-by-frame diagnostic curves of electron temperature and electron density in typical batches are presented, allowing for a direct observation of the drift patterns of plasma parameters between batches. The inverted electron temperature and electron density are then physically normalized component-by-component using the Boltzmann factor to the compressed characteristic matrix. After normalization, the characteristic distributions of similar samples from different batches show significant convergence.

[0118] Following step S04, the technicians calculate the plasma drift index for each frame. In a certain batch of measurements, a total of 1000 frames of spectrum were measured. The frame rate is approximately 12%. The frame rate is approximately 31%. The percentage of frames is approximately 57%, as shown in Table 1.

[0119] Table 1. Statistical table of plasma drift index distribution

[0120]

[0121] Following step S05, the normalized feature matrix, electron temperature, and electron density are input into the plasma parameters to decouple the dual-flow conditional graph network. The spectral graph construction flow uses a graph attention network to communicate with the feature spectral line nodes of seven types of controlled chemicals. After the conditional vector is incorporated into the attention scoring function, the network adaptively redistributes the intensity ratio of the same element's spectral lines under different matrices. The global spectral flow captures the broadband molecular emission residual structure through a five-layer one-dimensional convolutional neural network, such as... Figure 3As shown, Figure 3 Scatter plots of the distribution of seven categories of controlled chemicals in the final feature space of the global spectral flow are presented, demonstrating the clustering and separation effects of different categories in the feature space. After cross-attention fusion of the two feature paths, the classification head outputs the predicted probabilities of each category. Simultaneously, the adaptive plasma state-space iterative classification algorithm combines the Saha ionization equilibrium equation and the Boltzmann population equation, and after Newton iteration until the residual L2 norm is below the convergence threshold, constructs an element concentration vector based on the relative concentrations of each element.

[0122] The main components of the elemental concentration vectors of typical seven types of controlled chemical samples are shown in Table 2.

[0123] Table 2. Main components of elemental concentration vectors for typical samples

[0124]

[0125] In step S06, the Mahalanobis distance between the element concentration vector of each frame and the standard concentration vector library is calculated. Approximately 8% of the frames have Mahalanobis distances exceeding the discrimination threshold and are classified as mixture samples, triggering the Fredholm first-class integral equation regularization inversion spectral decomposition classification algorithm. A kernel matrix is ​​constructed using a pure component standard spectral library. Tikhorov 1 norm joint regularization is used to iterate for 40 steps using the alternating direction multiplier method to obtain sparse concentration vectors. Generalized cross-validation automatically selects the regularization parameters. The components corresponding to the non-zero support sets are the actual components of the mixture. After inputting into the decision tree classifier, the mixture component classification results are output.

[0126] Following step S07, the three classification results are fused using dynamic weighted voting. In this batch of measurements, after weighted fusion of the validation sets of the three algorithms, the proportion of frames with consistent results is 91%; inconsistent frames trigger a dual-pulse laser-induced breakdown spectral mode. After the first pulse preheats, the second pulse is triggered within the delay time interval corresponding to the self-absorption coefficient dropping below 20% under single-pulse conditions, and steps S03 to S06 are re-executed. Figure 4 As shown, Figure 4 The single-pulse and double-pulse modes are given. The curve showing the change of the self-absorption coefficient of the spectral line with the delay time intuitively demonstrates the suppression effect of the dual-pulse mode on the self-absorption effect.

[0127] It should be noted that the atomic spectral data used in this embodiment comes from the atomic spectral data database published by the Beijing Institute of Applied Physics and Computational Mathematics, which is released through the National Science and Technology Infrastructure Platform Project "Basic Science Data Sharing Network". Alternatively, the atomic spectral database (Atomic Spectra Database, ASD) published by the National Institute of Standards and Technology (NIST) in the United States can also be used.

[0128] Compared to traditional methods that directly input raw spectral line intensities into support vector machines or principal component analysis for classification, the advancements of this invention are reflected in the following principles. Traditional methods directly encode spectral line intensity changes caused by plasma parameter fluctuations into classification features, leading to significant shifts in feature vectors for the same type of sample under different batches, instruments, or matrix conditions, resulting in classification boundary failure. This invention decouples the modulation effect of plasma parameters at the data level through physical normalization, achieves adaptive redistribution at the network inference level by decoupling the two-current conditional graph network through plasma parameters, and replaces classification features with element concentration vectors independent of instruments and matrices at the physical inversion level through an adaptive plasma state-space iterative classification algorithm. The synergistic effect of these three physical constraints ensures the stability of classification features under a wide range of plasma parameter drift conditions, guaranteeing the cross-instrument and cross-matrix transferability of the classification model from a mechanistic perspective.

[0129] It should be noted that the variables involved in this invention are explained in detail in Tables 3 and 4.

[0130] Table 3. Variable Explanation Table (Part 1)

[0131]

[0132] Table 4. Variable Explanation Table (Part Two)

[0133]

[0134] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for constructing a model for classifying and identifying chemicals by rapid LIBS, characterized in that, Includes the following steps: Insert a mercury-argon reference lamp, acquire the reference lamp spectrum, update the pixel wavelength mapping polynomial, and output the original spectral dataset after wavelength calibration. Based on the atomic spectral database, the characteristic transition wavelength windows of the controlled elements are extracted, and the feature dimension is compressed to within the target dimension by a sparse autoencoder, and the compressed feature matrix is ​​output. The electron density is inverted using the hydrogen line Stark broadening, and the electron temperature is inverted using the Boltzmann plot slope. The compressed feature matrix is ​​then physically normalized, and the normalized feature matrix is ​​output. Calculate the plasma drift index and adjust the update strategy of the plasma parameter decoupling dual-flow condition graph network according to the range to which the plasma drift index belongs; The normalized feature matrix, electron temperature, and electron density are input into the plasma parameters and decoupled into the two-current conditional graph network. An adaptive plasma state-space iterative classification algorithm is executed synchronously, and the element concentration vector and category prediction probability are output. Calculate the Mahalanobis distance between the element concentration vector and the standard concentration vector library. Based on whether the Mahalanobis distance exceeds the discrimination threshold, decide whether to enable the Fredholm first-type integral equation regularization inversion spectral decomposition classification algorithm and output the mixture composition classification results. The three classification results are merged using a voting weighting fusion strategy to output the final classification label; if any two results are inconsistent, the dual-pulse laser-induced breakdown spectral mode is triggered, and the physical normalization to mixture component classification step is re-executed. The plasma drift index is calculated by summing the normalized deviations of electron temperature and electron density relative to the weighted average of the indexes of the previous several frames. When the plasma drift index is greater than or equal to the high drift threshold, the adaptive plasma state space iterative classification algorithm is triggered to re-iterate the entire algorithm and the condition vector input is forcibly updated. When the plasma drift index is between the low drift threshold and the high drift threshold, only the linear layer weights corresponding to the condition vector are updated. When the plasma drift index is less than the low drift threshold, the cached spectral intensity ratio matrix is ​​used directly for classification.

2. The method for constructing a rapid LIBS classification and identification model for controlled chemicals according to claim 1, characterized in that, The update of the pixel wavelength mapping polynomial is specifically achieved by acquiring the mercury-argon reference lamp spectrum before each laser trigger, fitting a third-order polynomial with the center pixel position of the known spectral line, writing the polynomial coefficients into the wavelength lookup table for the current measurement, and controlling the fitting error of the center pixel position within the fitting error threshold.

3. The method for constructing a rapid LIBS classification and identification model for controlled chemicals according to claim 2, characterized in that, The sparse autoencoder employs a three-layer fully connected structure with input dimensions of 512, 256, and target dimensions. The activation function is a linear rectified function, the decoding layer is symmetrically expanded, and the training objective is the sum of reconstruction error and sparsity penalty.

4. The method for constructing a rapid LIBS classification and identification model for controlled chemicals according to claim 3, characterized in that, The Boltzmann plot slope inversion of electron temperature is specifically achieved by selecting several weak spectral lines within the same excited state multivariate and fitting a straight line with the excitation energy as the horizontal axis and the logarithm of the normalized intensity as the vertical axis. The slope is inversely proportional to the electron temperature. The selection criteria for weak spectral lines are that the signal-to-noise ratio is greater than the signal-to-noise ratio threshold and the peak intensity is lower than the intensity ratio threshold of the strongest spectral line of the same element.

5. The method for constructing a rapid LIBS classification and identification model for controlled chemicals according to claim 4, characterized in that, The electron density inversion of the hydrogen line Stark broadening is specifically achieved by extracting the full width at half maximum (FWHM) of the spectral line, subtracting the instrument broadening to obtain the pure Stark broadening value, and then substituting it into the Stark broadening coefficient lookup table to calculate the electron density. The instrument broadening is obtained by calibration using the measured FWHM of a known narrow spectral line from a mercury-argon reference lamp.

6. The method for constructing a rapid LIBS classification and identification model for controlled chemicals according to claim 5, characterized in that, The physical normalization specifically involves using electron temperature and electron density to correct each spectral line intensity component in the compressed characteristic matrix component by Boltzmann factor, thereby eliminating the influence of plasma parameter fluctuations on the absolute intensity of the spectral lines.

7. The method for constructing a rapid LIBS classification and identification model for controlled chemicals according to claim 6, characterized in that, The first flow of the plasma parameter decoupled dual-flow conditional graph network is the spectral graph construction flow, which treats each detected spectral line as a graph node. The node feature vector consists of center wavelength, peak intensity, full width at half maximum (FWHM), integral area, and asymmetry. The graph neural network backbone adopts a graph attention network structure. When calculating the attention weight, the electron temperature and electron density are concatenated as a conditional vector and fused into the attention scoring function. The second flow is the global spectral flow, which inputs the normalized full spectrum after background subtraction into a one-dimensional convolutional neural network. The features of the two flows are fused through a cross-attention mechanism and then output as a category prediction probability by a classification head.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that, when executed in a computer, perform the method described in any one of claims 1-7.

9. A rapid LIBS classification and identification model construction system for controlled chemicals, characterized in that, The system comprises the computer-readable storage medium of claim 8, wherein the system is a computer, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes program instructions stored in the computer-readable storage medium.