A plasma spectral identification method, apparatus, electronic device, and storage medium

By using the LGBM spectral recognition model to train and optimize plasma spectral information, the problems of complex operation and high cost of plasma spectroscopy are solved, and accurate identification and classification of plasma spectra are achieved, reducing diagnostic costs.

CN116863182BActive Publication Date: 2026-05-26CHINA AGRI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA AGRI UNIV
Filing Date
2022-03-22
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing plasma spectroscopy methods are complex to operate and have high technical costs.

Method used

The LGBM spectral recognition model is used to train plasma spectral information. The raw data is processed through feature engineering and synthetic sampling methods. The model parameters are optimized by cross-validation and grid search to achieve accurate recognition and classification of plasma spectra.

Benefits of technology

This reduces the technical cost of plasma spectroscopy diagnostics and improves the ease of operation and accuracy of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116863182B_ABST
    Figure CN116863182B_ABST
Patent Text Reader

Abstract

This invention provides a plasma spectral identification method, apparatus, electronic device, and storage medium, comprising: inputting the plasma spectral information to be identified into a trained LGBM spectral identification model to obtain target classification spectral parameters corresponding to the plasma spectral information to be identified, and performing spectral wavelength feature analysis on the plasma spectral information to be identified based on the target classification spectral parameters; wherein the trained LGBM spectral identification model is obtained by training on plasma spectral information samples carrying real spectral parameter labels. The method of this invention automatically learns the implicit physical information between plasma spectral wavelength features by employing a machine learning LGBM algorithm, achieving accurate identification and classification of plasma spectra. It is easy to operate and can greatly reduce the technical cost of plasma spectral diagnostics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spectral detection technology, and in particular to a plasma spectral identification method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the continuous research in plasma science, plasma diagnostic technology has also developed. Plasma diagnostics is a scientific and technological field that uses appropriate methods and techniques to measure plasma parameters based on the understanding of plasma physical processes. Plasma diagnostics is an important foundation of astrometry and astrophysics.

[0003] Plasma spectroscopy is a common plasma diagnostic method that uses the emission or absorption spectra of plasma to diagnose parameters such as plasma temperature, density, and chemical composition. However, existing plasma spectroscopy analysis instruments and systems are complex to operate and costly. Summary of the Invention

[0004] This invention provides a plasma spectroscopy identification method, device, electronic device, and storage medium to address the shortcomings of existing plasma spectroscopy methods, which are characterized by complex operation processes and high technical costs.

[0005] This invention provides a plasma spectral identification method, comprising:

[0006] The spectral information of the plasma to be identified is input into the trained LGBM spectral recognition model to obtain the target classification spectral parameters corresponding to the spectral information of the plasma to be identified, so as to perform spectral wavelength feature analysis on the spectral information of the plasma to be identified based on the target classification spectral parameters.

[0007] The trained LGBM spectral recognition model is obtained by training plasma spectral information samples carrying real spectral parameter labels.

[0008] According to a plasma spectral identification method provided by the present invention, before inputting the plasma spectral information to be identified into the trained LGBM spectral identification model, the method further includes:

[0009] Obtain the raw plasma spectral information dataset;

[0010] The original plasma spectral information dataset is preprocessed using feature engineering methods and / or synthetic sampling methods to obtain a plasma spectral information sample dataset;

[0011] The LGBM spectral recognition model was trained using the plasma spectral information sample dataset through cross-validation.

[0012] According to a plasma spectral identification method provided by the present invention, the plasma spectral information sample includes 2048 characteristic wavelengths, the wavelength range of which is from 366.19 nm to 1051.14 nm.

[0013] According to the plasma spectral identification method provided by the present invention, after training the LGBM spectral identification model using the plasma spectral information sample dataset, the method includes:

[0014] The parameters of the LGBM spectral recognition model are optimized using a grid search method to obtain the optimal parameters of the LGBM spectral recognition model.

[0015] According to the present invention, a plasma spectral identification method is provided, wherein the method employs cross-validation to train an LGBM spectral identification model using the plasma spectral information sample dataset, comprising:

[0016] A cross-validation method is used to obtain a training sample set from the plasma spectral information sample dataset;

[0017] Multiple sets of training samples are obtained by taking the plasma spectral information samples in the training sample set and the real spectral parameter labels carried by the plasma spectral information samples as a set of training samples.

[0018] The LGBM spectral recognition model was trained using the aforementioned multiple sets of training samples.

[0019] According to the plasma spectral identification method provided by the present invention, the step of training an LGBM spectral identification model using the multiple sets of training samples includes:

[0020] For any set of training samples, the training samples are input into the LGBM spectral recognition model, and the predicted probability corresponding to the training samples is output; the predicted probability is used to determine the classification result of the training samples.

[0021] Using a preset loss function, the loss value is calculated based on the predicted probability corresponding to the training sample and the true spectral parameter label in the training sample;

[0022] If the loss value is less than a preset threshold, the LGBM spectral recognition model training is complete.

[0023] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the plasma spectral identification method as described above.

[0024] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the plasma spectral identification method as described above.

[0025] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the plasma spectral identification method as described above.

[0026] This invention provides a plasma spectral identification method, device, electronic device, and storage medium. It trains an LGBM model using plasma spectral information samples carrying real spectral parameter labels to obtain a trained LGBM spectral identification model. The plasma spectral information to be identified is then input into the trained LGBM spectral identification model to obtain the target classification spectral parameters corresponding to the plasma spectral information. Based on the target classification spectral parameters, spectral wavelength feature analysis is performed on the plasma spectral information to be identified. Thus, by employing a machine learning LGBM algorithm, the implicit physical information between plasma spectral wavelength features is automatically learned, achieving accurate identification and classification of plasma spectra. The method is convenient to operate and can greatly reduce the technical cost of plasma spectral diagnosis. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0028] Figure 1 This is a schematic flowchart of the plasma spectral identification method provided by the present invention;

[0029] Figure 2 This is a schematic diagram illustrating the feature importance output by the plasma spectral identification method provided by the present invention;

[0030] Figure 3 This is a schematic diagram of the model training and testing process for the plasma spectral identification method provided by the present invention;

[0031] Figure 4 This is a schematic diagram of the structure of the plasma spectral identification device provided by the present invention;

[0032] Figure 5 This is a schematic diagram of the physical structure of the electronic device provided by the present invention. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0034] The following is combined with Figures 1-5 This invention describes a plasma spectral identification method, apparatus, electronic device, and storage medium.

[0035] Figure 1 This is a schematic flowchart of the plasma spectral identification method provided by the present invention, as shown below. Figure 1 As shown, the method includes:

[0036] Step S1: Input the plasma spectral information to be identified into the trained LGBM spectral recognition model to obtain the target classification spectral parameters corresponding to the plasma spectral information to be identified, so as to perform spectral wavelength feature analysis on the plasma spectral information to be identified based on the target classification spectral parameters.

[0037] The trained LGBM spectral recognition model is obtained by training plasma spectral information samples carrying real spectral parameter labels.

[0038] It should be noted that plasma, also known as ionized plasma, is an ionized gaseous substance composed of positive and negative ions produced by the ionization of atoms and atomic groups after some electrons have been stripped away. It is a macroscopically neutral ionized gas with a scale greater than the Debye length. Its motion is mainly governed by electromagnetic forces and it exhibits significant collective behavior. It is widely present in the universe and is often regarded as the fourth state of matter, in addition to solids, liquids, and gases.

[0039] Plasma spectrum refers to the electromagnetic radiation spectrum emitted from within plasma, ranging from infrared to vacuum ultraviolet.

[0040] In this embodiment, the LGBM algorithm model is also called the LightGBM model. It is a fast, distributed, high-performance gradient boosting framework based on the decision tree algorithm, which can be used for ranking, classification, regression and many other machine learning tasks.

[0041] The LGBM model is based on the Gradient Boosting Decision Tree (GBDT) and XGboost boosting tree algorithms. It also employs histogram algorithms, histogram difference acceleration, depth-limited leaf-wise growth mechanism, gradient-based one-side sampling (GOSS), and exclusive feature bundling (EFB) methods, resulting in high classification accuracy and training speed.

[0042] The LGBM model employs an optimal leaf-wise growth strategy to split leaf nodes, abandoning the level-wise growth strategy used by most existing GBDT decision tree algorithms. Therefore, in the LightGBM algorithm, when growing to the same number of leaf nodes, the leaf-wise algorithm reduces loss more than the level-wise algorithm, thus achieving higher accuracy.

[0043] Specifically, the plasma spectral information to be identified described in this invention refers to the plasma spectral information that needs to be identified, which may be a plasma glow discharge spectrum, which can be obtained by a fiber optic spectrometer.

[0044] The target classification spectral parameters described in this invention refer to a class of parameters that can characterize the spectral characteristics of the plasma to be identified. Specifically, they may include at least one of the following parameters: ambient pressure of plasma glow discharge, discharge power, discharge location, etc.

[0045] The trained LGBM spectral recognition model described in this invention is trained on plasma spectral information samples carrying real spectral parameter labels. It is used to learn the implicit physical information between the wavelength features of the plasma spectrum to be identified, identify the input plasma spectrum to be identified, and output the target classification spectral parameters that can characterize the spectral features of the plasma to be identified. This enables the identification and classification of plasma spectral information under different target classification spectral parameters.

[0046] The model training samples can consist of multiple sets of plasma spectral information samples carrying labels of real spectral parameters.

[0047] The true spectral parameter labels described in this invention are predetermined based on plasma spectral information samples and correspond one-to-one with each sample. In other words, each plasma spectral information sample in the model training samples is pre-set to carry a corresponding spectral parameter as a true label.

[0048] In some embodiments, before inputting the plasma spectral information to be identified into the trained LGBM spectral recognition model in step S1, the method further includes:

[0049] Obtain the raw plasma spectral information dataset;

[0050] The original plasma spectral information dataset is preprocessed using feature engineering methods and / or synthetic sampling methods to obtain a plasma spectral information sample dataset;

[0051] The LGBM spectral recognition model was trained using a cross-validation method and a plasma spectral information sample dataset.

[0052] Specifically, the raw plasma spectral information dataset described in the embodiments of the present invention refers to a plasma spectral information dataset that has not undergone any data processing.

[0053] Before inputting the plasma spectral information to be identified into the trained LGBM spectral recognition model, the LGBM spectral recognition model also needs to be trained. The specific training process is as follows:

[0054] Complex plasmas were generated through radio frequency plasma discharge experiments. Raw plasma spectral information samples were collected using a fiber optic spectrometer, and each plasma spectral information sample was stored as a txt file.

[0055] In this embodiment, the plasma spectral information sample includes 2048 characteristic wavelengths, with the wavelength range being 366.19 nm to 1051.14 nm. More specifically, the spectral data in txt format can include two columns of data. The first column is the characteristic wavelength of the spectrum, which covers the range from the near-infrared band to the vacuum ultraviolet band, i.e., 366.19 nm to 1051.14 nm, totaling 2048 wavelength characteristics. The second column is the spectral intensity corresponding to the characteristic wavelength.

[0056] The method of this invention uses collected plasma spectral information samples, which contain 2048 characteristic wavelengths ranging from 366.19 nm to 1051.14 nm, to train an LGBM model. This facilitates the recognition and diagnosis of plasma spectra in the near-infrared to vacuum ultraviolet range by the trained LGBM spectral recognition model, thereby improving the applicability of the model.

[0057] Furthermore, in this embodiment, the collected plasma spectral information sample data can be read into a DataFrame data structure and stored as a comma-separated values ​​(CSV) file format using the Python language and the Numpy and Pandas data analysis libraries, thereby obtaining the original plasma spectral information dataset.

[0058] Furthermore, feature engineering methods and / or synthetic sampling methods are used to preprocess the original plasma spectral information dataset to obtain a plasma spectral information sample dataset. That is, in this embodiment, feature engineering methods, or synthetic sampling (SMOTE) methods, or both feature engineering methods and SMOTE methods can be used to preprocess the original plasma spectral information dataset to construct high-value feature data.

[0059] In this embodiment, the feature engineering methods used may specifically include feature derivation, feature discretization, feature normalization (standardization), and feature selection. Specifically, by performing feature derivation on the original plasma spectral information dataset, more features valuable for LGBM model classification can be generated. Feature discretization can extract the nonlinear laws of linear features, which helps to accelerate the convergence speed of the model, reduce the impact of outliers on the model, and thus improve the robustness of the algorithm. Feature normalization (standardization) is essentially a dimensionless process, which can improve the accuracy of the model and accelerate convergence. Feature selection reduces the dimensionality of the data, reduces the difficulty of model learning, improves accuracy, and accelerates convergence.

[0060] More specifically, in this implementation, the feature engineering method may include the following feature derivation methods: mathematical operations: addition, subtraction, multiplication, and division methods for continuous features; polynomial combination: constructing polynomial features to obtain cross-combinations and higher-order terms of features; Cartesian product: cross-combining discrete features; feature derivation can construct new linearly / nonlinearly related features based on the wavelength characteristics of the spectrum. Among these new features, there may be features that have a significant gain for model recognition, which can improve the model's recognition performance. Redundant features constructed can be eliminated in the feature selection process.

[0061] In this embodiment, the feature discretization method may include: ① Binarization: setting a threshold to discretize features into 0 and 1, transforming fine-grained features into coarse-grained features; ② Equal-frequency binning: dividing intervals into approximately the same number of samples; ③ Equal-distance binning: dividing intervals into intervals with the same value range; ④ Cluster binning: binning based on clustering algorithms, grouping samples within the same cluster into the same class. Discretized features accelerate model convergence while eliminating the influence of outliers in spectral data, improving the robustness of model recognition.

[0062] In this embodiment, feature scaling is essentially a dimensionless process that can improve model accuracy and accelerate convergence. Common methods include standardization (subtracting the mean and dividing by the standard deviation) and normalization (subtracting the minimum value and dividing by the difference in the data distribution range, i.e., the maximum value minus the minimum value). After scaling, the differences in spectral intensity distributions corresponding to different wavelengths are eliminated; that is, the spectral intensity corresponding to all wavelength features is scaled to the same range, eliminating the influence of dimensions.

[0063] The selection of a specific feature scaling scheme can refer to the following prior rules:

[0064] ① If there are requirements regarding the data range, normalization should be used;

[0065] ② The data contains outliers and noise; standardization should be used.

[0066] ③ In algorithms such as classification, clustering, and PCA, where distance is used to measure similarity, standardization yields better results;

[0067] ④ If the data does not conform to a normal distribution, normalization should be applied;

[0068] ⑤ Normalization is suitable for small samples, while standardization is suitable for large samples.

[0069] In this embodiment, the plasma spectrum has 2048 wavelengths. During training, some wavelength features have very limited gain for the model, while too many wavelengths can affect the model's convergence speed and fitting difficulty. By using feature selection, the dimensionality of the spectral data can be reduced, thereby reducing the difficulty of model learning, improving recognition accuracy, and accelerating convergence.

[0070] In this embodiment, the feature selection methods include Filter filtering, wrapper wrapping, and embedded embedding.

[0071] Among them, the Filter method selects features based on certain statistical test scores and correlation indicators; specifically, it can include: ① Variance selection method: set a variance threshold and select features with variance greater than the threshold; if the variance is too small, it means that the feature distribution difference is too small and has no effect on sample classification; ② Correlation coefficient method: set a correlation coefficient threshold, calculate the correlation coefficient between the feature variable and the target variable, and select variables with a correlation coefficient greater than the threshold; ③ IV value: set an IV threshold and select feature variables with an IV value greater than the threshold.

[0072] The wrapper method can specifically include: ① Stability method: Selecting features on different feature subsets and sample subsets, repeating and summarizing the feature selection results; for example, counting the frequency of a certain feature as an important feature and selecting the feature with the highest frequency; ② Recursive elimination method: Repeatedly building the model, selecting the best / worst features, placing the selected features and then filtering the remaining features until all features have been traversed; the order in which features are selected in this process is the feature sorting.

[0073] Embedded methods can specifically include: ① Penalty-based selection: Selecting models with regularization penalty terms to select features, such as L1 regularization in logistic regression, which produces sparse features; ② Feature importance-based selection: Selecting features using feature importance indices output by tree-based models (random forest, GBDT).

[0074] In this embodiment, the SMOTE synthetic sampling method is a method of artificially synthesizing new samples based on minority class samples and adding them to the sample dataset, specifically including steps 110, 120 and 130.

[0075] In step 110, multiple minority class samples are randomly selected from the original plasma spectral information dataset, and for each minority class sample, k nearest neighbor samples are calculated using Euclidean distance.

[0076] Step 120: Set the sampling factor N according to the sample imbalance ratio;

[0077] Step 130: Randomly select several samples from the k nearest neighbor samples, and calculate the new samples by comparing them with the original samples using the following formula:

[0078] New sample = nearest neighbor sample + (random number between 0 and 1) * (distance between original sample and nearest neighbor sample).

[0079] Therefore, the SMOTE algorithm can be used to synthesize and sample the original plasma spectral information dataset, thereby expanding the plasma spectral information sample dataset.

[0080] In this embodiment, the original plasma spectral information dataset is preprocessed using feature engineering methods and / or synthetic sampling methods to obtain a plasma spectral information sample dataset. Then, cross-validation is used to train and test the LGBM spectral recognition model using the plasma spectral information sample dataset.

[0081] In this embodiment, a five-fold cross-validation method can be adopted, dividing the data samples into a training sample set and a test sample set, with the ratio of the training sample set to the test sample set set being 4:1. The training sample set is used to train the model, and the test sample set is used to verify the training effect of the model. That is, the plasma spectral information sample dataset is divided into five parts, with four parts of the plasma spectral information sample dataset used to train the model and one part of the plasma spectral information sample dataset used to verify the training effect of the model.

[0082] Preferably, ten-fold cross-validation can also be adopted, dividing the data samples into a training sample set and a test sample set, with the ratio of the training sample set to the test sample set set being 9:1, which can help improve the model accuracy of the trained LGBM spectral recognition model.

[0083] In this embodiment, by employing feature engineering methods and / or synthetic sampling methods, the obtained original plasma spectral information dataset is preprocessed with data feature extraction and sample dataset expansion to obtain a plasma spectral information sample dataset. Cross-validation is then used to train the LGBM spectral recognition model using this sample dataset. This effectively accelerates the convergence speed of the LGBM model training, reduces the impact of outliers on the LGBM model, improves the algorithm's robustness, and significantly enhances the accuracy of the trained LGBM spectral recognition model. This, in turn, facilitates the accurate identification and classification of plasma spectra.

[0084] The plasma spectral identification method provided by this invention trains an LGBM model using plasma spectral information samples carrying real spectral parameter labels to obtain a trained LGBM spectral identification model. The plasma spectral information to be identified is then input into the trained LGBM spectral identification model to obtain the target classification spectral parameters corresponding to the plasma spectral information to be identified. Based on the target classification spectral parameters, spectral wavelength feature analysis is performed on the plasma spectral information to be identified. Thus, by using the machine learning LGBM algorithm, the implicit physical information between the plasma spectral wavelength features is automatically learned, achieving accurate identification and classification of plasma spectra. The method is easy to operate and can greatly reduce the technical cost of plasma spectral diagnosis.

[0085] In some embodiments, after training the LGBM spectral recognition model using the plasma spectral information sample dataset, the process includes:

[0086] Based on the grid search method, the parameters of the LGBM spectral recognition model are optimized to obtain the optimal parameters of the LGBM spectral recognition model.

[0087] Specifically, the optimal parameters described in the embodiments of the present invention refer to the set of model hyperparameters that best characterize the training effect of the LGBM spectral recognition model.

[0088] It should be noted that the grid search method is an exhaustive search method for specified parameter values. It optimizes the learning algorithm by using cross-validation to improve the parameters of the estimated function. Specifically, it involves permuting and combining all possible model parameter values, generating a "grid" of all possible combinations, then using each combination to train the model and evaluating its performance using cross-validation until the optimal parameter combination is found.

[0089] In this embodiment, a grid search method can be used to optimize the parameters of the LGBM spectral recognition model. Specifically, this may include adjusting parameters such as the maximum depth of the decision tree (max_depth), the minimum number of records a leaf may have (min_data_in_leaf), and the data ratio (bagging_fraction) used in each iteration. By traversing all possible model parameters, the optimal parameters of the LGBM spectral recognition model can be obtained.

[0090] In this embodiment, the parameters of the LGBM spectral recognition model are optimized using a grid search method. The optimized parameters may include at least one of the following: fitting parameters, sampling parameters, regularization parameters, and ensemble parameters.

[0091] The fitting parameters may include the maximum depth of the tree (max_depth), the minimum gain to split a node (min_gain_to_split), the number of branches (num_leaves), and the minimum number of data in a leaf node (min_data_in_leaf).

[0092] Sampling parameters may include: sample sampling ratio bagging_fraction, sample sampling frequency bagging_freq, and feature sampling ratio for each tree feature_fraction;

[0093] Regularization parameters may include: the coefficient of L2 regularization (Lambda_l2) in the loss function and the coefficient of L1 regularization (Lambda_l1) in the loss function;

[0094] Ensemble parameters can include the learning rate (learn_rate) and the maximum number of base models (n_estimators). The robustness of the model can be improved by optimizing the learning rate and reducing the weight of each tree.

[0095] Understandably, by using a grid search method to optimize the parameters of the LGBM spectral recognition model, the optimal parameters of the LGBM spectral recognition model can be determined, including the optimal fitting parameters, optimal sampling parameters, optimal regularization parameters, and optimal ensemble parameters.

[0096] More specifically, in this embodiment, the parameter optimization of the LGBM spectral recognition model can be achieved by the following optimization method to accelerate the model's training speed:

[0097] The sampling frequency is determined by the bagging_fraction and bagging_freq parameters, and the feature sampling frequency is determined by setting the feature_fraction parameter; reduce max_depth to reduce the depth of a single tree; reduce max_bin to control the number of feature bins.

[0098] In this embodiment, the following optimization methods can be used to accelerate the accuracy of the model and reduce model bias:

[0099] Use larger max_bin and num_leaves; use smaller learning_rate.

[0100] In this embodiment, the following optimization methods can be used to accelerate the accuracy of the model and reduce model overfitting:

[0101] Use smaller bin sizes (max_bin and num_leaves) for discretization based on feature values; use bagging by setting bagging_fraction and bagging_freq; use feature subsampling by setting feature_fraction; use more training data; use regularization with lambda_l1, lambda_l2, and minimum split gain (min_split_gain); try max_depth to avoid generating overly deep trees.

[0102] The method of this invention employs a grid search method to find the optimal parameters of the LGBM spectral recognition model, thereby further improving the convergence speed of LGBM model training, enhancing the robustness of the model algorithm, and increasing the accuracy of the trained LGBM spectral recognition model.

[0103] In some embodiments, a cross-validation method is used to train the LGBM spectral recognition model using a plasma spectral information sample dataset, including:

[0104] A cross-validation method was used to obtain a training sample set from the plasma spectral information sample dataset;

[0105] Multiple sets of training samples are obtained by taking the plasma spectral information samples and the real spectral parameter labels carried by the plasma spectral information samples in the training sample set as a set of training samples.

[0106] The LGBM spectral recognition model was trained using multiple sets of training samples.

[0107] Specifically, by adopting a five-fold cross-validation method or a ten-fold cross-validation method, the plasma spectral information sample dataset is divided into a training sample set and a test sample set. The plasma spectral information samples in the training sample set and the real spectral parameter labels carried by the plasma spectral information samples are used as a set of training samples. That is, each plasma spectral information sample with a real spectral parameter label is used as a set of training samples, thereby obtaining multiple sets of training samples.

[0108] In embodiments of the present invention, the plasma spectral information sample and the real spectral parameter label carried by the plasma spectral information sample are in one-to-one correspondence.

[0109] Then, after obtaining multiple sets of training samples, the multiple sets of training samples are input into the LGBM spectral recognition model in sequence. That is, the text sample sequence to be replied to and the real reply text sequence label in each set of training samples are input into the LGBM spectral recognition model at the same time. Based on each output result of the LGBM spectral recognition model, the model parameters in the LGBM spectral recognition model are adjusted by calculating the loss function value, and finally the training process of the LGBM spectral recognition model is completed.

[0110] The method of this invention uses plasma spectral information samples and the real spectral parameter labels carried by the plasma spectral information samples in the training sample set as a set of training samples to obtain multiple sets of training samples. By using multiple sets of training samples, the LGBM spectral recognition model can be effectively trained.

[0111] In some embodiments, the LGBM spectral recognition model is trained using multiple sets of training samples, including:

[0112] For any set of training samples, input the training samples into the LGBM spectral recognition model, and output the predicted probability corresponding to the training samples; the predicted probability is used to determine the classification result of the training samples.

[0113] Using a preset loss function, the loss value is calculated based on the predicted probability corresponding to the training sample and the true spectral parameter label in the training sample.

[0114] If the loss value is less than the preset threshold, the LGBM spectral recognition model training is complete.

[0115] Specifically, the preset loss function described in the embodiments of the present invention refers to the loss function pre-set in the LGBM spectral recognition model for model evaluation.

[0116] The preset threshold described in this embodiment of the invention refers to a threshold set in advance by the model to obtain the minimum loss value and complete model training.

[0117] After obtaining multiple sets of training samples, for any training sample, the plasma spectral information sample and the real spectral parameter label carried by the plasma spectral information sample are simultaneously input into the LGBM spectral recognition model, and the predicted probability corresponding to the training sample is output. The predicted probability refers to the predicted probability of the training sample for different spectral parameters, which can determine the classification result of the training sample.

[0118] Based on this, a pre-defined loss function is used to calculate the loss value according to the predicted probability corresponding to the training sample and the true spectral parameter labels in the training sample. The true spectral parameter labels can be represented as one-hot vectors, and the pre-defined loss function can be a squared loss function.

[0119] In the embodiments of the present invention, the representation of the true spectral parameter labels and the preset loss function can be set according to actual needs, and no specific limitation is made here.

[0120] After calculating the loss value, the current training process ends. The model parameters in the LGBM spectral recognition model are then updated before the next training iteration. During training, if the loss value calculated for a particular training sample is less than a preset threshold or the preset maximum number of iterations is reached, the LGBM spectral recognition model training is complete.

[0121] The method of this invention trains the LGBM spectral recognition model and controls the loss value of the LGBM spectral recognition model within a preset range, thereby improving the accuracy of the LGBM spectral recognition model in identifying the spectral parameter categories corresponding to plasma spectra.

[0122] In this embodiment, after training the LGBM spectral recognition model, precision, recall, and F1 score are used to evaluate the trained LGBM spectral recognition model. The formulas for calculating precision, recall, and F1 score are as follows:

[0123] Accuracy = (TP + TN) / (P + N);

[0124] Precision rate = TP / (TP + FP);

[0125] Recall rate = TP / (TP + FN);

[0126] F1 score = 2 * Precision * Recall / (Precision + Recall);

[0127] Where T represents the number of correctly classified samples, F represents the number of misclassified samples, P represents the number of positive samples, N represents the number of negative samples; TP represents the number of samples that are positive and the prediction result is positive, FP represents the number of samples that are negative and the prediction result is positive, TN represents the number of samples that are negative and the prediction result is negative, and FN represents the number of samples that are positive and the prediction result is negative.

[0128] Understandably, during model training, there is often a demand for superior model performance metrics such as accuracy and recall. However, in actual production, these metrics degrade as the model is deployed, necessitating rapid problem identification and repair. Therefore, understanding how the model operates and which features play a key role is of great significance. In this embodiment of the invention, we attempt to interpret the model by considering variable importance in conjunction with Shap values.

[0129] Shap is an abbreviation for Shapley Additive explanations. For each sample, the model produces a predicted value, and the Shap value is the numerical value assigned to each feature in that sample.

[0130] In embodiments of the present invention, based on the training and recognition performance of the LGBM spectral recognition model, and combined with the feature importance and Shap value output by the model itself, the contribution of a certain feature wavelength among multiple feature wavelengths is calculated to determine the importance of that feature wavelength. Correlation and feature importance analysis are performed on the plasma spectral wavelength features to identify significant variables affecting the model, and the results are visualized.

[0131] In this embodiment, the importance of a feature in the tree-based ensemble model is an average of the importance of that feature across all individual trees. The method for calculating the importance of a feature on an individual tree is to sum the reduction in squared loss after splitting based on that feature.

[0132] The Shap value is calculated by subtracting the payoff of a feature combination from the payoff of the combination without that feature. This gives the feature's contribution in that combination. Then, all combinations are calculated and weighted to get the overall contribution of the feature.

[0133] Based on the important features output by the feature importance and shap values ​​above, a correlation analysis is performed on the selected features.

[0134] Figure 2 This is a schematic diagram illustrating the feature importance output by the plasma spectral identification method provided by the present invention, such as... Figure 2 As shown, the horizontal axis represents the importance of features, i.e., the importance of the feature wavelength; the vertical axis represents multiple feature wavelengths. It can be understood that when the feature wavelength is 449.31nm, its feature importance weight value is the largest, which is 45. That is to say, in this embodiment, the feature wavelength of 449.31nm in the plasma spectrum contributes the most to completing the training of the LGBM spectral recognition model compared to feature wavelengths of 482.97nm, 480.87nm, 518.54nm, etc., and belongs to important features.

[0135] Figure 3 This is a schematic diagram illustrating the model training and testing process of the plasma spectral identification method provided by this invention, as shown below. Figure 3 As shown, the model training and testing of the plasma spectral identification method in this embodiment of the invention includes the following steps:

[0136] Step S310: Obtain raw plasma spectral data through discharge experiment, that is, generate complex plasma through radio frequency plasma discharge experiment and collect raw plasma spectral information samples using fiber optic spectrometer.

[0137] Step S320: Clean the raw data, construct the dataset and store it in CSV format. That is, store each raw plasma spectral information sample as a txt file. Use Python language and NumPy and Pandas data analysis libraries to read the collected plasma spectral information sample data into a DataFrame data structure and store it as a CSV file to obtain the raw plasma spectral information dataset.

[0138] Step S330: Using feature engineering and synthetic sampling methods, a plasma spectral information sample dataset is constructed. That is, the original plasma spectral information dataset is preprocessed using feature engineering and / or synthetic sampling methods to obtain the plasma spectral information sample dataset.

[0139] Step S340: Train the LGBM model through cross-validation and test the algorithm performance. That is, use the cross-validation method to train and test the LGBM spectral recognition model using a plasma spectral information sample dataset to obtain a trained LGBM spectral recognition model.

[0140] Then, it is divided into two parallel processing steps, namely step S3501 and step S3601, and step S3502 and step S3602.

[0141] In step S3501, the parameters of the LGBM spectral recognition model are optimized using a grid search method, that is, the parameters of the LGBM spectral recognition model are optimized based on the grid search method to obtain the optimal parameters of the LGBM spectral recognition model.

[0142] Step S3601: Evaluate the trained LGBM spectral recognition model using accuracy, precision, recall, and F1 score;

[0143] Step S3502: Output the model feature importance and Shap value, and select important features;

[0144] Step S3602: Compare the selected important features and peak features to verify the correlation.

[0145] Compared with existing technologies, the plasma spectral identification method provided by this invention has the following advantages:

[0146] This invention uses the LGBM machine learning model to identify and diagnose the complex plasma glow spectrum generated by radio frequency discharge, and uses multiple feature engineering methods and synthetic sampling methods to improve the model's recognition performance.

[0147] This invention performs spectral wavelength feature analysis on plasma glow spectrum based on the identification results of the LGBM model. Compared with the light peak features that traditional spectral analysis focuses on, it proposes a method for importance analysis of spectral wavelength features based on the LGBM machine learning model, and a correlation analysis method after selecting important features.

[0148] The plasma spectral identification method provided by this invention can be applied to the identification of plasma glow spectrum. It has good performance in terms of model identification accuracy, training speed, and algorithm stability, and can be applied to the identification of plasma glow spectrum.

[0149] The plasma spectral identification device provided by the present invention is described below. The plasma spectral identification device described below and the plasma spectral identification method described above can be referred to in correspondence.

[0150] Figure 4This is a schematic diagram of the structure of the plasma spectral identification device provided by the present invention, as shown below. Figure 4 As shown, it includes:

[0151] The identification module 410 is used to input the plasma spectral information to be identified into the trained LGBM spectral identification model to obtain the target classification spectral parameters corresponding to the plasma spectral information to be identified, so as to perform spectral wavelength feature analysis on the plasma spectral information to be identified based on the target classification spectral parameters.

[0152] The trained LGBM spectral recognition model is obtained by training plasma spectral information samples carrying real spectral parameter labels.

[0153] The plasma spectral identification device described in this embodiment can be used to execute the plasma spectral identification method embodiment described above. Its principle and technical effects are similar, and will not be repeated here.

[0154] The plasma spectral identification device provided by this invention trains an LGBM model using plasma spectral information samples carrying real spectral parameter labels to obtain a trained LGBM spectral identification model. The plasma spectral information to be identified is then input into the trained LGBM spectral identification model to obtain the target classification spectral parameters corresponding to the plasma spectral information to be identified. Based on the target classification spectral parameters, spectral wavelength feature analysis is performed on the plasma spectral information to be identified. Thus, by using the machine learning LGBM algorithm, the implicit physical information between the plasma spectral wavelength features is automatically learned, achieving accurate identification and classification of plasma spectra. The device is easy to operate and can greatly reduce the technical cost of plasma spectral diagnosis.

[0155] Figure 5 This is a schematic diagram of the physical structure of the electronic device provided by the present invention, such as... Figure 5 As shown, the electronic device may include a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute the plasma spectral recognition method provided by the above methods. The method includes: inputting the plasma spectral information to be identified into a trained LGBM spectral recognition model to obtain the target classification spectral parameters corresponding to the plasma spectral information to be identified, and performing spectral wavelength feature analysis on the plasma spectral information to be identified based on the target classification spectral parameters; wherein the trained LGBM spectral recognition model is obtained by training on plasma spectral information samples carrying real spectral parameter labels.

[0156] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0157] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the plasma spectral recognition method provided by the above methods. The method includes: inputting the plasma spectral information to be identified into a trained LGBM spectral recognition model to obtain the target classification spectral parameters corresponding to the plasma spectral information to be identified, so as to perform spectral wavelength feature analysis on the plasma spectral information to be identified based on the target classification spectral parameters; wherein the trained LGBM spectral recognition model is obtained by training on plasma spectral information samples carrying real spectral parameter labels.

[0158] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the plasma spectral identification method provided by the above methods. The method includes: inputting the plasma spectral information to be identified into a trained LGBM spectral identification model to obtain target classification spectral parameters corresponding to the plasma spectral information to be identified, and performing spectral wavelength feature analysis on the plasma spectral information to be identified based on the target classification spectral parameters; wherein the trained LGBM spectral identification model is obtained by training on plasma spectral information samples carrying real spectral parameter labels.

[0159] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0160] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A plasma spectral identification method, characterized in that, include: The spectral information of the plasma to be identified is input into the trained LGBM spectral recognition model to obtain the target classification spectral parameters corresponding to the spectral information of the plasma to be identified. Then, based on the target classification spectral parameters, the spectral wavelength feature analysis of the spectral information of the plasma to be identified is performed. The target classification spectral parameters include at least one type of parameter used to characterize the spectral characteristics of the plasma to be identified, namely, ambient pressure, discharge power, and discharge location. The trained LGBM spectral recognition model is obtained by training on plasma spectral information samples carrying real spectral parameter labels; the plasma spectral information samples include 2048 characteristic wavelengths, and the wavelength range of the characteristic wavelengths is from 366.19nm to 1051.14nm.

2. The plasma spectral identification method according to claim 1, characterized in that, Before inputting the plasma spectral information to be identified into the trained LGBM spectral recognition model, the following steps are also included: Obtain the raw plasma spectral information dataset; The original plasma spectral information dataset is preprocessed using feature engineering methods and / or synthetic sampling methods to obtain a plasma spectral information sample dataset; The LGBM spectral recognition model was trained using the plasma spectral information sample dataset through cross-validation.

3. The plasma spectral identification method according to claim 2, characterized in that, After training the LGBM spectral recognition model using the plasma spectral information sample dataset, the process includes: The parameters of the LGBM spectral recognition model are optimized using a grid search method to obtain the optimal parameters of the LGBM spectral recognition model.

4. The plasma spectral identification method according to claim 2, characterized in that, The method employing cross-validation, using the plasma spectral information sample dataset to train the LGBM spectral recognition model, includes: A cross-validation method is used to obtain a training sample set from the plasma spectral information sample dataset; Multiple sets of training samples are obtained by taking the plasma spectral information samples in the training sample set and the real spectral parameter labels carried by the plasma spectral information samples as a set of training samples. The LGBM spectral recognition model was trained using the aforementioned multiple sets of training samples.

5. The plasma spectral identification method according to claim 4, characterized in that, The process of training the LGBM spectral recognition model using the multiple sets of training samples includes: For any set of training samples, the training samples are input into the LGBM spectral recognition model, and the predicted probability corresponding to the training samples is output; the predicted probability is used to determine the classification result of the training samples. Using a preset loss function, the loss value is calculated based on the predicted probability corresponding to the training sample and the true spectral parameter label in the training sample; If the loss value is less than a preset threshold, the LGBM spectral recognition model training is complete.

6. A plasma spectral identification device, characterized in that, include: The identification module is used to input the plasma spectral information to be identified into the trained LGBM spectral identification model, obtain the target classification spectral parameters corresponding to the plasma spectral information to be identified, and then perform spectral wavelength feature analysis on the plasma spectral information to be identified based on the target classification spectral parameters. The target classification spectral parameters include at least one type of parameter used to characterize the spectral characteristics of the plasma to be identified, namely, ambient pressure, discharge power, and discharge location. The trained LGBM spectral recognition model is obtained by training on plasma spectral information samples carrying real spectral parameter labels; the plasma spectral information samples include 2048 characteristic wavelengths, and the wavelength range of the characteristic wavelengths is from 366.19nm to 1051.14nm.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the plasma spectral identification method as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the plasma spectral identification method as described in any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the plasma spectral identification method as described in any one of claims 1 to 5.