Dust covering target LIBS spectrum effectiveness identification method based on PHATE dimensionality reduction

The feature dimensionality reduction and K-means algorithm were performed through the PHATE algorithm to perform cluster analysis, which solved the problem of sand and dust interference in LIBS spectral data, and achieved accurate identification and distinction of LIBS spectra under Mars dust coverage.

CN120030385APending Publication Date: 2025-05-23SHANGHAI INSTITUTE OF TECHNICAL PHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510097519.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23

Smart Images

  • Figure CN120030385A_ABST
    Figure CN120030385A_ABST
Patent Text Reader

Abstract

The invention discloses a dust covering target LIBS spectrum effectiveness identification method based on PHATE dimensionality reduction, which combines a PHATE (affinity-based track embedding thermal diffusion potential) algorithm and a K-means algorithm to realize accurate identification of dust covering target LIBS spectrum effectiveness. The method comprises two stages of feature dimensionality reduction and clustering analysis: firstly, based on a PHATE algorithm, converting high-dimensional LIBS spectral data into low-dimensional feature vectors, and extracting spectral core features while retaining a data global structure; and then based on a K-means algorithm, carrying out clustering analysis on the low-dimensional spectral feature vectors so as to distinguish three types of different spectral data including a dust spectrum, a dust-matrix transition spectrum and a matrix spectrum. The method has the advantages of being simple and efficient in data processing, high in spectrum effectiveness recognition accuracy and high in robustness, is beneficial for screening out a real effective target spectrum from the dust covering target LIBS spectrum, and is suitable for rapidly analyzing LIBS spectrums collected in unclean scenes such as field exploration and deep space exploration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of laser spectral analysis, and in particular to a method for identifying the effectiveness of LIBS spectra of dust-covered targets based on PHATE dimension reduction. The method utilizes the PHATE algorithm to perform feature dimension reduction and the K-means algorithm to perform clustering, and can effectively distinguish the LIBS spectra of dust-covered targets, and preliminarily identify whether the target sample covered by dust has mineral components of research value, thereby accurately and efficiently screening out potential high-value spectra from a large number of LIBS spectra for subsequent further chemical composition analysis. Background Art

[0002] Laser-induced breakdown spectroscopy (LIBS) is a chemical composition detection technology based on laser-induced plasma radiation. It has the advantages of fast response, micro-damage to samples, simultaneous analysis of multiple elements, and remote detection. Therefore, it is widely used in environmental monitoring, industrial testing, planetary exploration and other fields. In particular, the special advantage of remote detection makes LIBS technology play an important role in the field of planetary exploration. At present, LIBS technology has been successfully applied to Mars exploration missions three times. NASA's "Curiosity" and "Perseverance", as well as my country's "Zhurong", all three Mars rovers are equipped with scientific payloads equipped with LIBS systems, namely the Chemical Camera (ChemCam), the Super Camera (SuperCam) and the Mars Surface Composition Detector (MarSCoDe). Using LIBS technology, these scientific payloads can detect and analyze the chemical element composition of soil, rocks and other materials on the surface of Mars.

[0003] However, compared to the clean environment of a conventional laboratory, the Martian atmosphere is full of micron-sized particles, namely Martian dust. These Martian dust covers the surface of LIBS detection target materials such as rocks, and the thickness of the dust layer can range from microns to meters, see reference [1]. Although high-energy laser pulses can remove part of the dust on the target surface during LIBS detection, they cannot completely remove it, see references [2,3], which will cause the LIBS spectrum to contain dust composition information to a certain extent, thus bringing significant interference to the analysis of the composition of the real target material.

[0004] In order to reduce the negative impact of dust on the target surface on the analysis of Mars LIBS spectral data, the international ChemCam and SuperCam payload teams have adopted a method of treating the first five LIBS spectra obtained in each detection as spectra derived from dust, and excluding these spectral data from the total data set for subsequent analysis. However, in fact, the sixth and subsequent LIBS spectra obtained in each detection may still contain information on the composition of dust, as detailed in reference [4]. Therefore, this overly simple judgment method cannot accurately determine the validity of the LIBS spectrum, that is, whether the spectrum is derived from the real base material under the dust layer.

[0005] In order to more accurately and reliably judge the effectiveness of LIBS spectra and identify whether the spectral data comes from surface dust or target substrate material, it is necessary to develop more reasonable spectral data clustering or classification methods. However, there are currently no research reports on this issue at home and abroad.

[0006] When clustering or classifying LIBS spectral data from different sources, the simplest and most direct way is to directly process the entire LIBS spectrum. However, a LIBS spectrum can usually contain thousands of data points (each original LIBS spectrum collected by MarSCoDe contains 5400 data points). Directly analyzing the entire spectrum is not only very time-consuming, but also increases the difficulty of clustering or classification due to the presence of a large number of redundant features.

[0007] In order to retain the main features of thousands of data points in each LIBS spectrum and remove redundant features, the spectral data can be first processed for feature dimension reduction, thereby reducing the difficulty of analysis and saving computing time. In this mode of first reducing the dimension and then analyzing, the rationality of the feature dimension reduction algorithm is one of the most critical factors affecting the accuracy of clustering or classification results. At present, the commonly used feature dimension reduction algorithms in data analysis are principal component analysis (PCA) and t-stochastic neighbor embedding (t-SNE), and their principles and disadvantages are as follows:

[0008] For the PCA algorithm, LIBS spectra are usually converted into a set of linearly unrelated feature vectors through orthogonal transformation. The features closer to the front in the feature vector show higher differences. Selecting the first N features (depending on the situation) as new features can achieve feature dimensionality reduction and retain most of the information of the LIBS spectrum. PCA is simple and feasible, and can well retain the global structure of the data. However, in terms of the interpretability of the dimensionality reduction results, the meaning of the principal component feature dimension in PCA is ambiguous and has poor interpretability; in terms of the degree of original information retention, the standard for PCA dimensionality reduction is to select the principal component that makes the original data have the largest variance on the new coordinate axis, but features with small variance are not necessarily unimportant. Such a standard will inevitably lose some effective information in the dimensionality reduction process, resulting in information loss; in addition, PCA is based on the linear assumption, so for data with nonlinear structure, PCA may find it difficult to extract its main features; finally, in terms of the robustness of the dimensionality reduction process, PCA may be overloaded by fluctuations or outliers in the data set and cannot accurately reflect the local structure of the data.

[0009] For the t-SNE algorithm, data dimensionality reduction is usually achieved by mapping the similarity between points in the high-dimensional space to the probability distribution in the low-dimensional space, and then minimizing the difference in the distribution between the two spaces. t-SNE works well in visualizing high-dimensional data and is very good at preserving the local structure in high-dimensional data. In the low-dimensional space after dimensionality reduction, similar data points will cluster together to form obvious clusters. However, in terms of model complexity, the computational complexity of t-SNE is relatively high, especially when dealing with large-scale data sets, it takes a long time to complete the dimensionality reduction task; in addition, when using t-SNE, it is necessary to try different initialization points and run the algorithm multiple times to select the best result to eliminate the influence of the local optimal solution; at the same time, t-SNE has poor robustness in the dimensionality reduction process, and has strict requirements on the setting of hyperparameters such as learning rate and number of iterations. Improper settings may lead to poor dimensionality reduction effects or even failure; finally, in terms of data structure information, t-SNE can only retain the local structure information of the data, but cannot well retain the global structure information.

[0010] References

[0011] [1] Newman CE, et al. 7.24-Martian Dust / / SHRODER J F. Treatise onGeomorphology (Second Edition). Oxford; Academic Press. 2022: 637-66.

[0012] [2]SKSharma,et al.Combined remote LIBS and Raman spectroscopy at8.6m of sulfur-containing minerals,and minerals coated with hematite or covered with basaltic dust.Spectrochimica Acta Part A 68(2007)1036–1045.

[0013] [3]Vasily N. Lednev, et al. Combining Raman and laser induced breakdownspectroscopy by double pulse lasing. Analytical and Bioanalytical Chemistry410 (2018) 277–286.

[0014] [4] Roger C. Wiens, et al. The ChemCam Instrument Suite on the MarsScience Laboratory (MSL) Rover: Body Unit and Combined System Tests. SpaceScience Reviews 170 (2012) 167–227. Summary of the invention

[0015] In view of the above background and the deficiencies of the prior art, the present invention proposes a method for identifying the effectiveness of LIBS spectra of dust-covered targets based on PHATE dimension reduction, which is used to effectively identify the spectra of surface dust and the spectra of covered scientific targets during the in-situ LIBS detection of Mars. This method utilizes the unique advantages of PHATE in feature dimension reduction and information extraction, and applies it to the effective feature extraction of Mars LIBS spectra. At the same time, it utilizes the high efficiency and fast clustering characteristics of the K-means algorithm and applies it to the clustering analysis of low-dimensional feature vectors to achieve effective distinction between three different types of spectral data: dust spectra, dust-base transition spectra, and base spectra, thereby providing important technical support and practical solutions for the subsequent efficient discrimination and screening of effective spectra, thereby accurately analyzing the real Mars environment LIBS spectra of Mars dust-covered samples.

[0016] The technical solution of the present invention is as follows:

[0017] The overall process of this technical solution can be divided into six steps, as shown in the attached manual. Figure 1The details are as follows:

[0018] S1. Preliminary preparation: Prepare the detection sample of laser induced breakdown spectroscopy (LIBS). The sample is a target with dust covered on the surface prepared by a specific process, that is, a dust-covered target, hereinafter referred to as a target. The name and composition information of the base material and the surface dust material of each target are counted and recorded. The target preparation process includes two parts. The first is to calculate the dust mass M corresponding to the dust sheet with a thickness of D based on the compaction density ρ of the dust material and the area S of the sheet. d , weighing mass is M d The dust material is placed on the powder plane and flattened in the tablet press mold; the base material is placed on the powder plane and flattened, and boric acid powder is placed on the outside of the mold to wrap the dust material and the base material. The pressure is set to 238MPa, the pressing time is 1.5 minutes, the mold diameter is 40mm, and the tablet press is started to complete the preparation of the target. When preparing each target, the mass of the dust, matrix and boric acid powder should be kept as the same as the previous target as much as possible, so that the prepared target has good consistency in shape, size and weight.

[0019] After the pressing is completed, check whether the target surface is completely covered by the dust cover layer. After the inspection is correct, put it into the sample bag for storage, name the sample as "dust name" + "target base material name", and record the material chemical composition information of the dust material and the target material.

[0020] S2. Use the LIBS spectrum detection device to collect the LIBS spectrum of the target. During the spectrum collection process of each target, the same position of the target needs to be continuously excited P times or more with a laser pulse to ensure that the dust covering layer on the target surface is broken down and the spectrum of the base material can be collected. The LIBS spectrum collected by the experiment is defined as the original spectrum data set. During the experiment, a LIBS device with strong laser energy and longitudinal profiling capability needs to be used to perform continuous LIBS excitation on the target to ensure that the laser pulse energy is sufficient to break through the dust layer at the excited position on the target surface after continuous excitation of more than P pulses. During the spectrum collection process, attention should be paid to the changes in the characteristic peaks of the elements to ensure that the surface dust covering layer can be broken down and the LIBS spectrum of the target base material can be collected.

[0021] S3. Preprocess the LIBS spectra in the original spectral dataset. Usually, the preprocessing steps include dark background subtraction, noise filtering, background baseline removal, wavelength calibration, invalid pixel screening, channel splicing, and sub-channel spectral normalization. Among them, the dark background subtraction operation refers to subtracting the dark background spectrum from the original LIBS spectrum to obtain the effective spectrum. The dark background spectrum refers to the spectrum of the spectrometer response when there is no light signal under the same instrument parameters as those for collecting the original spectrum, generally obtained by averaging 60 spectra; the noise filtering operation refers to using the adaptive wavelet denoising method to remove Gaussian noise in the effective spectrum to obtain the denoised spectrum; the background baseline removal operation refers to using the asymmetric least squares baseline correction method to remove the continuous baseline in the denoised spectrum; wavelength calibration refers to converting the spectrometer pixel number to the wavelength value through multivariate quadratic fitting; invalid pixel screening refers to removing the pixel response values outside the wavelength range in each band of the LIBS spectrum; channel splicing refers to splicing the LIBS spectra of multiple bands after invalid pixel screening into a whole according to the wavelength order; sub-channel spectral normalization refers to mapping each channel of the LIBS spectrum to the interval [0, 1] so that the sum of the LIBS spectral intensity values in each single channel is 1 and the sum of the LIBS spectral intensity values of the whole LIBS spectrum is 3. The LIBS spectrum after preprocessing is defined as the dust-covered target LIBS spectral dataset.

[0022] S4. In the dust-covered target LIBS spectral dataset, let the number of data points contained in each spectrum be M, and define M as the spectral dimension; perform feature dimensionality reduction operations on all the LIBS spectra of each target in turn, reducing the spectral dimension from M to N, where N < M, and the specific value of N is set according to the specific dataset and actual requirements. The feature dimensionality reduction operation is performed by the Potential of Heat-diffusion for Affinity-based Trajectory Embedding (PHATE) algorithm. After the feature dimensionality reduction operation, each LIBS spectrum is correspondingly transformed into an N-dimensional feature vector. As a non-linear dimensionality reduction algorithm, the PHATE algorithm is often used to process high-dimensional data to extract its low-dimensional representation, and is especially good at revealing the continuous structure of the data (such as the development trajectory). For the feature dimensionality reduction process, the input is all the LIBS spectra of a single target, and the output is a matrix composed of the low-dimensional feature vectors corresponding to all the spectra of this target, with the size of Q × N, where Q is the number of LIBS spectra of this target and N is the number of features after dimensionality reduction, depending on the actual situation. The relevant parameters of the PHATE algorithm include ndim, k, t_max, and mds_methods. Among them, ndim represents the feature dimension of the output vector, determining the number of features retained after dimensionality reduction; k represents the number of nearest neighbor points considered when constructing the graph, and this variable affects the local relationship modeling and can be adjusted according to the data density; t_max represents the maximum value of the optimal number of time steps for finding the Markov diffusion process; mds_methods represents the multi-dimensional scaling analysis method used to map the diffusion potential energy to the low-dimensional space.

[0023] When adjusting the parameters of the PHATE algorithm, it is necessary to make comprehensive considerations for various aspects. The default value of ndim is 2, that is, the output two-dimensional feature vector is used for two-dimensional visualization analysis. If the analysis result is not good, it can be set to a higher dimension (such as 3, 5); the default value of k is 5. If the input data is sparse, the k value needs to be appropriately increased. If the input data is dense, it needs to be appropriately reduced. t_max determines the maximum time step of the Markov diffusion process. Setting it to a smaller value will urge the algorithm to pay more attention to the local information of the input data, and setting it to a larger value will pay more attention to the global structure. mds_methods affects the calculation method of the final output low-dimensional feature vector. By default, metric multidimensional scaling analysis (Metric MDS) is used for calculation. The other two calculation methods are classical multidimensional scaling analysis (Classical MDS) and non-metric multidimensional scaling analysis (Non-metricMDS), which are respectively suitable for scenarios where the data strictly meets the Euclidean distance condition and scenarios where the data is a nonlinear relationship.

[0024] S5. Clustering operations are performed on all N-dimensional feature vectors of each target in turn, and the feature vectors are divided into three categories, which respectively indicate that the LIBS spectrum corresponding to the feature vector belongs to the surface dust layer spectrum, the dust-matrix transition zone spectrum and the matrix spectrum. The clustering operation is performed by the K-means clustering algorithm K-means. As a commonly used unsupervised learning clustering algorithm, the K-means algorithm is widely used due to its high efficiency and simplicity. Its core goal is to divide the input data into t clusters so that the samples in the same cluster are as similar (tight) as possible, and the samples between different clusters are as different (separated) as possible. For the clustering process, the input is all N-dimensional feature vectors and the number of categories corresponding to the LIBS spectrum of a single target, and the output is the category corresponding to each LIBS spectrum. The relevant parameters of the K-means algorithm include n_clusters, Distance, MaxIter and Start. Among them, n_clusters represents the number of clusters in the data, that is, the number of categories formed after clustering; Distance refers to the calculation method of the distance measurement between the cluster centroid and the sample point in the feature space; MaxIter refers to the maximum number of iterations of the algorithm; Start refers to the method of selecting the initial cluster centroid position.

[0025] When adjusting the parameters of the K-means algorithm, it is necessary to make comprehensive considerations for all aspects. n_clusters determines the number of categories, which directly affects the clustering results. It needs to be determined based on the prior knowledge of the input data and the specific application scenario before clustering. Distance uses Euclidean distance by default, but sometimes other distance metrics are more suitable for data characteristics. When using Euclidean distance, the data can be standardized or scaled to adapt to the assumptions of Euclidean clustering. MaxIter determines the maximum number of iterations for each run, which affects the speed and effect of algorithm convergence. For large-scale data sets or scenarios that may take more time to converge, it can be appropriately increased. If the amount of data is small, the number of iterations can be appropriately reduced to speed up the calculation. Start determines the position of the initial cluster centroid. By default, the heuristic method is used to repeatedly calculate and find the initial cluster centroid. This method can converge to the optimal value faster. Other methods include randomly selecting the initial cluster centroid from the data set.

[0026] S6. Evaluate the execution effect of steps S4 and S5 based on the spectral similarity index, and optimize the relevant parameters of the algorithms used in the feature dimensionality reduction process and the clustering process. The present invention introduces three evaluation indicators for evaluating the accuracy of spectral validity recognition, namely, the silhouette coefficient S, the spectral angle mapping θ, and the cluster centroid distance and d. Among them, the silhouette coefficient S is used as the core indicator for evaluating the performance of feature dimensionality reduction and clustering algorithms, and the other two indicators θ and d are used as auxiliary indicators; the silhouette coefficient S is directly normalized to obtain S norm , the spectrum angle plot θ is normalized after taking the inverse The cluster centroid distance and d are directly normalized to get d norm , S norm , and norm Sum to get Eval norm , with Eval norm As the evaluation index used in the subsequent parameter optimization process; in the process of optimizing the parameters related to steps S4 and S5, the scheme used is the grid optimization algorithm, which sets the grid value range and grid point spacing of each parameter to be optimized, tries every possible parameter combination scheme, and calculates the above evaluation index Eval norm Until all possible parameter combinations are traversed, the one that can make Eval norm The solution with the largest parameter combination is taken as the final solution, thus completing the optimization of relevant parameters.

[0027] The core indicator silhouette coefficient S can characterize whether each feature point is correctly assigned to the corresponding cluster. Therefore, when the algorithm and its parameters used in the clustering process are consistent, the silhouette coefficient S can clearly reflect the performance of the feature dimensionality reduction algorithm, that is, whether the feature points of the same category of LIBS spectra can be tightly aggregated after feature dimensionality reduction, and the feature points of different categories of LIBS spectra can be significantly separated. The silhouette coefficient S takes into account both the compactness within the cluster and the separation between clusters. The specific calculation method is shown in formula (1):

[0028]

[0029] Among them, S(i) is the silhouette coefficient of the ith feature point; a(i) represents the average distance between the ith feature point and other feature points in the cluster to which it belongs, called "cohesion", reflecting the compactness within the cluster; b(i) represents the average distance between the ith feature point and the feature point of the nearest cluster in the non-cluster to which it belongs, called "separation", reflecting the separation between clusters. When the silhouette coefficient is close to 1, it means that the ith feature point is compact within the cluster to which it belongs and is well separated from other clusters; when the silhouette coefficient is close to 0, it means that the feature point is at the cluster boundary, or the distance between points within the cluster is equivalent to the distance between clusters; when the silhouette coefficient is close to -1, it means that the feature point may be mistakenly classified into the wrong cluster. In order to evaluate the effect of the entire feature dimensionality reduction process, the average value of the silhouette coefficients of all feature points is usually calculated. This average value is used as the silhouette coefficient of the entire feature dimensionality reduction process for subsequent performance evaluation.

[0030] Auxiliary indicator spectral angle mapping θ is a common spectral similarity analysis method. This method evaluates the similarity between two spectra by calculating the angle between the sample spectrum and the reference spectrum. In order to verify the feature dimensionality reduction performance of PHATE and the clustering performance of the K-means algorithm, the present invention uses the LIBS spectrum corresponding to the feature point closest to the centroid of each cluster as the reference spectrum of the cluster, and the remaining spectra corresponding to the feature points of the cluster as sample spectra. The angle between each sample spectrum and the reference spectrum is calculated in turn and averaged, so as to judge the quality of feature dimensionality reduction and clustering effects according to the similarity of the LIBS spectra corresponding to the cluster. The specific calculation method is shown in formula (2):

[0031]

[0032] Among them, X * is the reference spectrum, X i is the sample spectrum.

[0033] The auxiliary index cluster centroid distance and d is an evaluation index for the accuracy of LIBS spectrum recognition of dust-covered targets proposed in the present invention. This index calculates the sum of the Euclidean distances between the centroid position of the cluster corresponding to the matrix LIBS spectrum and the centroid position of the cluster corresponding to the dust LIBS spectrum and the dust-matrix transition zone LIBS spectrum. The specific calculation method is shown in formula (3):

[0034]

[0035] Where d represents the cluster centroid distance and L subtrate , L dust , L dust-subtrate They represent the centroid position vectors of the clusters corresponding to the matrix LIBS spectrum, dust LIBS spectrum and dust-matrix transition region LIBS spectrum, respectively.

[0036] The mechanism of action of the present invention is as follows:

[0037] The present invention applies the PHATE algorithm to the dimensionality reduction of LIBS data of dust-covered targets. The PHATE algorithm encodes local data information through local similarity, and uses potential distance to encode the global relationship in the data, and finally embeds the potential distance information into a low dimension for visualization, thereby ultimately achieving the purpose of data dimensionality reduction. In actual LIBS data dimensionality reduction, the PHATE algorithm can well overcome the defects of conventional feature dimensionality reduction algorithms such as the PCA algorithm and the t-SNE algorithm, and accurately capture the true structure of high-dimensional data by denoising the spectral data, thereby robustly and clearly visualizing the spectral data and achieving data dimensionality reduction. The PHATE algorithm is applied to the LIBS spectral identification task. In terms of visualization results, the clustering effect of the data has been qualitatively improved compared with the above two conventional algorithms, showing a higher spectral discrimination.

[0038] Beneficial Effects

[0039] Compared with the prior art, the advantages of the present invention are:

[0040] 1. Compared with the scheme of directly identifying the validity of the entire LIBS spectrum without using the feature dimensionality reduction algorithm, the PHATE feature dimensionality reduction algorithm used in the present invention can significantly reduce the feature dimension and effectively reduce the influence of noise or redundant feature information in the spectrum on the subsequent clustering results, thereby improving the clustering accuracy, improving the calculation speed, and improving the interpretability of the results.

[0041] Second, compared with the LIBS spectrum validity identification scheme that uses the PCA algorithm as a feature dimension reduction algorithm, the PHATE algorithm used in the present invention can more accurately capture the true structure of high-dimensional LIBS data, and the visualization effect and interpretability are better. In addition, unlike the PCA algorithm that only focuses on the global structure of the data, the PHATE algorithm comprehensively retains the global and local structures of the data, can identify nonlinear effects in LIBS, and thus make the results more accurate, and is almost not affected by the fluctuations in the data set like the PCA algorithm.

[0042] 3. Compared with the LIBS spectrum validity identification scheme using the t-SNE algorithm as the feature dimension reduction algorithm, the PHATE algorithm used in the present invention requires relatively fewer computing resources and has a shorter running time when processing LIBS data, which can greatly save time costs. In addition, unlike the t-SNE algorithm that only focuses on the local structure information of the data, the PHATE algorithm can well obtain and retain the local and global structure information of the data, and is almost unaffected by hyperparameter restrictions and local optimal solutions.

[0043] In summary, the present invention uses the PHATE algorithm as a feature dimensionality reduction method in the dust-covered surface LIBS spectrum validity identification work, and combines the K-means algorithm to distinguish whether the target of LIBS laser pulse ablation is a matrix sample or the simulated Martian dust covering the surface. The PHATE algorithm can effectively transform high-dimensional LIBS spectral data into low-dimensional feature vectors, extracting the core features of the spectrum while retaining the global structure of the data; the K-means algorithm can effectively perform cluster analysis on the low-dimensional spectral feature vectors, thereby distinguishing three different types of spectral data: dust spectrum, dust-matrix transition spectrum, and matrix spectrum, and realizing accurate identification of the effectiveness of the LIBS spectrum of the dust-covered target. The advantages of the present invention include: simple and efficient data processing, high degree of feature dimensionality reduction, complete data structure information, high accuracy of spectral validity identification results, strong robustness, and strong interpretability. It is helpful to screen out truly effective target spectra from the LIBS spectra of dust-covered targets, and can better meet the application needs of LIBS spectrum effectiveness identification of unclean surface samples. It is suitable for rapid analysis of LIBS spectra collected in unclean scenes such as field exploration and deep space exploration. Therefore, it has important value in the field of laser spectral analysis technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a schematic diagram of the overall flow of the technical solution.

[0045] Figure 2 This is a comparison chart of the 1st, 40th, 80th, 120th, and 160th LIBS spectra of the "dolomite powder + basalt" target in continuous excitation LIBS spectra.

[0046] Figure 3The dimensionality reduction results of the LIBS spectra of dust-covered targets using the PHATE algorithm and two other common dimensionality reduction algorithms are as follows: (a) feature dimensionality reduction graph based on the PHATE algorithm, (b) feature dimensionality reduction graph based on the PCA algorithm, and (c) feature dimensionality reduction graph based on the t-SNE algorithm. The numbers in the color bar on the right represent the spectral numbers. DETAILED DESCRIPTION

[0047] The following is a specific experimental case to illustrate the application process of the method described in the summary of the invention:

[0048] 1. In this example, the detection sample prepared for laser induced breakdown spectroscopy (LIBS) is a dust-covered target prepared through a specific process, hereinafter referred to as the target. The dust material is the JMSS-1 model simulated Martian soil, which is basically consistent with the composition of the sand and dust material on the surface of Mars. The base material is selected from 10 potential water-bearing minerals on the surface of Mars, including chlorite, serpentine, calcite, talc, zeolite, muscovite, sodium montmorillonite, magnesite, kaolinite and kaolin. When preparing dust-covered targets, first calculate the dust mass required to press 0.6 mm dust flakes based on the compaction density of the dust material, weigh the corresponding mass of dust material and spread it flat in the tablet press mold; then place the base material on the powder plane and spread it flat, put boric acid powder on the outside and upper side of the mold to wrap the dust material and the base material, start the fully automatic fluorescent tablet press to press the powder sample into a cake shape, and set the parameters to an upper pressure limit of 30 T, a lower pressure limit of 27 T, a mold size of 40 mm, a mold pressure of 238 MPa, and a pressurization time of 2 minutes. After the pressing is completed, check whether the target surface is completely covered by the dust layer. After confirmation, make corresponding labels. The naming format is "dust name" + "target base material name", for example "JMSS-1 + chlorite".

[0049] 2. Use the LIBS spectrum detection device to collect the LIBS spectrum data of the target. In this example, the prepared target is placed in a vacuum chamber that simulates the atmospheric environment of the Martian surface, and the LIBS spectrum is collected in this environment. The detection device is a ground backup of the MarSCoDe payload of the Mars surface composition detector. The LIBS detection distance is 2m, the laser single pulse energy is about 9mJ, and the laser wavelength is 1064.4nm. One sampling position is selected for each sample, and 180 valid LIBS spectra are continuously collected at each position, totaling 1800 spectra, which constitute the original spectral data set. Instructions attached Figure 2The changes in the intensity of the characteristic peaks of the elements in the 240-340nm band of the 1st, 40th, 80th, 120th and 160th LIBS spectra of the basalt sample covered with dolomite powder are given when continuous LIBS spectra are collected. The results show that with the increase in the number of excitations, the intensity of the characteristic peaks of Mg and Ca elements representing dolomite powder gradually weakens, and the characteristic peaks of Si, Fe and Al elements representing basalt samples gradually increase from zero to one. This shows that the dust cover layer is broken down and the underlying substrate material is excited by LIBS, which is consistent with the experimental hypothesis.

[0050] 3. The LIBS spectra in the original spectral data set are preprocessed, including dark background subtraction, noise filtering, background baseline removal, wavelength calibration, invalid pixel screening, channel stitching and channel spectral normalization. The preprocessed LIBS spectra are defined as the dust cover target LIBS spectral data set, and each LIBS spectrum can be represented as a 1×4506 vector.

[0051] 4. Perform feature dimension reduction on all LIBS spectra of each target in turn to obtain a two-dimensional feature vector of the target LIBS spectrum. The feature dimension reduction process is performed by the affinity-based trajectory embedding thermal diffusion potential algorithm PHATE. According to the content of the invention, the relevant parameters of the PHATE algorithm are comprehensively adjusted. After the parameter optimization process in step 6, ndim is set to 2, that is, the two-dimensional visualization feature is output; k is set to 5, that is, the number of neighbor points considered when constructing the graph is 5; t_max is set to 100, that is, the maximum number of steps when finding the optimal time step of the Markov diffusion process is no more than 100; mds_methods is set to 'MetricMDS', that is, the diffusion potential is mapped to a low-dimensional space using metric multidimensional scaling analysis and a two-dimensional visualization feature vector is output.

[0052] After completing the algorithm parameter setting, all LIBS spectra of each target are sent to the PHATE algorithm for training and feature dimension reduction in turn, and finally the two-dimensional visual feature vector corresponding to the LIBS spectrum is output. This vector can be used to intuitively observe the difference between the surface dust layer spectrum, the dust-matrix transition zone spectrum and the matrix spectrum, and can also be used for subsequent clustering analysis. In order to demonstrate the superiority of the PHATE algorithm in feature dimension reduction, two other feature dimension reduction algorithms were used to perform feature dimension reduction operations on the LIBS spectral data set of dust-covered targets, and compared with the dimension reduction results of the PHATE algorithm. In this example, the three feature dimension reduction algorithms are the PHATE algorithm, the PCA algorithm, and the t-SNE algorithm. The feature dimension reduction results for all LIBS spectra of the "JMSS-1+chlorite" target are shown in the attached manual. Figure 3As shown in the figure. It is obvious from the figure that the two-dimensional visualization features reduced by the PHATE algorithm show clear class boundaries, and can easily separate the surface dust layer spectrum, the dust-matrix transition zone spectrum and the matrix spectrum. The two-dimensional visualization feature distributions output by the other two feature dimensionality reduction algorithms are overlapping or mixed, and it is difficult to distinguish the obvious boundaries of different categories or clusters through simple visual observation. They more reflect the continuous change of feature points and do not retain the boundary information of the original high-dimensional spectral data well. After completing the cluster analysis, we will further comprehensively compare the results of the three feature dimensionality reduction algorithms from multiple evaluation index dimensions.

[0053] 5. Perform clustering operations on all two-dimensional feature vectors of each target in turn, and divide the feature vectors into three categories, indicating that the LIBS spectrum corresponding to the feature vector belongs to the surface dust layer spectrum, the dust-matrix transition zone spectrum and the matrix spectrum, respectively. The clustering operation is performed based on the K-means clustering algorithm K-means. According to the content of the invention, the relevant parameters of the K-means algorithm are comprehensively adjusted. After the parameter optimization process in step 6, n_clusters is set to 3, indicating that all two-dimensional feature vectors of each target need to be divided into three categories; Distance is set to 'sqeuclidean', indicating that the K-means algorithm uses Euclidean distance to characterize when calculating the distance between the feature point and each cluster centroid; MaxIter is set to 100, indicating that the maximum number of iterations for each run is 100, which is sufficient for the clustering task of 180 feature points each time; Start is set to 'plus', and the initial cluster centroid is found by repeated calculation using a heuristic method, which can converge to the optimum faster.

[0054] After completing the algorithm parameter setting, all the two-dimensional feature vectors of each target are sent to the K-menas algorithm for training and cluster analysis, and finally the centroid position of each cluster and the cluster category corresponding to each two-dimensional feature vector are obtained. According to the clustering results, it can be identified whether each LIBS spectrum belongs to the surface dust layer spectrum, the dust-matrix transition zone spectrum or the matrix spectrum during the continuous excitation process, providing an effective technical solution for the subsequent identification and screening of potential mineral spectra of interest on Mars. We further evaluate the results of dimensionality reduction and cluster analysis of the LIBS spectrum features of each target in detail based on specific indicators.

[0055] 6. Evaluate the execution effects of Steps 4 and 5 based on relevant indicators, and comprehensively optimize the relevant parameters of the algorithms used in the feature dimension reduction process and the clustering process according to the described invention content, and finally obtain the optimal feature dimension reduction and clustering analysis results. The main evaluation indicators include the silhouette coefficient S, the spectral angle mapping θ, and the cluster centroid distance d. Among them, the silhouette coefficient S is the core indicator for evaluating the performance of the algorithms used in the feature dimension reduction and clustering processes, and the other two indicators are auxiliary indicators; the parameter optimization scheme is the grid search algorithm. In this example, we respectively used the PHATE algorithm, the PCA algorithm, and the t-SNE algorithm to perform feature dimension reduction on the LIBS spectra of 10 kinds of dust-covered target samples, and used the K-means algorithm to perform clustering analysis on the two-dimensional feature vectors obtained after dimension reduction to obtain the spectral categories to which each LIBS spectrum belongs. In order to evaluate the performance of different combinations of feature dimension reduction algorithms and the K-means algorithm for feature dimension reduction and clustering analysis, we calculated the above three evaluation indicators respectively for the optimal feature dimension reduction and clustering analysis results of each dust-covered target sample, and the calculation method is as described in S6 of the invention content.

[0056] The core indicator, the silhouette coefficient S, can characterize whether each feature point is correctly assigned to the corresponding cluster. Therefore, when the algorithms and their parameters used in the clustering process are the same, the silhouette coefficient S can clearly reflect the performance of the feature dimension reduction algorithm, that is, whether it can make the feature points of the LIBS spectra of the same category closely aggregated and the feature points of the LIBS spectra of different categories significantly separated after feature dimension reduction. In the calculation process, we calculate the silhouette coefficient for all the feature points obtained by dimension reduction using a certain feature dimension reduction algorithm for each target sample and take the average. It can be found that the average silhouette coefficient of the two-dimensional feature points of the LIBS spectra of 10 kinds of dust-covered target samples based on the PHATE algorithm reaches 0.78, while the average silhouette coefficients of the two-dimensional feature points based on the PCA algorithm and the t-SNE algorithm are 0.69 and 0.72 respectively, both lower than that of the PHATE algorithm. At the same time, the maximum value of the silhouette coefficient of the two-dimensional feature points obtained by the PHATE algorithm is 0.89, and the minimum value is 0.71, showing the highest silhouette coefficient on most target samples and only being equal to the other two algorithms on a few target samples. This shows that the feature dimension reduction algorithm based on the PHATE algorithm can better retain the boundary information of the target LIBS spectra after dimension reduction, facilitating the K-means algorithm to distinguish spectral categories. Table 1 lists the silhouette coefficient S of the LIBS spectra of each dust-covered target sample after dimension reduction using three feature dimension reduction algorithms and clustering using the K-means algorithm.

[0057] Table 1

[0058] Sample name <![CDATA[S PHATE ]]> <![CDATA[S PCA ]]> <![CDATA[S t-SNE ]]> JMSS-1+chlorite 0.79 0.56 0.66 JMSS-1+Serpentine 0.78 0.66 0.72 JMSS-1+calcite 0.79 0.68 0.73 JMSS-1+Talc 0.79 0.68 0.70 JMSS-1+zeolite 0.77 0.72 0.78 JMSS-1+Muscovite 0.89 0.75 0.64 JMSS-1+sodium montmorillonite 0.73 0.73 0.80 JMSS-1+magnesite 0.74 0.75 0.68 JMSS-1+Kaolinite 0.80 0.63 0.69 JMSS-1+Kaolin 0.71 0.69 0.76 average 0.78 0.69 0.72

[0059] The auxiliary indicator spectral angle mapping θ focuses on analyzing the effect of feature dimension reduction and cluster analysis from the LIBS spectrum level. First, the LIBS spectrum corresponding to the feature point closest to the centroid of each category is used as the reference spectrum of the category. The average spectral angle mapping of other LIBS spectra of the category and the reference spectrum is calculated as the spectral angle mapping of the category. The spectral angle mapping of the three categories is calculated in turn and then averaged to obtain the similarity of the spectra in each category in the LIBS spectrum of the target. The results show that after the LIBS spectrum of each target is reduced by three feature dimension reduction algorithms and clustered by K-means, the spectral angle mapping indicators of each combination method are basically close, which means that the dimensionality reduction clustering results obtained by the three combination methods can approximately characterize the category information of the original LIBS spectrum. Table 2 lists the average spectral angle mapping θ of the LIBS spectrum of each dust-covered target.

[0060] Table 2

[0061] Sample name <![CDATA[θ PHATE ]]> <![CDATA[θ PCA > <![CDATA[θ t-SNE ]]> JMSS-1+chlorite 5.98 6.90 6.12 JMSS-1+Serpentine 6.52 6.51 6.44 JMSS-1+calcite 6.68 6.47 6.07 JMSS-1+Talc 5.61 6.29 5.49 JMSS-1+zeolite 7.83 6.74 6.29 JMSS-1+Muscovite 8.81 8.54 8.55 JMSS-1+sodium montmorillonite 6.21 6.41 5.90 JMSS-1+magnesite 8.75 8.65 7.67 JMSS-1+Kaolinite 7.05 6.85 7.02 JMSS-1+Kaolin 6.97 6.12 6.85 average 7.04 6.95 6.64

[0062] The auxiliary index cluster centroid distance and d are used to evaluate the separation between the matrix spectrum category and the other two categories. The distance and distance between the matrix spectrum category centroid and the other two categories centroids are calculated to evaluate whether the feature dimensionality reduction algorithm can fully separate the three categories of LIBS spectra. Among them, after clustering, the two-dimensional feature points obtained by the three feature dimensionality reduction algorithms have an average cluster centroid distance and of 4.49, 4.47 and 4.79 for the 10 targets, respectively. The PHATE algorithm shows relatively high cluster centroid distance and on all 10 targets, while the PCA and t-SNE algorithms obtain results on "JMSS-1+chlorite" and "JMSS-1+calcite" that are much lower than other targets, much lower than the PHATE algorithm. This shows that after the feature points after PHATE algorithm dimension reduction are clustered by K-means, the clusters representing the matrix spectrum category are significantly different from the clusters representing the dust spectrum and the dust-target transition zone spectrum category, and similar results are obtained on different target LIBS spectra without abnormal points, indicating that the results of the PHATE algorithm are highly robust. Table 3 lists the cluster centroid distances and d of the LIBS spectra of each dust-covered target after dimension reduction by three feature dimension reduction algorithms and K-means clustering.

[0063] Table 3

[0064] Sample name <![CDATA[d PHATE ]]> <![CDATA[d PCA ]]> <![CDATA[d t-SNE ]]> JMSS-1+chlorite 4.46 3.68 3.60 JMSS-1+Serpentine 4.53 4.22 4.44 JMSS-1+calcite 4.36 4.05 4.08 JMSS-1+Talc 4.77 4.86 4.16 JMSS-1+zeolite 4.39 5.82 5.04 JMSS-1+Muscovite 4.60 4.52 5.65 JMSS-1+sodium montmorillonite 4.43 4.29 5.28 JMSS-1+magnesite 4.37 4.48 5.34 JMSS-1+Kaolinite 4.74 4.25 4.97 JMSS-1+Kaolin 4.24 4.50 5.33 average 4.49 4.47 4.79

[0065] From the results shown in this example, it can be seen that: in terms of the quantitative evaluation indicators of feature dimensionality reduction and clustering effect, for the core indicator silhouette coefficient S, the performance of the PHATE algorithm is better than that of the PCA algorithm and the t-SNE algorithm. For the auxiliary indicators spectral angle mapping θ and cluster centroid distance and d, the performance of the PHATE algorithm is similar to that of the PCA algorithm and the t-SNE algorithm, but the robustness of the PHATE algorithm results is stronger; in terms of the intuitive visualization of the clustering effect, the two-dimensional visualization feature points obtained based on the PHATE feature dimensionality reduction algorithm have significant intra-class compactness and inter-class separation performance, and well retain the important global features and boundary information of the high-dimensional LIBS spectrum, and the results have strong interpretability. In summary, the present invention has the advantages of high result accuracy, strong robustness, and strong interpretability in the application scenario of LIBS spectrum effectiveness identification of dust-covered targets, and is of great value for analyzing Mars exploration LIBS spectrum data.

[0066] The above specific embodiments are only explanations of the present invention, and they are not limitations of the present invention. After reading this specification, those skilled in the art can make modifications to the embodiments without creative contribution as needed. However, as long as they are within the scope of the claims of the present invention, they are protected by the patent law.

Claims

1. A LIBS spectrum validity identification method for dust-covered targets based on PHATE dimension reduction, characterized by It includes the following steps: S1. Preliminary preparation: Prepare a detection sample for Laser Induced Breakdown Spectroscopy (LIBS). The sample is a target with dust covering its surface prepared through a specific process, that is, a dust-covered target, hereinafter simply referred to as "target". Statistically record the names and composition information of the base material and surface dust material of each target; S2. Use the LIBS spectral detection device to collect the LIBS spectrum of the target. During the spectral collection process of each target, the same position of the target needs to be continuously excited by laser pulses more than P times to ensure that the dust covering layer on the target surface is broken down and the spectrum of the base material can be collected. The LIBS spectrum collected by the experiment is defined as the original spectral dataset; S3. Preprocess the LIBS spectra in the original spectral dataset. Usually, the preprocessing steps include dark background subtraction, noise filtering, background baseline removal, wavelength calibration, invalid pixel screening, channel splicing, and sub-channel spectral normalization. The LIBS spectrum after preprocessing is defined as the dust-covered target LIBS spectral dataset; S4. In the dust-covered target LIBS spectral dataset, assume that the number of data points contained in each spectrum is M, and M is defined as the spectral dimension; successively perform feature dimensionality reduction operations on all the LIBS spectra of each target, reducing the spectral dimension from M to N, where N < M, and the specific value of N is set according to the specific dataset and actual requirements. The feature dimensionality reduction operation is performed through the Potential of Heat-diffusion for Affinity-based Trajectory Embedding (PHATE) algorithm; after the feature dimensionality reduction operation, each LIBS spectrum is correspondingly transformed into an N-dimensional feature vector; S5. Successively perform clustering operations on all the N-dimensional feature vectors of each target, dividing the feature vectors into three categories, respectively indicating that the LIBS spectrum corresponding to the feature vector belongs to the surface dust layer spectrum, the dust-matrix transition zone spectrum, and the matrix spectrum. The clustering operation is performed through the K-means clustering algorithm; S6. Based on the spectral similarity index, evaluate the execution effects of steps S4 and S5, and optimize the relevant parameters of the algorithms used in the feature dimensionality reduction process and the clustering process.

2. The LIBS spectrum validity identification method for dust-covered surface targets based on PHATE dimension reduction as claimed in claim 1 is characterized in that: In step S1, the target preparation process includes two parts. First, the dust mass M corresponding to the dust sheet with a thickness of D is calculated according to the compaction density ρ of the dust material and the area S of the sheet. d , weighing mass is M d The first step is to prepare the dust material and flatten it in the tablet press mold; the second step is to place the base material on the powder plane and flatten it, put boric acid powder on the outside of the mold to wrap the dust material and the base material, and start the tablet press to complete the preparation of the target.

3. The LIBS spectrum validity identification method for dust-covered surface targets based on PHATE dimension reduction as claimed in claim 1 is characterized in that: In step S2, when using the LIBS spectral detection device for spectral collection, it is necessary to ensure that the laser pulse energy is sufficient to break down the dust layer at the excited position on the target surface after continuous excitation more than P times.

4. The LIBS spectrum validity identification method for dust-covered surface targets based on PHATE dimension reduction as claimed in claim 1 is characterized in that: In step S3, when preprocessing the LIBS spectra in the original spectral dataset, the dark background subtraction operation refers to subtracting the dark background spectrum from the original LIBS spectrum to obtain an effective spectrum. The dark background spectrum refers to the spectrum of the spectrometer response without light signal obtained under the same instrument parameters as those for collecting the original spectrum; The noise filtering operation refers to using the adaptive wavelet denoising method to remove Gaussian noise in the effective spectrum to obtain a denoised spectrum; The background baseline removal operation refers to using the asymmetric least squares baseline correction method to remove the continuous baseline in the denoised spectrum; Wavelength calibration refers to converting the pixel numbers of the spectrometer into wavelength values ​​by multivariate quadratic fitting; invalid pixel screening refers to removing the pixel response values ​​that exceed the wavelength range in each band of the LIBS spectrum; channel splicing refers to splicing the multiple band LIBS spectra that have been screened out of invalid pixels into a whole line in wavelength order; channel spectrum normalization refers to mapping the intensity values ​​of each channel LIBS spectrum to the interval [0,1] respectively, so that the sum of the LIBS spectrum intensity values ​​in each single channel is 1.

5. The method for identifying the effectiveness of LIBS spectra of dust-covered surface targets based on PHATE dimension reduction as claimed in claim 1, characterized in that: In step S4, a feature dimension reduction algorithm based on PHATE is designed, and relevant parameters of the PHATE algorithm are comprehensively adjusted, including ndim, k, t_max and mds_methods, where ndim represents the feature dimension of the output vector; k represents the number of neighboring points considered when constructing the graph. This variable affects the local relationship modeling and can be adjusted according to the data density; t_max represents the maximum value of the optimal time step for the Markov diffusion process; mds_methods represents a multidimensional scaling analysis method for mapping the diffusion potential energy to a low-dimensional space; for the feature dimension reduction process, the input is all LIBS spectra of a single target, and the output is a matrix composed of low-dimensional feature vectors corresponding to all spectra of the target, with a size of Q×N, where Q is the number of LIBS spectra of the target, and N is the number of features after dimension reduction.

6. The method for identifying the effectiveness of LIBS spectra of dust-covered surface targets based on PHATE dimension reduction as claimed in claim 1, characterized in that: In step S5, a K-means-based clustering algorithm, i.e., a K-means clustering algorithm, is designed, and the relevant parameters of the K-means algorithm are comprehensively adjusted, including n_clusters, Distance, MaxIter, and Start, wherein n_clusters represents the number of clusters in the data, i.e., the number of categories formed after clustering; Distance refers to the calculation method of the distance metric between the cluster centroid and the sample point in the feature space; MaxIter refers to the maximum number of iterations of the algorithm; Start refers to the method of selecting the initial cluster centroid position; for the clustering process, the input is all N-dimensional feature vectors and the number of categories corresponding to a single target LIBS spectrum, and the output is the category corresponding to each LIBS spectrum.

7. The method for identifying the effectiveness of LIBS spectra of dust-covered targets based on PHATE dimension reduction as claimed in claim 1 is characterized in that: In step S6, three evaluation indicators are introduced to evaluate the accuracy of spectral validity recognition, namely, the silhouette coefficient S, the spectral angle mapping θ, and the cluster centroid distance and d. Among them, the silhouette coefficient S is used as the core indicator for evaluating the performance of feature dimensionality reduction and clustering algorithms, and the other two indicators θ and d are used as auxiliary indicators. The three indicators are normalized and then summed to obtain Eval norm , with Eval norm As the evaluation index used in the subsequent parameter optimization process; in the process of optimizing the parameters related to steps S4 and S5, the scheme used is the grid optimization algorithm, which sets the grid value range and grid point spacing of each parameter to be optimized, tries every possible parameter combination scheme, and calculates the above evaluation index Eval norm Until all possible parameter combinations are traversed, the one that can make Eval norm The solution with the largest parameter combination is taken as the final solution, thus completing the optimization of relevant parameters.