A method for automatically calculating the layer position points of secondary ion mass spectrometry curves
Through automatic calculation methods, the hierarchical position points of the secondary ion mass spectrometer curve are automatically identified and calculated, which solves the problems of low efficiency, inconsistency and subjectivity of the manual stratification process in the prior art, and achieves more efficient and accurate data processing.
Patent Information
- Application Number
- CN202410497724.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-04-24
AI Technical Summary
The stratification process of the existing secondary ion mass spectrometer hierarchical position points depends on manual operation, and there are problems of low time efficiency, inconsistency, subjectivity and high technical threshold.
An automatic calculation method is proposed, including collecting and preprocessing the original data, identifying and extracting key feature position points, and calculating thresholds using automatic setting algorithms to determine the hierarchical position points.
This method can significantly reduce manual intervention, improve the accuracy and consistency of data processing, reduce errors, and improve work efficiency. It is suitable for processing large batches or complex data.
Smart Images

Figure CN118378020B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of semiconductor technology, and in particular to an automatic calculation method for layered position points of a secondary ion mass spectrum curve. Background Art
[0002] Secondary ion mass spectrometry (SIMS) is a highly sensitive and high-resolution chip detection method. The specific principle is that when the sample surface is bombarded by high-energy focused primary ions, the primary ions bombard the surface of the sample being analyzed, transfer kinetic energy to the solid atoms, and sputter out neutral particles and secondary ions with positive and negative charges through the cascade effect caused by the stacked collision. By collecting the mass signal of the sputtered secondary ions, the surface and internal element distribution characteristics of the bombarded sample can be analyzed. In the field of chips, secondary ion mass spectrometry is an extremely powerful surface analysis technology. SIMS has high sensitivity: it can detect extremely low concentrations of elements and molecules, which is crucial for the precise control of impurities and dopants in chip manufacturing. SIMS has high spatial resolution: this enables SIMS to analyze the composition of tiny areas, which is crucial for nano-scale chip manufacturing and research and development. SIMS also has deep analysis capabilities: SIMS can perform deep profiling and provide elemental composition information at different levels of the sample, which is crucial for understanding the multi-layer structure in chip manufacturing.
[0003] In secondary ion mass spectrometry (SIMS) data processing, the stratification of SIMS curves is a key step, which involves the analysis and visualization of raw data (usually about the element concentration on the surface or near-surface area of the sample). These curves are usually used to show the change of element concentration with the depth of the sample. At present, in many cases, the stratification of curves is performed manually. This method is performed by experienced engineers to visually observe and determine the location of the curve stratification. It is flexible and the operator can adjust it according to the specific situation and experience, especially when dealing with complex or non-standard samples. Manual operation also enables analysts to interact with the data directly, which helps to better understand the data characteristics. However, this method also has some disadvantages: 1. Low time efficiency: Manual stratification is a time-consuming process, especially when dealing with large amounts of data. 2. Consistency and repeatability issues: Different operators may have different stratification results, and even the same operator may get different results at different times. 3. Subjectivity: Manual stratification depends on the operator's experience and judgment, which may introduce subjectivity and affect the accuracy of the results. 4. Technical threshold: This method requires the operator to have certain professional knowledge and experience, which may be difficult for beginners to master. Summary of the invention
[0004] Based on this, the purpose of the present invention is to provide a method for automatically calculating the stratification position points of a secondary ion mass spectrum curve. Compared with the existing manual calculation method, the method is more convenient to operate and can quickly find the appropriate stratification position points.
[0005] The present invention is achieved through the following detailed technical solutions:
[0006] A method for automatically calculating layered position points of a secondary ion mass spectrometry curve comprises the following steps:
[0007] S1: collecting raw data of curves of samples measured by secondary ion mass spectrometry and preprocessing the raw data;
[0008] S2: Identify and extract key feature position points of the data preprocessed in step S1;
[0009] S3: Calculate a threshold using an automatic setting algorithm based on the key feature position points extracted in step S2, and the position points greater than the threshold are the stratification position points of the sample.
[0010] Compared with the prior art, the automatic calculation method of the stratification position points of the secondary ion mass spectrometry curve of the present invention can reduce human intervention and reduce difficulty. Even beginners can quickly master and use it. This automatic calculation method is more accurate in data processing and can significantly reduce errors. The stratification position point information calculated by the original manual calculation method will cause large errors due to too many human factors, while the automatic calculation method can ensure the consistency of the analysis results and greatly reduce errors. In addition, when processing large quantities or complex data, the automatic calculation method of the present invention can intelligently identify and quickly process large amounts of data without spending a lot of time and energy like manual stratification processing, thereby significantly improving work efficiency.
[0011] Furthermore, the raw data in step S1 are raw data collected from a SIMS experiment, including depth, concentration and measurement depth.
[0012] Furthermore, the preprocessing in step S1 includes data format conversion and curve depth distribution unification. The preprocessing not only facilitates calculation, but also aligns all curves for subsequent processing.
[0013] Furthermore, the preprocessing also includes noise processing, which includes at least one processing method of data cleaning, data standardization, baseline correction and smoothing. Noise processing can remove the influence of local measurement errors and improve the accuracy of the automatic calculation method.
[0014] Furthermore, the key feature position points in step S2 are one or more of the maximum point, the half-width height point, the slope change point, the mutation point and the periodic change point. By comprehensively utilizing these key feature position points, the characteristics of the curve can be analyzed more comprehensively, and the goal of automatically calculating the layered position points can be achieved.
[0015] Furthermore, the feature extraction method is a first-order derivative method. After the feature extraction, the feature importance needs to be evaluated by information gain and Gini index methods, and then a recursive feature elimination algorithm is applied to select the most effective feature subset.
[0016] Furthermore, the threshold in step S3 is automatically set using the Otsu method. The Otsu method is highly efficient and effective in automatically obtaining the threshold, and is very suitable for automatically calculating the layered position points of the secondary ion mass spectrometry curve.
[0017] Furthermore, the automatic calculation method also includes verification and adjustment of the automatic calculation results of the layered position points;
[0018] The verification method of the automatic calculation results of the stratified position points is to compare the automatic calculation results with the known sample results or the results of precise measurements to verify the accuracy of the automatic calculation. When the automatic calculation results are inconsistent with the known sample results or the results of precise measurements, the automatic calculation results of the stratified position points are adjusted. The adjustment method includes establishing a data set and adjusting parameters.
[0019] Furthermore, the data set is an expert data set with accurate measurement results, and the data set includes a training set, a validation set and a test set, and the data ratio of the training set, the validation set and the test set is 7:2:1.
[0020] Furthermore, the parameter adjustment includes model training, model verification and tuning, and model evaluation, wherein model training is to use samples in the training set to adjust the preprocessing steps, the method for extracting feature position points, or the method for automatically setting thresholds, until the stratification position points calculated by the adjusted automatic calculation method for the samples in the training set are consistent with the stratification position points of the training set itself; the model verification and tuning specifically is to use the verification set to test whether the stratification position points automatically calculated by the automatic calculation method after the model training for the samples in the verification set are accurate and consistent, and if not, continue to adjust parameters such as the preprocessing steps, the method for extracting feature position points, or the method for automatically setting thresholds; the model evaluation specifically is to use the test set to verify whether the stratification position points calculated by the automatic calculation method after tuning are accurate and consistent, and if not, repeat the steps of model training and model verification and tuning.
[0021] For better understanding and implementation, the present invention is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of the raw data of the secondary ion mass spectrometry curve;
[0023] Figure 2 It is a schematic diagram of the depth distribution of the secondary ion mass spectrum curve after unified processing;
[0024] Figure 3 It is a schematic diagram of the secondary ion mass spectrum curve after smoothing;
[0025] Figure 4 A schematic diagram of automatic calculation of the layer position points of the secondary ion mass spectrum curve;
[0026] Figure 5 Schematic diagram of the layering of secondary ion mass spectrum curves of the embodiment. DETAILED DESCRIPTION
[0027] The present invention provides a method for automatically calculating the stratified position points of a secondary ion mass spectrum curve. The steps of the method for automatically calculating the stratified position points of a secondary ion mass spectrum curve are detailed as follows:
[0028] S1: collecting raw data of curves of samples measured by secondary ion mass spectrometry and preprocessing the raw data;
[0029] The raw data is the raw data generated by the SIMS experiment, and includes element type, time and intensity, standard sample test data, measurement depth, data source, data type, measurement range and expected material layer structure. Figure 1 shown.
[0030] The element types include main elements and measured elements. The information of the main elements is used to determine the structure of the chip material layer, and the information of the measured elements is used to determine the doping information and isotope.
[0031] The time and intensity can be multiplied by the RSF coefficient to obtain the depth and concentration data of the chip elements.
[0032] The standard sample test data is used to calculate the RSF coefficient.
[0033] The measurement depth refers to the depth data of different material layers, which is used to determine the thickness of different material layers of the sample.
[0034] The data sources include SIMS data measured in the laboratory and SIMS data collected on site.
[0035] The data types include fixed samples, liquid samples, and multi-layer structure samples.
[0036] The calculation formulas for element concentration and depth are as follows:
[0037] C a =RSF×(I a / I m );
[0038] Among them, C a is the concentration of the measured element "a";
[0039] I a is the secondary ion strength of element "a";
[0040] RSF is relative sensitivity factor;
[0041] I m is the secondary ion intensity of the matrix element "m";
[0042] The relative sensitivity factor (RSF) is obtained by measuring a standard sample of the same matrix.
[0043] RSF coefficient = standard sample concentration × (I m / I a );
[0044] Depth = time (Time) * rate V;
[0045] The rate V = depth / measurement time;
[0046] The preprocessing includes data format conversion, curve depth distribution unification and noise processing.
[0047] The data format conversion converts the device-specific data structure into a universal numerical data structure, that is, converts the unique data structure into a universal depth distribution and concentration distribution. If the subsequent calculation is more complicated, the chart coordinates can be converted into numerical coordinates, which can ensure a clear data structure and facilitate subsequent calculations.
[0048] The curve depth distribution is unified to align the curves of different elements according to the structural information of the material layer, and all the curves that need to be layered can be processed uniformly later, which is more convenient. Figure 1 After the depth distribution is uniformly processed, the curve is as follows Figure 2 shown.
[0049] The noise processing includes data cleaning, data normalization, baseline correction and smoothing.
[0050] The data cleaning is to remove the erroneous data caused by instrument failure or operation error, the data standardization is to use the standardization method to standardize the data and scale all the data to a uniform range or distribution. The baseline correction is to use the mathematical model to remove the background signal and highlight the useful signal.
[0051] The standardization method is preferably a standard score standardization method, which is a process of standardizing data according to its mean (μ) and standard deviation (σ), so that the processed data conforms to the standard normal distribution, that is, the mean is 0 and the standard deviation is 1. This method is very useful in data preprocessing, especially in preparing data for machine learning and statistical analysis.
[0052] The smoothing process is to smooth the data after the format conversion to reduce the influence of noise. The smoothing process can select sliding average or median filter to process the data according to the actual situation of the data. The smoothed data helps to better identify the overall trend of the curve and avoid errors caused by excessive local noise. Figure 2 After smoothing, the curve is Figure 3 shown.
[0053] If the noise contained in the data after format conversion is mainly random and the overall trend of the data needs to be retained, choose to apply sliding average to smooth the data. Sliding average can retain the overall trend of the data by calculating the average value of the data points in a certain window, thereby reducing the impact of noise. The advantage of sliding average is that it is simple and easy to use and suitable for most smoothing needs. However, when there are many outliers or sudden noise in the data after format conversion, the data will be insufficiently smoothed, affecting the accuracy of the smoothing process.
[0054] If there are many outliers or sudden noise in the data after format conversion, and the original characteristics of the data need to be better preserved, median filtering is selected to smooth the data. Median filtering is a smoothing technique that replaces the original data points with the middle values of the data points. The advantage of median filtering is that it has a good effect on processing sudden noise because it is not affected by outliers and can better preserve the characteristics of the data.
[0055] S2: Identify and extract key feature location points of the data preprocessed by step S1.
[0056] The key feature position points include maximum value points, half-width height points, slope change points, mutation points and periodic change points.
[0057] The maximum point corresponds to a peak on the curve, i.e., a local maximum of the curve. The maximum point can be used to identify peak locations that indicate the presence of specific components in the sample and may correspond to locations on the curve where the rate of change of concentration is the greatest. By identifying the maximum point, key features that may be present in the curve can be determined, thereby aiding further analysis and understanding.
[0058] The half-width height point refers to the point where the curves on both sides of the peak point intersect with half of the peak height, and the half-width height point can provide information about the width of the peak. This width information can be used to evaluate the width of the peak, understand the distribution range of a specific component in the sample, or be used as one of the characteristics of the peak shape.
[0059] The specific identification and extraction method of the key feature position points is the first-order derivative method:
[0060] The first-order derivative method is to calculate the first-order derivative of each data point (the slope of the data point). The first-order derivative reflects the rate of change of the curve. The first-order derivative can be calculated by a numerical method (such as a difference method).
[0061] Maximum Points: First find the peak points of the curve by finding the zero crossing points of the first-order derivative (the slope changes from positive to negative) or by comparing the slopes of each point with its adjacent points.
[0062] Half-width height point: First record the height of the peak point as the reference height, then start from the peak point and move along both sides of the curve until a point with a height equal to half of the reference height is found. These two points are located on the left and right of the peak point, respectively. For the two half-height points found, linear interpolation is performed to calculate the exact position of the half-width point. Interpolation can use simple linear interpolation or more complex interpolation methods, such as cubic spline interpolation.
[0063] The maximum value is fmax, and the half-width can be calculated by finding two points x1 and x2 that satisfy the following conditions:
[0064]
[0065] Half-height width F WHM The calculation formula is:
[0066] F WHM =x2-x1
[0067] In samples with layers, the different layers in the samples represent the differences in the matrix and material. Therefore, when the secondary ion mass spectrometry is generated, the curve data seen will be obviously different. The concentration, depth and speed will be different. The maximum point and half-width height point can distinguish these layers well, thereby determining the curve stratification and facilitating subsequent data calculation and processing.
[0068] The slope change points are points on the curve where the slope changes significantly, and these points may correspond to the boundaries of different levels or components in the sample. The slope change points are calculated by calculating the slope between two points on the curve and taking out the two points with the maximum slope, that is, the sharp change.
[0069] k=(y1-y2) / (x1-x2)
[0070] Among them: x1, y1, x2, y2 are the horizontal and vertical coordinates of the two points respectively.
[0071] The mutation point is a point on the curve where a sharp change occurs, which may indicate a sudden change in the properties of the sample or the existence of a boundary. When a mutation point is found, the data point can be restored to a normal state by manually dragging it.
[0072] The periodic change points are characteristic points on the periodic change curve that can identify periodicity. These nodes can help understand the distribution of periodic structures in the sample.
[0073] The comprehensive use of these key feature positions can more comprehensively analyze the characteristics of the curve and assist in achieving the goal. The specific selection of features and how to combine these features for the automatic calculation of stratified position points requires the use of methods such as information gain and Gini index to evaluate the importance of features, and then apply algorithms such as recursive feature elimination (RFE) to select the most effective feature subset.
[0074] The information gain is based on the concept of entropy, which is a measure of the uncertainty of a data set. Information gain represents the amount by which the uncertainty of a data set is reduced given a given feature, i.e., the degree to which the feature makes the data set more ordered. In a decision tree, information gain is used to select the features that are most effective in classifying a data set.
[0075] Calculation method:
[0076] Given a data set D, its entropy H(D) is defined as:
[0077]
[0078] Where pi is the relative frequency of category i in dataset D and m is the number of categories.
[0079] If D is divided into several subsets D1, D2, ..., D according to different values of feature A, k , the information gain IG(D,A) of feature A to data set D is:
[0080]
[0081] Among them, feature A is the key feature position point in the optimal feature subset obtained in step S2.
[0082] The larger the information gain IG(D,A), the more the purity of the subset is improved after segmentation using feature A, that is, the more the uncertainty is reduced.
[0083] The Gini index is another indicator for measuring the impurity of a data set. The smaller the Gini index is, the higher the purity of the data set is.
[0084] Calculation method:
[0085] Given a data set D, its Gini index Gini(D) is defined as:
[0086]
[0087] Among them, p i is the relative frequency of category i in dataset D.
[0088] If D is divided into several subsets D1, D2, ..., D according to different values of feature A, k , the Gini index of feature A is:
[0089]
[0090] Among them, feature A is the key feature position point in the optimal feature subset obtained in step S2.
[0091] In this automatic calculation method, key feature position points that can minimize the Gini index Gini (D) are selected for segmentation. The purity of the segmented key feature position point subset is the highest, which can improve the accuracy of the automatic calculation method.
[0092] The recursive feature elimination (RFE) is a feature selection method that aims to find the features that are most helpful for model prediction by recursively reducing the size of the feature set. Specifically, it uses the weight coefficient or feature importance of the model itself to evaluate the importance of each feature and gradually removes the least important features, thereby optimizing the model performance.
[0093] The steps of the RFE algorithm are as follows:
[0094] 1. Train the model: Train the model using all available features.
[0095] 2. Evaluation of feature importance: Evaluate the importance of each feature based on the feature weight (or feature importance) given by the model. In some models, this may directly correspond to the size of the coefficient of the model parameter; in other models, additional methods may be needed to evaluate the importance of the feature.
[0096] 3. Remove the weakest features: Remove the least important features - that is, the features with the smallest weight.
[0097] 4. Recursive repetition: Repeat steps 1 to 3, retraining the model each iteration with a dataset with one less feature than the previous one, until the preset number of features is reached or a stopping criterion is met.
[0098] 5. Feature subset selection: Finally, the feature subset that performs best during the iteration process is selected as the final feature set.
[0099] The feature is a key feature position point in the optimal feature subset obtained in step S2.
[0100] By using three feature selection methods, namely information gain, Gini index and recursive feature elimination, the optimal feature subset can be determined among the above key feature locations.
[0101] S3: Calculate a threshold using an automatic setting algorithm based on the key feature position points extracted in step S2, and the position points greater than the threshold are the stratification position points of the sample.
[0102] Specifically, a threshold is first set, and the threshold is used to determine whether the rate of change in the first-order derivative is large enough. Only data points with a large enough rate of change can be identified as a stratification point, indicating that there is a significant concentration change or a change in curve characteristics near this point.
[0103] The delamination point is usually defined as a local maximum point on the curve, which represents a peak or maximum value in the curve, which may represent the presence of a certain component or feature in the sample. Therefore, the criterion for the delamination point is that the rate of change of the first-order derivative at this point is large enough to reach the set threshold, and the point is a maximum point on the curve.
[0104] In order to automatically select a suitable threshold, the present application adopts the Otsu method or the adaptive threshold method to automatically set the threshold, and the Otsu method includes the following steps:
[0105] 1. Calculate overall statistics:
[0106] Calculate the total average data intensity (μT) and total data variance (σ 2 T), the total number of data (N), and the frequency of occurrence of each intensity value.
[0107] 2. Iterate for each possible threshold:
[0108] For each possible threshold t, the data is divided into two classes: class 1 contains all points with intensity values less than or equal to t, while class 2 contains all points with intensity values greater than t.
[0109] Calculate the weights (w1 and w2) and average strengths (μ1 and μ2) of the two categories respectively. The weight can be obtained by dividing the number of points in the category by the total number of points (N), and the average strength can be obtained by averaging the strength values of all points in the category.
[0110] 3. Calculate the between-class variance:
[0111] Between-class variance (σ 2 B) can be calculated by the following formula:
[0112] σ 2 B(t)=w1(t)×w2(t)×[μ1(t)-μ2(t)]2
[0113] where w1(t) and w2(t) are the weights of class 1 and class 2 based on threshold t, and μ1(t) and μ2(t) are the average intensities of the corresponding classes.
[0114] 4. Choose the threshold that maximizes the between-class variance:
[0115] For all possible thresholds t, find the t that maximizes the between-class variance σ2B(t).
[0116] This t is the optimal threshold determined by the Otsu method.
[0117] According to the data processed by the optimal threshold determined by the Otsu method, different material layers or component areas are identified and marked, the depth distribution of each material layer and its corresponding element concentration changes are analyzed, and the automatic calculation of the curve layer position points is realized. The automatic calculation results are as follows Figure 4 shown.
[0118] Adaptive Threshold:
[0119] For a point x in a local region R, the adaptive threshold T(x) can be calculated based on local statistical characteristics, such as the local mean:
[0120] T(x)=α·mean(R)+β
[0121] Among them, α and β are adjustment parameters, mean(R) is the average value of point x in region R, and the point x is the key feature position point in the optimal feature subset obtained in step S2.
[0122] The Otsu method can automatically select the threshold value, which is suitable for the case where the image or data distribution has bimodal characteristics. The adaptive threshold dynamically adjusts the threshold value according to the local characteristics of the data, which is suitable for the case where the data characteristics vary significantly in different areas. These methods are suitable for automating the processing flow, reducing manual intervention, and improving efficiency and consistency.
[0123] In addition, a threshold may be set in combination with the slope change point. Specifically, a slope threshold is set. When the slope exceeds or falls below the threshold, it is considered that the curve has changed significantly, and corresponding processing or analysis is performed.
[0124] The setting of the threshold will also affect the calculation of the slope: for example, if the slope is defined as the rate of change between adjacent data points, then the choice of the threshold will directly affect which changes are considered significant. A larger threshold will result in only larger slope changes being considered significant, while smaller changes may be ignored; conversely, a smaller threshold may result in more changes being considered significant, including some minor changes.
[0125] In practical applications, the thresholds and criteria for determining the rate of change of the first-order derivative may vary depending on the specific data and analysis requirements. In order to ensure the accuracy and reliability of the stratification points, it is usually necessary to reasonably set the thresholds and criteria based on the actual situation, and further adjustments and optimizations may be required. Specifically, the following factors need to be considered:
[0126] 1. The noise level of the data: If there is a lot of noise or fluctuations in the data, you may need to choose a smaller threshold so that you can capture smaller changes. On the contrary, if the data is relatively clean, you can consider choosing a larger threshold.
[0127] 2. Curve variation range: If the curve has a large variation range, you may need to select a relatively large threshold value so that you can identify significant change points. If the variation range is small, you can select a smaller threshold value.
[0128] 3. Purpose of analysis: According to the purpose of analysis and the degree of attention paid to the curve characteristics, the choice of threshold can be adjusted. If you need to capture all possible change points, you can choose a smaller threshold; if you only focus on significant change points, you can choose a larger threshold.
[0129] 4. Resolution of the data: The higher the resolution of the data, the smaller the threshold may need to be selected to be able to identify more subtle changes. Conversely, if the data resolution is lower, a larger threshold may be selected to avoid over-identification.
[0130] S4: Verification and adjustment of the automatic calculation results of the stratified position points.
[0131] The verification method of the automatic calculation results of the stratified position points is to compare the automatic calculation results with the known sample results or the results of precise measurements to verify the accuracy of the automatic calculation. When the automatic calculation results are inconsistent with the known sample results or the results of precise measurements, the automatic calculation results of the stratified position points are adjusted. The adjustment method includes establishing a data set and adjusting parameters.
[0132] The data set is an expert data set with accurate measurement results. The data set includes a training set, a validation set and a test set. The data ratio of the training set, the validation set and the test set is 7:2:1.
[0133] The parameter adjustment includes model training, model verification and tuning, and model evaluation, wherein model training is to use samples in the training set to adjust parameters such as preprocessing steps, feature position point extraction methods, or automatic threshold setting methods, until the stratification position points calculated by the adjusted automatic calculation method for the samples in the training set are consistent with the stratification position points of the training set itself; the model verification and tuning specifically is to use the verification set to test whether the stratification position points automatically calculated by the automatic calculation method after the model training for the samples in the verification set are accurate and consistent, and if not, continue to adjust parameters such as preprocessing steps, feature position point extraction methods, or automatic threshold setting methods; the model evaluation specifically is to use the test set to verify whether the stratification position points calculated by the automatic calculation method after tuning are accurate and consistent, and if not, repeat the model training and model verification and tuning steps.
[0134] In order to further verify the accuracy of the stratified position points calculated by the automatic calculation method, this application processed the same set of data by the manual calculation method 6 times and the automatic calculation method 6 times. The stratification results of this set of data are as follows: Figure 5 As shown, the hierarchical location point information results obtained by the manual calculation method and the calculation method are shown in Table 1 below:
[0135] Table 1
[0136]
[0137] The distribution of layers can be observed through more precise instrumental analysis of the sample structure. The best stratification position point for this set of data is 1.294534538μm. By calculating the errors between the stratification positions obtained by the manual calculation method and the automatic calculation method and the best stratification point six times, it can be seen that the error of the stratification position points obtained by the manual calculation method is significantly greater than the error of the stratification position points obtained by the automatic calculation method.
[0138] Compared with the prior art, the automatic calculation method of the sample stratification position points based on the secondary ion mass spectrometry curve of the present invention can reduce manual intervention and reduce difficulty, and even beginners can quickly master and use it. This curve automatic calculation method is more accurate in data processing and can significantly reduce errors. The original manually selected stratified position information will cause large errors due to too many human factors, while the automatic calculation method can ensure the consistency of the analysis results and reduce errors. In addition, when processing large quantities or complex data, the automatic calculation method of the present invention can intelligently identify and quickly process large amounts of data, without spending a lot of time and energy like the manual calculation method, thereby significantly improving work efficiency.
[0139] The above-mentioned embodiments only express several implementation methods of the present invention, and the description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, and the present invention is also intended to include these modifications and modifications.
Claims
1. A method for automatically calculating sample stratification position points based on secondary ion mass spectrometry curves, characterized in that: The following steps are involved: S1: Collecting raw data of the curve of the sample measured by secondary ion mass spectrometry and preprocessing the raw data; the raw data is the raw data collected by the SIMS experiment, including element type, time and intensity, standard sample test data, and measurement depth; S2: Identify and extract key feature location points of the data preprocessed in step S1; S3: Calculate a threshold using an automatic setting algorithm based on the key feature position points extracted in step S2, and the position points greater than the threshold are the stratification position points of the sample; S4: Verification and adjustment of the automatic calculation results of the stratified position points; The verification method of the automatic calculation result of the stratified position point is to compare the automatic calculation result with the known sample result or the result of the precise measurement to verify the accuracy of the automatic calculation. When the automatic calculation result is inconsistent with the known sample result or the result of the precise measurement, the automatic calculation result of the stratified position point is adjusted, and the adjustment method includes establishing a data set and adjusting parameters; The data set is an expert data set with accurate measurement results, and the data set includes a training set, a validation set and a test set, and the data ratio of the training set, the validation set and the test set is 7:2:1; The parameter adjustment includes model training, model verification and tuning, and model evaluation, wherein model training is to use samples in the training set to adjust the preprocessing steps, the method for extracting feature position points, or the method for automatically setting thresholds, until the stratification position points calculated by the adjusted automatic calculation method for the samples in the training set are consistent with the stratification position points of the training set itself; the model verification and tuning specifically is to use the verification set to test whether the stratification position points automatically calculated by the automatic calculation method after the model training for the samples in the verification set are accurate and consistent, and if not, continue to adjust the parameters of the preprocessing steps, the method for extracting feature position points, or the method for automatically setting thresholds; the model evaluation specifically is to use the test set to verify whether the stratification position points calculated by the automatic calculation method after tuning are accurate and consistent, and if not, repeat the model training and model verification and tuning steps.
2. The method for automatically calculating sample stratification position points based on secondary ion mass spectrometry curve according to claim 1, characterized in that: The preprocessing in step S1 includes data format conversion and curve depth distribution unification.
3. The method for automatically calculating sample stratification position points based on secondary ion mass spectrometry curve according to claim 2, characterized in that: The preprocessing also includes noise processing, and the noise processing includes at least one processing method of data cleaning, data standardization, baseline correction and smoothing.
4. The method for automatically calculating sample stratification position points based on secondary ion mass spectrometry curve according to claim 3, characterized in that: The key feature position points in step S2 are one or more of the maximum point, the half-width height point, the slope change point, the mutation point and the periodic change point.
5. The method for automatically calculating sample stratification position points based on secondary ion mass spectrometry curve according to claim 4, characterized in that: The feature extraction method is the first-order derivative method. After the feature extraction, the feature importance needs to be evaluated by information gain and Gini index methods, and then the recursive feature elimination algorithm is applied to select the most effective feature subset.
6. The method for automatically calculating sample stratification position points based on secondary ion mass spectrometry curve according to claim 5, characterized in that: The threshold in step S3 is automatically set using the Otsu method.
Citation Information
Patent Citations
Mass spectrum data processing method and device, computer equipment and computer storage medium
CN109726667A
Mass spectrum data classification method based on support vector machine
CN117407779A