Nonlinear dynamic data anomaly detection method for multi-process stage of silicon single crystal growth

Through Bayesian online jump point detection and related vector machine model combined with dynamic threshold strategy, the abnormal detection difficulties in multi-process mode and dynamic environment during silicon single crystal growth in the prior art are solved, and nonlinear dynamic data detection with high accuracy and low error detection rate is achieved.

CN120296643AActive Publication Date: 2025-07-11XIAN UNIV OF TECH
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510784454.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The existing silicon single crystal growth anomaly detection methods cannot effectively adapt to multi-process modes and dynamic environments, and it is difficult to accurately detect nonlinear data anomalies, and the model training is high, resulting in frequent false alarms and missed alarms.

Method used

Bayesian online jump point detection and related vector machine model combined with dynamic threshold strategy are used to reconstruct the training set and test set for iterative training, predicted sequences and residual sequences are generated, and abnormal detection is performed using sliding windows and dynamic thresholds to achieve accurate detection of nonlinear dynamic data of silicon single crystal growth process.

Benefits of technology

It improves the accuracy and sensitivity of abnormal detection, reduces the error detection rate, can adapt to data fluctuations and nonlinear changes, and ensures the reliability and real-timeness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296643A_ABST
    Figure CN120296643A_ABST
Patent Text Reader

Abstract

The invention provides a nonlinear dynamic data anomaly detection method for multiple process stages of silicon single crystal growth, and belongs to the technical field of integrated circuit silicon single crystal preparation. Comprising the following steps: acquiring a historical data set and a monitoring data set in a silicon single crystal growth process, and constructing a reconstruction training set and a reconstruction test set; performing iterative training on a relevance vector machine model by using the reconstructed training set to obtain a relevance vector machine training model; performing single-step prediction on the monitoring data set to obtain a prediction sequence, and obtaining a residual sequence by using the prediction sequence; dividing the residual error sequence into a plurality of segmented windows, and sequentially traversing each segmented window by using all sliding windows to obtain a plurality of sliding window data sets; and a dynamic threshold strategy is introduced, anomaly detection is performed on the corresponding sliding windows according to the candidate threshold set generated by each sliding window data set, and all anomaly detection results are fused into a final anomaly detection result. According to the invention, the accuracy and reliability of abnormal detection of the nonlinear data can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of integrated circuit single crystal silicon preparation, and particularly to a method for detecting abnormal non-linear dynamic data in multiple process stages of single crystal silicon growth. Background Art

[0002] Single crystal silicon is a basic material for manufacturing integrated circuit chips. With the continuous improvement of the manufacturing process requirements for integrated circuit chips, more stringent standards are correspondingly put forward for the quality of single crystal silicon. During the preparation of single crystal silicon, due to factors such as the soft shaft shaking and mechanical vibration of the lifting system, the abnormal growth of single crystal silicon often occurs. How to timely detect the abnormal state during the growth process of single crystal silicon and take corresponding adjustments is a necessary measure to ensure the preparation of high-quality single crystal silicon.

[0003] Currently, the main process of the commonly used method for detecting abnormal growth conditions of single crystal silicon mainly includes: defining normal behavior, selecting an anomaly detection algorithm, training a model, setting a threshold, detecting anomalies, and result evaluation, etc. The specific process is as follows: First, define the range of normal behavior through statistical analysis, business knowledge, or expert opinions; then, select a suitable anomaly detection algorithm according to data characteristics and business requirements, such as statistical methods, machine learning models, or deep learning methods, etc.; then, use normal data to train the selected model to make it learn the characteristics of normal behavior; after that, set a reasonable threshold to distinguish normal data points from abnormal data points; finally, after the model training is completed, predict future data based on the existing historical data, and judge whether an anomaly occurs by comparing the relationship between the difference between the predicted value and the actual value and the reasonable threshold. However, the existing anomaly detection method has the following problems:

[0004] First, for the multiple process modes involved in the single crystal silicon growth process, the adopted model cannot well detect anomalies in different modes at the same time;

[0005] Second, using a fixed reasonable threshold cannot well adapt to the situation where data changes over time. Especially in a dynamic environment, it often cannot correctly capture data anomalies at a certain moment;

[0006] Finally, due to the higher requirements for model training for this non-linear data, simple models are difficult to adapt to complex modes.

[0007] Therefore, it is necessary to propose a solution to improve one or more problems existing in the above related technical solutions.

[0008] It should be noted that the information disclosed in the above background art section is only used to strengthen the understanding of the background of this application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0009] An embodiment of the present application provides a method for detecting abnormal non-linear dynamic data in multiple process stages of silicon single crystal growth. The method includes the following steps:

[0010] Obtain a historical data set and a monitoring data set during the silicon single crystal growth process, reconstruct the historical data set to obtain a reconstructed training set, reconstruct the monitoring data set to obtain a reconstructed test set, and perform Bayesian online change point detection on the monitoring data set to obtain multiple segmentation points in the monitoring data set. Among them, the reconstructed training set contains multiple training samples, and the reconstructed test set contains multiple test samples;

[0011] Construct a relevant vector machine model, and use the reconstructed training set to iteratively train the relevant vector machine model to obtain a relevant vector machine training model;

[0012] Use the reconstructed training set and the reconstructed test set in the relevant vector machine training model to perform single-step prediction on the monitoring data set to obtain a prediction sequence, and perform corresponding difference processing on the prediction sequence and the output target of each test sample to obtain a residual sequence. The residual sequence includes multiple residual values;

[0013] Use all the segmentation points to divide the residual sequence into multiple segmentation windows, and use all the sliding windows to sequentially traverse each segmentation window to obtain multiple sliding window data sets;

[0014] Introduce a dynamic threshold strategy, generate corresponding candidate threshold sets according to the statistical characteristics of all the sliding window data sets, use all the candidate threshold sets to perform anomaly detection on the corresponding sliding windows respectively, and fuse the anomaly detection results of all the sliding windows to obtain the final anomaly detection result.

[0015] Further, the step of obtaining a historical data set and a monitoring data set during the silicon single crystal growth process, reconstructing the historical data set to obtain a reconstructed training set, reconstructing the monitoring data set to obtain a reconstructed test set, and performing Bayesian online change point detection on the monitoring data set to obtain multiple segmentation points in the monitoring data set includes:

[0016] Obtain all the historical data during the silicon single crystal growth process to form the historical data set, and form the reconstructed training set with all the non-abnormal historical data in the historical data set;

[0017] The reconstructed training set is denoted by , , denotes the total length of all historical data in the reconstructed training set, denotes the The input vector of the th training sample, and the input vector of the th to the th consecutive pieces of the historical data make up, denotes the th piece of historical data in the reconstruction training set, denotes the th output target of the training sample, and, denotes the th piece of historical data in the reconstruction training set, denotes dimensional real space, denotes the th training sample, denotes the multiplication sign, denotes one-dimensional real space, denotes the th dimension;

[0018] Obtain all the monitoring data during the single-crystal silicon growth process to form the monitoring data set, and form the reconstruction test set with all the non-abnormal monitoring data in the monitoring data set;

[0019] Perform the Bayesian online change point detection on the monitoring data set to obtain multiple change points;

[0020] The reconstruction test set is denoted by , , denotes the total length of all the monitoring data in the reconstruction test set, denotes the th input vector of the test sample, and the th input vector of the test sample is composed of the th to the th consecutive pieces of the monitoring data, , denotes the th monitoring data in the reconstruction test set, denotes the th output target of the test sample, denotes the th test sample, denotes the th dimension;

[0021] The historical data set is denoted by , and the monitoring data set is denoted by .

[0022] Further, the steps of constructing the relevance vector machine model and iteratively training the relevance vector machine model with the reconstructed training set to obtain a relevance vector machine training model include:

[0023] In the process of iteratively training the relevance vector machine model with the reconstructed training set, all input vectors of the training samples are used as the input of the relevance vector machine model, all output targets of the corresponding training samples are used as the output of the relevance vector machine model, and a likelihood function of the reconstructed training set is constructed to obtain the relevance vector machine training model;

[0024] Prevent overfitting of parameters for the likelihood function of the reconstructed training set, find multiple hyperparameters, and alternately iterate and update all the hyperparameters until the change amplitude of all the hyperparameters is less than a preset amplitude threshold, completing the joint optimization of all the hyperparameters.

[0025] Further, the expression of the relevance vector machine training model is:

[0026] (1)

[0027] Wherein, represents the predicted output of the relevance vector machine training model, , represents the number of all training samples in the reconstructed training set, , represents the weight vector of the relevance vector machine training model, represents the th element of the weight vector of the relevance vector machine training model, represents the th input vector of the training sample, , represents the kernel function of the reconstructed training set, , represents the weight of the Gaussian kernel in the kernel function of the reconstructed training set, represents the width of the Gaussian kernel in the kernel function of the reconstructed training set, represents the transpose, represents the bias term of the polynomial kernel in the kernel function of the reconstructed training set, represents the power of the polynomial kernel in the kernel function of the reconstructed training set, represents the th Gaussian white noise corresponding to the training sample, and the Gaussian distribution of the Gaussian white noise is , represents the variance of the Gaussian white noise, represents the exponential function, represents the norm of a vector. The number of all elements of the weight vector is the same as the number of all training samples in the reconstruction training set, and for the th training sample in the reconstruction training set, there is a th element in the weight vector corresponding to it;

[0028] Using the Gaussian white noise corresponding to all the training samples, calculate the likelihood function of each training sample respectively, and using the mutual independence among all the training samples in the reconstruction training set, multiply the likelihood functions of all the training samples to obtain the likelihood function of the reconstruction training set;

[0029] The expression of the likelihood function of the training sample is:

[0030] P sample y h i w, σ 2 = 1 2π σ 2 exp { - [y h i -Y x h i ,w ] 2 2 σ 2 } (2)

[0031] where represents the likelihood function of the th training sample;

[0032] The expression of the likelihood function of the reconstruction training set is:

[0033] (3)

[0034] where represents the likelihood function of the reconstruction training set, represents the output targets of all training samples, represents the kernel matrix, , represents dimensional real space;

[0035] To prevent the weight vector from overfitting, assume the weight vector to be a Gaussian distribution with zero mean , and take as the prior distribution, , where represents the first hyperparameter, represents the Gaussian distribution, represents the reciprocal of the th value in the first hyperparameter;

[0036] According to Bayes' formula, the expression of the posterior distribution is:

[0037] (4)

[0038] Among them, represents the posterior distribution of the weight vector , represents the conditional probability of the weight vector , , , represents the second hyperparameter, represents the covariance matrix, , , represents the mean vector of the posterior distribution, , represents the marginal posterior probability of the weight vector ;

[0039] The expression of the first hyperparameter is:

[0040] (5)

[0041] Among them, represents the -th value of the first hyperparameter after the -th iteration update, represents the -th value of the effective degrees of freedom, , represents the -th value of the first hyperparameter after the -th iteration update, represents the element in the -th row and -th column of the covariance matrix, represents the -th value of the mean vector of the posterior distribution;

[0042] The expression of the second hyperparameter is:

[0043] (6)

[0044] Among them, represents the transpose of the all-ones column vector of , represents -dimensional real space, represents the effective degrees of freedom;

[0045] Using (5) and formula (6), the first hyperparameter and the second hyperparameter are alternately iteratively updated. When the change amplitudes of the first hyperparameter and the second hyperparameter are both smaller than the preset amplitude threshold, the joint optimization process ends.

[0046] Further, the step of using the reconstructed training set and the reconstructed test set in the relevant vector machine training model to perform single-step prediction on the monitoring data set to obtain a prediction sequence, and performing corresponding difference processing on the prediction sequence and the output target of each test sample to obtain a residual sequence includes:

[0047] Input both the reconstructed training set and the reconstructed test set into the relevant vector machine training model, and use the mean vector of the posterior distribution as the effective estimated value of the weight vector to perform the single-step prediction on the monitoring data set to obtain a prediction sequence;

[0048] Perform the corresponding difference processing on all the prediction data in the prediction sequence and the output target of the corresponding test sample in the reconstructed test set, and form all the obtained residual values into the residual sequence.

[0049] Further, the expression of the single-step prediction is:

[0050] (7)

[0051] where represents the prediction data at the th step, represents the th training sample, represents the kernel function of the reconstructed test set, , represents the weight of the Gaussian kernel in the kernel function of the reconstructed test set, represents the width of the Gaussian kernel in the kernel function of the reconstructed test set, represents the bias term of the polynomial kernel in the kernel function of the reconstructed test set, represents the power of the polynomial kernel in the kernel function of the reconstructed test set, represents the th value of the mean vector of the posterior distribution after the th iterative update;

[0052] The prediction sequence is expressed as: , where represents the prediction sequence, represents -dimensional real number space;

[0053] The residual sequence is expressed as: , represents the residual sequence, represents the output targets of all test samples.

[0054] Further, the step of dividing the residual sequence into a plurality of segmented windows by using all the segmentation points and sequentially traversing each of the segmented windows by using all the sliding windows to obtain a plurality of sliding window data sets includes:

[0055] Represent the set of all the segmentation points as: , where represents the -th segmentation point;

[0056] Divide the residual sequence into segmented windows by using all the segmentation points; represent the set of all the segmented windows as: , where represents all the residual values within the -th segmented window in the residual sequence, , represents the last residual value within the -th segmented window in the residual sequence;

[0057] Perform preset processing on all the sliding windows respectively, and sequentially traverse each of the segmented windows by using all the preset-processed sliding windows to obtain a plurality of the sliding window data sets.

[0058] Further, the step of introducing a dynamic threshold strategy, respectively generating corresponding candidate threshold sets according to the statistical characteristics of all the sliding window data sets, performing anomaly detection on the corresponding sliding windows by using all the candidate threshold sets, and fusing the anomaly detection results of all the sliding windows to obtain a final anomaly detection result includes:

[0059] Calculate the mean and standard deviation of all the window data in each of the sliding window data sets respectively, and make the value of each of the sliding window data sets take equally spaced values within a preset interval;

[0060] Calculate a plurality of candidate thresholds for the corresponding sliding window respectively according to the mean and standard deviation of all the window data in each of the sliding window data sets and the results of all the equally spaced value takings corresponding thereto, and form the candidate threshold set;

[0061] Calculate the value score corresponding to each of the candidate thresholds of all the sliding windows respectively;

[0062] Traverse all the value scores of each of the sliding windows respectively, and obtain the value corresponding to the maximum value score of the corresponding sliding window;

[0063] Utilize all the said maximum values corresponding to the values, and respectively calculate the optimal thresholds of the corresponding sliding windows;

[0064] Utilize all the said optimal thresholds to respectively perform anomaly detection on the corresponding sliding windows, and mark all the window data greater than or equal to the corresponding optimal thresholds in all the sliding windows as anomaly points, and mark all the remaining window data as normal points;

[0065] Introduce sliding window false alarm points, set a lower limit for the number of anomaly points, and respectively compare the number of all the anomaly points in each sliding window with the lower limit of the number of anomaly points;

[0066] If the number of all the anomaly points in the sliding window is greater than or equal to the lower limit of the number of anomaly points, it is considered that the sliding window has an anomaly, and fuse the anomaly detection results of all the sliding windows to obtain the final anomaly detection result;

[0067] If the number of all the anomaly points in the sliding window is less than the lower limit of the number of anomaly points, perform pruning processing on all the anomaly points, and fuse the anomaly detection results of all the sliding windows after the pruning processing to obtain the final anomaly detection result.

[0068] Furthermore, the expression of the candidate threshold is:

[0069] (8)

[0070] Wherein, represents the candidate threshold of the sliding window, represents the mean value of all the window data in the sliding window dataset, represents value, that is, the quantile in the standard normal distribution, represents the standard deviation of all the window data in the sliding window dataset, represents the number of all the window data in the sliding window dataset;

[0071] The expression of the value score is:

[0072] (9)

[0073] Wherein, represents value score, represents the offset of the mean value of all the anomaly points in the sliding window dataset, , represents the mean value of all normal points in the sliding window dataset, represents the number of all abnormal points in the sliding window dataset, , represents the number of all normal points in the sliding window dataset, , represents the offset of the standard deviation of all abnormal points in the sliding window dataset, , represents the standard deviation of all normal points in the sliding window dataset, represents the number of all abnormal points in the sliding window dataset after pruning, represents the number of all consecutive abnormal points in the sliding window dataset.

[0074] Furthermore, the expression of the optimal threshold is:

[0075] (10)

[0076] where, represents the optimal threshold of the sliding window, represents the maximum value score corresponding to value

[0077] The present application provides a method for detecting abnormal non - linear dynamic data in multiple process stages of single - crystal silicon growth, which has at least the following beneficial effects:

[0078] (1) By automatically segmenting the single - crystal silicon growth process and adopting independent dynamic thresholds within each segment, the present application gets rid of the limitations of manual segmentation and fixed window length, avoids false alarms and missed detections caused by process - stage differences, and improves the accuracy and sensitivity of abnormal detection;

[0079] (2) By using a relevant vector machine training model to perform single - step prediction on the monitoring dataset to obtain a prediction sequence, and using the prediction sequence to obtain a residual sequence, since the relevant vector machine training model can effectively process complex non - linear data, it can more accurately predict data changes in the single - crystal silicon growth process, effectively improving the accuracy of data prediction in the single - crystal silicon growth process;

[0080] (3) By introducing a dynamic threshold strategy, the present application uses all candidate threshold sets to perform anomaly detection on the corresponding sliding windows respectively, and fuses all the anomaly detection results into the final anomaly detection result, so as to adapt to data fluctuations in the anomaly detection process, effectively reduce false detections caused by factors such as noise, and improve the accuracy of anomaly detection; moreover, by adopting the dynamic threshold strategy, the window data of the sliding window can be automatically adjusted in real time, so as to better cope with data changes and ensure the accuracy and reliability of non-linear data anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0082] Figure 1 A schematic diagram showing the steps of a method for detecting non-linear dynamic data anomalies in multiple process stages of silicon single crystal growth in an exemplary embodiment of the present application;

[0083] Figure 2 A schematic flow diagram showing a method for detecting non-linear dynamic data anomalies in multiple process stages of silicon single crystal growth in an exemplary embodiment of the present application;

[0084] Figure 3 A schematic diagram showing multiple segmentation points detected in a method for detecting non-linear dynamic data anomalies in multiple process stages of silicon single crystal growth in an exemplary embodiment of the present application;

[0085] Figure 4 A schematic diagram showing the true value, predicted value, and abnormal conditions of the heater power in a simulation experiment of an exemplary embodiment of the present application;

[0086] Figure 5 A schematic diagram showing the anomaly detection result obtained by adopting the dynamic threshold strategy in a simulation experiment of an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0087] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.

[0088] In addition, the accompanying drawings are only schematic illustrations of the present application and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0089] Next, a method for detecting abnormal non - linear dynamic data in multiple process stages of silicon single crystal growth proposed in this exemplary embodiment will be described in more detail.

[0090] This exemplary embodiment provides a method for detecting abnormal non - linear dynamic data in multiple process stages of silicon single crystal growth. As Figure 1 and Figure 2 shown, the method may include the following steps:

[0091] Step S101: Obtain the historical data set and the monitoring data set during the silicon single crystal growth process, reconstruct the historical data set to obtain a reconstructed training set, reconstruct the monitoring data set to obtain a reconstructed test set, and perform Bayesian online change - point detection on the monitoring data set to obtain multiple segmentation points in the monitoring data set.

[0092] In step S101 of this embodiment, the reconstructed training set contains multiple training samples, and the reconstructed test set contains multiple test samples.

[0093] In step S101 of this embodiment, the historical data set is represented by and the monitoring data set is represented by .

[0094] Step S101 of this embodiment may include the following sub - steps:

[0095] Sub - step S1011: Obtain all the historical data during the silicon single crystal growth process to form a historical data set, and form a reconstructed training set from all the non - abnormal historical data in the historical data set.

[0096] In sub - step S1011, the reconstructed training set is represented by , , represents the total length of all historical data in the reconstructed training set, represents the input vector of the th training sample, and the input vector of the th training sample is composed of the th to the th consecutive historical data, , denotes the th historical data in the reconstruction training set, denotes the output target of the th training sample, denotes the th historical data in the reconstruction training set, denotes -dimensional real number space, denotes the th training sample, denotes the multiplication sign, denotes one-dimensional real number space, denotes the dimension.

[0097] Sub-step S1012: Obtain all the monitoring data during the single-crystal silicon growth process to form a monitoring data set, and form the reconstruction test set with all the monitoring data without anomalies in the monitoring data set.

[0098] In sub-step S1012, the reconstruction test set is denoted by ; denotes the total length of all the monitoring data in the reconstruction test set, denotes the input vector of the th test sample. The input vector of the th test sample is composed of the th to the th consecutive monitoring data, ; denotes the th monitoring data in the reconstruction test set, denotes the output target of the th test sample, denotes the th test sample, denotes the th dimension. ;

[0099] Sub-step S1013: Perform Bayesian online change point detection on the monitoring data set to obtain multiple segmentation points.

[0100] Step S102: Construct a relevance vector machine model, and iteratively train the relevance vector machine model using the reconstruction training set to obtain a relevance vector machine training model. Step S102 of this embodiment may include the following sub-steps:

[0101] Sub-step S1021: During the iterative training of the relevance vector machine model using the reconstruction training set, all input vectors of the training samples are used as the input of the relevance vector machine model, all output targets of the corresponding training samples are used as the output of the relevance vector machine model, and the likelihood function of the reconstruction training set is constructed to obtain the relevance vector machine training model.

[0102] In sub-step S1021, the relevance vector machine model is preferably a hybrid kernel relevance vector machine model.

[0103] In sub-step S1021, the expression of the relevance vector machine training model is:

[0104] (1)

[0105] Where, represents the predicted output of the relevance vector machine training model, , represents the number of all training samples in the reconstruction training set, , represents the weight vector of the relevance vector machine training model, represents the th element of the weight vector of the relevance vector machine training model, represents the th input vector of the training sample, , represents the kernel function of the reconstruction training set, , represents the weight of the Gaussian kernel in the kernel function of the reconstruction training set, represents the width of the Gaussian kernel in the kernel function of the reconstruction training set, represents the transpose, represents the bias term of the polynomial kernel in the kernel function of the reconstruction training set, represents the power of the polynomial kernel in the kernel function of the reconstruction training set, represents the th Gaussian white noise corresponding to the training sample, and the Gaussian distribution of the Gaussian white noise is , represents the variance of the Gaussian white noise, represents the exponential function, represents the norm of the vector. The number of all elements of the weight vector is the same as the number of all training samples in the reconstruction training set, and for the th training sample in the reconstruction training set, there is a th element in the weight vector corresponding to it.

[0106] In sub-step S1021, using the Gaussian white noise corresponding to all training samples, calculate the likelihood function of each training sample respectively, and utilize the mutual independence among all training samples in the reconstruction training set to multiply the likelihood functions of all training samples to obtain the likelihood function of the reconstruction training set.

[0107] Furthermore, the expression of the likelihood function of a training sample is:

[0108] P sample y h i w, σ 2 = 1 2π σ 2 exp { - [y h i -Y x h i ,w ] 2 2 σ 2 } (2)

[0109] Wherein, represents the likelihood function of the

[0110] th training sample.

[0111] (2)

[0112] Wherein, represents the likelihood function of the reconstruction training set, represents the output targets of all training samples, represents the kernel matrix, , represents the

[0113] dimensional real number space.

[0114] To prevent the weight vector from overfitting, assume the weight vector to be a Gaussian distribution with zero mean , and take as the prior distribution, , wherein, represents the first hyperparameter, represents the Gaussian distribution, represents the reciprocal of the th value in the first hyperparameter.

[0115] According to Bayes' formula, the expression of the posterior distribution is:

[0116] (4)

[0117] Among them, represents the posterior distribution of the weight vector , represents the conditional probability of the weight vector , , , represents the second hyperparameter, represents the covariance matrix, , , represents the mean vector of the posterior distribution, , represents the weight vector marginal posterior probability of. is a commonly used construction of a diagonal matrix in mathematics.

[0118] The expression for the first hyperparameter is:

[0119] (5)

[0120] Among them, represents the th value of the first hyperparameter after the th iteration update, represents the th value of the effective degrees of freedom, , represents the th value of the first hyperparameter after the th iteration update, represents the element in the th row and th column of the covariance matrix, represents the th value of the mean vector of the posterior distribution.

[0121] The expression for the second hyperparameter is:

[0122] (6)

[0123] Among them, represents transpose of the all-ones column vector of, represents dimensional real space, represents the effective degrees of freedom.

[0124] Using formulas (5) and (6), the first hyperparameter and the second hyperparameter are alternately iteratively updated. When the change amplitudes of both the first hyperparameter and the second hyperparameter are less than the preset amplitude threshold, the joint optimization process ends.

[0125] It should be noted here that the preset amplitude threshold is a parameter result continuously updated through experiments. After multiple iteration rounds, it basically remains below 0.01. Therefore, in this embodiment, the preset amplitude threshold is set to 0.01.

[0126] Step S103: Using the reconstructed training set and the reconstructed test set in the relevant vector machine training model, perform a single-step prediction on the monitoring data set to obtain a prediction sequence, and perform corresponding difference processing on the prediction sequence and the output target of each test sample respectively to obtain a residual sequence. Step S103 of this embodiment may include the following sub-steps:

[0127] Sub-step S1031: Input both the reconstructed training set and the reconstructed test set into the relevant vector machine training model, and use the mean vector of the posterior distribution as the effective estimated value of the weight vector to perform a single-step prediction on the monitoring data set to obtain a prediction sequence.

[0128] Sub-step S1032: As Figure 3 shown, perform corresponding difference processing on all the prediction data in the prediction sequence and the output target of the corresponding test sample in the reconstructed test set respectively, and form all the obtained residual values into a residual sequence. The residual value here refers to the absolute value of the corresponding difference obtained after the corresponding difference processing.

[0129] It can be seen from Figure 3 that in step S103 of this embodiment, the expression for single-step prediction is:

[0130] (7)

[0131] Among them, represents the prediction data of the th step, represents the th training sample, represents the kernel function of the reconstructed test set, , represents the weight of the Gaussian kernel in the kernel function of the reconstructed test set, represents the width of the Gaussian kernel in the kernel function of the reconstructed test set, represents the bias term of the polynomial kernel in the kernel function of the reconstructed test set, represents the power of the polynomial kernel in the kernel function of the reconstructed test set, represents after the The -th value of the mean vector of the posterior distribution after the

[0132] -th iteration update. The prediction sequence is expressed as: , where represents the prediction sequence, represents the

[0133] -dimensional real number space; The residual sequence is expressed as: , where

[0134] represents the residual sequence, and

[0135] represents the output targets of all test samples. Step S104: Use all the segmentation points to divide the residual sequence into multiple segmented windows, and use all the sliding windows to sequentially traverse each segmented window to obtain multiple sliding window datasets. Step S104 of this embodiment may include the following sub-steps: Sub-step S1041: Represent the set of all segmentation points as: , where

[0136] represents the -th segmentation point. Sub-step S1042: Use all the segmentation points to divide the residual sequence into segmented windows; the set of all segmented windows is represented as: , where represents all the residual values within the -th segmented window in the residual sequence, represents the last residual value within the

[0137] -th segmented window in the residual sequence.

[0138] Here, the preset processing means setting the size of all sliding windows to 150 and the step size of all sliding windows to 1.

[0139] Step S105: Introduce a dynamic threshold strategy, generate corresponding candidate threshold sets according to the statistical characteristics of all the sliding window datasets, use all the candidate threshold sets to perform anomaly detection on the corresponding sliding windows respectively, and fuse the anomaly detection results of all the sliding windows to obtain the final anomaly detection result. Step S105 of this embodiment may include the following sub-steps:

[0140] Sub-step S1051: Calculate the mean and standard deviation of all window data in each sliding window dataset respectively, and make the values take equally spaced values within a preset interval; in this embodiment, the range of the preset interval is [2.0,5.0] .

[0141] Sub-step S1052: Calculate multiple candidate thresholds for the corresponding sliding window respectively according to the mean and standard deviation of all window data in each sliding window dataset, and the results of all equally spaced values, and form a candidate threshold set.

[0142] In sub-step S1052, the expression of the candidate threshold is:

[0143] (8)

[0144] where, represents the candidate threshold of the sliding window, represents the mean of all window data in the sliding window dataset, represents value, the quantile in the standard normal distribution, represents the standard deviation of all window data in the sliding window dataset, represents the number of all window data in the sliding window dataset.

[0145] Sub-step S1053: Calculate the Z-value score corresponding to each candidate threshold of all sliding windows respectively.

[0146] In sub-step S1053, the expression of the value score is:

[0147] (9)

[0148] where, represents value score, represents the offset of the mean of all outliers in the sliding window dataset, , represents the mean of all normal points in the sliding window dataset, represents the number of all outliers in the sliding window dataset, , represents the number of all normal points in the sliding window dataset, , represents the offset of the standard deviation of all outliers in the sliding window dataset, , represents the standard deviation of all normal points in the sliding window dataset, represents the number of all outliers in the sliding window dataset after pruning, represents the number of all consecutive outliers in the sliding window dataset.

[0149] Sub-step S1054: Traverse all value scores of each sliding window to obtain the maximum value score corresponding to the value.

[0150] Sub-step S1055: Use all the maximum value scores corresponding to the values to calculate the optimal threshold for the corresponding sliding window respectively.

[0151] In sub-step S1055, the expression for the optimal threshold is:

[0152] (10)

[0153] where, represents the optimal threshold of the sliding window, represents the maximum value score corresponding to the value.

[0154] Sub-step S1056: Use all the optimal thresholds to perform outlier detection on the corresponding sliding windows respectively, and mark all window data greater than or equal to the corresponding optimal threshold in all sliding windows as outliers, and mark all the remaining window data as normal points.

[0155] Sub-step S1056: Introduce sliding window false alarm points, set the lower limit of the number of outliers, and compare the number of all outliers in each sliding window with the lower limit of the number of outliers respectively, including the following two cases:

[0156] The first case: If the number of all outliers in the sliding window is greater than or equal to the lower limit of the number of outliers, it is considered that the sliding window has outliers, and fuse the outlier detection results of all sliding windows to obtain the final outlier detection result.

[0157] The second case: If the number of all outliers in the sliding window is less than the lower limit of the number of outliers, perform pruning on all outliers, and fuse the outlier detection results of all sliding windows after pruning to obtain the final outlier detection result.

[0158] In sub-step S1056 of this embodiment, the lower limit of the number of outliers is set to 2.

[0159] It should be particularly noted that the false alarm points of the sliding window mentioned here refer to the situations where points are wrongly marked as abnormal points but actually belong to normal points.

[0160] In order to verify the excellent effect of a non - linear dynamic data anomaly detection method for multi - process stages of silicon single crystal growth proposed in this application, the following simulation experiments were carried out.

[0161] The data source of this simulation experiment: the National - Local Joint Engineering Research Center for Crystal Growth Equipment and System Integration (abbreviation: Crystal Center). Relevant simulation experiment data were collected from the operation process of a TDR135 silicon single crystal growth equipment in the Crystal Center.

[0162] The heater power data selected from the simulation experiment data were analyzed. As Figure 4 shown, the sampling frequency of the data is 0.2Hz, the data length includes 4800 sampling points, the sampling data set is the heater power data during the normal Czochralski silicon single crystal growth process, the sampling data set has anomalies in the sampling point interval of 1500 - 1600, and there is a process stage switch at the 3649th sampling point. It can be Figure 4 seen that the hybrid kernel relevance vector machine model can well predict the future changes of the data.

[0163] As Figure 5 shown, the dynamic threshold strategy can well adapt to data changes and can be automatically adjusted according to real - time data, only misdetection and missed detection occur at very few sampling points. For example, misdetection occurs at the 3794th sampling point and the 4392nd sampling point. Missed detection occurs at the 1542nd sampling point. However, from the whole detection process, adopting the dynamic threshold strategy can always maintain good and stable performance during the whole process of anomaly detection.

[0164] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this application, "a plurality" means two or more unless otherwise specifically defined.

[0165] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0166] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.

[0167] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed by the present application.

Claims

1. A method for detecting abnormal nonlinear dynamic data in multiple process stages of silicon single crystal growth, characterized in that The method includes the following steps: Obtain a historical data set and a monitoring data set during the single-crystal silicon growth process, reconstruct the historical data set to obtain a reconstructed training set, reconstruct the monitoring data set to obtain a reconstructed test set, and perform Bayesian online change point detection on the monitoring data set to obtain multiple segmentation points in the monitoring data set. Among them, the reconstructed training set contains multiple training samples, and the reconstructed test set contains multiple test samples; Construct a relevance vector machine model, and use the reconstructed training set to iteratively train the relevance vector machine model to obtain a relevance vector machine training model; Use the reconstructed training set and the reconstructed test set in the relevance vector machine training model to perform single-step prediction on the monitoring data set to obtain a prediction sequence, and perform corresponding difference processing on the prediction sequence and the output target of each test sample respectively to obtain a residual sequence, and the residual sequence includes multiple residual values; Use all the segmentation points to divide the residual sequence into multiple segmentation windows, and sequentially traverse each segmentation window using all sliding windows to obtain multiple sliding window data sets; Introduce a dynamic threshold strategy, generate corresponding candidate threshold sets according to the statistical characteristics of all the sliding window data sets, use all the candidate threshold sets to perform anomaly detection on the corresponding sliding windows respectively, and fuse the anomaly detection results of all the sliding windows to obtain a final anomaly detection result.

2. The method for detecting abnormal nonlinear dynamic data in multiple process stages of silicon single crystal growth according to claim 1, wherein The steps of obtaining a historical data set and a monitoring data set during the single-crystal silicon growth process, reconstructing the historical data set to obtain a reconstructed training set, reconstructing the monitoring data set to obtain a reconstructed test set, and performing Bayesian online change point detection on the monitoring data set to obtain multiple segmentation points in the monitoring data set include: Obtain all historical data during the single-crystal silicon growth process to form the historical data set, and form the reconstructed training set with all the non-abnormal historical data in the historical data set; The reconstructed training set is represented by and , represents the total length of all historical data in the reconstructed training set, represents the -th input vector of the training sample. The -th input vector of the training sample consists of the -th to the -th consecutive historical data. , represents the -th historical data in the reconstructed training set, represents the -th output target of the training sample, , represents the -th historical data in the reconstructed training set, represents -dimensional real number space, represents the -th training sample, represents the multiplication sign, represents one-dimensional real number space, represents the -th dimension; Obtain all monitoring data during the single-crystal silicon growth process to form the monitoring data set, and form the reconstructed test set with all the non-abnormal monitoring data in the monitoring data set; Perform the Bayesian online change point detection on the monitoring data set to obtain multiple segmentation points; The reconstructed test set is represented by , , which represents the total length of all monitoring data in the reconstructed test set, represents the input vector of the th test sample. The input vector of the th test sample consists of the th to the th consecutive monitoring data, , represents the th monitoring data in the reconstructed test set, represents the output target of the th test sample, represents the th test sample, represents the th dimension; The historical data set is represented by and the monitoring data set is represented by .

3. The method for detecting abnormal non-linear dynamic data in multiple process stages of silicon single crystal growth according to claim 2, wherein The steps of constructing a relevance vector machine model and using the reconstructed training set to iteratively train the relevance vector machine model to obtain a relevance vector machine training model include: During the iterative training of the relevance vector machine model using the reconstructed training set, use the input vectors of all the training samples as the input of the relevance vector machine model, use the output targets of all the corresponding training samples as the output of the relevance vector machine model, and construct the likelihood function of the reconstructed training set to obtain the relevance vector machine training model; Prevent overfitting of the likelihood function of the reconstructed training set, find multiple hyperparameters, and alternately iterate and update all the hyperparameters until the change amplitude of all the hyperparameters is less than a preset amplitude threshold, completing the joint optimization of all the hyperparameters.

4. The method for detecting abnormal non - linear dynamic data in multiple process stages of silicon single crystal growth according to claim 3, wherein, The expression of the relevant vector machine training model is: (1) Among them, represents the predicted output of the relevance vector machine training model, , represents the number of all training samples in the reconstruction training set, , represents the weight vector of the relevance vector machine training model, represents the th element of the weight vector of the relevance vector machine training model, represents the th input vector of the training sample, , represents the kernel function of the reconstruction training set, , represents the weight of the Gaussian kernel in the kernel function of the reconstruction training set, represents the width of the Gaussian kernel in the kernel function of the reconstruction training set, represents the transpose, represents the bias term of the polynomial kernel in the kernel function of the reconstruction training set, represents the power of the polynomial kernel in the kernel function of the reconstruction training set, represents the th Gaussian white noise corresponding to the training sample, and the Gaussian distribution of the Gaussian white noise is , represents the variance of the Gaussian white noise, represents the exponential function, represents the norm of the vector. The number of all the elements of the weight vector is the same as the number of all the training samples in the reconstruction training set, and for the th training sample in the reconstruction training set, there is a th element in the weight vector corresponding to it; Use the Gaussian white noise corresponding to all the training samples to calculate the likelihood function of each training sample respectively, and use the mutual independence between all the training samples in the reconstructed training set to multiply the likelihood functions of all the training samples to obtain the likelihood function of the reconstructed training set; The expression of the likelihood function of the training sample is: (2) Among them, represents the likelihood function of the th training sample; The expression of the likelihood function of the reconstructed training set is: (3) Among them, represents the likelihood function of the reconstructed training set, represents the output targets of all training samples, represents the kernel matrix, , represents d-dimensional real space; To prevent overfitting of the weight vector , assume that the weight vector follows a Gaussian distribution with zero mean , and take as the prior distribution , where represents the first hyperparameter represents the Gaussian distribution represents the reciprocal of the -th value in the first hyperparameter According to Bayes' formula, the expression of the posterior distribution is: (4) Among them, represents the posterior distribution of the weight vector represents the conditional probability of the weight vector , , represents the second hyperparameter, represents the covariance matrix, , , represents the mean vector of the posterior distribution, , represents the marginal posterior probability of the weight vector; The expression of the first hyperparameter is: (5) Among them, represents the -th value of the first hyperparameter after the -th iteration update, represents the -th value of the effective degrees of freedom, , represents the -th value of the first hyperparameter after the -th iteration update, represents the element in the -th row and -th column of the covariance matrix, represents the -th value of the mean vector of the posterior distribution; The expression of the second hyperparameter is: (6) Among them, denotes the transpose of the all-ones column vector of denotes the n-dimensional real space, denotes the effective degrees of freedom; Use formulas (5) and (6) to perform the alternate iterative update on the first hyperparameter and the second hyperparameter. When the change amplitudes of the first hyperparameter and the second hyperparameter are both less than the preset amplitude threshold, the joint optimization process ends.

5. The method for detecting abnormal non-linear dynamic data in multiple process stages of silicon single crystal growth according to claim 4, wherein The steps of using the reconstructed training set and the reconstructed test set in the relevant vector machine training model to perform single-step prediction on the monitoring data set to obtain a prediction sequence, and performing corresponding difference processing on the prediction sequence and the output target of each test sample to obtain a residual sequence include: Input both the reconstructed training set and the reconstructed test set into the relevant vector machine training model, and use the mean vector of the posterior distribution as the effective estimated value of the weight vector to perform the single-step prediction on the monitoring data set to obtain the prediction sequence; Perform the corresponding difference processing on all the prediction data in the prediction sequence and the output target of the corresponding test sample in the reconstructed test set respectively, and form the residual sequence with all the obtained residual values.

6. The method for detecting abnormal non - linear dynamic data in multiple process stages of silicon single crystal growth according to claim 5, wherein The expression of the single-step prediction is: (7) Among them, represents the prediction data of the step, represents the th training sample, represents the kernel function for reconstructing the test set, , represents the weight of the Gaussian kernel in the kernel function for reconstructing the test set, represents the width of the Gaussian kernel in the kernel function for reconstructing the test set, represents the bias term of the polynomial kernel in the kernel function for reconstructing the test set, represents the power of the polynomial kernel in the kernel function for reconstructing the test set, represents the th value of the mean vector of the posterior distribution after the th iteration update; The predicted sequence is expressed as: , where represents the predicted sequence, represents a real number space of dimension The residual sequence is expressed as: , represents the residual sequence, represents the output target of all test samples.

7. The method for detecting abnormal non - linear dynamic data in multiple process stages of silicon single crystal growth according to claim 1, wherein The steps of using all the segmentation points to divide the residual sequence into multiple segmentation windows and using all the sliding windows to sequentially traverse each segmentation window to obtain multiple sliding window data sets include: The set of all said segmentation points is represented as: , where represents the th segmentation point; Divide the residual sequence into such segmented windows by using all the segmentation points; the set of all the segmented windows is denoted as: , where represents all the residual values within the -th segmented window in the residual sequence, , represents the last residual value within the -th segmented window in the residual sequence; Perform preset processing on all the sliding windows respectively, and use all the preset processed sliding windows to sequentially traverse each segmentation window to obtain multiple sliding window data sets.

8. The method for detecting abnormal non-linear dynamic data in multiple process stages of silicon single crystal growth according to claim 7, characterized in that, The steps of introducing a dynamic threshold strategy, generating corresponding candidate threshold sets according to the statistical characteristics of all the sliding window data sets, using all the candidate threshold sets to perform anomaly detection on the corresponding sliding windows respectively, and fusing the anomaly detection results of all the sliding windows to obtain a final anomaly detection result include: Calculate the mean and standard deviation of all window data in each of the sliding window datasets respectively, and make the values take equally spaced values within a preset interval; Calculate multiple candidate thresholds of the corresponding sliding window respectively according to the mean and standard deviation of all the window data in each sliding window data set and the results of all the equidistant values, and form the candidate threshold set; Calculate the value scores corresponding to each of the candidate thresholds for all of the sliding windows respectively ; Traverse all the value scores of each of the sliding windows respectively, and obtain the maximum value score corresponding to the value; Using all of the said maximum values corresponding to the score values, respectively calculate the optimal thresholds of the corresponding sliding windows; Use all the optimal thresholds to perform anomaly detection on the corresponding sliding windows respectively, mark all the window data greater than or equal to the corresponding optimal threshold in all the sliding windows as anomaly points, and mark all the remaining window data as normal points; Introduce sliding window false alarm points, set the lower limit of the number of abnormal points, and compare the number of all the abnormal points in each of the sliding windows with the lower limit of the number of abnormal points respectively; If the number of all the abnormal points in the sliding window is greater than or equal to the lower limit of the number of abnormal points, it is considered that the sliding window has an abnormality, and the abnormal detection results of all the sliding windows are fused to obtain the final abnormal detection result; If the number of all the abnormal points in the sliding window is less than the lower limit of the number of abnormal points, pruning processing is performed on all the abnormal points, and the abnormal detection results of all the sliding windows after the pruning processing are fused to obtain the final abnormal detection result.

9. The method for detecting abnormal non-linear dynamic data in multiple process stages of silicon single crystal growth according to claim 8, wherein The expression of the candidate threshold is: (8) Among them, represents the candidate threshold of the sliding window, represents the mean value of all window data in the sliding window dataset, represents the value, that is, the quantile in the standard normal distribution, represents the standard deviation of all window data in the sliding window dataset, represents the number of all window data in the sliding window dataset; The expression of the value score is as follows: (9) Among them, denotes the value score, which represents the offset of the mean of all outliers in the sliding window dataset, , denotes the mean of all normal points in the sliding window dataset, denotes the number of all outliers in the sliding window dataset, , denotes the number of all normal points in the sliding window dataset, , denotes the offset of the standard deviation of all outliers in the sliding window dataset, , denotes the standard deviation of all normal points in the sliding window dataset, denotes the number of all outliers in the sliding window dataset after pruning, denotes the number of all consecutive outliers in the sliding window dataset.

10. The method for detecting abnormal non - linear dynamic data in multiple process stages of silicon single crystal growth according to claim 8, wherein, The expression of the optimal threshold is: (10) Among them, represents the optimal threshold of the sliding window, represents the maximum of the sliding window value score corresponding to value.

Citation Information

Patent Citations

  • Vibration monitoring data anomaly detection method based on multi-feature fusion and deep learning

    CN116502163A

  • Time sequence anomaly detection method based on neighborhood information fusion attention mechanism

    CN116680105A

  • Transform-based multivariable time sequence anomaly detection method

    CN116796272A

  • Time sequence anomaly detection method based on association difference and TimeBlock encoder

    CN117743899A

  • Silicon single crystal growth process fault detection method based on adaptive canonical variable analysis

    CN118171066A