Process module data prediction and adjustment method and device and storage medium

The BiGRU model optimized by kernel principal component analysis and artificial fish swarm algorithm solves the problems of computational complexity and accuracy of data-driven models in thin film deposition, realizes efficient prediction and adjustment of thin film deposition parameters, and reduces experimental costs.

CN120913677APending Publication Date: 2025-11-07PIOTECH (SHENYANG) SEMICONDUCTOR EQUIPMENT CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511023522.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing data-driven models for thin film deposition processes suffer from several problems: excessively high dimensionality of input parameters leading to increased computational complexity, prolonged model training and prediction time, overfitting affecting generalization ability, time-consuming manual selection of hyperparameters and difficulty in guaranteeing optimal accuracy, and reliance on extensive prior knowledge and computational resources for traditional mechanism modeling.

Method used

A bidirectional gated cyclic unit (BiGRU) model combining kernel principal component analysis (KPI) with a nested artificial fish swarm algorithm is used to construct a data-driven model for predicting and adjusting thin film deposition parameters through dimensionality reduction and parameter optimization.

Benefits of technology

It improves the prediction accuracy of thin film deposition process, reduces experimental costs and resource consumption, and optimizes the repeatability and stability of thin film deposition process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913677A_ABST
    Figure CN120913677A_ABST
Patent Text Reader

Abstract

The invention relates to a process module data prediction and adjustment method and device and a storage medium, and the method comprises the steps that multi-dimensional data of a process module is obtained, and the multi-dimensional data comprises parameters used for thin film deposition; a data driving model of a process module is used for predicting the multi-dimensional data, and the data driving model is obtained by training and constructing a BiGRU model embedded with an artificial fish swarm algorithm through a kernel principal component analysis algorithm; and parameters for thin film deposition are adjusted according to the prediction result, and the prediction result comprises the film thickness and the deposition rate. Therefore, the prediction precision of the deposition process is improved, the experiment cost and the resource consumption are reduced, and the thin film deposition process is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application mainly relates to the field of process modules, and particularly relates to a process module data prediction and adjustment method and device and a storage medium. BACKGROUND

[0002] With the rapid development of high-end manufacturing industries such as semiconductors, photovoltaics, and display panels, thin film deposition processes, as a core process link, directly determine the device performance and product yield. Since the new generation of devices has put forward more stringent requirements for the thickness, uniformity, composition, and microstructure of thin films, the traditional process development mode based on trial and error has been difficult to meet the demand. How to build an intelligent process model that can adapt to multiple devices and multiple formulations is crucial to improving the repeatability and stability of the deposition process. The existing model has the following problems:

[0003] 1) The thin film deposition process involves numerous input parameters, resulting in excessively high data dimensions. This not only increases the computational complexity but also significantly prolongs the model training and prediction time, which may cause overfitting and affect the generalization ability of the model. Therefore, there is an urgent need for effective dimension reduction techniques to improve computational efficiency and model performance.

[0004] 2) In modeling, manual selection of hyperparameters often relies on experience, which is time-consuming and difficult to ensure optimal model accuracy. Improper parameter selection may lead to insufficient training or failure to capture complex data patterns. Therefore, developing an automated parameter optimization method is crucial to improving model prediction accuracy.

[0005] 3) Traditional mechanism modeling relies on a large amount of prior knowledge and complex mathematical models, requiring researchers to have in-depth professional knowledge and consuming a large amount of computational resources. This limits the effectiveness of mechanism models in dealing with variable process conditions.

[0006] Therefore, seeking an effective modeling technique that does not require a large amount of prior knowledge and utilizes data-driven methods will be the key to promoting thin film deposition process research. SUMMARY

[0007] An object of the present application is to provide a process module data prediction and adjustment method, device and storage medium, which solves the problems of slow data-driven model operation speed caused by excessively large input parameter dimensions of the data-driven model, low model accuracy caused by manual selection of model parameters, and the need for a large amount of prior knowledge and complex algorithms in modeling.

[0008] According to one aspect of the present application, a process module data prediction and adjustment method is provided, which comprises:

[0009] acquiring multi-dimensional data of a process module, wherein the multi-dimensional data includes parameters for thin film deposition;

[0010] predicting the multidimensional data using a data-driven model of a process module, wherein the data-driven model is constructed by training a BiGRU model embedded with a kernel principal component analysis algorithm and an artificial fish swarm algorithm; and

[0011] adjusting parameters for thin film deposition according to the prediction results, wherein the prediction results include film thickness and deposition rate.

[0012] Optionally, the construction process of the data-driven model comprises the following steps:

[0013] performing dimension reduction processing on the multidimensional data of the original process module by kernel principal component analysis to obtain target dimensional data;

[0014] training a BiGRU model using the target dimensional data and initializing an artificial fish swarm algorithm using the target dimensional data during the training process;

[0015] performing parameter optimization on the BiGRU model using the initialized artificial fish swarm algorithm; and

[0016] inputting the optimal parameter combination obtained by the optimization into the trained BiGRU model to obtain a data-driven model, wherein the data-driven model is used for parameter prediction of the process module.

[0017] Optionally, the dimension reduction processing on the multidimensional data of the process module by kernel principal component analysis to obtain target dimensional data comprises:

[0018] constructing a sample matrix according to the multidimensional data of the process module;

[0019] determining a centralized kernel matrix according to the sample matrix and a Gaussian kernel function;

[0020] determining principal components according to the centralized kernel matrix and calculating the cumulative contribution rate of each principal component; and

[0021] selecting target dimensional data according to the target value of thin film deposition and the cumulative contribution rate.

[0022] Optionally, the parameter optimization on the BiGRU model using the initialized artificial fish swarm algorithm comprises:

[0023] defining a parameter combination of the BiGRU model, wherein the parameter combination comprises a learning rate, a number of hidden nodes, and L2 regularization;

[0024] using the parameter combination as the initial fish swarm position in the artificial fish swarm to perform iterative optimization on the BiGRU model and determine a fitness value; and

[0025] updating the optimal fish swarm position in the artificial fish swarm according to the fitness value, and obtaining the optimal parameter combination when the number of iterations reaches a set value.

[0026] Optionally, determining the fitness value comprises:

[0027] determining a fitness value function according to the true value of the target dimension data and the predicted value of the BiGRU model, and determining each fitness value according to the fitness value function, wherein the fitness value function satisfies the following formula:

[0028]

[0029] wherein F represents the fitness value function, N represents the number of iterations, o k represents the true value of the target dimension data, y k represents the predicted value of the BiGRU model.

[0030] Optionally, the method further comprises:

[0031] inputting a test set in the target dimension data into the data-driven model to verify whether the model precision meets the design requirement, if not, re-performing the construction of the data-driven model, and if yes, inputting the obtained multi-dimensional data of the process module into the data-driven model for prediction.

[0032] Optionally, inputting a test set in the target dimension data into the data-driven model to verify whether the model precision meets the design requirement comprises:

[0033] inputting a test set in the target dimension data into the data-driven model to determine the value of the evaluation index according to the output result, wherein the evaluation index comprises the mean absolute percentage error, the root mean square error and the mean absolute error;

[0034] verifying whether the model precision meets the design requirement according to the value of the evaluation index.

[0035] Optionally, the parameters for thin film deposition include any combination of in-cavity pressure, target-substrate distance, high frequency, low frequency, temperature, start frequency time, carbon dioxide, helium and silane.

[0036] According to yet another aspect of the present application, a device for process module data prediction and adjustment is also provided, which comprises:

[0037] one or more processors; and

[0038] a memory storing computer readable instructions which, when executed, cause the processor to perform the operations of the method as described above.

[0039] According to another aspect of the present application, a computer readable medium having computer instructions stored thereon is also provided, the computer instructions executable by a processor to implement the method as described above.

[0040] Compared with the prior art, the present application obtains multi-dimensional data of a process module, wherein the multi-dimensional data comprises parameters for thin film deposition; uses a data-driven model of the process module to predict the multi-dimensional data, wherein the data-driven model is constructed by training a BiGRU model embedded with a kernel principal component analysis algorithm and an artificial fish swarm algorithm; and adjusts the parameters for thin film deposition according to a prediction result, wherein the prediction result comprises film thickness and deposition rate. Thus, the prediction accuracy of the deposition process is improved, the experimental cost and resource consumption are reduced, and the thin film deposition process is optimized. BRIEF DESCRIPTION OF DRAWINGS

[0041] In order to make the above objectives, features and advantages of the present application more apparent and comprehensible, the specific embodiments of the present application are described in detail below with reference to the accompanying drawings, in which:

[0042] Figure 1 A process module data prediction and adjustment method flowchart according to an aspect of the present application is shown;

[0043] Figure 2 A data-driven model structure diagram in an embodiment of the present application is shown;

[0044] Figure 3 A data-driven model construction process diagram in an embodiment of the present application is shown;

[0045] Figure 4 A process parameter prediction flowchart in an embodiment of the present application is shown;

[0046] Figure 5 A KPCA dimensionality reduction processing data diagram in an embodiment of the present application is shown;

[0047] Figure 6 A fitness function value change diagram with iteration number in an embodiment of the present application is shown;

[0048] Figure 7 A deposition rate prediction result comparison diagram of each model in an embodiment of the present application is shown;

[0049] Figure 8 A film thickness prediction result comparison diagram of each model in an embodiment of the present application is shown;

[0050] Figure 9 A process module output parameter prediction evaluation index information diagram in an embodiment of the present application is shown;

[0051] Figure 10 FIG. 1 shows a schematic diagram of a framework of an apparatus according to an aspect of the present application.

[0052] The same or similar reference signs in the drawings represent the same or similar components. DETAILED DESCRIPTION

[0053] In order to make the above objectives, features and advantages of the present application more clear and easily understood, the specific embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0054] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. The present application, however, can be practiced in a variety of ways other than those specifically described herein, and the present application is not limited to the specific embodiments described herein.

[0055] As shown in the present application and claims, unless the context clearly indicates otherwise, the words "one", "an", "a", and / or "the" do not mean "only one", "single" or "just one", but can include a plurality or "one or more" than one. Generally, the terms "comprising" and "including" only indicate the inclusion of the steps and elements explicitly identified, and these steps and elements do not constitute an exclusive list, and the method or device can also include other steps or elements.

[0056] Figure 1 FIG. 1 shows a schematic diagram of a method flow of process module data prediction and adjustment according to an aspect of the present application, the method comprising steps S11-S13, wherein,

[0057] In step S11, multi-dimensional data of a process module is obtained, wherein the multi-dimensional data comprises parameters for thin film deposition.

[0058] The parameters of the process module include relevant parameters for the thin film deposition process, such as temperature, pressure, gas flow and ratio, time, etc. The obtained parameters of the process module in different dimensions can be used as input parameters for parameter prediction of the thin film deposition process.

[0059] In step S12, the multi-dimensional data is predicted using a data-driven model of the process module, wherein the data-driven model is constructed by training a BiGRU model embedded with a kernel principal component analysis algorithm and an artificial fish swarm algorithm.

[0060] The parameters used for thin film deposition are input into a pre-constructed data-driven model for prediction. This model is constructed using a Kernel Principal Components Analysis (KPCA) algorithm combined with a nested Artificial Fish Swarm Algorithm (AFSA) bidirectional gated recurrent unit (BiGRU) model. KPCA is used to reduce the dimensionality of the received data, while BiGRU consists of bidirectional gated recurrent units (GRUs). GRUs dynamically control information flow and alleviate gradient problems by introducing update and reset gates. Since thin film deposition is a dynamic temporal process, its state is constantly linked to the past and future. Using a unidirectional GRU would limit prediction performance because it only considers past information. Therefore, using a bidirectional gated recurrent unit (BiGRU) allows learning past and future contextual information for thin film deposition data-driven modeling. In this application, as... Figure 2 The schematic diagram of the model structure shown indicates that the data-driven model includes a KPCA module and a model module. The model module includes an AFSA module and a BiGRU neural network module. AFSA is nested in a bidirectional gated recurrent unit (BiGRU), so that AFSA can be used to optimize the BiGRU model and construct a data-driven model that meets the process requirements. This data-driven model is constructed based on KPCA-AFSA-BiGRU.

[0061] In one embodiment of this application, as Figure 3 As shown, the construction process of the data-driven model includes the following steps:

[0062] Step S121 involves dimensionality reduction of the multidimensional data of the original process module using kernel principal component analysis to obtain the target dimension data. Here, the multidimensional data of the original process module is obtained as the data foundation. After processing the data, training samples are obtained. The original process module refers to the process module before parameter adjustment, and the parameters of this process module are used to train the model. In one embodiment of this application, the multidimensional data may include any combination of several of the following: cavity pressure (P), target-base distance (Gap), high frequency (HF), low frequency (LF), temperature (Temp), dep time, carbon dioxide (CO2), helium (He), and silane (SiH4). In this example, parameters including the above nine dimensions are used as input data. After preprocessing the input data, training samples are obtained, and then the model is trained.

[0063] In the data preprocessing part, the dimensional data is converted into dimensionless data by normalization to solve the comparability problem between data; the multi-dimensional data of the original process module is processed by kernel principal component analysis (KPCA) for dimension reduction; for example, the original 9-dimensional input variables P (Torr), Gap (mm), HF (w), LF (w), Temp (℃), SiH4 (sccm), CO2 (sccm), He (sccm) and Dep time (s) of the process module are processed for dimension reduction. Under the premise of losing little accuracy, the original 9-dimensional input data is reduced to new 3-dimensional data containing the characteristics of the original 9-dimensional input parameters, greatly reducing the dimension of the input parameters.

[0064] Step S122, training the BiGRU model using the target dimension data, and initializing the artificial fish swarm algorithm using the target dimension data during the training process.

[0065] The total output value of the bidirectional gate recurrent unit (BiGRU) at any time is jointly determined by the output results of the forward GRU and the backward GRU at the time, satisfying the following formula:

[0066]

[0067] In the formula: and respectively represent the hidden layer output of the forward GRU and the backward GRU at time t; h t is the hidden layer output of the BiGRU at time t; α t and β t are the weights corresponding to the hidden layer states of the forward GRU and the backward GRU, respectively; b t is the corresponding bias coefficient.

[0068] The thin film deposition process is a typical dynamic time series process, and the quality of the current thin film depends not only on the current process parameters, but also on the historical parameter trajectory. The BiGRU bidirectional time series capture learns historical information and future information through the forward + backward GRU layer, for example, the current time thin film growth may be affected by the previous time pretreatment and the next time cooling process. A single GRU neural network can only learn forward data features, while BiGRU can learn forward + backward data information, so that BiGRU can learn more information from the same data set, and thus obtain a more accurate prediction model. Thin film deposition experiments are costly, and the amount of historical data is usually limited, and the data distribution of different devices and process formulations is quite different. BiGRU is suitable for small sample and transfer learning, and BiGRU has high parameter efficiency. GRU has one less gating unit than LSTM, and the structure is more concise, which is easier to train with a small amount of data, reduces the risk of overfitting, and is friendly to transfer learning. Using BiGRU for data-driven modeling of thin film deposition has unique advantages compared to traditional methods such as linear regression, PCA, static neural networks and other time series models, and compared to LSTM and unidirectional GRU, it is particularly suitable for modeling of thin film deposition, which is a multi-parameter, strong time series dependent and nonlinear dynamic process.

[0069] The BiGRU model can be pre-trained by the data after KPCA dimension reduction, and the initialized artificial fish swarm algorithm is used to optimize the parameters of the BiGRU model through step S123; then, the optimal parameter combination obtained by optimization is input into the trained BiGRU model to obtain a data-driven model, wherein the data-driven model is used for parameter prediction of a process module.

[0070] In data-driven modeling of thin film deposition processes, the artificial fish swarm algorithm (AFSA) is used to optimize the parameters of the BiGRU neural network, which has significant advantages and can effectively solve many challenges faced by traditional optimization methods. The performance of BiGRU highly depends on the reasonable configuration of hyperparameters such as the number of hidden nodes, the learning rate and the L2 regularization coefficient, while traditional grid search or random parameter tuning methods are inefficient and prone to local optimal solutions. AFSA realizes global optimization by simulating the foraging, grouping and tail chasing behaviors of fish schools, and its population parallel search mechanism and adaptive step size adjustment can significantly reduce the risk of falling into local optimum, which is particularly suitable for solving the gradient vanishing or explosion problem caused by the coupling of learning rate and network depth in BiGRU.

[0071] In the embodiments of the present application, the optimal learning rate, the number of hidden nodes and the L2 regularization parameter are found by AFSA, and the batch size, the number of training rounds and the number of units in each layer of BiGRU are fine-tuned on large-scale similar process data to quickly adapt to new equipment or new material systems.

[0072] Step S13, adjusting the parameters for thin film deposition according to the prediction results, wherein the prediction results include film thickness and deposition rate. Here, when using the data-driven model to make predictions for the process module, the input parameters are the parameters for thin film deposition, and the prediction results are film thickness and deposition rate. The input parameters are adjusted according to the predicted two-dimensional data, and the adjustment of the thin film deposition process is completed.

[0073] As shown in Figure 4 The obtained parameters for the process are subjected to data preprocessing, which includes dimensionality reduction by KPCA and normalization processing. Based on the processed data, a KPCA-AFSA-BiGRU-based process module data-driven model is trained. The data-driven model is a prediction model. It is determined whether the trained model meets the accuracy requirement. If not, further training is needed. If it does, the established prediction model is used for prediction, and the predicted value is output. The prediction model of the present application can improve the prediction accuracy of the deposition process, help researchers more accurately predict the properties and performance of thin films, reduce the dependence on a large number of experiments through computer simulation and data analysis, thereby reducing experimental costs and resource consumption, and the model can also adapt to different process conditions and material systems, enabling researchers to flexibly cope with various changes and optimize the thin film deposition process. The model construction and training process are specifically introduced as follows:

[0074] In an embodiment of the present application, in step S121, a sample matrix is constructed according to the multi-dimensional data of the process module; a centralized kernel matrix is determined according to the sample matrix and a Gaussian kernel function; each principal component is determined according to the central kernel matrix, and the cumulative contribution rate of each principal component is calculated; and target-dimensional data is selected according to the target value of the thin film deposition and the cumulative contribution rate.

[0075] The Gaussian kernel function KPCA is selected. The Gaussian kernel KPCA can flexibly capture the extremely complex nonlinear interaction effects and saturation effects between numerous process parameters, such as temperature, pressure, gas flow / ratio, time, etc. Moreover, thin film deposition usually has a narrow "process window", and a slight change in parameters within this window can have a significant impact on the performance of the thin film. For example, a few degrees of temperature fluctuation can change the crystal phase, or a slight change in gas ratio can affect the uniformity of the composition. The locality of the Gaussian kernel is its core advantage; it gives data points that are close in distance a higher similarity after kernel mapping in the feature space or the original space. This enables KPCA to have the following beneficial effects:

[0076] It can sensitively identify the slight deviation of process parameters and its influence on the results; it can accurately depict the process window boundary. In the principal component space, due to the high similarity of data points given by the Gaussian kernel, the identification of outliers becomes more obvious, and the samples within the process window may form relatively clustered but clear boundary clusters, so that in the data-driven modeling process, the neural network is easier to learn the data characteristics within the corresponding process window; it can detect subtle process drifts or outliers, which deviate from the cluster of normal process window in the principal component space.

[0077] KPCA maps data to high-dimensional space through kernel function, ensures linear separability, and extracts key features through principal component analysis for dimension reduction and feature extraction. Compared with traditional PCA, KPCA can effectively capture the nonlinear relationship of data, retain information and remove redundancy, improve data quality and reduce computational cost.

[0078] The specific steps of extracting principal component factors by KPCA method are as follows:

[0079] First, project the samples into the high-dimensional feature space. For the input sample matrix X = [x1, x2, …, xm] composed of various thin film deposition input parameters, where m is the number of samples, introduce the mapping to project the samples into the high-dimensional feature space, and form the new matrix m and satisfy

[0080] Next, select the Gaussian kernel function to calculate the kernel matrix R:

[0081]

[0082] where ‖x i -x j ‖ is the Euclidean distance between samples x i and x j , and σ is the bandwidth parameter of the Gaussian kernel.

[0083] Then, calculate the centralized kernel matrix R C using the kernel matrix: R C = R-I N R-KI N +I N KI N ; in the formula: I N represents an N × N matrix.

[0084] Calculate the eigenvalues and corresponding eigenvectors of the matrix, and sort them by size, where the largest eigenvalue is the first principal component. Calculate the cumulative contribution rates β1, β2, …, β n , which satisfy the following formula:

[0085] ​

[0086] The first h principal components selected on the basis of setting the target P according to the dimension reduction requirement are used as the input variables after dimension reduction, that is, the target dimension data is selected, wherein P = 99% is set in the thin film deposition data driven modeling. Figure 5 As shown in the figure, the original input parameter is 9-dimensional data, and through KPCA dimension reduction, the cumulative contribution rate reaches 99% at the principal component number 3, so the first 3 principal components are selected, thereby reducing the original 9-dimensional data to 3-dimensional data.

[0087] In an embodiment of the present application, in step S123, a parameter combination of the BiGRU model is defined, wherein the parameter combination includes a learning rate, a number of hidden nodes, and L2 regularization; the parameter combination is used as an initial fish population position in the artificial fish school to iteratively optimize the BiGRU model and determine a fitness value; the optimal fish population position in the artificial fish school is updated according to the fitness value, and when the number of iterations reaches a set value, the optimal parameter combination is obtained.

[0088] AFSA simulates the behavior characteristics of fish schools, moves to the position of individuals with higher food concentration in the neighborhood space, and realizes optimization through position updating, with characteristics such as fast search and global optimization. The state of an artificial fish individual is represented as a BiGRU hyperparameter vector θ, which contains a learning rate, a number of hidden nodes, and an L2 regularization parameter, and the hyperparameter vector is used as a variable to be optimized. The AFSA algorithm population is initialized, iteratively optimized according to the behavior of the artificial fish school, and the bulletin board, i.e., the hyperparameter vector of the BiGRU model, is updated; then, a fitness value function is determined, which can determine the fitness value function according to the true value of the target dimension data and the predicted value of the BiGRU model, and determine each fitness value according to the fitness value function, wherein the fitness value function satisfies the following formula:

[0089]

[0090] Wherein F represents the fitness value function, N represents the number of iterations, o k represents the true value of the target dimension data, y k represents the predicted value of the BiGRU model.

[0091] The initial parameters of the artificial fish swarm are used as the initial values ​​predicted by the BiGRU model. The root mean square error between the true values ​​of the training samples and the output values ​​of the BiGRU model is used as the food concentration at the location of the artificial fish individual, i.e., the fitness value Y. The target dimension data is split into a training set and a test set, with the data in the training set used as the training samples. The fitness value of the first-generation artificial fish swarm is determined according to the fitness value function. The positions of the artificial fish individuals are updated according to the three behaviors of the fish swarm, and the optimal fish swarm position in AFSA is updated. The iteration count is checked to see if the set value has been reached. If not, the fitness value function is redefined. If it has been completed, the optimal combination of learning rate, number of hidden nodes, and L2 regularization coefficient is obtained. In this embodiment, for the thin film deposition dataset, the obtained learning rate is 0.0091, the number of hidden nodes is 97, and the L2 regularization coefficient is 1.3214e. -4 Finally, a data-driven model for the process module was established.

[0092] More specifically, the update of the location of an artificial fish follows three behaviors:

[0093] 1) Foraging behavior, artificially raised fish individual θ i A state θ is randomly selected within its perceptual range visualization. j Calculate the food concentration value. If the moving condition is met, move one step in that direction; otherwise, move one step randomly.

[0094] 2) Grouping behavior, individual artificial fish θ i Find other individuals in the neighborhood and the center location θ c Calculate the food concentration value. If the conditions for moving forward are met, move one step towards the center. If not, perform foraging behavior and update the food concentration value.

[0095] 3) Tail-chasing behavior, artificial fish individual θ i Find the partner with the highest food concentration value θ in the neighborhood. j If the conditions for moving forward are met, move one step towards your partner; if not, perform foraging behavior and update the food concentration value.

[0096] like Figure 6 The graph shows the change in fitness function values ​​with the number of iterations. In the graph, A represents the fitness function value of the model without KPCA treatment, and B represents the change in fitness function value of the model with KPCA treatment. It can be seen that in the first 8 iterations, the fitness function values ​​of both types of models dropped sharply, and then stabilized at around 52, with minimal difference between the fitness function values ​​of the two types of models. When the number of iterations increases, exceeding 8, the fitness function value of the model with KPCA treatment is 52.08, and the fitness function value of the model without KPCA treatment is 51.77, a difference of 0.31.

[0097] In one embodiment of this application, the test set from the target dimension data can also be input into the data-driven model to verify whether the model accuracy meets the design requirements. If not, the data-driven model is reconstructed; if so, the multidimensional data of the acquired process modules is input into the data-driven model for prediction. Here, the target dimension data is divided into a training set and a test set. The training set is used to train the model, and the test set is input into the trained model to verify whether the model accuracy meets the design requirements. If not, the model needs to be reconstructed and trained again. If the design requirements are met, the data of the process modules to be tested is input into the model that meets the design requirements for prediction.

[0098] Specifically, the test set from the target dimension data is input into the data-driven model, and the values ​​of evaluation indicators are determined based on the output results. These evaluation indicators include mean absolute percentage error (MAPE), root mean square error (RMSE), and mean absolute error (MAE). The model accuracy is verified based on the values ​​of these evaluation indicators to determine whether it meets the design requirements. Here, mean absolute percentage error (MAPE), root mean square error (RMSE), and mean absolute error (MAE) are used to quantitatively evaluate the model and verify the accuracy of the established process module data-driven model.

[0099] like Figure 7 The comparison chart showing the prediction results of each model for deposition rate (DR) is shown below. Figure 8 The chart shows a comparison of the prediction results for membrane thickness by various models, comparing real data with the predictions of several models. KPCA-GRU refers to a model obtained by training a GRU after dimensionality reduction using KPCA; KPCA-LSTM is a model constructed by combining KPCA and LSTM; and KPCA-LSF is a model constructed by combining KPCA and LSF. The predictive performance of each model is evaluated according to evaluation metrics, including MAPE, MAE, and RMSE. Figure 9 As shown, the process module data-driven model proposed in this application, which integrates KPCA, AFSA, and BiGRU, can not only reduce the training time of the model, but also learn the patterns between historical operating data, thereby achieving the goal of accurately predicting various parameters.

[0100] Figure 10 The diagram shows a schematic frame of an apparatus according to another aspect of this application, which can be an apparatus for predicting and adjusting process module data, the apparatus including at least a processor 101 and a memory 102.

[0101] The processor 101 can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 101 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor 101 can also include a main processor and a coprocessor, the main processor being a processor for processing data in an awake state, also known as a CPU (Central Processing Unit), and the coprocessor being a low-power processor for processing data in a standby state. In some embodiments, the processor 101 can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing the content required to be displayed by the display screen. In some embodiments, the processor 101 can also include an AI (Artificial Intelligence) processor for processing machine learning-related computing operations.

[0102] The memory 102 can include one or more computer-readable storage media, which can be non-transitory. The memory 102 can also include a high-speed random access memory, and a nonvolatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 102 is used to store at least one instruction for being executed by the processor 101 to implement a process module data prediction and adjustment method provided by the method embodiment of the present application.

[0103] In some embodiments, the apparatus can also optionally include a peripheral device interface and at least one peripheral device. The processor 101, the memory 102, and the peripheral device interface can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface through a bus, a signal line, or a circuit board. Illustratively, the peripheral devices include, but are not limited to, radio frequency circuitry, a touch display screen, audio circuitry, and a power supply, etc.

[0104] Of course, the apparatus can also include fewer or more components, which are not limited in the present embodiment.

[0105] The present application also provides a computer-readable medium having computer instructions stored thereon, the computer instructions being executable by a processor to implement a process module data prediction and adjustment method as described above.

[0106] The method of process module data prediction and adjustment, when implemented as a computer program, can also be stored in a computer-readable storage medium as an article of manufacture. For example, the computer-readable storage medium can include, but is not limited to, magnetic storage devices (e.g., hard disk, floppy disk, magnetic strips), optical disks (e.g., compact disk (CD), digital versatile disk (DVD)), smart cards, and flash memory devices (e.g., electrically erasable programmable read only memory (EPROM), card, stick, key drive). Additionally, the various storage mediums described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" can include, without limitation, wireless channels and various other media (and / or storage media) that are capable of storing, containing, and / or carrying code and / or instructions and / or data.

[0107] It should be understood that the embodiments described above are only illustrative. The embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or any combination thereof. For a hardware implementation, the processors can be implemented within one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, and / or other electronic units designed to perform the functions described herein, or a combination thereof.

[0108] Some aspects of the present application can be performed entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. The above hardware or software can be referred to as a "block", "module", "engine", "unit", "component", or "system". The processor can be one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DAPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, or a combination thereof. In addition, aspects of the present application can be manifested as a computer product in a form of a computer-readable medium storing computer program codes. For example, the computer-readable medium can include, but is not limited to, magnetic storage devices (e.g., hard disk, floppy disk, magnetic tape...), optical disks (e.g., compact disk (CD), digital versatile disk (DVD)...), smart cards, and flash memory devices (e.g., card, stick, key drive...).

[0109] A computer readable medium can include a propagated data signal with computer program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any combination thereof. Computer readable media can be any media that can be accessed by a computer. By way of example, and not limitation, such computer readable media can comprise RAM, ROM, EEPROM, CD-ROM or any combination thereof. The computer program product can be tangibly embodied in an information carrier. The computer program product can also contain instructions that, when executed, perform one or more methods, such as those described above. The computer program product can be tangibly embodied in an information carrier code that can be accessed by a machine and that can cause the machine to perform a series of operations. The operations described above with reference to the methods discussed above can be stored as information in a computer readable medium. The computer program product can also contain appropriate distribution conditions.

[0110] The foregoing description of the exemplary embodiment has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the application to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. It is intended that the scope of the application be limited not with this detailed description, but rather by the claims appended hereto.

[0111] Also, the use of "a" or "an" to describe elements, instrumentation, etc., is merely for convenience and to give some the grammatical structure to this document. One of ordinary skill in the art will appreciate that "a" or "an" can be construed to be "one or more" or "at least one." Also, structures described herein are to be understood to refer to a single structure or a combination of structures.

[0112] In some embodiments, numerical descriptions of components, quantities of attributes, etc., are used. It should be understood that such numerical descriptions used in the description of embodiments can be modified by the modifying words "about," "approximately," or "generally" in some examples. Unless otherwise indicated, "about," "approximately," or "generally" indicates that the described numerical value allows for a ±20% variation. Accordingly, numerical values used in the specification and claims are approximations that can vary depending on the desired properties sought to be obtained in a particular embodiment. In some embodiments, numerical values should be considered in the context and be rounded to the appropriate significant figures. Notwithstanding that the numerical ranges and parameters setting forth the broadest scope of the application are approximations, the numerical values set forth in the specific examples are reported as precisely as possible.

Claims

1. A method of process module data prediction and adjustment, characterized by, The method comprises: acquiring multi-dimensional data of a process module, wherein the multi-dimensional data comprises parameters for thin film deposition; predicting the multi-dimensional data using a data-driven model of the process module, wherein the data-driven model is constructed by training a BiGRU model embedded with a kernel principal component analysis algorithm and an artificial fish swarm algorithm; and adjusting the parameters for thin film deposition according to a prediction result, wherein the prediction result comprises film thickness and deposition rate.

2. The method of claim 1, wherein, The construction process of the data-driven model comprises the following steps: performing dimension reduction processing on multi-dimensional data of an original process module by kernel principal component analysis to obtain target dimension data; training a BiGRU model using the target dimension data and initializing an artificial fish swarm algorithm using the target dimension data during the training process; performing parameter optimization on the BiGRU model using the initialized artificial fish swarm algorithm; and inputting an optimal parameter combination obtained by the optimization into the trained BiGRU model to obtain a data-driven model, wherein the data-driven model is used for parameter prediction of the process module.

3. The method of claim 2, wherein, performing dimension reduction processing on multi-dimensional data of a process module by kernel principal component analysis to obtain target dimension data, comprising: constructing a sample matrix according to the multi-dimensional data of the process module; determining a centralized kernel matrix according to the sample matrix and a Gaussian kernel function; determining principal components according to the centralized kernel matrix and calculating cumulative contribution rates of the principal components; and selecting target dimension data according to target values of thin film deposition and the cumulative contribution rates.

4. The method of claim 2, wherein, performing parameter optimization on the BiGRU model using the initialized artificial fish swarm algorithm, comprising: defining a parameter combination of the BiGRU model, wherein the parameter combination comprises a learning rate, a number of hidden nodes, and L2 regularization; inputting the parameter combination as an initial fish swarm position in the artificial fish swarm to iteratively optimize the BiGRU model and determine a fitness value; and updating an optimal fish swarm position in the artificial fish swarm according to the fitness value, and obtaining an optimal parameter combination when the number of iterations reaches a set value.

5. The method of claim 4, wherein, determining a fitness value comprises: determining a fitness value function according to true values of the target dimension data and prediction values of the BiGRU model, and determining each fitness value according to the fitness value function, wherein the fitness value function satisfies the following formula: Wherein, F represents the fitness value function, N represents the number of iterations, o k The real value of the target dimension data is represented by y k The predicted value of the BiGRU model is represented by 6. The method of claim 2, wherein, The method further comprises: inputting a test set in the target dimension data into the data-driven model to verify whether model accuracy meets design requirements, and if not, re-performing construction of the data-driven model, and if so, inputting the acquired multi-dimensional data of the process module into the data-driven model for prediction.

7. The method of claim 6, wherein, inputting a test set in the target dimension data into the data-driven model to verify whether model accuracy meets design requirements, comprising: inputting the test set in the target dimension data into the data-driven model to determine values of evaluation indexes according to output results, wherein the evaluation indexes comprise mean absolute percentage error, root mean square error, and mean absolute error; and verifying whether model accuracy meets design requirements according to the values of the evaluation indexes.

8. The method of claim 1, wherein, The parameters for thin film deposition include any combination of in-chamber pressure, target-to-substrate distance, high frequency, low frequency, temperature, on-time, carbon dioxide, helium, and silane.

9. An apparatus for process module data prediction and adjustment, comprising: The apparatus comprises: one or more processors; and memory having stored computer-readable instructions that, when executed, cause the processor to perform operations of the method of any of claims 1-8.

10. A computer-readable medium having stored thereon computer instructions executable by a processor to implement the method of any of claims 1-8.

Citation Information

Patent Citations

  • Photovoltaic power generation power prediction method based on gating circulation unit

    CN115579859A

  • Method and system for predicting thickness of deposited film of integrated circuit based on automatic machine learning

    CN116384247A

  • Integrated circuit deposition film thickness prediction method and system

    CN116431996A

  • CNN-XGBoost-GA model-based integrated circuit deposited film thickness prediction method

    CN116662809A

  • Lithium ion battery remaining service life prediction method based on AFSA-GRU

    CN116794547A