Film thickness prediction method and system based on self-supervision and characterization learning
By adopting self-supervised and characterization learning methods in film thickness prediction, using label-free data pre-trained encoder to build dynamic feature selectors and global prototype projection space, the problem of deep learning methods relying on labeled data and feature selection in film thickness prediction is solved, and higher prediction accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510593095.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-09
AI Technical Summary
Existing deep learning methods rely on a large amount of labeled data in film thickness prediction, feature selection is difficult, and feature representations have aliasing and localization problems, which affects the stability and accuracy of the prediction results.
Using a method based on self-supervisation and representation learning, the encoder is pre-trained with labelless data to reduce dependence on labeled data, and a dynamic feature selector is constructed through embedded feature selection, which optimizes the representation space, reduces aliasing and localization problems, and improves the reliability and adaptability of prediction through optimal transmission mechanism and orthogonality constraints.
Reliance on labeled data is reduced, the generalization ability and prediction accuracy of the model are improved, and the stability of the prediction results and the adaptability of the model are enhanced.
Smart Images

Figure CN120124013A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning and artificial intelligence, and particularly relates to a method and system for predicting film thickness based on self-supervised and representation learning. Background Art
[0002] In the field of semiconductor manufacturing, precise control of the production process is crucial for ensuring high-quality products and operational efficiency. Chemical Vapor Deposition (CVD) is one of the key processes for thin film preparation, and its process quality will directly affect the reliability of semiconductor devices. However, due to the diversity and complexity of process variables and possible performance drifts of the equipment itself, the uniformity of film thickness remains a very challenging problem to date. Traditional Virtual Metrology (VM) techniques mainly rely on prediction methods based on physical models or machine learning, and the existing deep learning prediction methods have the following three problems:
[0003] (1) Dependence on a large amount of labeled data: The accuracy of current deep learning methods for predicting film thickness highly depends on a large amount of labeled data. However, due to the time-consuming and costly measurement process, it is extremely difficult to obtain labeled data.
[0004] (2) Feature selection problem: Feature selection methods are mainly divided into three categories: filter method, wrapper method, and embedding method. The filter method is simple to calculate, but it does not consider the interaction between features and is difficult to apply to semiconductor manufacturing data with high complexity. The wrapper method selects features by training the model iteratively multiple times, with a high computational cost, suitable for traditional machine learning methods but not for deep learning. The existing deep learning feature screening methods lack in-depth research on the embedding method, that is, the lack of a dynamic feature screening mechanism, resulting in many deep learning models directly using the original sensor data for prediction without effectively screening key features. This may lead to redundant information interfering with the prediction results and reducing the generalization ability of the model.
[0005] (3) Representation aliasing and localization problems: The feature representations learned by existing deep learning methods usually have aliasing phenomena and lack clear separability, thus affecting the clarity and accuracy of decision-making. In addition, the feature representations of each sample are usually independent, ignoring the relevance of the global data structure, resulting in unstable prediction results. Summary of the Invention
[0006] In view of the above problems, the present invention proposes a thin film thickness prediction method and system based on self-supervised and representation learning. The encoder is pre-trained using unlabeled data to reduce the dependence on labeled data, enhance data features and improve the generalization ability of the model. In addition, a dynamic selector is constructed through embedded feature selection to improve the prediction accuracy. Furthermore, a global prototype projection space is constructed to optimize the representation space, reduce representation aliasing and localization problems, and make the prediction results more stable. With the help of the optimal transport mechanism and orthogonality constraints, the prediction reliability and the adaptability of the model to different data distributions are improved respectively.
[0007] The technical solution adopted by the present invention is as follows:
[0008] In a first aspect, the present invention proposes a thin film thickness prediction method based on self-supervised and representation learning, including the following steps:
[0009] (1) Collect the original process data of wafer manufacturing and the thin film thickness measurement data. The data pairs with successfully matched original process data and measurement data are used as the first training set after preprocessing, and the unmatched original process data are used as the second training set after preprocessing;
[0010] (2) Use the second training set to pre-train an encoder, and then use the first training set to fine-tune the encoder and synchronously train a dynamic feature selector and a regression head;
[0011] (3) Use the dynamic feature selector to reduce the dimension of the process data in the first training set, and use the fine-tuned encoder to initialize a global prototype projection space for the dimension-reduced process data;
[0012] (4) Use the latest updated encoder to generate the embedding representation of the dimension-reduced process data, calculate the projection coordinates of the embedding representation in the global prototype projection space through the feature representation head, calculate the projection loss based on the projection distribution, calculate the regression loss based on the regression head, and at the same time introduce the diversity loss and the global prototype orthogonality loss. The joint total loss is used to iteratively update the encoder, the regression head, the feature representation head, and synchronously update the global prototype projection space;
[0013] (5) Take the actual wafer manufacturing process data as the input, and use the dynamic feature selector, encoder and regression head to predict the thin film thickness.
[0014] Further, the preprocessing refers to the preliminary screening of the features in the process data. The screened features include deposition time, buffer pressure, wafer processing temperature, lateral and top silane gas reaction time, and top, bias and lateral radio frequency reflection power.
[0015] Further, when pre-training an encoder using the second training set, the K-Means clustering method is adopted to generate pseudo-measurement data of the second training set as pseudo-labels for training, and at the same time, the process data in the second training set is perturbed.
[0016] Further, the fine-tuning of the encoder using the first training set and the synchronous training of a dynamic feature selector and a regression head include:
[0017] Using the dynamic feature selector to screen the features of the process data in the first training set to obtain the process data after dimensionality reduction processing;
[0018] Using the pre-trained feature encoder to generate the embedded representation of the process data after dimensionality reduction processing;
[0019] Using the regression head to predict the film thickness according to the embedded representation;
[0020] Calculating the loss based on the predicted film thickness and the actual measurement data, and updating the parameters of the encoder, dynamic feature selector, and regression head.
[0021] Further, the dynamic feature selector is specifically:
[0022] Introduce a random gating variable , and the gating value is defined as:
[0023] ;
[0024] where is the j-th trainable parameter in the dynamic feature selector, j = 1, 2, …, d; d represents the number of features of the samples in the training set; is the independent sampling parameter of the j-th trainable parameter, and this parameter follows a Gaussian distribution; feature screening is achieved by calculating the Hadamard product of the process data and the random gating variable , and a regularization term is introduced in the training process of the dynamic feature selector.
[0025] Further, the initialization process of the global prototype projection space is:
[0026] Using the fine-tuned encoder to generate the embedded representation of the process data after dimensionality reduction processing;
[0027] Clustering the embedded representations, and regarding each clustering center as a prototype to obtain a global prototype projection space containing several prototypes.
[0028] Further, the projection coordinates of the embedded representation in the global prototype projection space are calculated through the feature representation head, which is expressed as:
[0029] ;
[0030] Embedded representation The projected coordinates of and the global prototype projection space satisfy the following relationship:
[0031] ;
[0032] wherein, represents the number of prototypes in the global prototype projection space, represents the k-th prototype in the global prototype projection space, represents the feature representation head.
[0033] Furthermore, the process of calculating the diversity loss is as follows:
[0034] Divide the range of the measurement data in the first training set into several sub-intervals, regard the process data in the same sub-interval as a pair of positive samples, regard the process data in different sub-intervals as a pair of negative samples, and calculate the contrastive learning loss of the projected coordinates according to the positive samples and negative samples.
[0035] Furthermore, the projection loss refers to the optimal transport distance loss between the distribution of the projected coordinates and the distribution of the embedded representation, and the global prototype orthogonality loss refers to that the global prototype projection space satisfies the orthogonality constraint.
[0036] In a second aspect, the present invention proposes a thin film thickness prediction system based on self-supervised and representation learning for implementing the above-mentioned thin film thickness prediction method based on self-supervised and representation learning.
[0037] The beneficial effects of the present invention are as follows:
[0038] The present invention first pre-trains the encoder using the wafer manufacturing process data without labels, which can not only reduce the dependence on a large amount of labeled data, but also enhance the data features, thereby improving the generalization ability of the model; adopts an embedded feature selection method to construct a dynamic feature selector, enabling the model to automatically screen the most representative wafer manufacturing process features, reducing the interference of redundant information, and effectively improving the prediction accuracy; in the optimization of the representation space, by constructing a global prototype projection space, it improves the separability of the embedded representation of the wafer manufacturing process data, reduces the problem of representation aliasing, and makes the thin film thickness prediction result more stable; with the help of the optimal transport mechanism to optimize the optimal transport loss, it prompts the model to more effectively learn the global structure information of the data, enhancing the reliability of the prediction; uses the orthogonality constraint method to achieve orthogonality constraint to enhance robustness, making the global prototypes more independent and enhancing the adaptability of the model to different data distributions. Description of the Drawings
[0039] Figure 1 is a flowchart of the thin film thickness prediction method based on self-supervised and representation learning;
[0040] Figure 2 It is a framework diagram of a thin film thickness prediction method based on self-supervised and representation learning;
[0041] Figure 3 It is a schematic diagram of the self-supervised pre-training stage;
[0042] Figure 4 It is a schematic diagram of the fine-tuning training stage;
[0043] Figure 5 It is a schematic diagram of the representation learning stage. Specific implementation manner
[0044] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention.
[0045] The accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0046] The flowcharts shown in the accompanying drawings are only exemplary illustrations and do not necessarily include all steps. For example, some steps can be decomposed, while some steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0047] A thin film thickness prediction method based on self-supervised and representation learning proposed by the present invention is named SGT_SRL (Self-Supervised and Representation-Aware Transformer). The overall framework structure is as Figure 2 shown, which combines the self-supervised pre-training, fine-tuning training stage and representation learning stage, and can effectively use the CVD process data to predict the thin film thickness. The training processes of the three stages will be described in detail below.
[0048] As Figure 1 shown, the implementation process of the present invention mainly includes the following steps:
[0049] S1. Obtain a first training set and a second training set for model training.
[0050] Collect real-time data of the fault detection and classification system (abbreviated as FDC system) in the wafer manufacturing process, including process data and measurement data of key semiconductor manufacturing steps. Regard the process data as process variables and the measurement data as target variables.
[0051] In a specific implementation, parameters such as the lateral / top silane gas flow rate, top / bias / lateral radio frequency reflection power, deposition time, and frontline pressure during the CVD process were collected as process data, and the final film thickness was used as measurement data. The pairs of data where the original process data and the measurement data were successfully matched were preprocessed and used as the first training set, and the original process data that was not successfully matched was preprocessed and used as the second training set. The two parts of the data set were used for subsequent SGT_SRL training. The SGT_SRL training process included 3 stages, namely the self-supervised pre-training stage, the fine-tuning training stage, and the representation learning stage. The model of the present invention was defined as
[0052] with parameters and and where consisted of a dynamic feature selector , an encoder , and a regression head .
[0053] S2, self-supervised pre-training stage: Use the second training set to pre-train an encoder .
[0054] As Figure 3 shown, in this stage, a pre-training task needs to be created from the unlabeled second training set. In actual situations, there is a high correlation between some input process data features and the original measurement data. For example, the original task is to use features TOP_MEAN_1, SID_MEAN_2, and SID_MEAN_3 to predict the film thickness BTH. And there is a high correlation between the deposition time TIME_1 and the film thickness BTH. In this way, a pre-training task can be created by using features TOP_MEAN_1, SID_MEAN_2, and SID_MEAN_3 to predict the deposition time TIME_1. To improve the diversity and rationality of the pseudo-labels, in this embodiment, the top 10 features with the highest correlation will be selected first for clustering to generate pseudo-labels.
[0055] First, define a binary mask matrix:
[0056]
[0057] where d is the feature dimension of the input data, which is 81, , indicates that the feature is selected, indicates that the feature is not selected. The features selected through the mask form a new input:
[0058]
[0059] Among them, represents element-wise multiplication, represents removing elements with a mask value of 0. Based on the selected features, K-Means clustering is run to generate task labels:
[0060]
[0061] Among them, represents the i-th original process data that fails to match, represents the binary mask matrix for preprocessing the original process data, represents the features of the preprocessed process data, with the feature dimension being w; represents the cluster center matrix, represents the pseudo-label of the i-th process data in the second training set, which is in the form of a one-hot encoding with a dimension of ks, and , is a vector of all 1s, represents the number of samples in the second training set, represents the square of the L2 norm. is the number of cluster centers, set to 100.
[0062] To prevent the model from directly memorizing the data, a random variable is introduced:
[0063]
[0064]
[0065] Among them, follows a Bernoulli distribution with a parameter of 0.5. The parameter 0.5 refers to the probability that the random variable takes the value 1; the superscript represents the transpose.
[0066] For the positions with a value of 1, the mean of the feature column is used to replace them, and thus the perturbed input is obtained. Finally, a pre-training task dataset is constructed and passed into the model for pre-training:
[0067]
[0068] The process of pre-training the model is as follows: Using the perturbed process data as the input of the encoder, a classification head is used to predict the pseudo-label of this process data, and the encoder and the classification head are updated according to the prediction result and the actual pseudo-label. Here, the encoder can adopt a Transformer network, and the classification head can be implemented by one or more linear layers. The classification head is only used in the training of this stage.
[0069] S3, Fine-tuning Training Stage: Fine-tune the encoder using the first training set and simultaneously train a dynamic feature selector and a regression head.
[0070] This stage is trained using the first training set.
[0071] As Figure 4 shown, at this time, initialize the parameters of the encoder using the encoder pre-trained in the previous stage. To perform feature selection, an embedded method is adopted. Specifically, introduce a random gating variable to achieve feature screening. For each feature perform independent sampling:
[0072]
[0073] where is the fixed noise intensity, and the gating value is defined as follows:
[0074]
[0075] where are the trainable parameters of the feature selector . When , the feature is completely retained, and when , the feature is completely suppressed. After the input process data passes through feature screening, it is passed to the encoder and the regression head to calculate the regression loss:
[0076]
[0077] To enhance feature sparsity, introduce a regularization term:
[0078]
[0079] where is the cumulative distribution function (CDF) of the standard normal distribution, which is used to adjust the weight of the regularization term in the overall loss. This regularization term encourages the of unimportant features, thereby closing the corresponding gates. According to the result of the loss function, update the model parameters and :
[0080]
[0081]
[0082] where and correspond to the encoder and regression head parameters with the feature selector of learning rate, repeat the above steps for iteration to complete the fine-tuning step.
[0083] S4, the representation learning stage.
[0084] As Figure 5 shown, this stage includes the initialization process of the global prototype projection space and the process of iteratively updating the encoder, regression head, and feature representation head.
[0085] First, in this stage, by constructing a Prototype-Based Projection Space (P-Space), the representation of data is optimized, the generalization ability of the model is improved, and the aliasing problem between different data features is reduced. After the training data is fine-tuned in step S3 and , the embedded representation is obtained, where represents the number of a batch of training samples, and the training samples refer to the samples in the first training set. Perform K-Means clustering on to obtain the clustering center matrix , and initialize the global prototype with it:
[0086]
[0087] Among them, is the number of clustering clusters, and the clustering center matrix of each cluster represents a vector of length s, denoted as a prototype; represents the k-th prototype. Define the prototype projection space as the projection space with the global prototype as the basis vector, where the embedded representation of each sample is formalized as , .
[0088] In this embodiment, the embedded representation of the i-th sample is used to calculate the projection coordinates in the prototype projection space through the feature representation head , and the feature representation head can be implemented by one or more linear layers.
[0089] The projection distributions of the embedded representation and the projection coordinates are respectively expressed as:
[0090]
[0091]
[0092] Among them, denotes the Dirac function of denotes the Dirac function of and are the projection distributions of the embedding representation and the projection coordinates, respectively.
[0093] To enable the prototype to capture the global data structure information, the optimal transport (OT) distance is used to calculate the projection loss:
[0094]
[0095] where N is the number of samples.
[0096] Considering that direct optimization may lead to coordinate collapse, constraints are introduced to distribute the projection coordinates to different regions. In this embodiment, the range of the measurement data is divided into t subintervals, where , is the batch size. Samples within the same partition are regarded as positive pairs, and samples across regions are negative pairs. The optimization objective is:
[0097]
[0098] where is the indicator function, which is 1 when belongs to the positive sample pair .
[0099] In addition, the global prototype projection space is composed of global prototypes. To make the prototypes independent, the orthogonality constraint needs to be satisfied:
[0100]
[0101] where , the first term forces M to be sparse, and the second term ensures . Combining the task loss with , and iteratively updating the parameters, the final model for prediction can be obtained.
[0102] It should be noted that the above only gives the calculation formula, while refers to the conventional regression loss, that is, the loss between the predicted thickness of the film and the actual measurement data.
[0103] S5. Using the actual wafer manufacturing process data as input, the dynamic feature selector, encoder, and regression head are used to predict the film thickness.
[0104] Here, the input process data first passes through a dynamic feature selector to screen the features of the process data, obtaining the process data after dimensionality reduction processing; then, a feature encoder is used to generate the embedding representation of the process data after dimensionality reduction processing; finally, a regression head is used to predict the film thickness based on the embedding representation.
[0105] To quantitatively evaluate the performance of the SGT_SRL model, the following experimental verification is carried out.
[0106] Baseline and evaluation metric selection: Representative methods in machine learning, deep learning, and self-supervised learning methods are selected for comparative evaluation, where the root mean square error (RMSE) and coefficient of determination (R 2 ) are used as evaluation metrics.
[0107] Deep learning methods include FT-Transforme, GNN, NODE, MDN, AutoInt, GANDALF, MLP, ResNet, TabDPT, STG, and TabTransformer. They have respectively achieved good performance on public datasets of tabular data, by adding self-attention mechanisms and stacking multiple layers of networks to achieve excellent non-linear modeling capabilities.
[0108] Machine learning methods include LightGBM, XGBoost, Random Forest, CatBoost, Extra-Trees, GPR, PLS, SVR, Elastic Net, and Linear Regression (LR). In traditional virtual metrology applications, these methods are often used, and they have the characteristics of less training time and relatively stable prediction performance.
[0109] Self-supervised learning methods include TabNet, VIME, SAINT, SCARF. They enhance the generalization ability of the model's feature representation by generating pseudo-labels on unlabeled data to generate pre-training tasks.
[0110] Data collection and preprocessing: 12-inch wafer production data is collected, including 81 feature variables such as reaction gas flow rate, RF power, temperature, and reaction chamber pressure. The data is processed for missing values and normalized. It is divided into four datasets CVD-1, CVD-2, CVD-3, and CVD-4 according to different machines, and the training set, validation set, and test set are divided and passed into the comparison models for evaluation respectively. The experimental results are shown in Table 1 and Table 2.
[0111] Table 1:
[0112]
[0113] Table 2:
[0114]
[0115] As can be seen from the comparison between Table 1 and Table 2, SGT_SRL achieved RMSE values of 18.796, 26.908, 19.710, and 21.730, as well as R 2 values of 0.769, 0.590, 0.649, and 0.640 on four different datasets. Except for slightly inferior performance on the CVD2 dataset, SGT_SRL outperformed all comparison methods in all other cases.
[0116] An ablation experiment was conducted on the present invention, and the results are shown in Table 3 and Table 4.
[0117] Table 3:
[0118]
[0119] Table 4:
[0120]
[0121] It can be found from Table 3 and Table 4 that the complete model (FTT_SG(SL, RL)) using SL+RL+feature selection performs best on all datasets. Among them, the representation learning RL is the key factor in improving the model performance, and the self-supervised learning SL contributes greatly on some datasets. Feature selection can improve the model performance, but it needs to be combined with RL and SL to maximize the effect.
[0122] The present invention also provides a thin film thickness prediction system based on self-supervised and representation learning, which is used to implement the above embodiments. Terms such as "module" and "unit" used hereinafter can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible.
[0123] A thin film thickness prediction system based on self-supervised and representation learning provided in this embodiment includes:
[0124] A training data preprocessing module, which is used to collect the original process data of wafer manufacturing and the thin film thickness measurement data, and use the data pairs with successfully matched original process data and measurement data as the first training set after preprocessing, and use the unmatched original process data after preprocessing as the second training set;
[0125] A pre-training module, which is used to pre-train an encoder using the second training set;
[0126] A fine-tuning module for fine-tuning the encoder obtained by the pre-trained module using the first training set and synchronously training a dynamic feature selector and a regression head;
[0127] A global prototype projection space initialization module for reducing the dimensionality of the process data in the first training set using the dynamic feature selector and initializing a global prototype projection space for the dimension-reduced process data using the fine-tuned encoder;
[0128] A representation learning module for generating an embedded representation of the dimension-reduced process data using the latest updated encoder, calculating the projection coordinates of the embedded representation in the global prototype projection space through a feature representation head, calculating a projection loss based on the optimal transport mechanism of the projection distribution, calculating a regression loss based on the regression head, introducing a diversity loss and a global prototype orthogonality loss at the same time, iteratively updating the encoder, the regression head, and the feature representation head for the combined total loss, and synchronously updating the global prototype projection space;
[0129] A film thickness prediction module for predicting the film thickness using the dynamic feature selector, the encoder, and the regression head with the actual wafer manufacturing process data as the input.
[0130] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment, and the implementation methods of the remaining modules will not be elaborated here. The system embodiments described above are merely illustrative, where the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0131] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The system embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking the software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation.
[0132] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. For those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A film thickness prediction method based on self-supervision and representation learning, characterized in that: The following steps are involved: (1) Collecting wafer manufacturing raw process data and film thickness measurement data, pre-processing the data pairs that successfully match the raw process data and the measurement data as the first training set, and pre-processing the raw process data that fail to match as the second training set; (2) Use the second training set to pre-train an encoder, then use the first training set to fine-tune the encoder and simultaneously train a dynamic feature selector and a regression head; (3) Using a dynamic feature selector to reduce the dimension of the process data in the first training set, and using the fine-tuned encoder to initialize a global prototype projection space for the process data after dimension reduction; (4) Use the latest updated encoder to generate an embedded representation of the process data after dimensionality reduction. Use the feature representation head to calculate the projection coordinates of the embedded representation in the global prototype projection space. Calculate the projection loss based on the optimal transmission mechanism of the projection distribution. Calculate the regression loss based on the regression head. At the same time, introduce diversity loss and global prototype orthogonal loss. Combine the total loss to iteratively update the encoder, regression head, and feature representation head, and simultaneously update the global prototype projection space. (5) Using actual wafer manufacturing process data as input, the film thickness is predicted using a dynamic feature selector, encoder, and regression head.
2. The thin film thickness prediction method based on self-supervision and representation learning according to claim 1, characterized in that: The preprocessing is to perform preliminary screening on the features in the process data, and the screened features include deposition time, buffer pressure, wafer processing temperature, lateral and top silane gas reaction time, and top, bias and lateral RF reflection power.
3. The thin film thickness prediction method based on self-supervision and representation learning according to claim 1, characterized in that: When the second training set is used to pre-train an encoder, a K-Means clustering method is used to generate pseudo measurement data of the second training set as pseudo labels for training, and the process data in the second training set is disturbed at the same time; The optimization goal of generating pseudo labels is: ; in, Indicates the i-th original process data that has not been successfully matched. represents the binary mask matrix used to preprocess the raw process data, Represents the features of the preprocessed process data, with a feature dimension of w; represents the cluster center matrix, represents the pseudo label of the i-th process data in the second training set, in the form of a one-hot encoding with dimension ks, and , is a vector of all 1s, represents the number of samples in the second training set, Represents the square of the L2 norm.
4. The thin film thickness prediction method based on self-supervision and representation learning according to claim 1, characterized in that: The method of reusing the first training set to fine-tune the encoder and synchronously training a dynamic feature selector and a regression head includes: Using a dynamic feature selector to filter the features of the process data of the first training set, and obtaining the process data after dimensionality reduction processing; Use the pre-trained feature encoder to generate an embedding representation of the process data after dimensionality reduction; Predicting film thickness based on the embedded representation using a regression head; The loss is calculated based on the predicted film thickness and the actual measurement data, and the encoder, dynamic feature selector and regression head parameters are updated.
5. The thin film thickness prediction method based on self-supervision and representation learning according to claim 4 is characterized in that: The dynamic feature selector is specifically: Introducing random gating variables , gate value Defined as: ; in, is the jth trainable parameter in the dynamic feature selector, j=1,2,…d; d represents the number of features of the samples in the training set; is the independent sampling parameter of the jth trainable parameter, which follows a Gaussian distribution; by calculating the process data and the random gating variable The Hadamard product is used to realize feature screening, and the regularization term is introduced into the training process of the dynamic feature selector.
6. The thin film thickness prediction method based on self-supervision and representation learning according to claim 1, characterized in that: The initialization process of the global prototype projection space is: Generate an embedding representation of the reduced-dimensionality process data using the fine-tuned encoder. The embedded representations are clustered, and each cluster center is regarded as a prototype, resulting in a global prototype projection space containing several prototypes.
7. The thin film thickness prediction method based on self-supervision and representation learning according to claim 6, characterized in that: The projection coordinates of the embedding representation in the global prototype projection space calculated by the feature representation head are expressed as: ; Embedding Representation The projection coordinates of The following relationship is satisfied with the global prototype projection space: ; in, represents the number of prototypes in the global prototype projection space, represents the kth prototype in the global prototype projection space, Indicates the feature of the head.
8. The thin film thickness prediction method based on self-supervision and representation learning according to claim 1 or 7, characterized in that: The diversity loss calculation process is as follows: The range of the measured data in the first training set is divided into several sub-intervals, the process data in the same sub-interval is regarded as a pair of positive samples, and the process data in different sub-intervals is regarded as a pair of negative samples, and the contrastive learning loss of the projection coordinates is calculated based on the positive samples and the negative samples.
9. The thin film thickness prediction method based on self-supervision and representation learning according to claim 1 or 7, characterized in that: The projection loss refers to the optimal transmission distance loss between the distribution of projection coordinates and the distribution of embedded representation, and the global prototype orthogonal loss refers to the global prototype projection space satisfying the orthogonality constraint.
10. A film thickness prediction system based on self-supervision and representation learning, used to implement the film thickness prediction method based on self-supervision and representation learning as claimed in claim 1, characterized in that: The system comprises: A training data preprocessing module, which is used to collect wafer manufacturing raw process data and film thickness measurement data, preprocess the data pairs that successfully match the raw process data and the measurement data as the first training set, and preprocess the raw process data that do not successfully match as the second training set; A pre-training module, configured to pre-train an encoder using a second training set; A fine-tuning module, which is used to fine-tune the encoder obtained by the pre-training module using the first training set and synchronously train a dynamic feature selector and a regression head; A global prototype projection space initialization module is used to reduce the dimension of the process data in the first training set by using a dynamic feature selector, and initialize a global prototype projection space for the process data after the dimension reduction by using a fine-tuned encoder; The representation learning module is used to generate an embedded representation of the process data after dimensionality reduction using the latest updated encoder, calculate the projection coordinates of the embedded representation in the global prototype projection space through the feature representation head, calculate the projection loss based on the optimal transmission mechanism of the projection distribution, calculate the regression loss based on the regression head, and introduce the diversity loss and the global prototype orthogonal loss at the same time. The encoder, regression head, and feature representation head are iteratively updated with the combined total loss, and the global prototype projection space is updated synchronously; The film thickness prediction module is used to predict the film thickness using the actual wafer manufacturing process data as input and using a dynamic feature selector, an encoder and a regression head.
Citation Information
Patent Citations
Self-supervision enhanced semi-supervised wafer failure prediction method and system
CN119557571A
System and method for user recognition using motion sensor data
US20240094828A1
Unsupervised domain adaptation of models with pseudo-label curation
US20240312197A1
Cited By
Weak supervision video anomaly detection method and system based on prototype orthogonality
CN121640198A