A semi-supervised learning method for label delay wafer yield prediction

CN122840152APending Publication Date: 2026-09-29ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611242846.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,在实际生产中,完整标注样本的获取受到测试周期和生产节拍限制,尤其在新产品、小批量或早期量产阶段,可用于训练的标注数据数量有限

Benefits of technology

[0044]第一,针对传统晶圆良率预测方法主要依赖已完成WAT且已返回CP良率标签的样本进行监督训练,导致在CP标签存在返回延迟时,大量已产生但尚未标注的WAT样本无法被有效利用的问题,本发明将未标注WAT样本纳入模型训练过程,通过联合利用已标注样本和标签延迟样本中包含的生产过程信息,提高了标签不完整条件下的数据利用率,使模型能够在CP标签尚未全部返回时完成训练或更新,从而缩短良率预测模型的等待周期,提高模型对实际生产节拍的适应能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840152A_ABST
    Figure CN122840152A_ABST
Patent Text Reader

Abstract

The application discloses a semi-supervised learning method for label-delay wafer yield prediction, and belongs to the field of semiconductor manufacturing quality prediction. According to wafer acceptance test data, an annotated data set and an unannotated data set are constructed; a prediction model containing an embedding layer, an encoder and a linear prediction head is constructed; for unannotated samples, weak and strong feature disturbances are applied to the encoder respectively, the pseudo-target mean value and uncertainty are obtained through multiple predictions under weak disturbance, and a confidence mask is generated; under strong disturbance, a single prediction is performed, the mask and the pseudo-target are combined to calculate a filtering consistency loss, the loss is further combined with a supervised regression loss calculated from annotated samples to obtain a total training loss containing a dynamic consistency weight, and the prediction model is trained. The method can utilize label-delay unannotated data without directly disturbing original wafer acceptance test measurement variables, and improve the stability of circuit probe yield prediction and yield risk monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of semiconductor manufacturing quality prediction and industrial artificial intelligence technology, specifically relating to a semi-supervised learning method for predicting the yield of labeled delay wafers. Background Technology

[0002] In the semiconductor manufacturing process, wafer acceptance test (WAT) data is usually available after the front-end manufacturing is completed, while yield tags obtained from circuit probe (CP) testing are only returned after subsequent testing stages. Therefore, during new product introduction or early production stages, the manufacturing system often accumulates a large amount of wafer acceptance test data, but the corresponding circuit probe yield tags are not yet complete, forming a typical tag-delay quality prediction scenario.

[0003] Existing wafer yield prediction methods mostly employ supervised learning, using paired wafer acceptance test data and circuit probe yield labels to train the prediction model. These methods can establish a mapping relationship between test parameters and yield when sufficient labeled data is available, enabling early assessment of wafer quality risks. However, in actual production, obtaining complete labeled samples is limited by testing cycles and production pace, especially in new products, small batches, or early mass production stages, where the amount of labeled data available for training is limited.

[0004] Furthermore, while general semi-supervised learning methods can utilize unlabeled samples, they typically construct enhanced samples through random masking, feature replacement, or input perturbation. For wafer acceptance test data, each feature usually corresponds to specific electrical test items or process-related measurement parameters. Directly perturbing the original input variables may change their physical meaning, leading to unreliable pseudo-labels or consistency constraints.

[0005] Therefore, a semi-supervised yield prediction scheme is needed that can utilize unlabeled WAT data while minimizing disruption of the original measurement semantics. Summary of the Invention

[0006] To address the technical challenge of training a yield prediction model using unlabeled WAT data in semiconductor manufacturing scenarios where CP yield labels are delayed and labeled WAT-CP samples are limited, while avoiding the disruption of physical measurement semantics by perturbations in the original WAT input, this invention provides a semi-supervised learning method for label-delayed wafer yield prediction.

[0007] This method acquires labeled and unlabeled data, preprocesses and embeds them, and then inputs them into a stacked adaptive learnable attention block encoder. It calculates the supervised regression loss for labeled samples, applies weak feature perturbations to unlabeled samples at early representation positions to estimate the pseudo-target mean and prediction uncertainty, constructs a confidence mask based on the prediction uncertainty, applies strong feature perturbations at later representation positions, and calculates the filtering consistency loss based on high-confidence pseudo-targets, and finally jointly optimizes the supervised loss and consistency loss by increasing the consistency weight.

[0008] The specific technical solution is as follows:

[0009] S1. Obtain test data from multiple wafers in the wafer acceptance test. Wafers that have obtained yield labels are used as labeled datasets, and wafers that have not obtained yield labels are used as unlabeled datasets.

[0010] S2, construct a prediction model including an embedding layer, an encoder, and a linear prediction head. The embedding layer is used to map the input test data to an initial embedding representation. The encoder is formed by stacking multiple adaptive learnable attention blocks and is used to output a final hierarchical representation based on the initial embedding representation. The linear prediction head is used to output a yield prediction value based on the final hierarchical representation.

[0011] S3, apply weak feature perturbation and strong feature perturbation to the encoder of the prediction model respectively, use the prediction model with weak feature perturbation to make multiple predictions on the unlabeled dataset, calculate the pseudo target mean and prediction uncertainty, and further construct a confidence mask;

[0012] The prediction model with strong feature perturbation is used to make a prediction on the unlabeled dataset, and the filtering consistency loss is calculated by combining the pseudo-target mean and confidence mask.

[0013] S4. Use the prediction model to predict the labeled dataset, calculate the supervised regression loss, and calculate the total training loss including dynamic consistency weights together with the filter consistency loss. Then train the prediction model to obtain the trained prediction model.

[0014] S5. Using the trained prediction model, obtain the predicted yield value of the wafer to be inspected.

[0015] Furthermore, in S2, the adaptive learnable attention block comprises a normalization layer, a multi-head self-attention layer, a gated interactive selection layer, a linear projection layer, and an attention-based learnable residual aggregation layer connected in sequence.

[0016] Furthermore, in the encoder, the first... The process of outputting hierarchical representations for each adaptive learnable attention block is as follows:

[0017] S201, receiving the first The hierarchical representation of the output of each adaptive learnable attention block is normalized and multi-head self-attention operations are performed sequentially in the normalization layer and the multi-head self-attention layer to obtain the attention representation.

[0018] S202, in the gated interaction selection layer, the gate factor is obtained based on the hierarchical representation output by the previous adaptive learnable attention block; the gate factor and the attention representation are combined to obtain the gated feature representation, which is then passed through the linear projection layer to obtain the transformation representation;

[0019] S203, the initial embedding representation and the previous The transformation representation of each adaptive learnable attention block is input to the attention-based learnable residual aggregation layer, and weighted summation is performed by combining the residual aggregation weights to obtain the hierarchical representation.

[0020] Furthermore, in S203, the residual aggregation weights are learnable parameters, specifically:

[0021] Initialize a residual pair consisting of a learnable query vector and a residual candidate representation for the initial embedding representation and all transformed representations respectively; during training, use the initial embedding representation and all transformed representations obtained in the previous training round as the residual candidate representations used in the current training round.

[0022] The residual aggregation weight refers to the ratio of the attention score of the initial embedding representation and the attention score of all transformed representations to the total attention score;

[0023] For each residual pair, the root mean square normalized result of the residual candidate representation is multiplied by the transpose of the learnable query vector, and then an exponential operation is performed to obtain the attention score of the initial embedded representation or transformed representation corresponding to the residual pair; the total attention score is the sum of the attention score of the initial embedded representation and the attention scores of all transformed representations.

[0024] Furthermore, in step S3, the process of performing multiple predictions using a prediction model with applied weak feature perturbations and obtaining the pseudo-target mean, prediction uncertainty, and confidence mask is as follows:

[0025] S301, the test data of the unlabeled dataset is passed through the embedding layer of the prediction model to obtain the initial embedding representation;

[0026] S302, input the initial embedded representation into the encoder, add the hierarchical representation obtained from the first adaptive learnable attention block to the weak feature perturbation to obtain the hierarchical representation after weak perturbation;

[0027] S303, input the hierarchical representation after weak perturbation into the next adaptive learnable attention block, until all adaptive learnable attention blocks in the encoder are traversed, input the hierarchical representation output by the last adaptive learnable attention block into the linear prediction head to obtain the yield prediction value.

[0028] S304, repeat S302-S303 until the specified number of perturbations is reached;

[0029] S305, taking each wafer in the unlabeled dataset as a unit, averages the predicted yield values ​​of the wafer under all weak feature perturbations to obtain the pseudo-target mean, and further calculates the prediction uncertainty of each wafer based on the predicted yield value and the pseudo-target mean.

[0030] S306. Based on the prediction uncertainty, a confidence mask is constructed. If the prediction uncertainty is less than or equal to the threshold, the confidence mask corresponding to the wafer in the unlabeled dataset is set to 1, otherwise it is set to 0.

[0031] Furthermore, in S302, the weak feature perturbation adopts an additive Gaussian perturbation, and the weak perturbation noise corresponding to the weak feature perturbation is sampled from a Gaussian distribution with a mean of 0 and a covariance equal to the square of the weak perturbation noise scale multiplied by the identity matrix.

[0032] Furthermore, in S3, the process of performing a prediction using a prediction model with applied strong feature perturbation and calculating the filtering consistency loss is as follows:

[0033] S311, the test data of the unlabeled dataset is passed through the embedding layer of the prediction model to obtain the initial embedding representation;

[0034] S312, the hierarchical representation output from the last adaptive learnable attention block is added to the initial embedded representation input to the encoder, and the strong feature perturbation is added to obtain the hierarchical representation after strong perturbation.

[0035] S313, input the hierarchical representation after strong perturbation into the linear prediction head to obtain the yield prediction value;

[0036] S314: Using the pseudo-target mean as the true value, calculate the mean square error between the predicted yield value and the true value; combine the mean square error and the confidence mask to calculate the filtering consistency loss.

[0037] Furthermore, in S314, the formula for calculating the consistency loss is as follows:

[0038] ;

[0039] in, To filter out consistency loss; This represents the total number of unlabeled samples in a training batch, i.e., the number of wafers in a training batch that did not obtain a yield label. Let be the confidence mask for the j-th unlabeled sample; This represents the predicted yield value of the j-th unlabeled sample after strong feature perturbation. Let be the pseudo-target mean of the j-th unlabeled sample; To prevent extremely small constants with a denominator of zero, Let be the mean square error function.

[0040] Furthermore, in S312, the strong feature perturbation adopts an additive Gaussian perturbation, and the strong perturbation noise corresponding to the strong feature perturbation is sampled from a Gaussian distribution with a mean of 0 and a covariance equal to the square of the strong perturbation noise scale multiplied by the identity matrix.

[0041] Furthermore, in S5, the total training loss is calculated by multiplying the filtering consistency loss by the consistency weight and then adding the supervised regression loss;

[0042] The consistency weight is calculated as follows: First, divide the index value of the current training round by the preset consistency weight increment, compare the calculation result with 1, and multiply the smaller value by the preset maximum consistency weight to obtain the consistency weight.

[0043] The beneficial effects of this invention are:

[0044] First, addressing the issue that traditional wafer yield prediction methods primarily rely on samples that have completed WAT (Wafer Attainment) and returned CP (Cumulative Production) yield labels for supervised training, resulting in a large number of generated but unlabeled WAT samples being unable to be effectively utilized when there is a delay in the return of CP labels, this invention incorporates unlabeled WAT samples into the model training process. By jointly utilizing the production process information contained in labeled samples and label-delayed samples, it improves the data utilization rate under incomplete label conditions, enabling the model to complete training or updates before all CP labels have been returned, thereby shortening the waiting period of the yield prediction model and improving the model's adaptability to actual production cycles.

[0045] Second, addressing the problem that traditional semi-supervised methods often construct enhanced samples in the original WAT input space using random masking or random replacement, which can easily alter the actual numerical relationships of electrical measurement variables and destroy their corresponding physical meanings, this invention constructs weak and strong perturbations in the latent representation space generated by the encoder. Since the latent representation is a high-level abstraction of the original WAT measurement information, this invention can form differentiated training views while maintaining the basic physical semantics of the original electrical measurement data. This reduces the risk of directly modifying the original measurement variables and causing non-physical samples or abnormal data distributions, thus improving the stability and effectiveness of the consistency learning process.

[0046] Third, traditional semi-supervised learning schemes typically use single prediction results directly as the supervision target for unlabeled samples, which can easily lead to unreliable predictions from the model being continuously fed back into the training process, causing error accumulation. This invention obtains multiple sets of prediction results based on multiple weakly perturbated views of the same unlabeled sample and estimates the prediction uncertainty based on the dispersion of these multiple sets of prediction results. On this basis, only unlabeled samples with stable predictions and confidence levels meeting preset conditions are selected to participate in consistency constraints, thereby suppressing the interference of high-uncertainty samples and erroneous false targets on model parameter updates, reducing confirmation bias and error propagation risks, and improving the robustness of semi-supervised training and the reliability of yield prediction results.

[0047] Fourth, addressing the shortcomings of traditional multilayer perceptrons, fixed residual connections, or simple feature splicing methods in fully characterizing the complex nonlinear relationships between high-dimensional WAT variables, and the tendency for effective low-level measurement information to decay with increasing network layers, this invention constructs an adaptive learnable attention encoder. This encoder adaptively filters different features and their interactions through a gating interaction selection mechanism, and assigns corresponding weights to the outputs of different encoding layers through an attention-based learnable residual aggregation mechanism. Thus, it can dynamically retain yield-related measurement information based on specific samples, suppress redundant features and ineffective interaction information, enhance the nonlinear modeling capability of high-dimensional WAT features, reduce the loss of useful information during deep feature propagation, and improve the model's feature representation capability and prediction accuracy.

[0048] Fifth, traditional wafer yield prediction models typically require sufficient labeled CP data before retraining, resulting in significant lag in model updates when new batches, new process stages, or changes in production conditions occur. This invention addresses this by employing latent space weak-strong perturbations, uncertainty screening, and joint optimization with labeled and unlabeled samples. This allows the model to dynamically adjust the consistency learning contributions of labeled and unlabeled samples, reducing reliance on immediate CP labels while improving the model's adaptability to changes in production data distribution. Consequently, it provides more timely and stable technical support for online yield prediction, early risk identification, and production decision-making in the wafer manufacturing process. Attached Figure Description

[0049] Figure 1 This is an overall framework diagram of the method of the present invention.

[0050] Figure 2 This is a flowchart illustrating the method of the present invention.

[0051] Figure 3 This is a schematic diagram of the adaptive learnable attention block of the present invention. Detailed Implementation

[0052] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in the various embodiments of the present invention can be combined accordingly without mutual conflict.

[0053] The overall process of the semi-supervised learning method for predicting the yield of tag-delayed wafers provided by this invention is as follows: Figure 1 and Figure 2 As shown, this method uses wafer acceptance test data from the semiconductor manufacturing process as input, and combines labeled samples with obtained circuit probe yield labels and unlabeled samples without obtained circuit probe yield labels to train a prediction model that outputs wafer yield predictions or yield risk monitoring results before the circuit probe results are completed.

[0054] The specific steps are as follows:

[0055] Step 1: Obtain wafer acceptance test data and construct labeled and unlabeled datasets.

[0056] This step is used to obtain test data from wafer acceptance testing during the semiconductor manufacturing process and to build a dataset.

[0057] For wafers that have completed circuit probe testing and obtained yield labels during wafer acceptance testing, the voltage, current, capacitance, and inductance test data measured during the wafer acceptance testing are used to form the wafer acceptance testing feature vector, which is then combined with the corresponding circuit probe yield labels to form a labeled dataset. , This represents the wafer acceptance test feature vector of the i-th labeled wafer sample. The label indicates the yield of the circuit probe corresponding to the sample, including qualified and unqualified labels; This indicates the total number of labeled wafers.

[0058] For wafers that have not yet returned circuit probe yield labels during wafer acceptance testing, the voltage, current, capacitance, and inductance test data measured during the wafer acceptance test are used to construct the wafer acceptance test feature vector, which is then used to form an unlabeled dataset. , This represents the wafer acceptance test feature vector of the j-th unlabeled wafer sample. This indicates that the number of samples was not labeled.

[0059] The goal of the prediction model constructed in this invention is to utilize and Training prediction model This enables the output circuit probe yield prediction or yield risk monitoring results based on the wafer acceptance test feature vector. The prediction model constructed in this invention includes an embedding layer, an adaptive learnable attention encoder, and a linear prediction head, which is composed of multiple adaptive learnable attention blocks stacked together.

[0060] Step 2: Preprocess and embed the feature vectors of wafer acceptance testing.

[0061] This step is used to perform missing value imputation, categorical variable encoding, and numerical standardization on the wafer acceptance test feature vector.

[0062] Missing numerical features can be filled in using the median of the labeled dataset; categorical variables (such as test item grouping, wafer batch number, etc.) can be encoded using ordinal or one-hot encoding based on the cardinality of the variable; and numerical features can be standardized using the mean and standard deviation obtained from the labeled dataset.

[0063] Subsequently, the preprocessed test feature vectors are passed through the embedding layer of the prediction model and mapped to the initial embedding representation. This serves as the input for the subsequent adaptive learnable attention encoder.

[0064] Step 3: Extract hierarchical wafer acceptance test characterization using an adaptive learnable attention encoder.

[0065] This step involves the initial embedding characterization of all wafers. The input consists of an adaptive learnable attention encoder formed by stacking multiple adaptive learnable attention blocks, which is used to obtain the wafer acceptance test characterization.

[0066] Adaptive learnable attention blocks are used to model nonlinear interactions between high-dimensional wafer acceptance test variables and retain useful original measurement information, with a structure such as... Figure 3 As shown, each adaptive learnable attention block includes a normalization layer, a multi-head self-attention layer, a gated interaction selection layer, a linear projection layer, and an attention-based learnable residual aggregation layer. The output of the previous adaptive learnable attention block is used as input, and the result is used as input for the next adaptive learnable attention block. The input of the first adaptive learnable attention block is the initial embedding representation. .

[0067] No. The specific process for an adaptive, learnable attention block is as follows:

[0068] In the In each adaptive learnable attention block, the representation output from the previous adaptive learnable attention block is first received. After normalization in the normalization layer, multi-head self-attention operation is performed through the multi-head self-attention layer to obtain the attention representation. The calculation method is as follows:

[0069]

[0070] in, For the first Hierarchical representation of the output of an adaptive learnable attention block For the adaptive learnable attention encoder, the first An adaptive, learnable attention block; This is a normalization operation used to stabilize the feature distribution. This is a multi-head self-attention operation used to model the dependencies between different wafer acceptance test features from multiple attention heads; For the first Attention representations obtained from adaptive learnable attention blocks.

[0071] Subsequently, the attention representation is filtered through a gating interaction selection layer, calculated as follows:

[0072]

[0073] in, For the first Feature representation after adaptive learnable attention block gating; This represents element-wise multiplication; This is the Sigmoid activation function, used to generate gate weights with values ​​ranging from 0 to 1; Learnable weight matrix for gating interaction selection layer; Learnable biases for gating interaction selection layers; For use according to the A hierarchical representation of an adaptive learnable attention block is used to generate a gating factor.

[0074] This gating mechanism is used to suppress weakly correlated or noisy feature interactions and improve the stability of feature interaction selection.

[0075] The gated feature representation is input into the linear projection layer to obtain the transformed representation of the current adaptive learnable attention block. The calculation method is as follows:

[0076]

[0077] in, For the first Transformation representation of the output of an adaptive learnable attention block; This is a linear projection operation; For gating representation Perform normalization.

[0078] To avoid losing the original wafer acceptance test measurement semantics during deep feature transformation, this invention further incorporates an attention-based learnable residual aggregation layer into the adaptive learnable attention block. This layer employs an attention-based learnable residual aggregation mechanism to integrate the original embedded representation. and present and past Transformation representation of an adaptive learnable attention block Adaptive aggregation is performed, and the calculation method is as follows:

[0079]

[0080] in, For the first Hierarchical representation of the output of an adaptive learnable attention block; This is the initial embedding representation; For the first Transformation representations generated by adaptive learnable attention blocks; Assigned to the initial embedding representation during aggregation The residual aggregation weights; To be assigned to the Transformation representation The residual aggregation weights; This indicates a summation operation.

[0081] Initialization phase, and Obtained from random initialization, in subsequent iterations, and The calculation process is as follows: The initial embedding representation generated by the embedding layer in the previous iteration and the transformed representation generated by the adaptive learnable attention block are used to calculate the following:

[0082] The weights for each residual aggregation are obtained by normalizing the attention score, and the calculation method is as follows:

[0083]

[0084]

[0085]

[0086] Where L is the total number of layers of adaptive learnable attention blocks stacked in the adaptive learnable attention encoder, that is, the total number of adaptive learnable attention blocks; For root mean square normalization operation, , σ is the mean calculated from all candidate residual representations, and σ is the variance calculated from all candidate residual representations. To calculate the compatibility score; It is an exponential function;

[0087] when hour, The learnable query vector is the initial embedded representation, and the initial value is obtained by random initialization; The residual candidate representation is the initial embedding representation, and its value is the initial embedding representation output by the embedding layer in the previous iteration. ; The residual aggregation weights of the initial embedding representation after Softmax normalization;

[0088] when for hour, For the first The learnable query vector corresponding to the residual aggregation of each adaptive learnable attention block. For the first The residual candidate representation of each adaptive learnable attention block is taken as the transformed representation of the adaptive learnable attention block output in the previous iteration, i.e. ; The th after Softmax normalization Residual aggregate weights of an adaptive learnable attention block.

[0089] Through this mechanism, the encoder can adaptively choose to retain the original measurement information or utilize intermediate predictive related representations.

[0090] Step 4: Calculate the supervised regression loss based on the labeled samples.

[0091] In this step, for the labeled dataset Input the labeled samples into the prediction model After passing through the embedding layer, adaptive learnable attention encoder, and linear prediction head, the output yield prediction value of the circuit probe is obtained, and the supervised regression loss is calculated by comparing it with the true yield label.

[0092]

[0093]

[0094] in, To monitor regression losses; The number of labeled samples in a training batch; Let the mean square error function be used. The predicted yield value of the circuit probe for the i-th labeled sample; This is the yield label for the actual circuit probe corresponding to this sample; For parameters The prediction model.

[0095] Step 5: Perform weak feature perturbation on the unlabeled samples and generate pseudo-targets.

[0096] For unlabeled samples Weak feature perturbations are applied to the early layer representations at the first representation position of the encoder to construct multiple weak views. The first representation position is set after the first adaptive learnable attention block to take advantage of the measurement-level information retained by the earlier layer representations.

[0097] The weak feature perturbation is an additive Gaussian perturbation in the latent characterization space, rather than a perturbation of the original wafer acceptance test variables. The corresponding latent characterization calculation method is as follows:

[0098]

[0099] in, H represents the latent representation after weak perturbation, and H is the latent representation calculated by the adaptive learnable attention block located before the first representation position. This is a preset weak disturbance noise; This indicates that the mean is 0 and the covariance is... Gaussian distribution; The preset weak disturbance noise scale; It is an identity matrix.

[0100] The perturbated latent representation is input into the next adaptive learnable attention block until the final latent representation is obtained, and the yield prediction result is obtained through the linear prediction head.

[0101] Repeat the above process, performing N weak-view random forward predictions on the same unlabeled sample, calculated as follows:

[0102]

[0103] in, is the prediction result of the j-th unlabeled sample in the nth weak view forward propagation; N is the number of random predictions for the weak view. This is the nth weak feature perturbation operation; Let be the wafer acceptance test feature vector of the j-th unlabeled sample.

[0104] Calculate the pseudo-target mean and prediction uncertainty based on the N weak view prediction results:

[0105]

[0106]

[0107] in, Let be the pseudo-target mean of the j-th unlabeled sample; The prediction uncertainty for the j-th unlabeled sample; This is a square root operation. Prediction uncertainty is used to measure the consistency of prediction results from multiple weak view analyses.

[0108] Step 6: Filter high-confidence unlabeled samples based on prediction uncertainty.

[0109] Based on the prediction uncertainty in step five Construct a confidence mask:

[0110]

[0111] in, Let be the confidence mask for the j-th unlabeled sample; This is an indicator function that takes the value 1 if the condition within the parentheses is true, and 0 otherwise. This is a preset uncertainty threshold.

[0112] This setting can effectively achieve the following: Not greater than If so, the corresponding unlabeled sample is identified as a high-confidence unlabeled sample and participates in consistency training; if Greater than This reduces or eliminates the contribution of that sample to the consistency loss.

[0113] Step 7: Apply strong feature perturbation to the high-confidence unlabeled samples and calculate the filtering consistency loss.

[0114] For the same unlabeled sample, a strong feature perturbation is applied to the later-level representation at the second representation position of the encoder to obtain a strong view prediction. The second representation position is set after the last adaptive learnable attention block to perturb the task-relevant representation closer to the prediction head and improve the model's robustness.

[0115] Strong feature perturbation is also an additive Gaussian perturbation in the latent characterization space, rather than perturbing the original wafer acceptance test variables. The corresponding latent characterization calculation method is as follows:

[0116]

[0117] in, This represents the potential characteristics following a strong perturbation. This is strong disturbance noise; This indicates that the mean is 0 and the covariance is... Gaussian distribution; I represents the preset scale for strong disturbance noise; I is the identity matrix. This represents the latent representation of the output of the last adaptive learnable attention block.

[0118] Will By directly inputting the linear prediction head, the strong view prediction result is obtained, expressed by the following formula:

[0119]

[0120] in, The strong view prediction result for the j-th unlabeled sample; This is a strong feature perturbation operation.

[0121] because Less than Weak views are used to generate relatively stable pseudo-targets, while strong views are used to enhance the robustness of the prediction model to potential representational perturbations.

[0122] Define the weak view pseudo-target as And calculate the filtering consistency loss based on the confidence mask:

[0123]

[0124] in, Filtering consistency loss to mitigate uncertainty; The number of unlabeled samples in a training batch; For confidence mask; For strong view prediction results; The target is a pseudo-target obtained from the mean of the weak view prediction; To prevent extremely small constants with a denominator of zero.

[0125] This loss method allows for the use of only high-confidence unlabeled samples with low uncertainty in consistency training.

[0126] Step 8: Jointly optimize supervised regression loss and uncertainty filtering consistency loss.

[0127] Set a dynamic consistency weight that gradually increases with each training round:

[0128]

[0129]

[0130] in, Here, represents the consistency weight for the e-th training round; e is the index corresponding to the current training round. The consistency weight is incremented gradually according to the preset parameters; The maximum consistency weight is predefined. To select the smaller value function; L is the total training loss; To monitor regression losses; Filter out consistency loss for uncertainty.

[0131] In the early stages of training, the model mainly relies on the supervision signals of labeled samples; as training progresses, the consistency learning contribution of high-confidence unlabeled samples is gradually increased, thereby improving the stability of yield prediction in label delay scenarios.

[0132] The total training loss is used to update all parameters in the prediction model. The iteration stops when the preset number of training rounds is reached or the total training loss converges, and the completed prediction model is obtained.

[0133] In one alternative implementation, the encoder may include 1 to 16 stacked adaptive learnable attention blocks, the hidden dimension may be set to 16 to 1024, the number of heads for multi-head self-attention may be set to 1 to 100, the number of weak view random predictions N may be set to 2 to 20, and an uncertainty threshold. It can be set to 0 to 1, which represents the scale of weak feature noise. It can be set to 0 to 1, strong characteristic noise scale It can be set to 0 to 1, the maximum consistency weight. It can be set from 0 to 1, gradually increasing in size. It can be set to a positive integer according to the training rounds.

[0134] Step 9: Deploy the trained prediction model and output the yield prediction results.

[0135] After training, only the shared prediction model is retained during production deployment. For newly generated wafer acceptance test data, the system performs the same preprocessing as during training, inputs the preprocessed wafer acceptance test data into the trained prediction model, performs forward inference, and outputs the circuit probe yield prediction value. The residual aggregation weights are calculated using the initial embedding representation and transformation representation after the last round of training.

[0136] Weak feature perturbation, strong feature perturbation, multi-weak view uncertainty estimation, and confidence mask filtering are only used in the training phase and are not executed in the deployment phase, so they do not add inference branches in the deployment phase.

[0137] To verify the effectiveness of the method of the present invention, further experiments were conducted on five semiconductor manufacturing datasets from different real production lines. The specific experimental setup is as follows:

[0138] All datasets include test data from wafer acceptance testing and corresponding circuit probe yield labels, and label delay scenarios are constructed according to the time order of the circuit probe labels. The earliest 25% of wafer acceptance test-circuit probe samples that obtained labels during the training phase are designated as labeled samples, and the remaining training samples are designated as unlabeled samples whose labels have not yet been returned. The validation set and test set are used only for model selection and final performance evaluation.

[0139] The encoder consists of six adaptive learnable attention blocks, with a hidden dimension of 128, 8 attention heads, and N = 5. It is 0.2. It is 0.01. It is 0.05. It is 0.2. There are 50 training rounds, for a total of 100 training rounds.

[0140] The method of this invention is compared with methods such as Random Forest, XGBoost, CatBoost, TabTransformer, FT-Transformer, PTaRL, TabPFN, VIME-self, SCARF, SubTab, SAINT, CARTE, VIME-semi, and GFTab. Evaluation metrics include the coefficient of determination R0. 2 Mean absolute error (MAE) and root mean square error (RMSE), where MAE and RMSE are expressed in percentage points.

[0141] The experimental results are shown in Table 1. The average experimental results on the five datasets show that the method of the present invention achieves better overall prediction performance under the label delay setting, with an average R0. 2 The mean R² is 0.628, the mean MAE is 1.956, and the mean RMSE is 2.955. Compared to the best-performing supervised learning baseline TabPFN, its mean R² is significantly lower. 2 The R² value increased from 0.499 to 0.628, a relative improvement of 25.9%; the MAE decreased from 2.516 to 1.956, a relative decrease of 22.3%; and the RMSE decreased from 3.423 to 2.955, a relative decrease of 13.7%. Compared to the best-performing semi-supervised learning baseline GFTab, its average R² value decreased significantly. 2 The MAE increased from 0.481 to 0.628, a relative increase of 30.6%; the MAE decreased from 2.620 to 1.956, a relative decrease of 25.3%; and the RMSE decreased from 3.484 to 2.955, a relative decrease of 15.2%.

[0142] Table 1. Average performance comparison under five dataset label delay settings.

[0143]

[0144] The above results demonstrate that the present invention, by performing weak and strong perturbations in the potential characterization space, using the uncertainty of multiple weak views to screen high-confidence unlabeled samples, and performing filtering consistency learning, can effectively utilize unlabeled data with label delays without directly perturbing the original wafer acceptance test measurement variables, thereby improving the stability of circuit probe yield prediction or yield risk monitoring.

[0145] The embodiments described above are merely preferred embodiments of the present invention, and are not intended to limit the invention. The experimental results described above are only used to illustrate the effects that the technical solutions of the present invention can produce, and do not constitute a limitation on the scope of protection of the present invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by equivalent substitution or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A semi-supervised learning method for predicting the yield of tag-delayed wafers, characterized in that, include: S1. Obtain test data from multiple wafers in the wafer acceptance test. Wafers that have obtained yield labels are used as labeled datasets, and wafers that have not obtained yield labels are used as unlabeled datasets. S2, Construct a prediction model that includes an embedding layer, an encoder, and a linear prediction head. The embedding layer is used to map the input test data to an initial embedding representation. The encoder is formed by stacking multiple adaptive learnable attention blocks and is used to output a final hierarchical representation based on the initial embedding representation. The linear prediction head is used to output a yield prediction value based on the final level representation; S3, apply weak feature perturbation and strong feature perturbation to the encoder of the prediction model respectively, use the prediction model with weak feature perturbation to make multiple predictions on the unlabeled dataset, calculate the pseudo target mean and prediction uncertainty, and further construct a confidence mask; The prediction model with strong feature perturbation is used to make a prediction on the unlabeled dataset, and the filtering consistency loss is calculated by combining the pseudo-target mean and confidence mask. S4. Use the prediction model to predict the labeled dataset, calculate the supervised regression loss, and calculate the total training loss including dynamic consistency weights together with the filter consistency loss. Then train the prediction model to obtain the trained prediction model. S5. Using the trained prediction model, obtain the predicted yield value of the wafer to be inspected.

2. The semi-supervised learning method for predicting the yield of tag-delayed wafers according to claim 1, characterized in that, In S2, the adaptive learnable attention block includes a normalization layer, a multi-head self-attention layer, a gated interactive selection layer, a linear projection layer, and an attention-based learnable residual aggregation layer connected in sequence.

3. The semi-supervised learning method for predicting the yield of tag-delayed wafers according to claim 2, characterized in that, In the encoder, the first The process of outputting hierarchical representations for each adaptive learnable attention block is as follows: S201, receiving the first The hierarchical representation of the output of each adaptive learnable attention block is normalized and multi-head self-attention is performed sequentially in the normalization layer and the multi-head self-attention layer to obtain the attention representation; S202, In the gated interaction selection layer, the gate factor is obtained based on the hierarchical representation of the output of the previous adaptive learnable attention block; By combining the gating factor and the attention representation, the gated feature representation is obtained, which is then passed through a linear projection layer to obtain the transformed representation. S203, the initial embedding representation and the previous The transformation representation of each adaptive learnable attention block is input to the attention-based learnable residual aggregation layer, and weighted summation is performed by combining the residual aggregation weights to obtain the hierarchical representation.

4. The semi-supervised learning method for predicting the yield of tag-delayed wafers according to claim 3, characterized in that, In S203, the residual aggregation weight is a learnable parameter, specifically: Initialize a residual pair consisting of a learnable query vector and a residual candidate representation for the initial embedding representation and all transformed representations respectively; during training, use the initial embedding representation and all transformed representations obtained in the previous training round as the residual candidate representations used in the current training round. The residual aggregation weight refers to the ratio of the attention score of the initial embedding representation and the attention score of all transformed representations to the total attention score; For each residual pair, the root mean square normalized result of the residual candidate representation is multiplied by the transpose of the learnable query vector, and then an exponential operation is performed to obtain the attention score of the initial embedded representation or transformed representation corresponding to the residual pair. The total attention score is the sum of the attention score of the initial embedding representation and the attention scores of all transformed representations.

5. A semi-supervised learning method for predicting the yield of tag-delayed wafers according to claim 1, characterized in that, In step S3, the process of performing multiple predictions using a prediction model with applied weak feature perturbations and obtaining the pseudo-target mean, prediction uncertainty, and confidence mask is as follows: S301, the test data of the unlabeled dataset is passed through the embedding layer of the prediction model to obtain the initial embedding representation; S302, input the initial embedded representation into the encoder, add the hierarchical representation obtained from the first adaptive learnable attention block to the weak feature perturbation to obtain the hierarchical representation after weak perturbation; S303, input the hierarchical representation after weak perturbation into the next adaptive learnable attention block, until all adaptive learnable attention blocks in the encoder are traversed, input the hierarchical representation output by the last adaptive learnable attention block into the linear prediction head to obtain the yield prediction value. S304, repeat S302-S303 until the specified number of perturbations is reached; S305, taking each wafer in the unlabeled dataset as a unit, averages the predicted yield values ​​of the wafer under all weak feature perturbations to obtain the pseudo-target mean, and further calculates the prediction uncertainty of each wafer based on the predicted yield value and the pseudo-target mean. S306. Based on the prediction uncertainty, a confidence mask is constructed. If the prediction uncertainty is less than or equal to the threshold, the confidence mask corresponding to the wafer in the unlabeled dataset is set to 1, otherwise it is set to 0.

6. A semi-supervised learning method for predicting the yield of tag-delayed wafers according to claim 5, characterized in that, In S302, the weak feature perturbation adopts additive Gaussian perturbation, and the weak perturbation noise corresponding to the weak feature perturbation is sampled from a Gaussian distribution with a mean of 0 and a covariance equal to the square of the weak perturbation noise scale multiplied by the identity matrix.

7. A semi-supervised learning method for predicting the yield of tag-delayed wafers according to claim 1, characterized in that, In step S3, the process of using a prediction model with applied strong feature perturbation to perform a prediction and calculating the filtering consistency loss is as follows: S311, the test data of the unlabeled dataset is passed through the embedding layer of the prediction model to obtain the initial embedding representation; S312, the hierarchical representation output from the last adaptive learnable attention block is added to the initial embedded representation input to the encoder, and the strong feature perturbation is added to obtain the hierarchical representation after strong perturbation. S313, input the hierarchical representation after strong perturbation into the linear prediction head to obtain the yield prediction value; S314: Using the pseudo-target mean as the true value, calculate the mean square error between the predicted yield value and the true value; combine the mean square error and the confidence mask to calculate the filtering consistency loss.

8. A semi-supervised learning method for predicting the yield of tag-delayed wafers according to claim 7, characterized in that, In S314, the formula for calculating the consistency loss is as follows: ; in, To filter out consistency loss; This represents the total number of unlabeled samples in a training batch, i.e., the number of wafers in a training batch that did not obtain a yield label. Let be the confidence mask for the j-th unlabeled sample; This represents the predicted yield value of the j-th unlabeled sample after strong feature perturbation. Let be the pseudo-target mean of the j-th unlabeled sample; To prevent extremely small constants with a denominator of zero, Let be the mean square error function.

9. A semi-supervised learning method for predicting the yield of tag-delayed wafers according to claim 7, characterized in that, In S312, the strong feature perturbation adopts an additive Gaussian perturbation, and the strong perturbation noise corresponding to the strong feature perturbation is sampled from a Gaussian distribution with a mean of 0 and a covariance equal to the square of the strong perturbation noise scale multiplied by the identity matrix.

10. A semi-supervised learning method for predicting the yield of tag-delayed wafers according to claim 1, characterized in that, In S5, the total training loss is calculated by multiplying the filtering consistency loss by the consistency weight and then adding the supervised regression loss. The consistency weight is calculated as follows: First, divide the index value of the current training round by the preset consistency weight increment, compare the calculation result with 1, and multiply the smaller value by the preset maximum consistency weight to obtain the consistency weight.