A wafer defect detection method, device, equipment and medium
By acquiring the multimodal dataset of wafers, using lightweight alignment and open-world wildcard methods to generate pseudo-label defect datasets, fine-tuning the pre-trained model, solving the problem of identifying unknown types in wafer defect detection, and improving detection accuracy and efficiency.
Patent Information
- Application Number
- CN202510513147.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Existing wafer defect detection methods require a large amount of labeled data, making it difficult to effectively identify and detect unknown wafer defect types, resulting in low detection accuracy.
By acquiring the wafer multimodal data set, using the lightweight alignment method for spatial alignment, generating the target wafer multimodal data set, using the open-world wildcard method to create pseudo-label defect data sets, and using the pseudo-label defect data set to pre-train and fine-tune the preset supervision model to obtain the object detection model.
It reduces the cost of manually labeling data, improves the accuracy of wafer defect detection, can identify unknown wafer defect types, and enhances the generalization ability and robustness of the model.
Smart Images

Figure CN120031883B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semiconductor defect detection, and in particular, to a wafer defect detection method, device, equipment and medium. Background Art
[0002] The rapid development of semiconductor technology has attracted wide attention in today's society, promoting the continuous progress in fields such as communication, information technology, and embedded systems. During the semiconductor chip processing, many wafer defects will occur, and the detection and analysis of wafer defects are key links to ensure product quality and yield. These defects may come from various factors in the production process, such as material problems, improper process parameter settings, or equipment failures.
[0003] However, traditional wafer defect detection methods need to use a large amount of labeled data to train a relatively stable initial detection model, but it may be very difficult and costly to obtain sufficient labeled data, and only a few types of wafer defects can be learned, and unknown wafer defect types cannot be effectively identified and detected, resulting in low detection accuracy for unknown wafer defect types. Therefore, how to effectively identify and detect unknown wafer defect types to improve the detection accuracy of wafer defects is a technical problem to be solved urgently. Summary of the Invention
[0004] Based on this, in view of the above technical problems, embodiments of the present invention provide a wafer defect detection method, device, equipment and medium, which can effectively identify and detect unknown wafer defect types, thereby improving the detection accuracy of wafer defects.
[0005] The first aspect of the embodiments of the present application provides a wafer defect detection method, and the wafer defect detection method includes:
[0006] Obtain a wafer multimodal data set;
[0007] Based on a lightweight alignment method, perform spatial alignment on the wafer multimodal data set to generate a target wafer multimodal data set;
[0008] Use the open-world wildcard method to make pseudo-labels for the unlabeled wafer data in the target wafer multimodal data set to generate multiple pseudo-label defect data sets;
[0009] Use the multiple pseudo-label defect data sets to pre-train a preset supervised model, and use a preset fine-tuning algorithm to fine-tune the initial detection model obtained after pre-training to obtain a target detection model;
[0010] Input the wafer image to be measured into the target detection model for defect detection to obtain a wafer defect detection result.
[0011] In a second aspect of the embodiments of the present application, a wafer defect detection device is provided. The wafer defect detection device includes:
[0012] An acquisition module, configured to acquire a wafer multimodal dataset;
[0013] An alignment module, configured to perform spatial alignment on the wafer multimodal dataset based on a lightweight alignment method to generate a target wafer multimodal dataset;
[0014] A generation module, configured to use an open-world wildcard method to make pseudo-labels for unlabeled wafer data in the target wafer multimodal dataset, generating a plurality of pseudo-label defect datasets;
[0015] A training module, configured to pre-train a preset supervised model using the plurality of pseudo-label defect datasets, and perform model fine-tuning on the initial detection model obtained after pre-training using a preset fine-tuning algorithm to obtain a target detection model;
[0016] A detection module, configured to input a to-be-detected wafer image into the target detection model for defect detection to obtain a wafer defect detection result.
[0017] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the wafer defect detection method described in the first aspect is implemented.
[0018] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the wafer defect detection method described in the first aspect is implemented.
[0019] In summary, the present invention provides a wafer defect detection method, apparatus, device and medium. A wafer multi-modal data set is acquired, and based on a lightweight alignment method, spatial alignment is performed on the wafer multi-modal data set to generate a target wafer multi-modal data set. The open-world wildcard method is used to create pseudo-labels for the unlabeled wafer data in the target wafer multi-modal data set, generating multiple pseudo-label defect data sets. The multiple pseudo-label defect data sets are used to pre-train a preset supervised model, and a preset fine-tuning algorithm is used to fine-tune the initial detection model obtained after pre-training to obtain a target detection model. The image of the wafer to be tested is input into the target detection model for defect detection to obtain the wafer defect detection result. It can be seen that in this application, by using the open-world wildcard method to create pseudo-labels for the unlabeled wafer data in the target wafer multi-modal data set, generating multiple pseudo-label defect data sets, and pre-training a preset supervised model according to the multiple pseudo-label defect data sets, the problem of lack of training data is overcome, the cost of manually labeled data is greatly reduced, and then a preset fine-tuning algorithm is used to fine-tune the initial detection model obtained after pre-training, so that the obtained target detection model can learn different wafer defect features, identify unknown wafer defect types, and improve the detection accuracy of wafer defects. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0021] Figure 1 is a flowchart of a wafer defect detection method provided by an embodiment of the present invention;
[0022] Figure 2 is a structural diagram of a wafer defect detection device provided by an embodiment of the present invention;
[0023] Figure 3 is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the protection scope of the present invention.
[0025] It should be understood that, as used in the specification of the present invention and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their combinations.
[0026] It should also be understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0027] As used in the specification of the present invention and the appended claims, the term "if" can be interpreted as "when" or "once" or "in response to determining" according to the context. Similarly, the phrases "if determined" or "if compared to [the described condition or event]" can be interpreted as meaning "once determined" or "in response to determining" or "once compared to [the described condition or event]" or "in response to comparing to [the described condition or event]" according to the context.
[0028] In addition, in the description of the specification of the present invention and the appended claims, the terms "first", "second", "third", etc. are only used for differentiating descriptions and cannot be understood as indicating or implying relative importance.
[0029] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present invention means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0030] It should be understood that the magnitude of the sequence numbers of the steps in the following embodiments does not mean the order of execution is prior or subsequent, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0031] To illustrate the technical solution of the present invention, the following specific embodiments are used for illustration.
[0032] See Figure 1 , which is a schematic flow chart of a wafer defect detection method provided by an embodiment of the present invention. As Figure 1 shown, the wafer defect detection method can be implemented through the following steps.
[0033] S101: Obtain a wafer multi-modal dataset.
[0034] In step S101, by obtaining a wafer surface defect dataset and various ultra-large-scale multi-modal datasets, for example, the CC12M dataset, the GQA dataset, the Flickr30K dataset, and the Open Images V7 dataset, etc., this application does not make any limitations. These datasets cover various types of images and labels, including visual descriptions of image content, relationship graphs of questions and answers, etc. multi-modal data. Among them, the wafer surface defect dataset is collected by deploying high-precision industrial cameras on the semiconductor wafer production line. These cameras can capture defect images on the wafer surface in real time, covering different types of surface defects (such as scratches, cracks, bubbles, etc.); The CC12M (Conceptual 12m) dataset contains 12 million pairs of image-text pairing data. The images are two-dimensional images, and the texts are natural language descriptions related to the images, usually involving detailed descriptions of the content, scene, or objects in the images; The GQA (Graph Question Answering) dataset consists of a series of images and their related questions and answers, focusing on semantic information in the images and the relationships between image elements. Each image is accompanied by a set of natural language questions designed based on the image content. The questions involve objects, scenes, and their relationships in the images, and each question has a corresponding answer; The Flickr30K dataset contains 30,000 pictures and their corresponding natural language descriptions. Each picture is accompanied by 5 different text descriptions, and the content of the descriptions covers multi-dimensional information such as scenes, objects, and activities in the images; The Open Images V7 dataset contains approximately 9 million images and is annotated with object detection boxes for up to 6,000 categories, mainly applicable to supervised object detection and image segmentation and other deep learning task training.
[0035] Since the obtained wafer surface defect dataset and various ultra-large-scale multi-modal datasets are represented in multiple formats and there will be problems such as inconsistent data formats, units, and dimensions. In addition, some data has missing values, outliers, and error values, and these problems will affect the training and performance of subsequent models. Therefore, the wafer surface defect dataset and various ultra-large-scale multi-modal datasets can be preprocessed to unify the data into the same format and standard to ensure data consistency, and then the preprocessed wafer surface defect dataset and various ultra-large-scale multi-modal datasets can be integrated to obtain a wafer multi-modal dataset.
[0036] In the embodiments of the present application, by preprocessing and integrating these data sets, the consistency of the wafer multi-modal data set is ensured, so that the subsequent model can be trained on multiple different data sources, overcoming the problem of lack of training data for the detection model and enhancing its generalization ability in unknown defect categories and complex working conditions.
[0037] S102: Based on a lightweight alignment method, perform spatial alignment on the wafer multi-modal data set to generate a target wafer multi-modal data set.
[0038] In step S102, in multi-modal tasks (such as tasks that simultaneously process images and text), alignment pre-training is a very important step. Its goal is to make the image and its related text description have similar representations in the embedding space, so as to better understand the relationship between the image and the text, and thus be able to effectively perform cross-modal understanding and reasoning in multi-modal tasks. In traditional detection models, the alignment of image and text modalities usually relies on complex fusion operations (such as CLIP), which require a large amount of computing resources and significantly reduce the model efficiency in the inference stage. The lightweight alignment method reduces the computational cost when achieving modality alignment while maintaining the improvement of the model's performance. Therefore, by using the lightweight alignment method, perform spatial alignment on the wafer multi-modal data set to generate a target wafer multi-modal data set.
[0039] In an embodiment of the invention, the step of performing spatial alignment on the wafer multi-modal data set based on the lightweight alignment method to generate a target wafer multi-modal data set includes:
[0040] Based on a preset low-rank adaptation method and multi-scale feature fusion method, process the wafer text data and the wafer image data to determine the target interaction information between the wafer text data and the wafer image data;
[0041] According to the target interaction information between the wafer text data and the wafer image data, use the negative sample label method and the alignment method to perform image-text association on the wafer text data and the wafer image data to generate multiple image-text data pairs, where each image-text data pair corresponds to a respective description text;
[0042] For each image-text data pair, calculate the similarity between the description text and the image in the image-text data pair, and generate an image-text pre-training data pair based on the image-text data pair, the description text, and the similarity;
[0043] Generate a target wafer multi-modal data set according to each of the image-text pre-training data pairs.
[0044] Specifically, the wafer multi-modal dataset includes wafer text data and wafer image data. By introducing low-rank adaptation into the CLIP text encoder to eliminate complex fusion operations, the CLIP text encoder originally had some fixed parameters (weights). Now, a low-rank matrix is added to these parameters to process the wafer text data and wafer image data and determine the initial interaction information between the wafer text data and the wafer image data. That is, by using the low-rank adaptation method, low-rank matrices are introduced into all queries, keys, values, and output projections of the CLIP text encoder. The specific formula is:
[0045] ;
[0046] where, represents the pre-trained weights of the CLIP text encoder, represents the product of two low-rank matrices, with the input and output being x and h respectively. The rank of the low-rank matrix is set to a value much smaller than the model feature dimension. Through the low-rank matrix, the model dynamically stores information related to cross-modal interactions during training, thereby improving the adaptability of the decision boundary. This method ensures that the parameters of the pre-trained text encoder remain unchanged, while the low-rank matrix dynamically stores cross-modal interaction information. And in the inference stage, the calibrated text embeddings can be pre-computed and stored offline, thus avoiding the computational cost of the text encoder, which makes the model more efficient during inference. Using the RT-UTER model as an efficient detector, the classification head in the RT-UETR model is adapted through multi-modal dual-head matching. In region-text contrastive learning, the region embeddings of the two heads are refined by aligning with the shared and semantically rich text representation, realizing end-to-end training and inference.
[0047] Thus, after obtaining the initial interaction information between the wafer text data and the wafer image data, although the low-rank adaptation method can achieve text-image interaction, there are still key issues such as temporal synchronization in the currently acquired images (such as image drift caused by dynamic process parameter fluctuations), which will lead to a decline in the quality after multi-modal data alignment. Therefore, it is necessary to use the multi-scale feature fusion method to process the initial interaction information between the wafer text data and the wafer image data, and then obtain the target interaction information between the wafer text data and the wafer image data to extract the wafer texture and frequency domain features and enhance the robustness of spatial alignment. Among them, the multi-scale feature fusion method includes the gray-level co-occurrence matrix (GLCM) and the Fourier spectrum cross-correlation method. The gray-level co-occurrence matrix is a method of characterizing texture features by statistically analyzing the gray-scale spatial distribution of pixel pairs in an image. In our semiconductor defect detection products, it is mainly used to identify texture anomalies such as scratches on the wafer surface, thereby assisting in label establishment. The Fourier spectrum cross-correlation is a defect detection method based on frequency domain analysis. In our semiconductor defect detection products, it is mainly used for periodic pollution anomalies such as particles formed on the wafer surface, thereby assisting in label establishment. It can be seen that by accurately determining the target interaction information, the accuracy of wafer defect detection and classification can be improved, the situations of false detection and missed detection can be reduced, and the quality and efficiency of semiconductor manufacturing can be improved. Furthermore, according to the target interaction information between the wafer text data and the wafer image data, the negative sample label method and the alignment method are used to perform image-text association on the wafer text data and the wafer image data to generate multiple image-text data pairs. Among them, the descriptive text corresponding to each image-text data pair. In multi-modal learning, especially in the task of image and text alignment, the situation where images and texts do not strictly correspond one-to-one is usually encountered. For example, an image may contain multiple objects, while the text description may only cover a part of them, or a text description may apply to multiple images. This non-strict correspondence relationship will cause noise in the traditional global instance-level alignment objective (that is, assuming that each image and text pair is one-to-one), thus affecting the training effect of the model. It is adjusted by softening the negative sample label method, and this adjustment process is divided into several stages:
[0048] a. Initial stage (all-1 labels): In the early stage of training, the labels of negative samples are set to all 1, which means that the model will regard all negative samples as equally important as positive samples in the initial stage, thus avoiding premature strict discrimination of negative samples.
[0049] b. Intermediate stage (softened labels): As training progresses, the labels of negative samples gradually transition from all 1 to softened labels. The value of the softened label is between 0 and 1, indicating that the similarity between negative samples and positive samples gradually decreases. This softening process can help the model gradually distinguish negative samples from positive samples while reducing the impact of noise.
[0050] c. Late stage (softer labels): In the late stage of training, the labels of negative samples are further softened and approach 0, indicating that the model's ability to distinguish negative samples is gradually enhanced, and it can more accurately identify the differences between negative samples and positive samples.
[0051] This alignment method is divided into a multi-head alignment method and a single-head alignment method. In region-text contrastive learning, through a consistent double-alignment strategy, the decision boundaries of the two classification heads are made more consistent. The specific formula is expressed as:
[0052] ;
[0053] where u represents the IoU value between the predicted box and the ground truth box, α and β represent the classification heads, and s represents the classification score obtained through multimodal information. The calculation formula for this classification score is:
[0054] ;
[0055] where sim represents the cosine similarity, T represents the text embedding, and I represents the pixel-level features of the image. To ensure the consistency of the supervision signals of the two heads in multimodal dual-head matching, a consistent setting is adopted, which is specifically expressed as:
[0056] .
[0057] This enables the one-to-one head to effectively learn the same supervision signals as the one-to-many heads, thereby improving the performance of the model. The core idea of the one-to-one head strategy is to assign a unique text category to each image region, ensuring a one-to-one correspondence between image features and text features.
[0058] Furthermore, if there is exactly one dropped object defect on a detection sub - graph in our scenario, the one - to - one reasoning efficiency is relatively high at this time. The core idea of the one - to - multiple - heads strategy is to allow an image region to match multiple text categories. This strategy is more flexible and can handle complex scenarios, but the computational cost is relatively high. If there is one dirt defect on a detection sub - graph in our scenario, and the dirt usually consists of multiple particle defects, the one - to - many reasoning efficiency is relatively high at this time. That is, for each image - text data pair, by introducing the open - source CLIP model, this CLIP model is first trained on large images with captions to learn the representations of images and texts in the joint embedding space. Based on the assumption that "in this space, the distance between the image embedding and its corresponding caption embedding is relatively close, while the distance between unrelated image embeddings and caption embeddings is farther", the CLIP model can extract text from the image and can compare the obtained text with the given text to generate a semantic relevance score between the image and the text. According to the semantic relevance score between the image and the text, the similarity between the description text and the image in the image - text data pair is calculated respectively, and based on the image - text data pair, the description text, and the similarity, an image - text pre - training data pair is generated, so as to generate a target wafer multi - modal data set according to each of the image - text pre - training data pairs. It can be seen that by using the low - rank adaptation method to reduce the feature dimension, the target interaction information between the wafer text data and the image data can be effectively captured. By introducing negative samples (mismatched text and image pairs), it helps the model better learn the discriminative features between text and images, can automatically generate high - quality image - text data pairs, reduces the dependence on a large amount of manually labeled data, significantly reduces the cost and time of data annotation, and at the same time ensures the quality of the data.
[0059] In this embodiment, by using the lightweight alignment method, the spatial alignment of the wafer multi - modal data set is performed to generate the target wafer multi - modal data set, thereby reducing the computational cost during multi - modal data alignment, facilitating the subsequent improvement of the efficiency and accuracy of wafer defect detection, and at the same time reducing the costs of data annotation and model training.
[0060] S103: Use the open - world wildcard method to create pseudo - labels for the unlabeled wafer data in the target wafer multi - modal data set, generating multiple pseudo - label defect data sets.
[0061] In step S103, for the wafer defect detection scenario, we do not want the model to have a low confidence score for the defect category, while also hoping to expand the detection of defects outside the defined categories. The open-world wildcard is designed to enable the model to detect objects that do not exist in the predefined vocabulary and label them as "unknown". This method is achieved by using a wildcard embedding that can capture unknown objects in the scenario in a zero-shot manner. Specifically, all wildcard embeddings are initialized from the text features of a general text (such as "object"), which are extracted by a calibrated text encoder. At this time, all unknown objects are initially mapped to a general category, providing a basis for subsequent learning. Since the target wafer multi-modal dataset contains a small amount of labeled wafer detection data and a large amount of pseudo-labeled wafer detection data, the wildcard embedding is fine-tuned using the pre-training dataset, treating all real instances as the same "object" category. This fine-tuning enables the embedding to capture richer semantic information, thereby enhancing the model's ability to identify objects not covered by the predefined specific categories, so as to distinguish wafer detection data of different defect types, and thus create pseudo-labels for a large amount of unlabeled wafer detection data to generate multiple pseudo-label defect datasets. And for some ultra-small defects (5x5, pixel unit), a super-resolution network can be used to increase the defect size, thereby improving the accuracy of the label box.
[0062] In an embodiment of the invention, the open-world wildcard method is used to create pseudo-labels for the unlabeled wafer data in the target wafer multi-modal dataset, generating multiple pseudo-label defect datasets, including:
[0063] Based on the open-world wildcard method, learn the semantic defect information between modalities from the target wafer multi-modal dataset;
[0064] According to the semantic defect information, use multiple autoencoders to predict the labels of the unlabeled wafer data in the target wafer multi-modal dataset respectively, obtaining multiple prediction results;
[0065] Judge whether the multiple prediction results meet the preset IoU threshold and score threshold;
[0066] If the multiple prediction results meet the preset IoU threshold and score threshold, determine the pseudo-labels corresponding to the unlabeled wafer data according to the multiple prediction results, generating multiple pseudo-label defect datasets.
[0067] Specifically, by using the open-world wildcard method, semantic defect information between modalities is learned from the target wafer multi-modal dataset, that is, initialized from the text features of general text (such as "defect") through wildcard embedding. This method allows the model to still understand and infer the characteristics of defects when facing unseen defect types. The features extracted by the calibrated text encoder ensure that the wildcard embedding is not only general but also can capture the semantic information of potential defects, which may include key features such as the category, location, and size of wafer defects. The open-world wildcard method may involve using wildcards to represent potential associations or similarities between different modalities in wafer data, which may require the use of multi-modal information fusion technologies, such as physical layer fusion, feature layer fusion, decision layer fusion, etc.
[0068] Furthermore, after determining the semantic defect information, multiple autoencoders are trained for the semantic defect information. Each autoencoder is responsible for extracting features from a specific data modality and attempting to reconstruct or predict labels. The training of the autoencoder can be based on unsupervised learning or self-supervised learning methods and fine-tuned using labeled data (if available). The unlabeled wafer data is input into the trained multiple autoencoders to obtain multiple prediction results. Each autoencoder will output a prediction about the wafer defect, including information such as the location and category of the defect. Calculate the intersection over union (IoU) and score (usually the confidence or probability of the prediction result) of each prediction result, and compare the calculated IoU and score with the preset IoU threshold and score threshold. If the IoU and score of a certain prediction result both meet the threshold requirements, then this prediction result is considered valid. By setting the IoU threshold (o1 = 0.5) and score threshold (o2 = 0.01) to select pseudo-labels for training the "unknown" wildcard. For multiple prediction results that meet the threshold requirements, the pseudo-labels corresponding to the unlabeled wafer data are determined according to the multiple prediction results and the prediction accuracy corresponding to their respective encoders to generate multiple pseudo-label defect datasets. It can be seen that by making pseudo-labels for unlabeled wafer data, the cost and time of manual annotation can be reduced, so as to make full use of these unlabeled data to train the detection model subsequently, thereby improving the prediction accuracy. And by setting the IoU threshold and score threshold, the prediction results can be flexibly controlled and screened, thus ensuring the quality of the generated pseudo-label dataset.
[0069] In this embodiment, by using the open-world wildcard method to create pseudo-labels for the unlabeled wafer data in the target wafer multi-modal dataset, the labeling cost can be significantly reduced, so that the detection model trained by multiple pseudo-label defect datasets can still understand and infer the characteristics of defects when facing unseen defect types, thereby improving the efficiency of wafer defect detection.
[0070] S104: Use the multiple pseudo-label defect datasets to pre-train a preset supervised model, and use a preset fine-tuning algorithm to fine-tune the initial detection model obtained after pre-training to obtain a target detection model.
[0071] In step S104, after using the open-world wildcard method to create pseudo-labels for the unlabeled wafer data in the target wafer multi-modal dataset and generating multiple pseudo-label defect datasets, the generated pseudo-label defect datasets are then merged with the dataset with real labels to form a labeled defect dataset. Select a suitable supervised learning model as the preset model, which should have the ability to handle wafer detection tasks. Use the labeled defect dataset to pre-train the preset supervised model. During the pre-training process, the model will learn the feature representations in the labeled defect dataset, including real-label data and pseudo-label data. Select a preset fine-tuning algorithm, which should be applicable to the target detection task and can effectively utilize the pseudo-label data. Apply the fine-tuning algorithm to the initial detection model obtained after pre-training to obtain a target detection model. During the fine-tuning process, the model will be further optimized for the specific wafer detection task to improve the detection performance.
[0072] In an embodiment of the invention, based on the open-world wildcard method, pre-training a preset supervised model according to the target wafer multi-modal dataset includes:
[0073] Perform data processing on the multiple pseudo-label defect datasets to obtain the processed multiple pseudo-label defect datasets;
[0074] Merge the processed multiple pseudo-label defect datasets and the real-label defect datasets to form a target labeled defect dataset.
[0075] Iteratively train the detection model according to the target labeled defect dataset until the trained detection model meets the expected performance to obtain an initial detection model.
[0076] Specifically, perform data processing on multiple pseudo-label defect data sets to obtain multiple processed pseudo-label defect data sets. Concatenate the processed pseudo-label defect data sets and the true-label defect data sets to form a target label defect data set. When concatenating, it is necessary to ensure the consistency of the data set format and labels. To avoid overfitting of the model to a certain type of sample, the target label defect data set can be balanced. For example, through oversampling or undersampling techniques, the number of samples in different categories can be kept balanced. Select a supervised learning model suitable for the wafer detection task, such as a convolutional neural network (CNN), a support vector machine (SVM), etc., and initialize the parameters of the model. Random initialization or pre-trained weights can be used. Use the target label defect data set to iteratively train the model. In each iteration, divide the data set into a training set and a validation set. Use the training set to update the model parameters and use the validation set to evaluate the model performance. According to the performance feedback of the validation set, adjust the model parameters and training strategy until the trained detection model meets the expected performance. When the performance of the detection model reaches the preset standard, use it as the initial detection model. This initial detection model can be used for subsequent wafer defect detection tasks and can be further optimized and improved according to actual needs. It can be seen that by combining the processed pseudo-label defect data sets and the true-label defect data sets, the scale of the data set is further increased, the training effect of the model is improved, and the initial detection model obtained by training can better learn the characteristics and rules of different wafer defects, thereby enhancing the generalization ability of the model and further improving the detection ability and robustness for unknown defect categories.
[0077] In an embodiment of the invention, performing data processing on multiple pseudo-label defect data sets to obtain multiple processed pseudo-label defect data sets includes:
[0078] Performing duplicate filtering on the multiple pseudo-label defect data sets to obtain multiple filtered pseudo-label defect data sets;
[0079] Performing vocabulary expansion processing on the multiple filtered pseudo-label defect data sets to obtain multiple expanded pseudo-label defect data sets;
[0080] Based on a preset confidence threshold, judging the confidence information of the multiple expanded pseudo-label defect data sets to obtain multiple judgment results;
[0081] In response to the multiple judgment results respectively meeting the first judgment condition, adjusting the confidence information in the pseudo-label defect data groups corresponding to the judgment results that meet the conditions to obtain multiple processed pseudo-label defect data sets.
[0082] Specifically, a hash algorithm, similarity calculation (such as cosine similarity, Jaccard similarity, etc.) or clustering algorithm (such as K-means, DBSCAN, etc.) is used to filter and remove duplicate data items from multiple pseudo-label defect data sets, and the filtered multiple pseudo-label defect data sets are obtained. Each data set is de-duplicated. That is, during the inference process, a simple unknown filtering strategy is used to remove unknown class predictions that highly overlap with known class predictions (IoU threshold τ = 0.99) to reduce duplicates. This strategy helps reduce false detections and redundant predictions, ensuring that the detection results of each "unknown" class have sufficient distinctiveness and credibility. Through this step, the model can be prevented from generating excessive duplicate detections, optimizing the inference efficiency and accuracy. Based on the existing label vocabulary, methods such as synonym replacement, near-synonym expansion, and context reasoning are used to perform vocabulary expansion processing on the filtered multiple pseudo-label defect data sets, and the expanded multiple pseudo-label defect data sets are obtained. Each data set contains more diverse label vocabularies. That is, by discovering new classes from "unknown" class predictions and adding their class names to the vocabulary, known classes are provided for the next iteration, thus realizing dynamic vocabulary expansion. Dynamically expanding the vocabulary not only improves the generalization ability of the model but also enables the model to continuously adapt to newly emerging defect classes, with stronger adaptability and long-term application capabilities. Each newly added class is evaluated and verified to ensure its accuracy and representativeness, avoiding the addition of irrelevant classes.
[0083] Furthermore, by evaluating the confidence of each data item according to a preset confidence threshold, multiple judgment results are obtained. Each result indicates whether the confidence of the corresponding data item meets the threshold requirement. The confidence can be based on the probability predicted by the model, similarity score, or other measurement criteria. In response to the judgment result meeting the first judgment condition (i.e., the confidence does not meet the threshold requirement), the confidence information in the corresponding pseudo-label defect data group is adjusted. The adjustment method can be recalculating the confidence, using a more accurate prediction model, or introducing external verification data. The first judgment condition can be that the confidence is lower than a specific value, or the confidence differs too much from the confidence of other data items, etc. The present application does not make any limitations in this regard. Then, the processed multiple pseudo-label defect data sets are obtained, and the confidence of each data set has been appropriately adjusted. It can be seen that through filtering duplicate processing and vocabulary expansion processing, redundant data can be removed and the richness of label vocabularies can be increased, thereby improving the quality and diversity of the data sets. By processing and adjusting the pseudo-label defect data sets, the dependence on manual annotation can be reduced, and the annotation cost and time can be lowered. By setting the confidence threshold and adjusting the data items that do not meet the requirements, it can be ensured that the confidence of each data item in the data set is more accurate and reliable.
[0084] In an embodiment of the invention, the initial detection model obtained after pre-training is fine-tuned by using a preset fine-tuning algorithm to obtain a target detection model, including:
[0085] Select a part of the pseudo-labeled defect datasets from the processed multiple pseudo-labeled defect datasets in advance, and perform manual annotation on the part of the pseudo-labeled defect datasets to obtain the manual annotation data result;
[0086] Use a preset similarity method to match the manual annotation data result with another part of the pseudo-labeled defect datasets to obtain a label matching result;
[0087] According to the label matching result, use a parameter-efficient fine-tuning strategy to fine-tune the initial detection model obtained after pre-training to obtain a target detection model.
[0088] Specifically, based on the wildcard method, the defect categories have been pseudo-labeled, that is, all defects are collectively referred to as "object". To avoid the negative impact of errors in the pseudo-labels on the training process, a threshold is set, and multiple prediction results with confidence higher than the threshold are used as candidate pseudo-labels, rather than only selecting the label with the highest confidence. This method can effectively reduce the bias caused by a single high-confidence prediction and increase the diversity of the pseudo-labels. Therefore, to further ensure the correctness of defect classification, from the processed multiple pseudo-labeled defect datasets, according to factors such as data quality and diversity, select a part of the pseudo-labeled defect datasets, perform manual verification and annotation on the selected pseudo-labeled defect datasets to ensure the accuracy and consistency of the annotation results, and obtain the manual annotation data result. This part of the data will be used as the "gold standard" for subsequent matching and model fine-tuning. Use a preset similarity method (such as cosine similarity, Euclidean distance, Jaccard similarity, etc.) to match the manual annotation data result with another part of the pseudo-labeled defect datasets to obtain a label matching result. The label matching result includes the pseudo-labeled defect datasets similar or related to the manual annotation data result.
[0089] Furthermore, by adopting parameter-efficient fine-tuning strategies (such as Fine-tuning, weight transfer, feature extraction, etc.), the pseudo-labeled defect dataset obtained by matching is used as training data to fine-tune the initial detection model obtained after pre-training, thereby obtaining the target detection model. For example, randomly select 20% of the labels from all pseudo-labels for manual annotation. During the manual annotation process, each label box is reviewed by at least 2 experts to ensure the accuracy and consistency of the labels. In addition, the cosine similarity method is used to match the 20% of the manually annotated labels with the remaining 80% of the pseudo-label information to ensure that there is unique and accurate label information for all pseudo-labels. This step not only improves the quality of the pseudo-labels but also further enhances the reliability of the labels through the expert review mechanism. In the model fine-tuning stage, considering only modifying the defect category information without affecting the defect feature information, the RT-UTER classification head and the network layer above it are frozen. The AdamW optimizer is used, and the learning rate is set to 0.0002, and the model is fine-tuned 50,000 times. This parameter-efficient fine-tuning strategy not only reduces the consumption of computing resources but also avoids over-adjusting the learned features by freezing specific layers, thereby further optimizing the performance of defect classification while maintaining the generalization ability of the model. It can be seen that through manual annotation and matching screening, the annotation cost and time can be significantly reduced, ensuring the data quality and annotation accuracy for model fine-tuning, thereby improving the accuracy of obtaining the target detection model and further enhancing the model's detection ability for specific wafer defects.
[0090] In this embodiment, by using multiple pseudo-labeled defect datasets to pre-train a preset supervised model, the problem of lack of model training data is overcome, so that the model can learn more wafer defect features and representations, providing a good foundation for the subsequent fine-tuning stage. At the same time, using a preset fine-tuning algorithm to fine-tune the initial detection model obtained after pre-training can further improve the accuracy and robustness of the model. It can be seen that by combining pre-training and model fine-tuning, while maintaining the model performance, the dependence on a large amount of labeled data can be reduced, thereby reducing the annotation cost and time.
[0091] S105: Input the wafer image to be measured into the target detection model for defect detection to obtain the wafer defect detection result.
[0092] In step S105, it mainly realizes defect detection based on the target detection model under actual application scenario data. By inputting the image of the wafer to be tested into the target detection model for defect detection, the wafer defect detection result is obtained. The wafer defect detection result includes the wafer defect type judgment result (category label), the wafer defect position (bounding box coordinates), and the corresponding confidence score, etc. And post-process the detection result output by the model, which may include techniques such as non-maximum suppression (NMS) to filter duplicate detection boxes. Then evaluate the accuracy of the detection result, such as calculating metrics like precision and recall. If the corresponding metrics are met, visualize the detection result and store the detection result for subsequent analysis or tracking.
[0093] In an embodiment of the invention, the image of the wafer to be tested is input into the target detection model for defect detection to obtain the wafer defect detection result, including:
[0094] After correcting the angle of the image of the wafer to be tested, crop the image of the wafer to be tested based on the set picture size to obtain the cropped image of the wafer to be tested;
[0095] Perform Gaussian filtering on the cropped image of the wafer to be tested to obtain the filtered image of the wafer to be tested;
[0096] Perform image enhancement on the filtered image of the wafer to be tested to obtain the enhanced image of the wafer to be tested;
[0097] Input the enhanced image of the wafer to be tested into the target detection model for forward inference to obtain the wafer defect detection result.
[0098] Specifically, image processing techniques (such as Hough transform, edge detection, etc.) are used to identify specific patterns or edges in the wafer image, thereby determining the principal axis direction of the wafer. Then, the image is rotated according to the principal axis direction to make the wafer image reach a standard angle (such as horizontal or vertical). According to the preset wafer image size (such as width and height), the angle-corrected wafer image to be measured is cropped to remove unnecessary parts in the image. A Gaussian filter is applied to the cropped wafer image to smooth the image and reduce noise, obtaining the filtered wafer image to be measured. The filtered wafer image to be measured is subjected to image enhancement processing (such as contrast stretching, histogram equalization, etc.) to obtain the enhanced wafer image to be measured. The enhanced wafer image to be measured is input into the target detection model for forward inference to obtain the wafer defect detection result. That is, load the trained target detection model and fix its model weights, turn off the gradient derivation. Secondly, after angle-correcting the wafer image to be measured collected by the camera and cropping it to meet the model input size according to the requirements, a sub-image is obtained. Then, the sub-image is preprocessed such as Gaussian filtering and image enhancement. Finally, the processed sub-image is input into the target detection model for forward inference. Through forward inference, the model will analyze the defect area in the image based on the learned feature representation and label information and generate the defect detection result. It can be seen that through preprocessing steps such as angle correction, cropping, Gaussian filtering, and image enhancement, the accuracy and efficiency of wafer defect detection can be effectively improved, which helps to enhance the generalization ability of the deep learning model.
[0099] In this embodiment, by inputting the wafer image to be measured into the target detection model for defect detection, the problem of lack of training data for the wafer epitaxial surface defect detection model is overcome, the computational complexity of data processing is reduced, defects other than the defined defect categories can be detected, so as to identify various types of defects, and the efficiency and accuracy of wafer defect detection are improved.
[0100] In summary, the present invention provides a wafer defect detection method, device, equipment and medium. A wafer multimodal dataset is obtained, and based on a lightweight alignment method, spatial alignment is performed on the wafer multimodal dataset to generate a target wafer multimodal dataset. The open-world wildcard method is used to create pseudo-labels for the unlabeled wafer data in the target wafer multimodal dataset, generating multiple pseudo-label defect datasets. The multiple pseudo-label defect datasets are used to pre-train a preset supervised model, and a preset fine-tuning algorithm is used to fine-tune the initial detection model obtained after pre-training to obtain a target detection model. The image of the wafer to be tested is input into the target detection model for defect detection to obtain the wafer defect detection result. It can be seen that in this application, by using the open-world wildcard method to create pseudo-labels for the unlabeled wafer data in the target wafer multimodal dataset, generating multiple pseudo-label defect datasets, and pre-training a preset supervised model according to the multiple pseudo-label defect datasets, the problem of lack of training data is overcome, the cost of manually labeled data is greatly reduced, and then a preset fine-tuning algorithm is used to fine-tune the initial detection model obtained after pre-training, so that the obtained target detection model can learn different wafer defect features, identify unknown wafer defect types, and improve the detection accuracy of wafer defects.
[0101] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of the wafer defect detection device provided by an embodiment of the present invention. This wafer defect detection device corresponds one-to-one with the wafer defect detection method in the above embodiment. For specific details, please refer to Figure 1 and Figure 1 the relevant descriptions in the corresponding embodiments. For the sake of convenience of description, only the parts related to this embodiment are shown. Refer to Figure 2 , the wafer defect detection device 20 includes: an acquisition module 21, an alignment module 22, a generation module 23, a training module 24, and a detection module 25.
[0102] The acquisition module 21 is used to acquire a wafer multimodal dataset;
[0103] The alignment module 22 is used to perform spatial alignment on the wafer multimodal dataset based on a lightweight alignment method to generate a target wafer multimodal dataset;
[0104] The generation module 23 is used to create pseudo-labels for the unlabeled wafer data in the target wafer multimodal dataset by using the open-world wildcard method, generating multiple pseudo-label defect datasets;
[0105] The training module 24 is used to pre-train a preset supervised model by using the multiple pseudo-label defect datasets, and use a preset fine-tuning algorithm to fine-tune the initial detection model obtained after pre-training to obtain a target detection model;
[0106] The detection module 25 is configured to input the image of the wafer to be tested into the target detection model for defect detection, so as to obtain the wafer defect detection result.
[0107] Optionally, the above alignment module 22 is specifically configured to:
[0108] Process the wafer text data and the wafer image data based on a preset low-rank adaptation method and a multi-scale feature fusion method to determine the target interaction information between the wafer text data and the wafer image data;
[0109] According to the target interaction information between the wafer text data and the wafer image data, use the negative sample label method and the alignment method to perform image-text association on the wafer text data and the wafer image data to generate a plurality of image-text data pairs, wherein each of the image-text data pairs corresponds to a respective description text;
[0110] For each of the image-text data pairs, calculate the similarity between the description text and the image in the image-text data pair, and generate an image-text pre-training data pair based on the image-text data pair, the description text, and the similarity;
[0111] Generate a target wafer multi-modal dataset according to each of the image-text pre-training data pairs.
[0112] Optionally, the above generation module 23 is specifically configured to:
[0113] Learn the semantic defect information between modalities from the target wafer multi-modal dataset based on the open-world wildcard method;
[0114] According to the semantic defect information, use a plurality of autoencoders to predict the labels of the unlabeled wafer data in the target wafer multi-modal dataset respectively to obtain a plurality of prediction results;
[0115] Determine whether the plurality of prediction results meet a preset IoU threshold and score threshold;
[0116] If the plurality of prediction results meet the preset IoU threshold and score threshold, determine the pseudo-labels corresponding to the unlabeled wafer data according to the plurality of prediction results, and generate a plurality of pseudo-label defect datasets.
[0117] Optionally, the above training module 24 is specifically configured to:
[0118] Perform data processing on the plurality of pseudo-label defect datasets to obtain a plurality of processed pseudo-label defect datasets;
[0119] Merge the processed multiple pseudo-label defect data sets and the real-label defect data sets to form a target label defect data set.
[0120] Iteratively train the detection model according to the target label defect data set until the trained detection model meets the expected performance, and obtain an initial detection model.
[0121] Optionally, the above training module 24 is further configured to:
[0122] Filter and duplicate the multiple pseudo-label defect data sets to obtain multiple filtered pseudo-label defect data sets;
[0123] Perform vocabulary expansion processing on the multiple filtered pseudo-label defect data sets to obtain multiple expanded pseudo-label defect data sets;
[0124] Based on a preset confidence threshold, judge the confidence information of the multiple expanded pseudo-label defect data sets to obtain multiple judgment results;
[0125] In response to the multiple judgment results respectively meeting the first judgment condition, adjust the confidence information in the pseudo-label defect data group corresponding to the judgment result that meets the condition to obtain multiple processed pseudo-label defect data sets.
[0126] Optionally, the above training module 24 is further configured to:
[0127] Pre-select a part of the pseudo-label defect data sets from the processed multiple pseudo-label defect data sets, and perform manual annotation on the part of the pseudo-label defect data sets to obtain a manual annotation data result;
[0128] Use a preset similarity method to match the manual annotation data result with another part of the pseudo-label defect data sets to obtain a label matching result;
[0129] According to the label matching result, use the parameter-efficient fine-tuning strategy to fine-tune the pre-trained initial detection model to obtain a target detection model.
[0130] Optionally, the above detection module 25 is specifically configured to:
[0131] After performing angle correction on the wafer image to be measured, crop the wafer image to be measured based on the set picture size to obtain a cropped wafer image to be measured;
[0132] Perform Gaussian filtering on the cropped wafer image to be measured to obtain a filtered wafer image to be measured;
[0133] Perform image enhancement on the filtered wafer image to be measured to obtain an enhanced wafer image to be measured;
[0134] Input the enhanced image of the wafer to be measured into the target detection model for forward inference to obtain the wafer defect detection result.
[0135] It should be noted that for the information interaction, execution process, etc. between the above units, since they are based on the same concept as the method embodiment of the present invention, for their specific functions and the technical effects brought, please refer to the method embodiment section for details, and will not be elaborated here.
[0136] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 3 shown, the electronic device of this embodiment includes: at least one processor ( Figure 3 only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, the steps in any of the above-mentioned wafer defect detection method embodiments are implemented.
[0137] The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 3 merely examples of electronic devices, and do not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include a network interface, a display screen, and an input system, etc.
[0138] In one embodiment, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by the processor in the electronic device, the electronic device can execute each step of any of the wafer defect detection methods disclosed in the present invention, which will not be repeated here. The computer-readable storage medium may be non-volatile or volatile.
[0139] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0140] The memory includes a readable storage medium, an internal memory, etc. Among them, the internal memory can be the memory of an electronic device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of the electronic device, and in some other embodiments, it can also be an external storage device of the electronic device. For example, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device. Further, the memory can also include both the internal storage unit of the electronic device and the external storage device. The memory is used to store the operating system, cooperative applications, a BootLoader, data, and other programs, such as the program code of a computer program. The memory can also be used to temporarily store the data that has been output or will be output.
[0141] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory, storage, database, or other media used in the embodiments provided in the present application can include non-volatile and / or volatile memories. The non-volatile memory can include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can include a random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0142] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working processes of the units and modules in the above-mentioned system can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0143] The above-mentioned embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.
Claims
1. A wafer defect detection method, characterized in that, Including: Obtain a wafer multi-modal dataset; Based on a lightweight alignment method, perform spatial alignment on the wafer multi-modal dataset to generate a target wafer multi-modal dataset; Use the open-world wildcard method to create pseudo-labels for the unlabeled wafer data in the target wafer multi-modal dataset, generating multiple pseudo-label defect datasets; Use the multiple pseudo-label defect datasets to pre-train a preset supervised model, and use a preset fine-tuning algorithm to fine-tune the initial detection model obtained after pre-training to obtain a target detection model; Input the wafer image to be measured into the target detection model for defect detection to obtain a wafer defect detection result; Among them, the wafer multi-modal dataset includes wafer text data and wafer image data. The method of performing spatial alignment on the wafer multi-modal dataset based on the lightweight alignment method to generate a target wafer multi-modal dataset includes: Based on a preset low-rank adaptation method and multi-scale feature fusion method, process the wafer text data and the wafer image data to determine the target interaction information between the wafer text data and the wafer image data; According to the target interaction information between the wafer text data and the wafer image data, use the negative sample label method and the alignment method to perform image-text association on the wafer text data and the wafer image data to generate multiple image-text data pairs, where each image-text data pair corresponds to a respective description text; For each image-text data pair, calculate the similarity between the description text and the image in the image-text data pair, and generate an image-text pre-training data pair based on the image-text data pair, the description text, and the similarity; Generate a target wafer multi-modal dataset according to each image-text pre-training data pair.
2. The wafer defect detection method according to claim 1, wherein, The method of using the open-world wildcard method to create pseudo-labels for the unlabeled wafer data in the target wafer multi-modal dataset, generating multiple pseudo-label defect datasets, includes: Based on the open-world wildcard method, learn the semantic defect information between modalities from the target wafer multi-modal dataset; According to the semantic defect information, use multiple autoencoders to predict the labels of the unlabeled wafer data in the target wafer multi-modal dataset respectively, obtaining multiple prediction results; Judge whether the multiple prediction results meet the preset IoU threshold and score threshold; If the multiple prediction results meet the preset IoU threshold and score threshold, determine the pseudo-labels corresponding to the unlabeled wafer data according to the multiple prediction results, generating multiple pseudo-label defect datasets.
3. The wafer defect detection method according to claim 1, characterized in that, The method of using the multiple pseudo-label defect datasets to pre-train a preset supervised model includes: Perform data processing on the multiple pseudo-label defect datasets to obtain processed multiple pseudo-label defect datasets; Merge the processed multiple pseudo-label defect datasets and the real-label defect datasets to form a target label defect dataset; Iteratively train the detection model according to the target label defect dataset until the trained detection model meets the expected performance, obtaining an initial detection model.
4. The wafer defect detection method according to claim 3, characterized in that Performing data processing on the multiple pseudo-label defect data sets to obtain the processed multiple pseudo-label defect data sets includes: Performing duplicate filtering on the multiple pseudo-label defect data sets to obtain the filtered multiple pseudo-label defect data sets; Performing vocabulary expansion processing on the filtered multiple pseudo-label defect data sets to obtain the expanded multiple pseudo-label defect data sets; Based on a preset confidence threshold, judging the confidence information of the expanded multiple pseudo-label defect data sets to obtain multiple judgment results; In response to the multiple judgment results respectively meeting the first judgment condition, adjusting the confidence information in the pseudo-label defect data groups corresponding to the judgment results that meet the conditions to obtain the processed multiple pseudo-label defect data sets.
5. The wafer defect detection method according to claim 1, characterized in that, Utilizing a preset fine-tuning algorithm to perform model fine-tuning on the initially obtained detection model after pre-training to obtain the target detection model includes: Pre-selecting a part of the pseudo-label defect data sets from the processed multiple pseudo-label defect data sets and performing manual annotation on the part of the pseudo-label defect data sets to obtain the manual annotation data results; Utilizing a preset similarity method to match the manual annotation data results with another part of the pseudo-label defect data sets to obtain the label matching results; According to the label matching results, utilizing a parameter-efficient fine-tuning strategy to perform model fine-tuning on the initially obtained detection model after pre-training to obtain the target detection model.
6. The wafer defect detection method according to claim 1, characterized in that Inputting the wafer image to be measured into the target detection model for defect detection to obtain the wafer defect detection result includes: After performing angle correction on the wafer image to be measured, cropping the wafer image to be measured based on the set picture size to obtain the cropped wafer image to be measured; Performing Gaussian filtering on the cropped wafer image to be measured to obtain the filtered wafer image to be measured; Performing image enhancement on the filtered wafer image to be measured to obtain the enhanced wafer image to be measured; Inputting the enhanced wafer image to be measured into the target detection model for forward inference to obtain the wafer defect detection result.
7. A wafer defect detection device, characterized in that, Including: An acquisition module for acquiring a wafer multi-modal data set; An alignment module for spatially aligning the wafer multi-modal data set based on a lightweight alignment method to generate a target wafer multi-modal data set; A generation module for using the open-world wildcard method to make pseudo-labels for the unlabeled wafer data in the target wafer multi-modal data set to generate multiple pseudo-label defect data sets; A training module for pre-training a preset supervised model using the multiple pseudo-label defect data sets and performing model fine-tuning on the initially obtained detection model after pre-training using a preset fine-tuning algorithm to obtain the target detection model; A detection module for inputting the wafer image to be measured into the target detection model for defect detection to obtain the wafer defect detection result; Wherein, the wafer multi-modal data set includes wafer text data and wafer image data, and the spatially aligning the wafer multi-modal data set based on the lightweight alignment method to generate the target wafer multi-modal data set includes: Based on a preset low-rank adaptation method and a multi-scale feature fusion method, process the wafer text data and the wafer image data to determine the target interaction information between the wafer text data and the wafer image data; According to the target interaction information between the wafer text data and the wafer image data, use the negative sample label method and the alignment method to perform image-text association on the wafer text data and the wafer image data, generating multiple image-text data pairs, where each of the image-text data pairs corresponds to a respective description text; For each of the image-text data pairs, calculate the similarity between the description text and the image in the image-text data pair, and based on the image-text data pair, the description text, and the similarity, generate an image-text pre-training data pair; Generate a target wafer multi-modal dataset according to each of the image-text pre-training data pairs.
8. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the wafer defect detection method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the wafer defect detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Unsupervised wafer defect detection method, device and equipment and storage medium
CN114998330A
Wire rod surface defect detection method based on semi-supervised learning
CN118447322A