Wafer defect detection method, device, equipment and medium

By acquiring and aligning wafer multimodal datasets, using open-world wildcard methods to create pseudo-labels, pre-training and fine-tuning of models, the problem that traditional detection methods are difficult to identify unknown defects is solved, and higher detection accuracy and lower labeling costs are achieved.

CN120031883AActive Publication Date: 2025-05-23深圳市壹倍科技有限公司

Patent Information

Application Number
CN202510513147.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

Traditional wafer defect detection methods are difficult to effectively identify and detect unknown wafer defect types, resulting in low detection accuracy.

Method used

By acquiring the multimodal data set of wafers, using a lightweight alignment method to perform spatial alignment, the target wafer multimodal data set is generated. Then, using the open-world wildcard method to create pseudo-labels for labelless data, generating multiple pseudo-label defect data sets. These data sets are used to pre-train the preset supervised model and optimize the initial detection model through a fine-tuning algorithm to obtain the object detection model.

Benefits of technology

This method can effectively identify and detect unknown types of wafer defects, improve the accuracy of wafer defect detection, reduce the cost of manual labeling of data, and enhance the model's learning ability of different defect characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031883A_ABST
    Figure CN120031883A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of semiconductor defect detection, in particular to a wafer defect detection method, device, equipment and medium, and the method comprises the steps: obtaining a wafer multi-modal data set, carrying out the space alignment of the wafer multi-modal data set based on a lightweight alignment method, and generating a target wafer multi-modal data set; making pseudo labels for the label-free wafer data in the multi-modal data set of the target wafer by using an open world wildcard character method, generating a plurality of pseudo label defect data sets, and pre-training a preset supervision model by using the plurality of pseudo label defect data sets, and performing model fine tuning on the initial detection model obtained after pre-training by using a preset fine tuning algorithm to obtain a target detection model, and inputting the to-be-detected wafer image into the target detection model for defect detection to obtain a wafer defect detection result. Therefore, the obtained target detection model can learn different wafer defect features and identify unknown wafer defect types, and the wafer defect detection precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semiconductor defect detection technology, and in particular to a wafer defect detection method, device, equipment and medium. Background Art

[0002] The rapid development of semiconductor technology has attracted widespread attention in today's society and promoted the continuous progress in the fields of communications, information technology, embedded systems, etc. During the processing of semiconductor chips, many wafer defects will be generated, and the detection and analysis of wafer defects is a key link to ensure product quality and yield. These defects may come from a variety of factors in the production process, such as material problems, improper process parameter settings, or equipment failures.

[0003] However, traditional wafer defect detection methods require a large amount of labeled data to train a relatively stable initial detection model, but obtaining sufficient labeled data may be very difficult and costly, and only a few types of wafer defects can be learned, and unknown types of wafer defects cannot be effectively identified and detected, resulting in low detection accuracy in unknown types of wafer defects. Therefore, how to effectively identify and detect unknown types of wafer defects to improve the detection accuracy of wafer defects is a technical problem that needs to be solved urgently. Summary of the invention

[0004] Based on this, it is necessary to address the above technical problems. The embodiments of the present invention provide a wafer defect detection method, device, equipment and medium, which can effectively identify and detect unknown wafer defect types, thereby improving the detection accuracy of wafer defects.

[0005] A first aspect of an embodiment of the present application provides a wafer defect detection method, the wafer defect detection method comprising: Obtain wafer multimodal dataset; Based on a lightweight alignment method, spatially aligning the wafer multimodal dataset to generate a target wafer multimodal dataset; Using an open-world wildcard method to create pseudo labels for unlabeled wafer data in the target wafer multimodal dataset, generating multiple pseudo-label defect datasets; Pre-training a preset supervision model using the multiple pseudo-label defect data sets, and fine-tuning an initial detection model obtained after pre-training using a preset fine-tuning algorithm to obtain a target detection model; The wafer image to be tested is input into the target detection model for defect detection to obtain a wafer defect detection result.

[0006] A second aspect of an embodiment of the present application provides a wafer defect detection device, the wafer defect detection device comprising: An acquisition module, used for acquiring a wafer multimodal data set; An alignment module, used for spatially aligning the wafer multimodal dataset based on a lightweight alignment method to generate a target wafer multimodal dataset; A generation module, used for making pseudo labels for unlabeled wafer data in the target wafer multimodal data set by using an open-world wildcard method, and generating multiple pseudo-label defect data sets; A training module, used to pre-train a preset supervision model using the multiple pseudo-label defect data sets, and to fine-tune the initial detection model obtained after the pre-training using a preset fine-tuning algorithm to obtain a target detection model; The detection module is used to input the wafer image to be tested into the target detection model for defect detection to obtain a wafer defect detection result.

[0007] In a third aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the wafer defect detection method as described in the first aspect when executing the computer program.

[0008] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the wafer defect detection method as described in the first aspect is implemented.

[0009] In summary, the present invention provides a wafer defect detection method, device, equipment and medium, which obtain a wafer multimodal dataset, spatially align the wafer multimodal dataset based on a lightweight alignment method, generate a target wafer multimodal dataset, use an open world wildcard method to create pseudo labels for unlabeled wafer data in the target wafer multimodal dataset, generate multiple pseudo-label defect datasets, use the multiple pseudo-label defect datasets to pre-train a preset supervision model, and use a preset fine-tuning algorithm to fine-tune the initial detection model obtained after pre-training to obtain a target detection model, input the wafer image to be tested into the target detection model for defect detection, and obtain a wafer defect detection result. It can be seen that the present application utilizes the open-world wildcard method to create pseudo-labels for the unlabeled wafer data in the target wafer multimodal dataset, generates multiple pseudo-label defect datasets, and pre-trains the preset supervision model based on the multiple pseudo-label defect datasets, thereby overcoming the problem of lack of training data and greatly reducing the cost of manually annotating data. The preset fine-tuning algorithm is then used to fine-tune the initial detection model obtained after pre-training, so that the obtained target detection model can learn different wafer defect characteristics, identify unknown wafer defect types, and improve the detection accuracy of wafer defects. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0011] Figure 1 It is a schematic flow chart of a wafer defect detection method provided by an embodiment of the present invention; Figure 2 It is a structural schematic diagram of a wafer defect detection device provided by an embodiment of the present invention; Figure 3 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0012] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technical personnel in this field without creative work are within the scope of protection of the present invention.

[0013] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0014] It should also be understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0015] As used in the present specification and the appended claims, the term "if" can be interpreted as "when" or "uponce" or "in response to a determination" depending on the context. Similarly, the phrase "if it is determined" or "if matched to [described condition or event]" can be interpreted as meaning "uponce determined" or "in response to a determination" or "uponce matched to [described condition or event]" or "in response to matching to [described condition or event]" depending on the context.

[0016] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0017] References to "one embodiment" or "some embodiments" etc. described in the present specification mean that one or more embodiments of the present invention include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0018] It should be understood that the order of execution of the steps in the following embodiments does not imply a precedence of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0019] In order to illustrate the technical solution of the present invention, specific embodiments are provided below for illustration.

[0020] See also Figure 1 , is a flow chart of a wafer defect detection method provided by an embodiment of the present invention, such as Figure 1 As shown, the wafer defect detection method can be implemented by the following steps.

[0021] S101: Acquire a wafer multimodal dataset.

[0022] In step S101, by acquiring a wafer surface defect dataset and various ultra-large-scale multimodal datasets, such as the CC12M dataset, the GQA dataset, the Flickr30K dataset, and the Open Images V7 dataset, etc., the present application does not impose any limitation on this. These datasets cover various types of images and labels, including multimodal data such as visual descriptions of image content, relationship graphs of questions and answers, etc. Among them, the wafer surface defect dataset is collected by deploying high-precision industrial cameras on a semiconductor wafer production line. These cameras can capture defect images of the wafer surface in real time, covering different types of surface defects (such as scratches, cracks, bubbles, etc.); the CC12M (Conceptual 12m) dataset contains 12 million pairs of image-text pairing data, where the image is a two-dimensional image and the text is a natural language description related to the image, usually involving a detailed description of the content, scene or object of the image; the GQA (Graph Question Answering) dataset consists of a series of images and their related questions and answers, focusing on the semantic information in the image and the relationship between image elements. Each image is accompanied by a set of natural language questions designed based on the image content. The questions involve objects, scenes and their relationships in the image, and each question has a corresponding answer; the Flickr30K dataset contains 30,000 images and their corresponding natural language descriptions. Each image is accompanied by 5 different text descriptions, and the descriptions cover multi-dimensional information such as scenes, objects, and activities in the image; the Open lmages V7 dataset contains about 9 million images and is annotated with object detection boxes of up to 6,000 categories. It is mainly suitable for supervised target detection and image segmentation and other deep learning task training.

[0023] Since the wafer surface defect datasets and various ultra-large-scale multimodal datasets obtained above are represented in multiple formats, and there will be inconsistencies in data formats, units, dimensions, etc., in addition, some data have missing values, outliers, and erroneous values, and these problems will affect the training and performance of subsequent models. Therefore, the wafer surface defect datasets and various ultra-large-scale multimodal datasets can be preprocessed, and the data can be unified to the same format and standard to ensure data consistency, so as to integrate the preprocessed wafer surface defect datasets and various ultra-large-scale multimodal datasets to obtain a wafer multimodal dataset.

[0024] In an embodiment of the present application, these data sets are preprocessed and integrated to ensure the consistency of the wafer multimodal data set, so that subsequent models can be trained on a variety of different data sources, overcoming the problem of lack of training data for the detection model and enhancing its generalization ability under unknown defect categories and complex working conditions.

[0025] S102: Based on a lightweight alignment method, spatially align the wafer multimodal dataset to generate a target wafer multimodal dataset.

[0026] In step S102, in multimodal tasks (such as tasks that process images and texts simultaneously), alignment pre-training is a very important step, and its goal is to make the image and the text description associated therewith have similar representations in the embedding space, so that the relationship between the image and the text can be better understood, so that cross-modal understanding and reasoning can be effectively performed in multimodal tasks. In traditional detection models, the alignment of image and text modalities usually relies on complex fusion operations (such as CLIP), which require a lot of computing resources and significantly reduce the efficiency of the model in the reasoning stage. The lightweight alignment method reduces the computing cost when implementing modal alignment while maintaining the performance of the improved model. Therefore, by using the lightweight alignment method, the wafer multimodal dataset is spatially aligned to generate a target wafer multimodal dataset.

[0027] In an embodiment of the invention, the method of spatially aligning the wafer multimodal dataset based on a lightweight alignment method to generate a target wafer multimodal dataset includes: Based on a preset low-rank adaptation method and a multi-scale feature fusion method, the wafer text data and the wafer image data are processed to determine target interaction information between the wafer text data and the wafer image data; According to the target interaction information between the wafer text data and the wafer image data, the wafer text data and the wafer image data are image-text associated by using a negative sample labeling method and an alignment method to generate a plurality of image-text data pairs, wherein each of the image-text data pairs has a corresponding description text; For each of the image-text data pairs, respectively calculate the similarity between the description text and the image in the image-text data pair, and generate an image-text pre-training data pair based on the image-text data pair, the description text and the similarity; A target wafer multimodal dataset is generated based on each of the image-text pre-training data pairs.

[0028] Specifically, the wafer multimodal dataset includes wafer text data and wafer image data. By introducing low-rank adaptation to eliminate complex fusion operations in the CLIP text encoder, the CLIP text encoder originally has some fixed parameters (weights). Now a low-rank matrix is ​​added on the basis of these parameters to process the wafer text data and wafer image data and determine the initial interaction information between the wafer text data and the wafer image data. That is, by using the low-rank adaptation method, a low-rank matrix is ​​introduced in all queries, keys, values, and output projections of the CLIP text encoder. The specific formula is: ; in, Denoted as the pre-trained weights of the CLIP text encoder, It is represented as the product of two low-rank matrices, with input and output x and h respectively. The rank of the low-rank matrix is ​​set to a value much smaller than the model feature dimension. Through the low-rank matrix, the model dynamically stores information related to cross-modal interactions during training, thereby improving the adaptability of the decision boundary. This method ensures that the pre-trained text encoder parameters remain unchanged, while the low-rank matrix dynamically stores information about cross-modal interactions. In the inference stage, the calibrated text embeddings can be pre-computed and stored offline, thereby avoiding the computational cost of the text encoder, which makes the model more efficient during inference. Using the RT-UTER model as an efficient detector, the classification head in the RT-UETR model is adapted through multimodal dual-head matching. In region-text contrastive learning, the region embeddings of the two heads are refined by aligning with the shared, semantically rich text representation, achieving end-to-end training and inference.

[0029] Therefore, after obtaining the initial interaction information between the wafer text data and the wafer image data, although the low-rank adaptation method can realize text-image interaction, there are still key problems such as time synchronization in the current acquired images (such as image drift caused by dynamic process parameter fluctuations), which will lead to a decrease in the quality of multimodal data after alignment. Therefore, it is necessary to use a multi-scale feature fusion method to process the initial interaction information between the wafer text data and the wafer image data, and then obtain the target interaction information between the wafer text data and the wafer image data to extract the wafer texture and frequency domain features and enhance the robustness of spatial alignment. Among them, the multi-scale feature fusion method includes the grayscale co-occurrence matrix (GLCM) and the Fourier spectrum cross-correlation method. The grayscale co-occurrence matrix is ​​a method for characterizing texture features by statistically analyzing the grayscale spatial distribution of pixel pairs in an image. In our semiconductor defect detection products, it is mainly used to identify texture anomalies such as scratches on the wafer surface, thereby assisting in label establishment. Fourier spectrum cross-correlation is a defect detection method based on frequency domain analysis. In our semiconductor defect detection products, it is mainly used for periodic contamination anomalies such as particles on the wafer surface, thereby assisting in label establishment. It can be seen that accurate target interaction information can be determined to improve the accuracy of wafer defect detection and classification, reduce false detection and missed detection, and improve the quality and efficiency of semiconductor manufacturing. Then, according to the target interaction information between wafer text data and wafer image data, the negative sample label method and alignment method are used to associate the wafer text data and wafer image data with images to generate multiple image-text data pairs, where each image-text data pair has its own corresponding description text. In multimodal learning, especially in the task of image and text alignment, it is often encountered that images and texts are not strictly one-to-one corresponding. For example, an image may contain multiple objects, and the text description may only cover part of them, or a text description may apply to multiple images. This non-strict correspondence will cause noise in the traditional global instance-level alignment target (i.e., assuming that each image and text pair is one-to-one corresponding), thereby affecting the training effect of the model. Adjustment is made by softening the negative sample label method, and the adjustment process is divided into several stages: a. Initial stage (all 1 labels): In the early stages of training, the labels of negative samples are set to all 1, which means that the model will regard all negative samples as equally important as positive samples in the initial stage, thus avoiding strict distinction of negative samples too early.

[0030] b. Intermediate stage (softened label): As training progresses, the labels of negative samples gradually transition from all 1 to softened labels. The values ​​of softened labels are between 0 and 1, indicating that the similarity between negative samples and positive samples gradually decreases. This softening process can help the model gradually distinguish between negative and positive samples while reducing the impact of noise.

[0031] c. Later stages (softer labels): In the later stages of training, the labels of negative samples are further softened and close to 0, which means that the model's ability to distinguish negative samples is gradually enhanced and it can more accurately identify the difference between negative and positive samples.

[0032] The alignment method is divided into multi-head alignment method and single-head alignment method. In region-text contrastive learning, a consistent double alignment strategy is used to make the decision boundaries of the two classification heads more consistent. The specific formula is expressed as: ; Among them, u represents the loU value between the predicted box and the true box, α and β represent the classification head, and s represents the classification score obtained through multimodal information. The calculation formula of the classification score is: ; Among them, sim represents cosine similarity, T represents text embedding, and I represents the pixel-level features of the image. In order to ensure that the supervision signals of the two heads in multimodal dual-head matching are consistent, a consistent setting is adopted, which is specifically expressed as follows: .

[0033] This enables the one-to-one head to effectively learn supervisory signals consistent with the one-to-many heads, thereby improving the performance of the model. The core idea of ​​the one-to-one strategy is to assign a unique text category to each image region, ensuring a one-to-one correspondence between image features and text features.

[0034] Furthermore, if there is only one drop defect on a detection sub-image in our scenario, then the one-to-one reasoning efficiency is higher. The core idea of ​​the one-to-many strategy is to allow one image region to match multiple text categories. This strategy is more flexible and can handle complex scenarios, but the computational cost is relatively high. If there is one dirt defect on a detection sub-image in our scenario, dirt is usually composed of multiple particle defects, then the one-to-many reasoning efficiency is higher. That is, for each image-text data pair, by introducing the open source CLIP model, the CLIP model is first trained on a large image with a caption to learn the representation of images and texts in the joint embedding space. Based on the assumption that "in this space, the distance between the image embedding and its corresponding title embedding is close, while the distance between the unrelated image embedding and title embedding is farther", the CLIP model can extract text from the image, and can compare the obtained text with the given text to generate a semantic relevance score between the image and the text. According to the semantic relevance score between the image and the text, the similarity between the description text and the image in the image-text data pair is calculated respectively, and based on the image-text data pair, the description text and the similarity, the image-text pre-training data pair is generated, so as to generate the target wafer multimodal data set according to each of the image-text pre-training data pairs. It can be seen that by reducing the feature dimension through the low-rank adaptation method, the target interaction information between the wafer text data and the image data can be effectively captured. By introducing negative samples (unmatched text and image pairs), the model can be helped to better learn the discriminative features between text and image, and can automatically generate high-quality image-text data pairs, reducing the dependence on a large amount of manually annotated data, significantly reducing the cost and time of data annotation, while ensuring the quality of the data.

[0035] In this embodiment, a lightweight alignment method is used to spatially align wafer multimodal datasets to generate a target wafer multimodal dataset, thereby reducing computational costs during multimodal data alignment, thereby subsequently improving the efficiency and accuracy of wafer defect detection, and reducing the cost of data annotation and model training.

[0036] S103: Using an open-world wildcard method to create pseudo labels for the unlabeled wafer data in the target wafer multimodal dataset, and generating multiple pseudo-label defect datasets.

[0037] In step S103, for the wafer defect detection scenario, we do not want the model to have a low confidence score for the defect category, but at the same time we want to be able to expand the detection of defects outside the defined categories. Open-world wildcards are designed to enable the model to detect objects that do not exist in the predefined vocabulary and mark them as "unknown". This method is achieved by using a wildcard embedding that can capture unknown objects in the scene in zero-shot situations. Specifically, all wildcard embeddings are initialized from text features of a general text (such as "object") extracted by a calibrated text encoder. At this point, all unknown objects will be preliminarily mapped to a general category, providing a basis for subsequent learning. Since the target wafer multimodal dataset contains a small amount of labeled wafer inspection data and a large amount of pseudo-labeled wafer inspection data, the wildcard embedding is fine-tuned using the pre-training dataset, treating all real instances as the same-"object" category. This fine-tuning enables the embedding to capture richer semantic information, thereby enhancing the model's ability to identify objects that are not covered by predefined specific categories, so as to distinguish wafer inspection data of different defect types, thereby making pseudo labels for a large amount of unlabeled wafer inspection data, so as to generate multiple pseudo-labeled defect datasets. And for some ultra-small defects (5x5, pixel units), a super-resolution network can be used to increase the defect size, thereby improving the accuracy of the label box.

[0038] In one embodiment of the invention, an open-world wildcard method is used to create pseudo labels for unlabeled wafer data in a target wafer multimodal data set, generating multiple pseudo-label defect data sets, including: Based on an open-world wildcard approach, semantic defect information between modalities is learned from the target wafer multimodal dataset; According to the semantic defect information, using multiple autoencoders to respectively predict labels of unlabeled wafer data in the target wafer multimodal data set to obtain multiple prediction results; Determine whether the multiple prediction results meet the preset IoU threshold and score threshold; If the multiple prediction results meet the preset IoU threshold and score threshold, the pseudo labels corresponding to the unlabeled wafer data are determined according to the multiple prediction results to generate multiple pseudo-label defect data sets.

[0039] Specifically, by utilizing the open-world wildcard method, semantic defect information between modalities is learned from the target wafer multimodal dataset, that is, the wildcard embedding is initialized from the text features of the general text (such as "defect"). This method allows the model to understand and infer the characteristics of defects when facing unseen defect types. The features extracted by the calibrated text encoder ensure that the wildcard embedding is not only universal but also captures the semantic information of potential defects. This information may include key features such as the category, location, and size of wafer defects. The open-world wildcard method may involve the use of wildcards to represent the potential association or similarity between different modalities in the wafer data, which may require the use of multimodal information fusion techniques, such as physical layer fusion, feature layer fusion, decision layer fusion, etc.

[0040] Furthermore, after determining the semantic defect information, multiple autoencoders are trained for the semantic defect information. Each autoencoder is responsible for extracting features from a specific data modality and attempting to reconstruct or predict labels. The training of the autoencoder can be based on unsupervised learning or self-supervised learning methods, using labeled data (if any) for fine-tuning. The unlabeled wafer data is input into the trained multiple autoencoders to obtain multiple prediction results. Each autoencoder outputs a prediction about the wafer defect, including the location, category and other information of the defect. Calculate the intersection over union (IoU) and score (usually the confidence or probability of the prediction result) of each prediction result, compare the calculated IoU and score with the preset IoU threshold and score threshold, if the IoU and score of a prediction result meet the threshold requirements, then the prediction result is considered valid, and select pseudo labels by setting the IoU threshold (o1=0.5) and score threshold (o2=0.01) for training the "unknown" wildcard. For multiple prediction results that meet the threshold requirements, determine the pseudo labels corresponding to the unlabeled wafer data according to the multiple prediction results and the prediction accuracy corresponding to the respective encoders, and generate multiple pseudo-label defect data sets. It can be seen that by making pseudo labels for unlabeled wafer data, the cost and time of manual labeling can be reduced, so that these unlabeled data can be fully utilized to train the detection model in the future, thereby improving the accuracy of the prediction, and by setting the IoU threshold and score threshold, the prediction results can be flexibly controlled and screened, thereby ensuring the quality of the generated pseudo-label data set.

[0041] In this embodiment, by using the open-world wildcard method to create pseudo-labels for the unlabeled wafer data in the target wafer multimodal dataset, the labeling cost can be significantly reduced so that the detection model subsequently trained with multiple pseudo-label defect datasets can still understand and infer the characteristics of the defects when faced with unseen defect types, thereby improving the efficiency of wafer defect detection.

[0042] S104: Pre-training a preset supervision model using the multiple pseudo-label defect data sets, and fine-tuning an initial detection model obtained after the pre-training using a preset fine-tuning algorithm to obtain a target detection model.

[0043] In step S104, an open-world wildcard method is used to create pseudo labels for unlabeled wafer data in the target wafer multimodal data set. After generating multiple pseudo-label defect data sets, the generated pseudo-label defect data sets are merged with the data sets of real labels to form a labeled defect data set. A suitable supervised learning model is selected as a preset model. The model should have the ability to handle wafer detection tasks. The preset supervised model is pre-trained using the labeled defect data set. During the pre-training process, the model will learn the feature representation in the labeled defect data set, including real labeled data and pseudo-label data. A preset fine-tuning algorithm is selected. The algorithm should be suitable for target detection tasks and can effectively utilize pseudo-label data. The fine-tuning algorithm is applied to the initial detection model obtained after pre-training to obtain a target detection model. During the fine-tuning process, the model will be further optimized for specific wafer detection tasks to improve detection performance.

[0044] In one embodiment of the invention, based on an open world wildcard method, a preset supervision model is pre-trained according to the target wafer multimodal dataset, including: Performing data processing on the multiple pseudo-label defect data sets to obtain multiple processed pseudo-label defect data sets; The processed multiple pseudo-label defect data sets and the real label defect data set are merged to form a target label defect data set.

[0045] The detection model is iteratively trained according to the target label defect data set until the trained detection model meets the expected performance, thereby obtaining an initial detection model.

[0046] Specifically, multiple pseudo-label defect data sets are processed to obtain multiple processed pseudo-label defect data sets, and the processed pseudo-label defect data sets and the real label defect data sets are spliced ​​to form a target label defect data set. When splicing, it is necessary to ensure the consistency of the format and label of the data set. In order to avoid the overfitting of the model to a certain type of sample, the target label defect data set can be balanced. For example, through oversampling or undersampling technology, the number of samples of different categories is kept balanced. Select a supervised learning model suitable for wafer inspection tasks, such as convolutional neural network (CNN), support vector machine (SVM), etc., initialize the parameters of the model, and use random initialization or pre-trained weights. The model is iteratively trained using the target label defect data set. In each iteration, the data set is divided into a training set and a validation set, and the model parameters are updated using the training set. The model performance is evaluated using the validation set. According to the performance feedback of the validation set, the model parameters and training strategy are adjusted until the trained detection model meets the expected performance. When the performance of the detection model reaches the preset standard, it is used as the initial detection model. The initial detection model can be used for subsequent wafer defect detection tasks and can be further optimized and improved according to actual needs. It can be seen that by merging the processed pseudo-label defect dataset and the real label defect dataset, the size of the dataset is further increased and the training effect of the model is improved, so that the trained initial detection model can better learn the characteristics and laws of different wafer defects, thereby enhancing the generalization ability of the model and further improving the detection capability and robustness of unknown defect categories.

[0047] In one embodiment of the invention, data processing is performed on a plurality of pseudo-label defect data sets to obtain a plurality of processed pseudo-label defect data sets, including: Repeating filtering on the multiple pseudo-label defect data sets to obtain multiple filtered pseudo-label defect data sets; Performing vocabulary expansion processing on the filtered multiple pseudo-label defect data sets to obtain multiple expanded pseudo-label defect data sets; Based on a preset confidence threshold, the confidence information of the expanded multiple pseudo-label defect data sets is judged to obtain multiple judgment results; In response to the multiple judgment results respectively meeting the first judgment condition, the confidence information in the pseudo-label defect data set corresponding to the judgment results meeting the condition is adjusted to obtain a plurality of processed pseudo-label defect data sets.

[0048] Specifically, hashing algorithms, similarity calculations (such as cosine similarity, Jaccard similarity, etc.) or clustering algorithms (such as K-means, DBSCAN, etc.) are used to filter and duplicate multiple pseudo-label defect data sets, identify and remove duplicate data items, and obtain multiple filtered pseudo-label defect data sets, where each data set is deduplicated, that is, during the inference process, a simple unknown filtering strategy is used to remove unknown category predictions that are highly overlapped with known category predictions (loU threshold τ=0.99) to reduce duplication. This strategy helps to reduce false detections and redundant predictions, ensuring that the detection results of each "unknown" category have sufficient discrimination and credibility. Through this step, the model can be prevented from generating too many duplicate detections, optimizing inference efficiency and accuracy. Based on the existing label vocabulary, the filtered multiple pseudo-label defect data sets are expanded using methods such as synonym replacement, near-synonym expansion, and context reasoning to obtain multiple expanded pseudo-label defect data sets, each of which contains a more diverse label vocabulary, that is, by discovering new categories from "unknown" category predictions and adding their class names to the vocabulary, known categories are provided for the next iteration, thereby achieving dynamic vocabulary expansion. The dynamically expanded vocabulary not only improves the generalization ability of the model, but also enables the model to continuously adapt to newly emerging defect categories, with stronger adaptability and long-term application capabilities. Each newly added category is evaluated and verified to ensure its accuracy and representativeness, and avoid the addition of irrelevant categories.

[0049] Furthermore, by evaluating the confidence of each data item according to a preset confidence threshold, a plurality of judgment results are obtained, each of which indicates whether the confidence of the corresponding data item meets the threshold requirement, wherein the confidence can be a probability, a similarity score or other metric based on model prediction. In response to the judgment result meeting the first judgment condition (i.e., the confidence does not meet the threshold requirement), the confidence information in the corresponding pseudo-label defect data group is adjusted, and the adjustment method can be to recalculate the confidence, use a more accurate prediction model or introduce external verification data. Among them, the first judgment condition can be that the confidence is lower than a specific value, or the confidence is too different from the confidence of other data items, etc., and the present application does not make any limitation on this, thereby obtaining a plurality of processed pseudo-label defect data sets, wherein the confidence of each data set has been appropriately adjusted. It can be seen that by filtering duplicate processing and vocabulary expansion processing, redundant data can be removed and the richness of label vocabulary can be increased, thereby improving the quality and diversity of the data set. By processing and adjusting the pseudo-label defective data set, the dependence on manual labeling can be reduced, and the labeling cost and time can be reduced. By setting the confidence threshold and adjusting the data items that do not meet the requirements, the confidence of each data item in the data set can be ensured to be more accurate and reliable.

[0050] In one embodiment of the invention, a preset fine-tuning algorithm is used to fine-tune the initial detection model obtained after pre-training to obtain a target detection model, including: Pre-selecting a portion of the pseudo-label defect data set from the processed multiple pseudo-label defect data sets, and manually annotating the portion of the pseudo-label defect data set to obtain a manually annotated data result; Using the preset similarity method, the manually labeled data results are matched with another part of the pseudo-label defect data set to obtain the label matching results; According to the label matching results, an efficient parameter fine-tuning strategy is used to fine-tune the initial detection model obtained after pre-training to obtain a target detection model.

[0051] Specifically, the defect categories have been annotated with pseudo labels based on the wildcard method, that is, all defects are collectively referred to as "objects". In order to avoid the negative impact of errors in pseudo labels on the training process, a threshold is set, and multiple prediction results with confidence levels higher than the threshold are used as candidate pseudo labels, rather than only selecting the label with the highest confidence level. This method can effectively reduce the deviation caused by a single high-confidence prediction and increase the diversity of pseudo labels. Therefore, in order to further ensure the correctness of defect classification, a part of the pseudo-label defect data set is selected from the processed multiple pseudo-label defect data sets based on factors such as data quality and diversity. The selected pseudo-label defect data sets are manually checked and annotated to ensure the accuracy and consistency of the annotation results, and the manually annotated data results are obtained. This part of the data will be used as the "gold standard" for subsequent matching and model fine-tuning. The manually annotated data results are matched with another part of the pseudo-label defect data set using a preset similarity method (such as cosine similarity, Euclidean distance, Jaccard similarity, etc.) to obtain a label matching result, which includes a pseudo-label defect data set that is similar or related to the manually annotated data result.

[0052] Furthermore, by adopting efficient parameter fine-tuning strategies (such as Fine-tuning, weight transfer, feature extraction, etc.), the matched pseudo-label defect data set is used as training data, and the initial detection model obtained after pre-training is fine-tuned to obtain the target detection model. For example, 20% of the labels are randomly selected from all pseudo-labels for manual labeling. During the manual labeling process, each label box is reviewed by at least 2 experts to ensure the accuracy and consistency of the label. In addition, the 20% manually labeled labels are matched with the remaining 80% of the pseudo-label information through the cosine similarity method to ensure that all pseudo-labels have unique and accurate label information. This step not only improves the quality of pseudo-labels, but also further enhances the reliability of labels through the expert review mechanism. In the model fine-tuning stage, considering that only the defect category information is modified without affecting the defect feature information, the RT-UTER classification head and the network layer above it are frozen. The AdamW optimizer was used, the learning rate was set to 0.0002, and the model was fine-tuned 50,000 times. This efficient parameter fine-tuning strategy not only reduced the consumption of computing resources, but also avoided over-adjustment of learned features by freezing specific layers, thereby further optimizing the performance of defect classification while maintaining the generalization ability of the model. It can be seen that through manual labeling and matching screening, the labeling cost and time can be significantly reduced, the data quality and labeling accuracy used for model fine-tuning can be ensured, thereby improving the accuracy of the target detection model and further improving the model's detection ability for specific wafer defects.

[0053] In this embodiment, by using multiple pseudo-label defect data sets to pre-train the preset supervision model, the problem of lack of model training data is overcome, so that the model can learn more wafer defect characteristics and representations, providing a good foundation for the subsequent fine-tuning stage. At the same time, the initial detection model obtained after pre-training is fine-tuned using a preset fine-tuning algorithm, which can further improve the accuracy and robustness of the model. It can be seen that by combining pre-training and model fine-tuning, the dependence on a large amount of labeled data can be reduced while maintaining the model performance, thereby reducing the labeling cost and time.

[0054] S105: Inputting the wafer image to be tested into the target detection model for defect detection to obtain a wafer defect detection result.

[0055] In step S105, defect detection is mainly implemented based on the target detection model under the actual application scenario data. By inputting the wafer image to be tested into the target detection model for defect detection, the wafer defect detection result is obtained, which includes the wafer defect type judgment result (category label), wafer defect position (bounding box coordinates) and corresponding confidence scores, etc. The detection results output by the model are post-processed, which may include non-maximum suppression (NMS) and other technologies to filter duplicate detection frames. Then evaluate the accuracy of the detection results, such as calculating accuracy, recall rate and other indicators. If the corresponding indicators are met, the detection results are visualized and stored for subsequent analysis or tracking.

[0056] In one embodiment of the invention, inputting the wafer image to be tested into the target detection model for defect detection to obtain a wafer defect detection result includes: After performing angle correction on the wafer image to be tested, cropping the wafer image to be tested based on a set picture size to obtain a cropped wafer image to be tested; Performing Gaussian filtering on the cropped wafer image to obtain a filtered wafer image to be tested; Performing image enhancement processing on the filtered wafer image to obtain an enhanced wafer image to be tested; The enhanced wafer image to be tested is input into the target detection model for forward reasoning to obtain a wafer defect detection result.

[0057] Specifically, image processing techniques (such as Hough transform, edge detection, etc.) are used to identify specific patterns or edges in the wafer image, so as to determine the main axis direction of the wafer. The image is then rotated according to the main axis direction so that the wafer image reaches a standard angle (such as horizontal or vertical). According to a preset wafer image size (such as width and height), the angle-corrected wafer image to be tested is cropped to remove unnecessary parts of the image. A Gaussian filter is applied to the cropped wafer image to smooth the image and reduce noise to obtain a filtered wafer image to be tested. The filtered wafer image to be tested is subjected to image enhancement processing (such as contrast stretching, histogram equalization, etc.) to obtain an enhanced wafer image to be tested. The enhanced wafer image to be tested is input into a target detection model for forward reasoning to obtain wafer defect detection results. That is, load the trained target detection model and fix its model weights, turn off gradient derivation, and then perform angle correction on the wafer image to be tested collected by the camera and crop it to the model input size as required to obtain a sub-image. Then, perform Gaussian filtering, image enhancement and other pre-processing on the sub-image. Finally, input the processed sub-image into the target detection model for forward reasoning. Through forward reasoning, the model will analyze the defective area in the image based on the learned feature representation and label information, and generate defect detection results. It can be seen that through pre-processing steps such as angle correction, cropping, Gaussian filtering and image enhancement, the accuracy and efficiency of wafer defect detection can be effectively improved, which helps to enhance the generalization ability of deep learning models.

[0058] In this embodiment, by inputting the image of the wafer to be tested into the target detection model for defect detection, the problem of lack of training data for the wafer epitaxial surface defect detection model is overcome, the computational complexity of data processing is reduced, and defects outside the defined defect categories can be detected, so as to identify multiple types of defects, thereby improving the efficiency and accuracy of wafer defect detection.

[0059] In summary, the present invention provides a wafer defect detection method, device, equipment and medium, which obtain a wafer multimodal dataset, spatially align the wafer multimodal dataset based on a lightweight alignment method, generate a target wafer multimodal dataset, use an open world wildcard method to create pseudo labels for unlabeled wafer data in the target wafer multimodal dataset, generate multiple pseudo-label defect datasets, use the multiple pseudo-label defect datasets to pre-train a preset supervision model, and use a preset fine-tuning algorithm to fine-tune the initial detection model obtained after pre-training to obtain a target detection model, input the wafer image to be tested into the target detection model for defect detection, and obtain a wafer defect detection result. It can be seen that the present application utilizes the open-world wildcard method to create pseudo-labels for the unlabeled wafer data in the target wafer multimodal dataset, generates multiple pseudo-label defect datasets, and pre-trains the preset supervision model based on the multiple pseudo-label defect datasets, thereby overcoming the problem of lack of training data and greatly reducing the cost of manually annotating data. The preset fine-tuning algorithm is then used to fine-tune the initial detection model obtained after pre-training, so that the obtained target detection model can learn different wafer defect characteristics, identify unknown wafer defect types, and improve the detection accuracy of wafer defects.

[0060] See also Figure 2 , Figure 2 is a schematic diagram of the structure of a wafer defect detection device provided by an embodiment of the present invention. The wafer defect detection device corresponds one-to-one to the wafer defect detection method in the above embodiment. Figure 1 as well as Figure 1 For the convenience of explanation, only the parts related to this embodiment are shown. Figure 2 The wafer defect detection device 20 includes: an acquisition module 21, an alignment module 22, a generation module 23, a training module 24, and a detection module 25.

[0061] An acquisition module 21 is used to acquire a wafer multimodal data set; An alignment module 22, configured to perform spatial alignment on the wafer multimodal dataset based on a lightweight alignment method to generate a target wafer multimodal dataset; A generating module 23 is used to generate a plurality of pseudo-label defect data sets by using an open-world wildcard method to generate pseudo-labels for the unlabeled wafer data in the target wafer multimodal data set; A training module 24 is used to pre-train a preset supervision model using the multiple pseudo-label defect data sets, and to fine-tune the initial detection model obtained after the pre-training using a preset fine-tuning algorithm to obtain a target detection model; The detection module 25 is used to input the image of the wafer to be tested into the target detection model to perform defect detection to obtain a wafer defect detection result.

[0062] Optionally, the alignment module 22 is specifically used for: Based on a preset low-rank adaptation method and a multi-scale feature fusion method, the wafer text data and the wafer image data are processed to determine target interaction information between the wafer text data and the wafer image data; According to the target interaction information between the wafer text data and the wafer image data, the wafer text data and the wafer image data are image-text associated by using a negative sample labeling method and an alignment method to generate a plurality of image-text data pairs, wherein each of the image-text data pairs has a corresponding description text; For each of the image-text data pairs, respectively calculate the similarity between the description text and the image in the image-text data pair, and generate an image-text pre-training data pair based on the image-text data pair, the description text and the similarity; A target wafer multimodal dataset is generated based on each of the image-text pre-training data pairs.

[0063] Optionally, the generating module 23 is specifically used for: Based on an open-world wildcard approach, semantic defect information between modalities is learned from the target wafer multimodal dataset; According to the semantic defect information, using multiple autoencoders to respectively predict labels of unlabeled wafer data in the target wafer multimodal data set to obtain multiple prediction results; Determine whether the multiple prediction results meet the preset IoU threshold and score threshold; If the multiple prediction results meet the preset IoU threshold and score threshold, the pseudo labels corresponding to the unlabeled wafer data are determined according to the multiple prediction results to generate multiple pseudo-label defect data sets.

[0064] Optionally, the training module 24 is specifically used for: Performing data processing on the multiple pseudo-label defect data sets to obtain multiple processed pseudo-label defect data sets; The processed multiple pseudo-label defect data sets and the real label defect data set are merged to form a target label defect data set.

[0065] The detection model is iteratively trained according to the target label defect data set until the trained detection model meets the expected performance, thereby obtaining an initial detection model.

[0066] Optionally, the training module 24 is further used for: Repeating filtering on the multiple pseudo-label defect data sets to obtain multiple filtered pseudo-label defect data sets; Performing vocabulary expansion processing on the filtered multiple pseudo-label defect data sets to obtain multiple expanded pseudo-label defect data sets; Based on a preset confidence threshold, the confidence information of the expanded multiple pseudo-label defect data sets is judged to obtain multiple judgment results; In response to the multiple judgment results respectively meeting the first judgment condition, the confidence information in the pseudo-label defect data set corresponding to the judgment results meeting the condition is adjusted to obtain a plurality of processed pseudo-label defect data sets.

[0067] Optionally, the training module 24 is further used for: Pre-selecting a portion of the pseudo-label defect data sets from the processed multiple pseudo-label defect data sets, and manually annotating the portion of the pseudo-label defect data sets to obtain manually annotated data results; Using the preset similarity method, the manually annotated data results are matched with another part of the pseudo-label defect data set to obtain the label matching results; According to the label matching results, an efficient parameter fine-tuning strategy is used to fine-tune the initial detection model obtained after pre-training to obtain a target detection model.

[0068] Optionally, the detection module 25 is specifically used for: After performing angle correction on the wafer image to be tested, cropping the wafer image to be tested based on a set picture size to obtain a cropped wafer image to be tested; Performing Gaussian filtering on the cropped wafer image to obtain a filtered wafer image to be tested; Performing image enhancement processing on the filtered wafer image to obtain an enhanced wafer image to be tested; The enhanced wafer image to be tested is input into the target detection model for forward reasoning to obtain a wafer defect detection result.

[0069] It should be noted that the information interaction, execution process and other contents between the above-mentioned units are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0070] Figure 3 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 3 As shown, the electronic device of this embodiment includes: at least one processor ( Figure 3Only one is shown), a memory, and a computer program stored in the memory and executable on at least one processor, wherein when the processor executes the computer program, the steps in any of the above-mentioned wafer defect detection method embodiments are implemented.

[0071] The electronic device may include, but is not limited to, a processor and a memory. It can be understood by those skilled in the art that Figure 3 These are merely examples of electronic devices and do not constitute limitations on the electronic devices. The electronic devices may include more or fewer components than those shown in the figures, or a combination of certain components, or different components. For example, they may also include a network interface, a display screen, and an input system.

[0072] In one embodiment, a computer-readable storage medium is provided, and when the instructions in the computer-readable storage medium are executed by a processor in an electronic device, the electronic device can perform the steps of any embodiment of a wafer defect detection method disclosed in the present invention, which will not be repeated here. The computer-readable storage medium can be non-volatile or volatile.

[0073] The processor may be a CPU, or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0074] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory may be the memory of an electronic device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be a hard disk of an electronic device, and in other embodiments, it may also be an external storage device of the electronic device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Further, the memory may also include both an internal storage unit of the electronic device and an external storage device. The memory is used to store an operating system, a cooperative application, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program, etc. The memory may also be used to temporarily store data that has been output or is to be output.

[0075] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0076] The technical business in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0077] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A wafer defect detection method, characterized in that: include: Obtain wafer multimodal dataset; Based on a lightweight alignment method, spatially aligning the wafer multimodal dataset to generate a target wafer multimodal dataset; Using an open-world wildcard method to create pseudo labels for unlabeled wafer data in the target wafer multimodal dataset, generating multiple pseudo-label defect datasets; Pre-training a preset supervision model using the multiple pseudo-label defect data sets, and fine-tuning an initial detection model obtained after pre-training using a preset fine-tuning algorithm to obtain a target detection model; The wafer image to be tested is input into the target detection model for defect detection to obtain a wafer defect detection result.

2. The wafer defect detection method according to claim 1, characterized in that: The wafer multimodal dataset includes wafer text data and wafer image data. The method of spatially aligning the wafer multimodal dataset based on a lightweight alignment method to generate a target wafer multimodal dataset includes: Based on a preset low-rank adaptation method and a multi-scale feature fusion method, the wafer text data and the wafer image data are processed to determine target interaction information between the wafer text data and the wafer image data; According to the target interaction information between the wafer text data and the wafer image data, the wafer text data and the wafer image data are image-text associated by using a negative sample labeling method and an alignment method to generate a plurality of image-text data pairs, wherein each of the image-text data pairs has a corresponding description text; For each of the image-text data pairs, respectively calculate the similarity between the description text and the image in the image-text data pair, and generate an image-text pre-training data pair based on the image-text data pair, the description text and the similarity; A target wafer multimodal dataset is generated based on each of the image-text pre-training data pairs.

3. The wafer defect detection method according to claim 1, characterized in that: The method of using the open world wildcard method to create pseudo labels for the unlabeled wafer data in the target wafer multimodal data set to generate multiple pseudo label defect data sets includes: Based on an open-world wildcard approach, semantic defect information between modalities is learned from the target wafer multimodal dataset; According to the semantic defect information, using multiple autoencoders to respectively predict labels of unlabeled wafer data in the target wafer multimodal data set to obtain multiple prediction results; Determine whether the multiple prediction results meet the preset IoU threshold and score threshold; If the multiple prediction results meet the preset IoU threshold and score threshold, the pseudo labels corresponding to the unlabeled wafer data are determined according to the multiple prediction results to generate multiple pseudo-label defect data sets.

4. The wafer defect detection method according to claim 1, characterized in that: The pre-training of a preset supervision model using the multiple pseudo-label defect data sets includes: Performing data processing on the multiple pseudo-label defect data sets to obtain multiple processed pseudo-label defect data sets; Merging the processed multiple pseudo-label defect data sets and the real label defect data set to form a target label defect data set; The detection model is iteratively trained according to the target label defect data set until the trained detection model meets the expected performance, thereby obtaining an initial detection model.

5. The wafer defect detection method according to claim 4, characterized in that: The performing data processing on the multiple pseudo-label defect data sets to obtain multiple processed pseudo-label defect data sets includes: Repeating filtering on the multiple pseudo-label defect data sets to obtain multiple filtered pseudo-label defect data sets; Performing vocabulary expansion processing on the filtered multiple pseudo-label defect data sets to obtain multiple expanded pseudo-label defect data sets; Based on a preset confidence threshold, the confidence information of the expanded multiple pseudo-label defect data sets is judged to obtain multiple judgment results; In response to the multiple judgment results respectively meeting the first judgment condition, the confidence information in the pseudo-label defect data set corresponding to the judgment results meeting the condition is adjusted to obtain a plurality of processed pseudo-label defect data sets.

6. The wafer defect detection method according to claim 1, characterized in that: The method of fine-tuning the initial detection model obtained after pre-training by using a preset fine-tuning algorithm to obtain a target detection model includes: Pre-selecting a portion of the pseudo-label defect data set from the processed multiple pseudo-label defect data sets, and manually annotating the portion of the pseudo-label defect data set to obtain a manually annotated data result; Using the preset similarity method, the manually labeled data results are matched with another part of the pseudo-label defect data set to obtain the label matching results; According to the label matching results, an efficient parameter fine-tuning strategy is used to fine-tune the initial detection model obtained after pre-training to obtain a target detection model.

7. The wafer defect detection method according to claim 1, characterized in that: The step of inputting the wafer image to be tested into the target detection model for defect detection to obtain a wafer defect detection result includes: After performing angle correction on the wafer image to be tested, cropping the wafer image to be tested based on a set picture size to obtain a cropped wafer image to be tested; Performing Gaussian filtering on the cropped wafer image to obtain a filtered wafer image to be tested; Performing image enhancement processing on the filtered wafer image to obtain an enhanced wafer image to be tested; The enhanced wafer image to be tested is input into the target detection model for forward reasoning to obtain a wafer defect detection result.

8. A wafer defect detection device, characterized in that: include: An acquisition module, used for acquiring a wafer multimodal data set; An alignment module, used for spatially aligning the wafer multimodal dataset based on a lightweight alignment method to generate a target wafer multimodal dataset; A generation module, used for making pseudo labels for unlabeled wafer data in the target wafer multimodal data set by using an open-world wildcard method, and generating multiple pseudo-label defect data sets; A training module, used to pre-train a preset supervision model using the multiple pseudo-label defect data sets, and to fine-tune the initial detection model obtained after the pre-training using a preset fine-tuning algorithm to obtain a target detection model; The detection module is used to input the wafer image to be tested into the target detection model for defect detection to obtain a wafer defect detection result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the wafer defect detection method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the wafer defect detection method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Unsupervised wafer defect detection method, device and equipment and storage medium

    CN114998330A

  • Wire rod surface defect detection method based on semi-supervised learning

    CN118447322A

Cited By

  • Self-supervised defect detection method based on thermodynamic diagram pseudo defects

    CN120931608A

  • Wafer defect detection method and device based on self-supervised learning algorithm

    CN121414709A

  • Chip end face crack detection method, system and equipment based on multi-mode large model fine tuning and medium

    CN121563948A

  • Placental lesion tissue detection method and placental lesion tissue detection system based on artificial intelligence

    CN121563999A

  • Wafer yield analysis model screening method and system, medium and program product

    CN121743806A