Semiconductor wafer defect detection method, device and equipment and storage medium

By combining the preprocessing and classification model of semiconductor wafer images with class activation mapping technology, the thermal map is generated to identify defect areas, solving the problem of high manual labeling cost in the existing methods, and achieving efficient and accurate defect detection.

CN120259191APending Publication Date: 2025-07-04SKYVERSE TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510257812.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the existing semiconductor wafer defect detection methods, the object detection algorithm requires complicated manual labeling of data, which leads to high cost and inconvenient maintenance.

Method used

Preprocessing technology is used to preprocess wafer images, and heat maps are generated using pre-trained classification models and class activation mapping methods to identify and label defect areas to reduce dependence on manual annotation.

Benefits of technology

Improve detection efficiency, reduce manual labeling costs, and simplify model maintenance process while maintaining high detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259191A_ABST
    Figure CN120259191A_ABST
Patent Text Reader

Abstract

The invention discloses a semiconductor wafer defect detection method, device and equipment and a storage medium, and the method comprises the steps: preprocessing an obtained to-be-detected image of a semiconductor wafer, and obtaining an input image corresponding to the to-be-detected image; inputting the input image into the classification model to obtain all target defect categories reaching a confidence threshold and confidence values; generating a thermodynamic diagram corresponding to each target defect category based on a category activation mapping mode; respectively processing all the thermodynamic diagrams to obtain an effective area with defects on the input image; target areas on the to-be-detected image are determined based on the effective areas, and after target defect categories and confidence values corresponding to the effective areas are marked for each target area, the marked to-be-detected image is output. According to the method, the target detection effect is achieved through the classification model and the class activation mapping mode to replace a target detection algorithm, so that the target detection algorithm does not need to be trained by tedious data labeling, and the manual labeling cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of wafer detection, and particularly to a semiconductor wafer defect detection method, device, equipment and storage medium. Background Art

[0002] Wafer defect detection is a crucial part in semiconductor manufacturing. Wafer defect detection technology helps to improve the yield of chip manufacturing processes, increase production efficiency and reduce production costs. During the production process, wafers are prone to various defects such as particles, scratches, cracks, contamination, pits, protrusions, etc. These defects will affect the functions and performance of chips and even cause chip failures. Therefore, it is very necessary to effectively detect and analyze wafer defects.

[0003] Due to the complex production process of semiconductor wafers, there are also many types of defects generated (generally up to hundreds), and the characteristics of different defects are often not very clear, resulting in the inability of quality inspectors to handle manual rejudgment effectively. At the same time, due to the subjectivity and fatigue-prone characteristics of people, the final rejudgment effect is not very good. Therefore, methods based on deep learning are often used for automatic discrimination.

[0004] Currently, object detection algorithms are mainly used in the field of semiconductor wafer defect detection. This algorithm can effectively identify the categories of defects and provide relatively accurate positioning. However, such algorithms must use specific tools to perform complex manual annotation on training data, which also brings great inconvenience to the subsequent maintenance of the algorithm and directly affects the popularization of such detection algorithms. Summary of the Invention

[0005] In view of this, the present application provides a semiconductor wafer defect detection method, device, equipment and storage medium to solve the problems of high cost and inconvenient maintenance of manual annotation data in the existing object detection algorithm during the automatic detection process of wafers.

[0006] To solve the above technical problems, a technical solution adopted by the present application is: to provide a semiconductor wafer defect detection method, which includes: preprocessing the to-be-detected image of the semiconductor wafer obtained to obtain an input image corresponding to the to-be-detected image; inputting the input image into a pre-trained classification model to obtain all target defect categories reaching the confidence threshold and the confidence value corresponding to each target defect category; generating a heat map corresponding to each target defect category based on the class activation mapping method; respectively processing all heat maps to obtain the effective regions where defects exist on the input image; determining the target regions on the to-be-detected image based on the effective regions, and after labeling each target region with the target defect category and confidence value corresponding to the effective region, outputting the labeled to-be-detected image.

[0007] As a further improvement of the present application, preprocess the acquired image to be measured of the semiconductor wafer to obtain an input image corresponding to the image to be measured, including: obtaining a template image corresponding to the image to be measured of the semiconductor wafer, where the template image is a defect-free image stored in advance and corresponding to the image to be measured; superimposing the image to be measured and the template image to generate an input image.

[0008] As a further improvement of the present application, preprocess the acquired image to be measured of the semiconductor wafer to obtain an input image corresponding to the image to be measured, including: grayscale the image to be measured of the semiconductor wafer to obtain a grayscale image of the image to be measured; obtain a template image corresponding to the image to be measured, where the template image is a defect-free image stored in advance and corresponding to the image to be measured; calculate the absolute difference between the grayscale image and the template image to generate a difference image; perform denoising processing on the difference image, and then superimpose it with the image to be measured to generate an input image.

[0009] As a further improvement of the present application, generate a heat map corresponding to each target defect category based on the class activation mapping method, including: obtaining the feature map of a pre-selected target layer in the classification model; calculating the gradient of each channel of the feature map based on the confidence value corresponding to each target defect category; performing global average pooling operation on the gradient in the spatial dimension to obtain the weight value of each channel; performing weighted summation of the feature map and the weight value, and using an activation function for activation to obtain a heat map.

[0010] As a further improvement of the present application, the classification model is constructed based on a convolutional neural network, and the target layer includes the last convolutional layer of the classification model.

[0011] As a further improvement of the present application, process all heat maps respectively to obtain the effective regions with defects on the input image, including: performing binarization processing on the heat map to obtain a binarized image, where the binarized image includes at least one significant region; performing morphological operations on the binarized image; based on the pre-set filtering parameters corresponding to each defect category, combine the parameters of each significant region to filter the significant regions in the binarized image after morphological operations to obtain effective regions.

[0012] As a further improvement of the present application, the class activation mapping method includes any one of the CAM algorithm, Grad-CAM algorithm, Grad-CAM++ algorithm, Score-CAM algorithm, Layer-CAM algorithm, and Ablation-CAM algorithm.

[0013] To solve the above technical problems, another technical solution adopted by this application is: to provide a semiconductor wafer defect detection device, which includes: a preprocessing module for preprocessing the to-be-tested image of the semiconductor wafer obtained to obtain an input image corresponding to the to-be-tested image; a prediction module for inputting the input image into a pre-trained classification model to obtain all target defect categories reaching the confidence threshold and the confidence value corresponding to each target defect category; a generation module for generating a heat map corresponding to each target defect category based on the class activation mapping method; a processing module for processing all the heat maps respectively to obtain the effective regions with defects on the input image; and an output module for respectively bounding all the effective regions on the to-be-tested image, labeling the corresponding target defect category and confidence value for each effective region, and then outputting the labeled to-be-tested image.

[0014] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer device, which includes a processor and a memory coupled to the processor. Program instructions are stored in the memory. When the program instructions are executed by the processor, the processor is caused to execute the steps of the semiconductor wafer defect detection method as described in any one of the above.

[0015] To solve the above technical problems, another technical solution adopted by this application is: to provide a storage medium storing program instructions capable of implementing the semiconductor wafer defect detection method as described in any one of the above.

[0016] The beneficial effects of this application are as follows: The semiconductor wafer defect detection method of this application classifies the to-be-tested image of the semiconductor wafer by using a pre-trained classification model to confirm whether there are defects in the to-be-tested image. When there are defects, the heat map corresponding to each defect category is generated in combination with the class activation mapping method, the position where the defect is located is confirmed from the heat map, and then the defect is bounded on the to-be-tested image and the category and confidence of the defect are labeled, and then the labeled to-be-tested image is output. It can achieve the detection effect of the target detection algorithm by using the classification model plus the class activation mapping method, without going through the complicated data annotation work of the target detection algorithm, greatly improving the work efficiency, and also bringing great convenience to the later maintenance of the model. In addition, implementing the class activation mapping technology has less modification to the structure of the classification model based on deep learning, has a lower impact on the inference efficiency of the classification model, and does not occupy too much system resources. Description of the Drawings

[0017] Figure 1 is a flowchart of the semiconductor wafer defect detection method according to an embodiment of the present invention;

[0018] Figure 2 is a schematic diagram of the to-be-tested image and the heat map of the semiconductor wafer defect detection method according to an embodiment of the present invention;

[0019] Figure 3 is another schematic flowchart of the semiconductor wafer defect detection method according to an embodiment of the present invention;

[0020] Figure 4 is a schematic diagram of a to-be-detected image and a template image of the semiconductor wafer defect detection method according to an embodiment of the present invention;

[0021] Figure 5 is yet another schematic flowchart of the semiconductor wafer defect detection method according to an embodiment of the present invention;

[0022] Figure 6 is a schematic diagram of functional modules of the semiconductor wafer defect detection device according to an embodiment of the present invention;

[0023] Figure 7 is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0024] Figure 8 is a schematic diagram of the structure of a storage medium according to an embodiment of the present invention. Detailed Embodiments

[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of the present application.

[0026] The terms "first", "second", and "third" in the present application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined. All directional indications (such as up, down, left, right, front, back...) in the embodiments of the present application are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0027] References to "embodiments" in this specification mean that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0028] Figure 1 is a schematic flowchart of a semiconductor wafer defect detection method according to an embodiment of the present invention. It should be noted that if there are substantially the same results, the method of the present invention is not limited to Figure 1 the process sequence shown. As Figure 1 shown, the semiconductor wafer defect detection method includes the steps of:

[0029] Step S1: Preprocess the to-be-tested image of the semiconductor wafer to obtain an input image corresponding to the to-be-tested image.

[0030] Specifically, after obtaining the to-be-tested image of the semiconductor wafer, preprocess the to-be-tested image to obtain a preprocessed image. The preprocessing process includes but is not limited to data cleaning and denoising, size adjustment, image cropping and rotation, image pixel value normalization, standardization of the image dataset, image color adjustment, etc. By preprocessing the to-be-tested image, the prediction performance and stability of the classification model can be improved.

[0031] Step S2: Input the input image into a pre-trained classification model to obtain all target defect categories that reach the confidence threshold, and the confidence value corresponding to each target defect category.

[0032] It should be noted that the classification model includes but is not limited to a convolutional neural network, and other classification algorithms capable of classifying defect categories in the to-be-tested image can also be used. In this embodiment, the classification model is pre-trained, and the training process is as follows:

[0033] 1. Data collection and annotation: According to the classification problem to be solved, collect relevant data from various channels as data samples, and then label each data sample with a corresponding category label. The data samples include an image set of semiconductor wafers, and the image set includes normal images and images with various defects. The category labels include all preset defect categories. Finally, the collected data is divided into a training set, a validation set, and a test set.

[0034] 2. Input the training set data into the classification model to be trained. Calculate the prediction results of the model through forward propagation, and then calculate the loss value based on the prediction results and the true labels. Then, use the optimization algorithm to calculate the gradients of the loss function with respect to the model parameters through backpropagation, and update the model parameters according to the gradients. Repeat the above process until the loss function converges or reaches the preset number of training epochs.

[0035] 3. Use the validation set to evaluate the performance of the model during training. Through the evaluation on the validation set, it can be known whether the model is overfitting or underfitting, and then adjust the hyperparameters of the model according to the evaluation results of the validation set.

[0036] 4. Use the test set to perform a final evaluation on the tuned model to obtain the true performance of the model on unseen data.

[0037] Specifically, after obtaining the input image, input the input image into the pre-trained classification model for prediction to obtain the prediction results. The prediction results include the target defect categories whose prediction scores exceed the score threshold. The prediction score is the confidence value, the score threshold is the confidence threshold, and the prediction score is used to represent the probability that the target defect category exists in the input image.

[0038] Furthermore, it should be noted that compared with existing open-source datasets, the dataset for wafer defect detection has the following three important differences:

[0039] 1. Wafer defects are sensitive to defect size. That is, for the same defect morphology, if the defect size is greater than a certain threshold, it is judged as a defect, and if it is less than this threshold, it is judged as acceptable, that is, no defect. And the function related to scaling commonly used in model training under the PyTorch framework is RandomResizedCrop(). This function randomly crops a smaller region from the input image and scales this region to the specified size. The key parameter of this function is scale, which controls the ratio of the area of the transformed image to the area of the original image, set as ( β is generally taken as 1.0). Assuming that both the original image and the transformed image are square, and the side lengths are m and n respectively, the actual scaling range of the image side length is Assuming that after testing, it is found that the safe side length scaling range is [γ, δ], then the corresponding Substitute to obtain the scaling coefficient of the RandomResizedCrop() function, so as to control the image scaling within a controllable range.

[0040] 2. Wafer defects are sensitive to color. For example, color deviation defects (such as oxidation defects) will show some color abnormalities. At this time, it is necessary to pay attention to controlling the amplitude of image color transformation during model training (especially the hue value). In the pytorch framework, the commonly used color-related function during model training is ColorJitter(), which increases the generalization ability of the model by randomly adjusting the brightness, contrast, saturation, and hue of the image. The hue factor in the function is randomly sampled from [-hue, hue]. It is necessary to determine the maximum hue value within the controllable color range before and after the transformation through experiments. Let it be ρ, then the transformation range of the hue factor is [-ρ, v].

[0041] 3. Wafer defects are sensitive to position. Wafer defects will preferentially appear in the middle area of the image. Therefore, it is necessary to control the amplitude of operations such as image translation during model training. In the pytorch framework, the commonly used position-related function during model training is Translate(), which is used to move all pixels in the image horizontally and vertically according to the specified pixel values. Assuming that the input size of the model is square and the side length is N, then when using the Translate() function, the safe movement range in the horizontal and vertical directions is [-N / 2, N / 2]. Additionally, considering that the defect position during model training may also be affected by other transformation operations, it is recommended to introduce a safety factor ε (ε ∈ [0, 1.0], and generally ε ∈ [0, 0.70]) on the basis of the above range. Therefore, when finally using the Translate() function, the movement range in the horizontal and vertical directions is

[0042] Based on the above differences, the heatmaps generated by adopting the class activation mapping method are highly adaptable to the size, color, and position sensitivity of wafer defects, as follows:

[0043] 1. Size sensitivity and the dynamic focusing ability of the heatmap. The determination of wafer defects highly depends on whether the defect size exceeds the preset threshold. Traditional object detection requires manual annotation of the defect boundary to identify the size, while the heatmap dynamically focuses on the region that contributes the most to classification through the activation intensity of the feature layer. As Figure 2 shown, Figure 2In (a), it represents the image to be measured, and in (b), it represents the heat map. The significant area (the darkest part in color) of the heat map directly reflects the core area of the defect, regardless of the defect size. When the defect approaches the size threshold, the coverage range of the heat map can be accurately extracted through morphological processing, and combined with area filtering conditions (such as the area of the minimum bounding rectangle), it can automatically determine whether it exceeds the limit, avoiding the subjective error of manual annotation. In addition, during the training stage, the scaling range is controlled by the RandomResizedCrop() function (such as the scaling factor is limited to [γ, ε]), ensuring the sensitivity of the model to defects of different sizes, and the generation process of the heat map naturally inherits this scaling constraint, so as to accurately adapt to the size criterion during the inference stage.

[0044] 2. Color sensitivity and color correlation of the heat map. Wafer defects (such as oxidation defects) are often accompanied by color abnormalities. Traditional methods need to rely on complex color enhancement strategies or manual experience for judgment. The generation mechanism of the heat map is deeply coupled with color feature extraction: the model automatically learns local patterns of color abnormalities (such as hue shift) through convolutional layers, and the heat map visualizes these patterns through gradient weighting. For example, in the ColorJitter() function, the hue_factor is restricted to [-v, ρ]. While suppressing irrelevant color interference, the significant area of the heat map can still accurately capture subtle color differences (such as Figure 2 the boxed area shown in (b) in the figure). This heat response mechanism based on color sensitivity enables the model to achieve defect localization without relying on additional color annotations, significantly superior to the bounding box annotation method of traditional object detection.

[0045] 3. Position sensitivity and spatial attention guidance of the heat map. Wafer defects are mostly concentrated in the central area of the image. Traditional methods need to set translation limits (such as the translation range is set by the Translate() function to be ) to prevent defects from moving out of the field of view. As Figure 2 shown in the example, the significant area of the heat map is concentrated in the center of the image, which is consistent with the actual distribution law of wafer defects. In addition, the generation process of the heat map and the translation constraint work together to ensure that even if there is a slight position offset, the model can still correct the positioning result through attention weights, thereby improving the detection robustness.

[0046] In summary, using the heat map can dynamically combine the feature extraction ability of the classification model with the defect criterion: in terms of size determination, it is automatically completed through the area of the heat map region without bounding box annotation; in terms of color abnormality, it is directly visualized through the gradient response of the heat map, avoiding reliance on manual experience; in terms of position sensitivity, it is double-verified through the heat map and the spatial attention module to ensure accurate positioning. Therefore, the present invention improves the detection of wafer defects through heat map technology to improve the final detection accuracy.

[0047] Step S3: Generate a heatmap corresponding to each target defect category based on the class activation mapping method.

[0048] It should be noted that Class Activation Mapping (CAM for short) is an interpretability technique that can generate a heatmap to show the regions in the image that play a key role in the prediction of a specific category.

[0049] In this embodiment, the class activation mapping method includes any one of the CAM algorithm, Grad-CAM algorithm, Grad-CAM++ algorithm, Score-CAM algorithm, Layer-CAM algorithm, and Ablation-CAM algorithm.

[0050] Further, step S3 specifically includes:

[0051] 1. Obtain the feature map of the target layer pre-selected in the classification model.

[0052] In this embodiment, the classification model is constructed based on a Convolutional Neural Network (CNN), and the target layer includes the last convolutional layer of the classification model. Specifically, in a convolutional neural network, as the network depth increases, the features extracted by the convolutional layer gradually transition from low-level simple features such as edges and textures to high-level semantic features. For example, in an image classification task, shallower convolutional layers may detect basic elements such as lines and corners in the image, while deeper convolutional layers, especially the last convolutional layer, can capture high-level semantic information closely related to the category. For instance, in a cat-dog classification, the last convolutional layer may extract discriminative complex feature combinations such as cat faces and dog ears. Moreover, the feature map of the last convolutional layer integrates the global and local information of the image to a certain extent. It retains the feature details of each local region in the image and, through the sliding of the convolutional kernel and the combination of multiple convolutional layers, integrates these local information to form an abstract representation of the entire image. This fusion of global and local features enables the CAM generated based on the last convolutional layer to accurately locate the regions in the image that play a key role in classification, rather than being limited to small local features.

[0053] 2. Calculate the gradient of each channel of the feature map based on the confidence value corresponding to each target defect category.

[0054] Specifically, in order to calculate the gradient of each channel of the target layer feature map with respect to the confidence value of the target category, the backpropagation algorithm needs to be used. First, create a one-hot encoding vector with the same size as the model output, set the value at the position corresponding to the target defect category to 1, and set the values at the remaining positions to 0. Then, use this one-hot encoding vector as the gradient parameter for the backpropagation operation. During the backpropagation process, calculate the gradients from the output layer to each layer, including the gradient of the feature map of the target layer with respect to the confidence value of the target defect category.

[0055] 3. Perform global average pooling operation on the gradient in the spatial dimension to obtain the weight value of each channel.

[0056] Specifically, for each channel, calculate the average value of all elements in the gradient map of this channel. For example, for the gradient map of the i-th channel, add up all the height*width (height and width refer to the spatial dimensions) elements in this map, and then divide by height*width to obtain a scalar value. This scalar value is the weight value of the i-th channel, which reflects the relative importance of this channel in the classification decision of the target defect category. In this way, a weight value is calculated for each channel of the feature map.

[0057] 4. Perform weighted summation of the feature map and the weight value, and use an activation function for activation to obtain a heat map.

[0058] Specifically, for each channel of the feature map, multiply all the elements of this channel by the corresponding weight value, and then add up the results of all channels to obtain a two-dimensional tensor, which reflects the comprehensive importance of different positions in the image for the target defect category. In this embodiment, the activation function uses the ReLU (Rectified Linear Unit) function. The ReLU function sets all negative values to 0 and only retains positive values, which can enhance the regions that make positive contributions to the target defect category and suppress the regions that make small or negative contributions to the target defect category. After ReLU activation, the obtained two-dimensional tensor is the heat map.

[0059] Step S4: Process all the heat maps respectively to obtain the effective regions with defects on the input image.

[0060] Specifically, after obtaining the heat map corresponding to each target defect category, analyze the heat map to confirm the effective region corresponding to the target defect category in the heat map and the coordinate information of the effective region.

[0061] Furthermore, step S4 specifically includes:

[0062] 1. Binarize the heat map to obtain a binarized image, which includes at least one significant region.

[0063] Specifically, after obtaining the heat map, binarizing it can convert the image into a binary image containing only two colors (usually black and white), thereby highlighting the significant regions. When performing binarization, first determine the binarization threshold. Pixels with values higher than this binarization threshold will be set to a fixed value (usually 255, representing white), while pixels lower than this threshold will be set to another fixed value (usually 0, representing black). Then, based on the selected threshold, judge each pixel of the heat map and convert it to 0 or 255. In the binarized image, the white region (pixel value of 255) usually represents the significant region.

[0064] 2. Perform morphological operations on the binarized image.

[0065] Specifically, morphological operations include erosion, dilation, opening operation, closing operation, etc. By performing morphological operations, the noise points in the binarized image are removed.

[0066] 3. Based on the filtering parameters corresponding to each preset defect category, combined with the parameters of each significant region, filter the significant regions in the binarized image after morphological operations to obtain effective regions.

[0067] Specifically, when screening the significant regions, each filtering parameter is related to a specific defect category and needs to be preset in advance. The parameters of the significant regions usually include information such as the area, length, and width of the significant regions. Based on the filtering parameters corresponding to each target defect category, filter the significant regions to obtain effective regions that meet the requirements.

[0068] Step S5: Based on the effective regions, determine the target regions on the image to be measured, and after labeling each target region with the corresponding target defect category and confidence value of the effective region, output the labeled image to be measured.

[0069] It should be noted that in this embodiment, the size of the obtained heat map is the same as the size of the image to be measured, and the coordinate positions of the effective regions on the heat map are the same as those on the image to be measured. Specifically, after obtaining the effective regions in the heat map, calculate the minimum bounding rectangle of each effective region, and then use this minimum bounding rectangle to frame the corresponding positions on the image to be measured. Mark the corresponding target defect category and confidence value on the minimum bounding rectangle frame, and then output the labeled image to be measured. The user can then see the defective regions, the types of defects, and the credibility from the image to be measured.

[0070] The semiconductor wafer defect detection method of this embodiment classifies the image to be measured of the semiconductor wafer by using a pre-trained classification model to confirm whether there are defects in the image to be measured. When there are defects, the class activation mapping method is combined to generate a heat map corresponding to each defect category. The location of the defect is confirmed from the heat map, and then the defect is framed in the image to be measured and the category and confidence of the defect are marked, and then the marked image to be measured is output. It can achieve the detection effect of the target detection algorithm by using the classification model plus the class activation mapping method, without going through the complicated data annotation work of the target detection algorithm, greatly improving the work efficiency, and also bringing great convenience to the later maintenance of the model. In addition, the implementation of the class activation mapping technology has little modification to the structure of the classification model based on deep learning, has a low impact on the inference efficiency of the classification model, and does not occupy too many system resources.

[0071] Further, in order to improve the detection accuracy of the classification model for defects on the wafer, in some embodiments, as Figure 3 shown, step S1 specifically includes:

[0072] Step S101: Obtain a template image corresponding to the image to be measured of the semiconductor wafer. The template image is a defect-free image stored in advance corresponding to the image to be measured.

[0073] Specifically, in this embodiment, the template image needs to be obtained and stored in advance. Please refer to Figure 4 shown, Figure 4 In (a), it represents the image to be measured, and in (b), it represents the template image. There are defects in the image to be measured (the circular black block circled by the red dotted line in the figure), and there are no defects in the template image. When there is no template image as a reference, this defect is likely to be misidentified as a normal bump structure, resulting in missed detection.

[0074] Step S102: Superimpose the image to be measured and the template image to generate an input image.

[0075] Specifically, the image to be measured is a three-channel image. The image to be measured and the template image are superimposed to form a four-channel image, and this four-channel image is used as the input image.

[0076] This embodiment further enriches the information of the input image of the classification model by introducing a template image corresponding to the image to be measured, thereby improving the detection ability of the classification model and further enhancing the positioning accuracy of the subsequent generated heat map.

[0077] Further, in order to improve the detection accuracy of the classification model for defects on the wafer, in some embodiments, as Figure 5 shown, step S1 specifically includes:

[0078] Step S111: Grayscale the image to be measured of the semiconductor wafer to obtain a grayscale image of the image to be measured.

[0079] Step S112: Obtain a template image corresponding to the image to be measured. The template image is a defect-free image corresponding to the image to be measured and stored in advance.

[0080] Specifically, in this embodiment, the template image needs to be obtained and stored in advance. Please refer to Figure 4 as shown Figure 4 In (a) represents the image to be measured, and (b) represents the template image. There are defects in the image to be measured (the circular black block circled by the dotted line in the figure), while there are no defects in the template image. When there is no template image as a reference, this defect is likely to be misidentified as a normal bump structure, resulting in missed detection.

[0081] Step S113: Calculate the absolute difference between the grayscale image and the template image to generate a difference image.

[0082] Step S114: Denoise the difference image, and then superimpose it on the image to be measured to generate an input image.

[0083] Specifically, the image to be measured is a three-channel image. The image to be measured and the difference image are superimposed to form a four-channel image, and this four-channel image is used as the input image.

[0084] In this embodiment, by introducing a template image corresponding to the image to be measured, the information of the input image of the classification model is further enriched, thereby improving the detection ability of the classification model, and further improving the positioning accuracy of the subsequent generated heat map.

[0085] Figure 6 is a schematic diagram of the functional modules of the semiconductor wafer defect detection device according to an embodiment of the present invention. As Figure 6 shown, the semiconductor wafer defect detection device 20 includes: a preprocessing module 21, a prediction module 22, a generation module 23, a processing module 24, and an output module 25.

[0086] The preprocessing module 21 is configured to preprocess the image to be measured of the semiconductor wafer obtained to obtain an input image corresponding to the image to be measured;

[0087] The prediction module 22 is configured to input the input image into a pre-trained classification model to obtain all target defect categories that reach the confidence threshold, and the confidence value corresponding to each target defect category;

[0088] The generation module 23 is configured to generate a heat map corresponding to each target defect category based on the class activation mapping method;

[0089] The processing module 24 is used to process all the heat maps respectively to obtain the effective regions with defects on the input image;

[0090] The output module 25 is used to determine the target regions on the image to be tested based on the effective regions, and after labeling each target region with the corresponding target defect category and confidence value of the effective region, output the labeled image to be tested.

[0091] Optionally, the operation of the preprocessing module 21 to perform preprocessing on the image to be tested of the semiconductor wafer to obtain the input image corresponding to the image to be tested specifically includes: obtaining a template image corresponding to the image to be tested of the semiconductor wafer, where the template image is a defect-free image corresponding to the image to be tested stored in advance; superimposing the image to be tested and the template image to generate the input image.

[0092] Optionally, the operation of the preprocessing module 21 to perform preprocessing on the image to be tested of the semiconductor wafer to obtain the input image corresponding to the image to be tested specifically includes: graying the image to be tested of the semiconductor wafer to obtain the grayscale image of the image to be tested; obtaining a template image corresponding to the image to be tested, where the template image is a defect-free image corresponding to the image to be tested stored in advance; calculating the absolute difference between the grayscale image and the template image to generate a difference image; performing denoising processing on the difference image, and then superimposing it with the image to be tested to generate the input image.

[0093] Optionally, the operation of the generating module 23 to generate heat maps corresponding to each target defect category based on the class activation mapping specifically includes: obtaining the feature map of the pre-selected target layer in the classification model; calculating the gradient of each channel of the feature map based on the confidence value corresponding to each target defect category; performing global average pooling operation on the gradient in the spatial dimension to obtain the weight value of each channel; performing weighted summation of the feature map and the weight value, and using an activation function for activation to obtain the heat map.

[0094] Optionally, the classification model is constructed based on a convolutional neural network, and the target layer includes the last convolutional layer of the classification model.

[0095] Optionally, the operation of the processing module 24 to process all the heat maps respectively to obtain the effective regions with defects on the input image specifically includes: performing binarization processing on the heat map to obtain a binarized image, where the binarized image includes at least one significant region; performing morphological operations on the binarized image; filtering the significant regions in the binarized image after morphological operations based on the pre-set filtering parameters corresponding to each defect category, combined with the parameters of each significant region, to obtain the effective regions.

[0096] Optionally, the class activation mapping method includes any one of the CAM algorithm, Grad-CAM algorithm, Grad-CAM++ algorithm, Score-CAM algorithm, Layer-CAM algorithm, and Ablation-CAM algorithm.

[0097] For other details of the implementation technical solutions of each module in the semiconductor wafer defect detection device in the above embodiments, reference can be made to the description in the semiconductor wafer defect detection method in the above embodiments, which will not be elaborated here.

[0098] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0099] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of the computer device according to an embodiment of the present invention. As Figure 7 shown, the computer device 30 includes a processor 31 and a memory 32 coupled to the processor 31. Program instructions are stored in the memory 32. When the program instructions are executed by the processor 31, the processor 31 executes the steps of the semiconductor wafer defect detection method described in any of the above embodiments.

[0100] Among them, the processor 31 can also be referred to as a resource (Central Processing Unit, CPU). The processor 31 may be an integrated circuit chip with signal processing capabilities. The processor 31 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0101] Refer to Figure 8 , Figure 8Schematic diagram of the structure of the storage medium according to an embodiment of the present invention. The storage medium according to the embodiment of the present invention stores program instructions 41 capable of implementing the above semiconductor wafer defect detection method. Among them, the program instructions 41 can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or computer devices such as computers, servers, mobile phones, and tablets.

[0102] In several embodiments provided in the present application, it should be understood that the disclosed computer devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.

[0103] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. The above is only the embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A semiconductor wafer defect detection method, characterized in that, It includes: Preprocessing the to-be-tested image of the semiconductor wafer to obtain an input image corresponding to the to-be-tested image; Inputting the input image into a pre-trained classification model to obtain all target defect categories reaching the confidence threshold and the confidence value corresponding to each target defect category; Generating a heat map corresponding to each target defect category based on the class activation mapping method; Processing all the heat maps respectively to obtain the effective regions with defects on the input image; Determining the target regions on the to-be-tested image based on the effective regions, and outputting the labeled to-be-tested image after labeling each target region with the corresponding target defect category and confidence value of the effective region.

2. The semiconductor wafer defect detection method according to claim 1, wherein The preprocessing of obtaining the to-be-tested image of the semiconductor wafer to obtain an input image corresponding to the to-be-tested image includes: Obtaining a template image corresponding to the to-be-tested image of the semiconductor wafer, where the template image is a defect-free image stored in advance corresponding to the to-be-tested image; Overlaying the to-be-tested image and the template image to generate the input image.

3. The semiconductor wafer defect detection method according to claim 1, wherein The preprocessing of obtaining the to-be-tested image of the semiconductor wafer to obtain an input image corresponding to the to-be-tested image includes: Gray-scaling the to-be-tested image of the semiconductor wafer to obtain the gray-scale image of the to-be-tested image; Obtaining a template image corresponding to the to-be-tested image, where the template image is a defect-free image stored in advance corresponding to the to-be-tested image; Calculating the absolute difference between the gray-scale image and the template image to generate a difference image; Performing denoising processing on the difference image and then overlaying it with the to-be-tested image to generate the input image.

4. The semiconductor wafer defect detection method according to claim 1, wherein, The generating of a heat map corresponding to each target defect category based on the class activation mapping method includes: Obtaining the feature map of a pre-selected target layer in the classification model; Calculating the gradient of each channel of the feature map based on the confidence value corresponding to each target defect category; Performing global average pooling operation on the gradient in the spatial dimension to obtain the weight value of each channel; Performing weighted summation on the feature map and the weight value and activating it using an activation function to obtain the heat map.

5. The semiconductor wafer defect detection method according to claim 4, characterized in that, The classification model is constructed based on a convolutional neural network, and the target layer includes the last convolutional layer of the classification model.

6. The semiconductor wafer defect detection method according to claim 1, characterized in that, The processing of all the heat maps respectively to obtain the effective regions with defects on the input image includes: Performing binarization processing on the heat map to obtain a binarized image, where the binarized image includes at least one significant region; Performing morphological operations on the binarized image; Filtering the significant regions in the binarized image after morphological operations based on the filtering parameters corresponding to each defect category set in advance and combining the parameters of each significant region to obtain the effective regions.

7. The semiconductor wafer defect detection method according to claim 1, characterized in that, The class activation mapping method includes any one of the CAM algorithm, Grad-CAM algorithm, Grad-CAM++ algorithm, Score-CAM algorithm, Layer-CAM algorithm, and Ablation-CAM algorithm.

8. A semiconductor wafer defect detection device, characterized in that, It includes: A preprocessing module, configured to preprocess a to-be-tested image of a semiconductor wafer obtained, to obtain an input image corresponding to the to-be-tested image; A prediction module, configured to input the input image into a pre-trained classification model, to obtain all target defect categories reaching a confidence threshold, and a confidence value corresponding to each target defect category; A generation module, configured to generate a heat map corresponding to each target defect category based on a class activation mapping method; A processing module, configured to process all the heat maps respectively, to obtain valid regions with defects on the input image; An output module, configured to determine a target region on the to-be-tested image based on the valid regions, and after labeling each target region with a target defect category and a confidence value corresponding to the valid region, output the labeled to-be-tested image.

9. A computer device, characterized in that, The computer device includes a processor and a memory coupled to the processor, and program instructions are stored in the memory. When the program instructions are executed by the processor, the processor executes the steps of the semiconductor wafer defect detection method according to any one of claims 1-7.

10. A storage medium, characterized in that, Program instructions capable of implementing the semiconductor wafer defect detection method according to any one of claims 1-7 are stored.

Citation Information

Cited By

  • Mask-based wafer defect classification system and method

    CN120431093A

  • Chip batch detection system and method

    CN120807517A

  • A chip batch detection system and method

    CN120807517B