Wafer defect detection method, electronic device, storage medium and product

CN122199540BActive Publication Date: 2026-09-29SHENZHEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610661807.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-14
Publication Date
2026-09-29
Estimated Expiration
2046-05-14

AI Technical Summary

Technical Problem

[0004]本申请的主要目的在于提供一种晶圆缺陷检测方法、电子设备、存储介质与产品,旨在解决晶圆表面缺陷检测效果差的技术问题

Benefits of technology

[0016]本申请提供了一种晶圆缺陷检测方法,晶圆缺陷检测方法包括:获取待检测晶圆表面缺陷图像;基于预设特征提取网络对待检测晶圆表面缺陷图像进行特征提取,输出待检测晶圆表面缺陷图像对应的多尺度特征;基于预设解码器对多尺度特征进行融合,得到融合特征,基于融合特征并行进行频域增强处理以及防爆注意力处理得到输出特征;基于输出特征生成包含预设缺陷类型的晶圆缺陷分割图像。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122199540B_ABST
    Figure CN122199540B_ABST
Patent Text Reader

Abstract

The application discloses a wafer defect detection method, an electronic device, a storage medium and a product, relates to the technical field of computer vision and deep learning, and comprises the following steps: acquiring a wafer surface defect image to be detected; performing feature extraction on the wafer surface defect image to be detected based on a preset feature extraction network, and outputting multi-scale features corresponding to the wafer surface defect image to be detected; fusing the multi-scale features based on a preset decoder to obtain fused features; performing frequency domain enhancement processing and explosion-proof attention processing in parallel based on the fused features to obtain output features; and generating a wafer defect segmentation image containing a preset defect type based on the output features. The application solves the technical problem of poor wafer surface defect detection effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer vision and deep learning technology, and in particular to a wafer defect detection method, electronic device, storage medium and product. Background Technology

[0002] With the continuous advancement of semiconductor manufacturing processes, accurate detection of wafer surface defects is crucial for improving chip yield and process stability. Currently, traditional non-destructive testing methods, such as manual visual inspection and machine vision, struggle to handle the complex periodic textures and noise interference in the wafer background. Therefore, traditional non-destructive testing methods suffer from poor wafer surface defect detection performance.

[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this application is to provide a wafer defect detection method, electronic device, storage medium, and product, aiming to solve the technical problem of poor wafer surface defect detection performance.

[0005] To achieve the above objectives, this application proposes a wafer defect detection method, which includes: Acquire images of surface defects on the wafer to be inspected; Based on a preset feature extraction network, feature extraction is performed on the surface defect image of the wafer to be detected, and multi-scale features corresponding to the surface defect image of the wafer to be detected are output. The multi-scale features are fused based on a preset decoder to obtain fused features. Frequency domain enhancement processing and explosion-proof attention processing are performed in parallel based on the fused features to obtain output features. Based on the output features, a wafer defect segmentation image containing preset defect types is generated.

[0006] In one embodiment, the wafer defect detection method further includes: Obtain a preset set of wafer surface defect images. For any preset wafer surface defect image in the preset set of wafer surface defect images, use the preset wafer surface defect image and its corresponding image label as training samples. The training samples are input into a preset pyramid network for forward propagation to obtain the layer features corresponding to the output of each network layer of the preset pyramid network. The preset pyramid network includes shallow layers and deep layers. The weight parameters of the shallow layers remain unchanged, while the weight parameters of the deep layers are updated during the training process of the preset pyramid network. The network loss is calculated based on the hierarchical features corresponding to each network layer and the image label, and the weight parameters of the deeper layers of the network are updated in reverse based on the network loss until the preset pyramid network meets the preset training conditions, thus obtaining the preset feature extraction network.

[0007] In one embodiment, the step of acquiring a preset set of wafer surface defect images includes: Obtain a preset basic defect image set, wherein the preset basic defect image set includes basic defect images of each preset defect type, and each preset defect type includes at least one of center circle defect, edge defect, random defect and scratch defect; For any basic defect image in the preset basic defect image set, perform a random horizontal flip operation, a random vertical flip operation, and a random rotation operation with a preset probability on the basic defect image to obtain an initial transformed image. The initial transformation images containing different preset defect types are merged to obtain a preset wafer surface defect image, and the preset defect types contained in the preset wafer surface defect image are recorded, wherein the preset wafer surface defect image contains at least one preset defect type; The preset wafer surface defect images are merged into a preset wafer surface defect image set.

[0008] In one embodiment, the step of fusing the multi-scale features based on a preset decoder to obtain fused features, and performing frequency domain enhancement processing and explosion-proof attention processing in parallel based on the fused features to obtain output features includes: The multi-scale features are input into a preset decoder, and the multi-scale features are fused based on the first fusion module and the second fusion module in the preset decoder to obtain fused features. The first fusion module is composed of a preset residual direct connection channel and a series channel connected in parallel. The series channel is composed of a first preset convolutional layer, a first preset layer normalization layer and a first preset linear rectified activation layer connected in series. The second fusion module is composed of a second preset convolutional layer, a second preset layer normalization layer and a second preset linear rectified activation layer stacked alternately. The fused features are input in parallel into the wavelet transform branch and the attention branch of the preset decoder. In the wavelet transform branch, the fused features are subjected to frequency domain enhancement processing to obtain frequency domain enhanced features. In the attention branch, the fused features are subjected to explosion-proof attention processing to obtain spatial attention features. The frequency domain enhancement feature is fused with the spatial attention feature to obtain the output feature.

[0009] In one embodiment, the step of performing frequency domain enhancement processing on the fused features in the wavelet transform branch to obtain frequency domain enhanced features includes: A two-dimensional discrete wavelet transform is performed on the fused features to obtain low-frequency components and high-frequency components in each preset orientation, wherein the preset orientation includes the horizontal direction, the vertical direction and the diagonal direction; A first dynamic threshold interval is generated based on the historical statistics of the fusion features. The low-frequency components are numerically hard-truncated through the first dynamic threshold interval to obtain the first truncated components. The high-frequency components are numerically hard-truncated through the first dynamic threshold interval to obtain the second truncated components. Each of the second truncated components is added element by element to generate a high-frequency comprehensive feature; Based on the edge response intensity of each high-frequency component to each preset defect type, dynamic weighting coefficients are assigned to the high-frequency integrated features, and the high-frequency integrated features are amplified based on the dynamic weighting coefficients to obtain amplified features. The amplification feature is added to the first truncated component to obtain the frequency domain enhancement feature.

[0010] In one embodiment, the step of performing explosion-proof attention processing on the fused features in the attention branch to obtain spatial attention features includes: Based on the fusion features, a query matrix, a key matrix, and a value matrix are generated. The attention score matrix is ​​obtained by multiplying the query matrix with the transpose of the key matrix. Based on a preset piecewise linear truncation function, each matrix element in the attention score matrix is ​​truncated to a second dynamic threshold interval to obtain an attention weight matrix, wherein the second dynamic threshold interval is determined by learnable parameters, and the learnable parameters are updated by the attention branch; The attention weight matrix is ​​multiplied by the value matrix to obtain the attention enhancement feature, and the attention enhancement feature is concatenated with the spatial feature of the preset decoder to obtain the spatial attention feature.

[0011] In one embodiment, the step of generating a wafer defect segmentation image containing a preset defect type based on the output features includes: The output features are input to a preset segmentation head, and the output features are upsampled in terms of spatial resolution based on the preset segmentation head to restore the spatial size of the output features to the same size as the image of the surface defect of the wafer to be detected, so as to obtain a restored feature map. For any pixel in the restored feature map, classify the pixel into categories and determine the target defect type corresponding to the pixel; Based on the recovered feature map and the target defect type, a wafer defect segmentation image is output, wherein different colors in the wafer defect segmentation image represent different preset defect types, and the preset defect types include the target defect type.

[0012] Furthermore, to achieve the above objectives, this application also proposes a wafer defect detection system, which includes: The image acquisition module is used to acquire images of surface defects on the wafer to be inspected. The feature extraction module is used to extract features from the surface defect image of the wafer to be detected based on a preset feature extraction network, and output the multi-scale features corresponding to the surface defect image of the wafer to be detected. The feature fusion module is used to fuse the multi-scale features based on a preset decoder to obtain fused features, and to perform frequency domain enhancement processing and explosion-proof attention processing in parallel based on the fused features to obtain output features; The defect determination module is used to generate a wafer defect segmentation image containing preset defect types based on the output features.

[0013] In addition, to achieve the above objectives, this application also proposes an electronic device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the wafer defect detection method as described above.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the wafer defect detection method described above.

[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the wafer defect detection method described above.

[0016] This application provides a wafer defect detection method, which includes: acquiring an image of a wafer surface defect to be detected; extracting features from the image of the wafer surface defect to be detected based on a preset feature extraction network, and outputting multi-scale features corresponding to the image of the wafer surface defect to be detected; fusing the multi-scale features based on a preset decoder to obtain fused features; performing frequency domain enhancement processing and explosion-proof attention processing in parallel based on the fused features to obtain output features; and generating a wafer defect segmentation image containing a preset defect type based on the output features.

[0017] This application acquires images of surface defects on a wafer to be inspected, eliminating the subjectivity and visual fatigue of manual inspection. Based on a pre-defined feature extraction network, it extracts features from the surface defect images and outputs corresponding multi-scale features. The network's multi-stage structure can capture different levels of information, from high-resolution details to low-resolution semantics, effectively addressing complex periodic textures and noise interference in the wafer background. This solves the problem that traditional machine vision cannot adaptively extract discriminative features, thus enhancing feature robustness and representational ability. Based on a pre-defined decoder, the multi-scale features are fused to obtain output features, integrating spatial and frequency domain information at different levels. This solves the problem that single-scale features cannot simultaneously take into account global context and local edge details, improving the integrity and boundary fineness of the defect area. Based on the output features, a wafer defect segmentation image containing pre-defined defect types is generated, achieving pixel-level defect category prediction and directly outputting a visualized segmentation mask. This solves the problem that traditional methods can only determine the presence or absence of defects but cannot locate the defect type and contour, thus providing high-precision and interpretable defect detection results. Compared to current traditional non-destructive testing methods that struggle to handle complex periodic textures and noise interference in wafer backgrounds, this application achieves end-to-end automated detection from the original image to the defect mask through deep learning-driven multi-scale feature extraction, fusion, and pixel-level segmentation. This significantly improves the accuracy and robustness of detection, overcoming the shortcomings of traditional methods that rely on human experience and have weak noise resistance. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating an embodiment of the wafer defect detection method of this application. Figure 2 This is a schematic diagram of wafer surface defects provided in Embodiment 1 of this application; Figure 3 This is a schematic diagram of a wafer defect segmentation image provided in Embodiment 1 of this application; Figure 4 This is a flowchart illustrating Embodiment 2 of the wafer defect detection method of this application; Figure 5This is a schematic diagram of the network architecture of the wafer defect detection method provided in Embodiment 2 of this application; Figure 6 This is a schematic diagram of the module structure of the wafer defect detection system according to an embodiment of this application; Figure 7 This is a schematic diagram of the equipment structure of the hardware operating environment involved in the wafer defect detection method in this application embodiment.

[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0022] It should be understood that the first embodiment described herein is merely used to explain the technical solution of this application and is not intended to limit this application.

[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0024] The main solution of the first embodiment of this application is: to acquire an image of a wafer surface defect to be detected; to extract features from the image of the wafer surface defect to be detected based on a preset feature extraction network, and output multi-scale features corresponding to the image of the wafer surface defect to be detected; to fuse the multi-scale features based on a preset decoder to obtain fused features; to perform frequency domain enhancement processing and explosion-proof attention processing in parallel based on the fused features to obtain output features; and to generate a wafer defect segmentation image containing a preset defect type based on the output features.

[0025] In the first embodiment, for ease of description, the following description uses a wafer defect detection system as the execution subject.

[0026] In semiconductor manufacturing, wafer surfaces are highly susceptible to various defects (such as center circle defects, edge defects, random noise, and scratches). Accurately detecting and segmenting these defective regions is crucial for improving chip yield and analyzing production line faults. Traditional non-destructive testing methods (such as manual visual inspection or traditional machine vision) are inefficient and struggle to adapt to complex background noise. In recent years, deep learning-based segmentation models (such as U-Net (a U-shaped network, a semantic segmentation network with an encoder-decoder structure, named for its U-shaped shape, which fuses shallow details and deep semantics through skip connections) and Deep-Lab (a series of semantic segmentation models that use dilated convolutions to expand the receptive field and often combine conditional random fields for post-processing)) have been introduced into wafer defect detection. However, these CNN (Convolutional Neural Network) models have limitations in capturing global contextual information. For example, although the Transformer architecture (a deep learning architecture based on self-attention mechanism) performs well in global modeling, it has two major pain points when applied to small sample datasets in industrial manufacturing: First, the deep attention mechanism is prone to numerical overflow during computation (gradient explosion leads to NaN (Not a Number, indicating that the calculation result is outside the range of numerical representation or invalid, usually caused by gradient explosion, division by zero or abnormal input of exponential function)); Second, as a low-pass filter, the Transformer itself is prone to ignoring high-frequency edge details such as wafer scratches.

[0027] This application provides a solution that, by acquiring images of surface defects on a wafer to be inspected, eliminates the subjectivity and visual fatigue of manual inspection. Based on a pre-defined feature extraction network, features are extracted from the surface defect images, outputting corresponding multi-scale features. The network's multi-stage structure can capture different levels of information, from high-resolution details to low-resolution semantics, effectively addressing complex periodic textures and noise interference in the wafer background. This solves the problem of traditional machine vision's difficulty in adaptively extracting discriminative features, enhancing feature robustness and representational ability. Based on a pre-defined decoder, multi-scale features are fused to obtain output features, integrating spatial and frequency domain information at different levels. This solves the problem that single-scale features cannot simultaneously take into account global context and local edge details, improving the integrity and boundary refinement of defect areas. Based on the output features, a wafer defect segmentation image containing pre-defined defect types is generated, achieving pixel-level defect category prediction and directly outputting a visualized segmentation mask. This solves the problem that traditional methods can only determine the presence or absence of defects but cannot locate defect types and contours, providing high-precision and interpretable defect detection results. Compared to current traditional non-destructive testing methods that struggle to handle complex periodic textures and noise interference in wafer backgrounds, this application achieves end-to-end automated detection from the original image to the defect mask through deep learning-driven multi-scale feature extraction, fusion, and pixel-level segmentation. This significantly improves the accuracy and robustness of detection, overcoming the shortcomings of traditional methods that rely on human experience and have weak noise resistance.

[0028] It should be noted that the executing entity of the first embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, or other electronic device, or a system, application, or program capable of implementing the above functions. The first embodiment and the following embodiments will be described using a wafer defect detection system as an example.

[0029] All actions involving the acquisition of signals, information, or data in this application are carried out in accordance with the relevant data protection laws and policies of the country where the application is located, and with the authorization of the owner of the relevant device.

[0030] Based on this, embodiments of this application provide a wafer defect detection method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the wafer defect detection method of this application.

[0031] In this embodiment, the wafer defect detection method includes steps S01 to S04: Step S01: Obtain an image of the surface defects of the wafer to be inspected; It should be noted that the image of the wafer surface defect to be detected is a raw digital image of the wafer acquired by an image acquisition device (such as a camera). This image may contain wafer surface defect anomalies, such as scratches, damage, clumps, etc.

[0032] Understandably, since traditional nondestructive testing relies on human experience for image interpretation, step S01, by directly acquiring the image to be inspected, avoids the subjective bias and inefficiency caused by manual sampling. This achieves the effect of improving the automation level and data consistency of the inspection process from the source, and solves the problems of non-standard image acquisition and inability to be directly adapted to subsequent intelligent analysis in traditional manual visual inspection or machine vision. It provides a standardized input interface for automated inspection.

[0033] Step S02: Based on the preset feature extraction network, feature extraction is performed on the image of the wafer surface defect to be detected, and the multi-scale features corresponding to the image of the wafer surface defect to be detected are output. It should be noted that the preset feature extraction network is a hierarchical deep learning network model (such as a pyramid network) used to extract features of different levels of abstraction from the input image stage by stage. This network contains multiple downsampling stages, each outputting a feature map with a corresponding spatial resolution. Feature extraction is the forward propagation process of the input image through the preset feature extraction network. Each layer of the network performs convolution, pooling, or attention operations on the image, gradually transforming the pixel-level raw data into a set of feature maps with semantic information. Multi-scale features are the set of feature maps output by different stages of the feature extraction network. These feature maps have different spatial resolutions and channel dimensions. For example, low-level features (closer to the input) have high resolution and weak semantics, preserving edge and texture details; high-level features (closer to the output) have low resolution and strong semantics, encoding defect categories and global context.

[0034] For example, the preprocessed image of the wafer surface defects to be detected is input into a preset feature extraction network, such as a pyramid hierarchical backbone network (which could be PVTv2 (Pyramid Vision Transformer version 2, a hierarchical vision transformer that generates multi-scale feature maps through progressive downsampling)). This network can contain four feature extraction stages with different spatial resolutions: the first stage is used to extract high-resolution edges and texture details of the wafer surface defect image; the second stage is used to identify local spatial patterns of the wafer surface defect image; the third stage is used to encode the mid-level semantic relationships of the wafer surface defect image; and the fourth stage is used to perceive the global defect context of the wafer surface defect image. The final output consists of four hierarchical feature maps (i.e., multi-scale features) with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image, respectively.

[0035] Understandably, since existing deep learning models either lose global context or weaken high-frequency details, step S02, by outputting hierarchical features at different resolutions, allows the network to simultaneously retain high-resolution scratch edge details and low-resolution global semantic information. This solves the limitations of traditional CNN models in capturing global context information and the problem that Transformer tends to ignore high-frequency edge details on industrial small sample datasets. It achieves the effect of adaptively extracting discriminative features from local texture to global layout without the need for manual feature design.

[0036] Step S03: Based on the preset decoder, multi-scale features are fused to obtain fused features. Based on the fused features, frequency domain enhancement processing and explosion-proof attention processing are performed in parallel to obtain output features. It should be noted that the pre-defined decoder is a neural network module used to fuse multi-scale features. It includes operations such as upsampling, feature concatenation, convolution, and attention mechanisms, aiming to integrate multi-scale features of different scales into a unified representation. The output feature is the feature map finally generated by the pre-defined decoder after multi-level fusion, which retains both high-resolution spatial information for accurate localization and strong semantic information for distinguishing defect categories.

[0037] In one feasible implementation, step S03 includes steps A01 to A03: Step A01: Input the multi-scale features into the preset decoder, and fuse the multi-scale features based on the first fusion module and the second fusion module in the preset decoder to obtain fused features. The first fusion module is composed of a preset residual direct connection channel and a series channel connected in parallel. The series channel is composed of a first preset convolutional layer, a first preset layer normalization layer and a first preset linear rectified activation layer connected in series. The second fusion module is composed of a second preset convolutional layer, a second preset layer normalization layer and a second preset linear rectified activation layer stacked alternately.

[0038] It should be noted that the first fusion module is the basic module used for the initial fusion of multi-scale features. It consists of a series channel and a preset residual direct connection channel connected in parallel. The series channel sequentially includes a first preset convolutional layer, a first preset layer normalization layer, and a first preset linear rectified activation layer. The preset residual direct connection channel is the connection path that directly passes the input to the output in the parallel structure of the first fusion module, allowing gradients to be directly backpropagated, alleviating the gradient vanishing problem in deep network training, and preserving the original input information. The first preset convolutional layer is a two-dimensional convolutional layer located in the series channel. It is used to perform local spatial filtering on the input features (such as multi-scale features) to extract higher-level abstract features and is used to adjust the number of channels or spatial size. The first preset layer normalization layer is used to normalize the features output by the first preset convolutional layer along the channel dimension, making the feature distribution mean zero and variance one, thereby stabilizing the training process and accelerating convergence. The first preset linear rectified activation layer uses a linear rectified function as the activation function, setting negative values ​​to zero and keeping positive values ​​unchanged, thus introducing nonlinear transformation capability to the network. It can be a ReLU (Rectified Linear Unit), and the activation function is defined as follows: The first fusion module has the advantages of simple computation and mitigating gradient vanishing. The second fusion module is similar to the first, also used for initial fusion of multi-scale features, but its structure consists of two alternately stacked layers: a second preset convolutional layer, a second preset normalization layer, and a second preset linear rectified activation layer are repeated twice. It does not include residual direct connections. This module enhances the expressive power of features through deeper nonlinear transformations. The second preset convolutional layer, second preset normalization layer, and second preset linear rectified activation layer are similar to the first preset convolutional layer, first preset normalization layer, and first preset linear rectified activation layer, forming two sets of identical structural layers in the second fusion module. The first set performs convolution-normalization-activation, and the second set repeats the same operation. The difference from the first fusion module is the absence of residual direct connections and the use of two stacked layers to enhance the depth of feature transformation. The fused feature is a comprehensive feature map obtained by processing multi-scale features (such as X1, X2, X3, X4) sequentially through the first and second fusion modules, which integrates semantic and detailed information at different resolutions.

[0039] For example, multi-scale features at different resolutions are input into a preset decoder and initially fused through a first fusion module. The first fusion module can be composed of a preset residual direct connection channel and a serial channel connected in parallel. The serial channel sequentially includes a first preset convolutional layer, a first preset layer normalization layer, and a first preset linear rectified activation layer. The input features are processed simultaneously by the residual direct connection (identity mapping) and the serial channel, and their outputs are added together, thereby introducing a nonlinear transformation while preserving the original information. The fused features are then input into a second fusion module, which consists of a second preset convolutional layer, a second preset layer normalization layer, and a second preset linear rectified activation layer stacked alternately twice. This module further extracts deep semantics and eliminates statistical fluctuations caused by different defect scales, ultimately outputting fused features that integrate multi-scale information, providing a unified input for subsequent frequency domain enhancement and explosion-proof attention processing.

[0040] Step A02: Input the fused features in parallel into the wavelet transform branch and attention branch of the preset decoder. In the wavelet transform branch, perform frequency domain enhancement processing on the fused features to obtain frequency domain enhanced features. In the attention branch, perform explosion-proof attention processing on the fused features to obtain spatial attention features. It should be noted that the wavelet transform branch is one of the parallel processing branches in the decoder, used to perform frequency domain enhancement processing on the fused features to enhance the edge information of the image of the wafer surface defects to be detected. Frequency domain enhancement processing is a series of operations performed in the wavelet transform branch to highlight the high-frequency edge details of defects in the image of the wafer surface defects to be detected. The frequency domain enhanced feature is the feature map output after frequency domain enhancement processing by the wavelet transform branch. In this feature, the edge response of small defects such as scratches is significantly amplified, while high-frequency background noise is suppressed. The attention branch is another parallel processing branch in the decoder, used to perform self-attention-based global context modeling on the fused features. This branch introduces an attention anti-explosion mechanism to dynamically truncate the attention score of the fused features to prevent numerical overflow. Explosion-proof attention processing is a series of operations performed in the attention branch to avoid exponential overflow caused by excessive feature variance in the attention mechanism. The spatial attention feature is the feature obtained after explosion-proof attention processing, which simultaneously possesses global semantic dependence on fine spatial location information.

[0041] In one feasible implementation, step A02, which involves performing frequency domain enhancement processing on the fused features in the wavelet transform branch to obtain frequency domain enhanced features, includes steps A11 to A15: Step A11: Perform a two-dimensional discrete wavelet transform on the fused features to obtain low-frequency components and high-frequency components in each preset orientation, wherein the preset orientation includes the horizontal direction, the vertical direction and the diagonal direction. It should be noted that the 2-Dimensional Discrete Wavelet Transform (DWT2d) is a frequency domain analysis operation used to decompose an input feature map (e.g., a fused feature) into a low-frequency approximation component and high-frequency detail components at preset orientations (e.g., three). This transform decomposes the feature map layer by layer using a filter bank, separating information from different frequency bands. The low-frequency component (LL) is the approximation component obtained after the 2-dimensional discrete wavelet transform, representing regions with gradual changes in the fused feature (such as the smooth body of a defect or the background), containing the main energy and global structural information of the image. The preset orientations are predefined spatial directions used to distinguish the orientation of the high-frequency components, including horizontal, vertical, and diagonal directions. The high-frequency components are the detail components obtained after the 2-dimensional discrete wavelet transform, corresponding to the horizontal (LH), vertical (HL), and diagonal (HH) directions, respectively, representing regions with drastic gray-level changes in the fused feature (such as high-frequency information like defect edges and background textures). The horizontal direction, i.e., the x-axis of the image coordinate system, corresponds to the LH component, which reflects abrupt changes in grayscale along the horizontal direction (such as vertical edges). The vertical direction, i.e., the y-axis of the image coordinate system, corresponds to the HL component, which reflects abrupt changes in grayscale along the vertical direction (such as horizontal edges). The diagonal direction is at a 45° angle to both the horizontal and vertical directions, and the HH component reflects abrupt changes in grayscale along the diagonal direction (such as oblique edges or noise).

[0042] For example, a two-dimensional discrete wavelet transform is performed on the fused feature, decomposing it into a low-frequency component and three high-frequency components in preset orientations. The low-frequency component (LL) preserves the approximate contour and smooth regions of the fused feature, reflecting the overall distribution of defects; the high-frequency components correspond to the horizontal direction (LH, responding to horizontal edges), the vertical direction (HL, responding to vertical edges), and the diagonal direction (HH, responding to oblique edges), respectively. For instance, for a wafer image containing scratch defects, oblique scratches will produce a strong response in the diagonal high-frequency component, while the horizontal and vertical high-frequency components will have weaker responses.

[0043] Step A12: Generate a first dynamic threshold interval based on the historical statistics of the fusion features, perform numerical hard truncation on the low-frequency components through the first dynamic threshold interval to obtain the first truncated components, and perform numerical hard truncation on each high-frequency component through the first dynamic threshold interval to obtain each second truncated component. It should be noted that historical statistics, such as mean and standard deviation, are statistical indicators calculated based on historical fusion characteristics. The first dynamic threshold interval is a variable numerical range whose boundaries are dynamically calculated from historical statistics. It is used to truncate low-frequency and high-frequency components to suppress extreme outliers. Numerical hard truncation is a non-linear operation used to restrict the values ​​of low-frequency components to a specified interval. For example, elements below the lower limit of the first dynamic threshold interval are set to the lower limit, and elements above the upper limit are set to the upper limit, while the rest remain unchanged. The first truncated component is the output of the low-frequency components after numerical hard truncation, retaining the main structural information and removing outliers in the low-frequency domain. The second truncated component is the output of each high-frequency component (including LH, HL, and HH) after numerical hard truncation. The high-frequency components in each direction are truncated independently, removing extreme high-frequency peaks caused by noise or acquisition anomalies.

[0044] Optionally, since the characteristic distributions of different wafer batches or different defect types vary significantly, a fixed threshold is difficult to adapt to diverse data statistical characteristics. By dynamically defining the cutoff boundary, extreme high-frequency noise can be adaptively suppressed while retaining effective defect edge information. Therefore, the first dynamic threshold interval can be generated based on the historical statistics of the fused features by calculating the mean and standard deviation of the current fused features, and setting the boundary of the first dynamic threshold interval as the mean plus or minus a preset multiple (e.g., 3 times) of the standard deviation to obtain a numerical range that is adaptive to the data distribution. For example, if the fused features follow a standard normal distribution, the threshold interval can be set to the mean ± three times the standard deviation.

[0045] For example, based on a first dynamic threshold interval, each element in the component is compared with the upper and lower bounds of the interval. Elements smaller than the lower bound are assigned the lower bound, elements larger than the upper bound are assigned the upper bound, and elements within the interval remain unchanged. After truncation, the low-frequency component becomes the first truncated component, and the high-frequency components in the horizontal, vertical, and diagonal directions become the corresponding second truncated components. For example, if the threshold interval is [-3, 3], and an element of a high-frequency component has a value of 5, it is truncated to 3.

[0046] It is understandable that industrial wafer images often contain extreme high-frequency outliers caused by sensor noise or lattice reflection. If these outliers are not suppressed, subsequent high-frequency amplification operations will excessively amplify the noise, masking the true defect edges. Therefore, by forcibly limiting outliers to a reasonable range through a first dynamic threshold interval while preserving effective edge responses, hard truncation, compared to existing schemes that use soft thresholding or global normalization, does not require the introduction of additional learnable parameters or complex transformations, resulting in low computational overhead. Furthermore, when combined with adaptive dynamic thresholding, it can adapt to changes in data distribution across different batches of wafers, avoiding edge response attenuation or smoothing distortion problems that may be caused by soft thresholding.

[0047] Step A13: Add each of the second truncated components element by element to generate high-frequency comprehensive features; It should be noted that the high-frequency integrated feature is a single feature map obtained by adding the second truncated components element by element. This feature integrates high-frequency information from multiple directions and serves as the input for subsequent attention weighting.

[0048] Step A14: Based on the edge response intensity of each high-frequency component to each preset defect type, assign dynamic weight coefficients to the high-frequency integrated features, and amplify the high-frequency integrated features based on the dynamic weight coefficients to obtain amplified features; It should be noted that edge response intensity refers to the sensitivity of different high-frequency components to defect edges in the current input feature. For example, horizontal high-frequency components respond strongly to vertical scratch edges but weakly to horizontal edges. Dynamic weighting coefficients are variable weight values ​​generated by the attention mechanism based on the edge response intensity of each high-frequency component; different high-frequency components can be assigned different weights. Feature amplification multiplies the dynamic weighting coefficients with the corresponding high-frequency composite features, enhancing the frequency band information sensitive to defect edges while suppressing irrelevant background noise. Amplified features are the high-frequency enhanced features obtained after feature amplification, where the edge responses of weak defects such as scratches are significantly improved.

[0049] Optionally, the steps for generating dynamic weight coefficients can be as follows: extract global information from each high-frequency component (e.g., global average pooling), generate a weight vector related to the defect type through a fully connected layer and activation function via a wavelet transform branch, with each weight coefficient corresponding to a high-frequency direction or an overall high-frequency comprehensive feature. For example, for scratch defects, the high-frequency component in the diagonal direction has the strongest response, so a weight coefficient greater than one is assigned to amplify its edge; for random noise, the response in each direction is uniform, so a weight less than one is assigned to suppress interference.

[0050] It is understandable that different preset defect types (such as scratches and edge defects) have anisotropic edge direction characteristics. If each high-frequency component is added with equal weight, the information of the dominant direction will be diluted, while irrelevant noise will be amplified. Therefore, by using an adaptive frequency channel attention mechanism, weights are dynamically allocated according to the response intensity of each high-frequency component to the current defect edge. This prioritizes the enhancement of directions that contribute the most, suppresses noise directions, and significantly improves the salience of high-frequency details such as weak scratches. Compared with existing schemes that use fixed gain coefficients or global normalization, this mechanism has directional adaptability and can dynamically adjust according to input characteristics, avoiding suboptimal enhancement caused by using a uniform amplification strategy for all defect types.

[0051] Step A15: Add the amplified feature to the first truncated component to obtain the frequency domain enhancement feature.

[0052] In one feasible implementation, step A02, the step of performing explosion-proof attention processing on the fused features in the attention branch to obtain spatial attention features, includes steps A21-A23: Step A21: Generate a query matrix, a key matrix, and a value matrix based on the fusion features. Multiply the query matrix and the transpose of the key matrix to obtain the attention score matrix. It should be noted that the query matrix, key matrix, and value matrix are three sets of tensors generated by performing three different linear transformations (such as fully connected layers or convolutional layers) on the fused features. The query matrix is ​​used to match the correlations between fused features, the key matrix is ​​used for matrix query matching, and the value matrix stores the content features to be weighted. For example, for a wafer image containing scratches and defects, the query matrix represents the query for the global context at each location, the key matrix represents the context identifier provided by each location, and the value matrix represents the feature content at each location. The attention score matrix is ​​a square matrix obtained by multiplying the query matrix and the transpose of the key matrix. Each element represents the correlation strength between corresponding locations, with larger element values ​​indicating a stronger association between the two locations.

[0053] Step A22: Based on the preset piecewise linear truncation function, each matrix element in the attention score matrix is ​​truncated to the second dynamic threshold interval to obtain the attention weight matrix. The second dynamic threshold interval is determined by the learnable parameters, which are updated by the attention branches. It should be noted that the preset piecewise linear truncation function is a numerical processing function, which can be expressed as:

[0054] in, The input elements, such as matrix elements, are individual values ​​in the attention score matrix, representing the original relevance score between a pair of locations. The second dynamic threshold interval is defined by the learnable parameters. Defined numerical range [- , Learnable parameters An initial value (e.g., 10) is assigned during attention branch initialization, and then updated with gradient descent during training, enabling the model to adapt to the feature variance of different wafer batches. This function forces the input values ​​to be bound to [-]. , Within the interval, the portion exceeding the limit is truncated as boundary values.

[0055] For example, for any matrix element among all matrix elements, if its value is greater than Then adjust to If less than Then adjust to Otherwise, it remains unchanged. After this hard truncation operation, all matrix elements are forcibly confined to the interval [- , The attention score matrix is ​​obtained by inputting the truncated matrix into the Softmax function (a normalization exponential function that converts the input vector into a probability distribution with each element value between (0,1) and summing to 1) to normalize it, so that the sum of the elements in each row is 1 and all are positive. The final output is the attention weight matrix.

[0056] Step A23: Multiply the attention weight matrix and the value matrix to obtain the attention enhancement feature, and concatenate the attention enhancement feature with the spatial features of the preset decoder to obtain the spatial attention feature.

[0057] It should be noted that the attention weight matrix is ​​the probability distribution matrix output by inputting the truncated attention score matrix into the Softmax function. Its row sum is 1, and each element represents the normalized attention weight of each key position at a given query position. The attention enhancement feature is the feature obtained by multiplying the attention weight matrix by the value matrix. This feature is a weighted aggregation of the global context, and the new feature at each position is the weighted sum of all positional features according to their attention weights. Spatial features are feature maps provided by the decoder that retain high-resolution spatial location information. These can be features received by the decoder from shallow stages (such as the first or second stage) of the preset feature extraction network via skip connections, or raw features from intermediate layers within the decoder that have not undergone attention compression. Spatial attention features are the final feature output by concatenating the attention enhancement feature and spatial features along the channel dimension, followed by value fusion (such as dimensionality reduction via convolutional layers). This feature possesses both global semantic dependence (from attention) and local spatial accuracy (from spatial features), for example, it can accurately locate the continuous edges of scratch defects and distinguish them from background texture.

[0058] Understandably, by performing two-dimensional discrete wavelet transform, dynamic threshold truncation, high-frequency synthesis, and adaptive frequency channel attention amplification on the fused features through wavelet transform branches, the high-frequency edge response of minor defects such as scratches is effectively enhanced while suppressing background noise. However, this frequency domain enhancement process artificially amplifies high-frequency features, leading to a surge in the variance of the query matrix and key matrix in the subsequent attention branch. Directly performing normalization calculations can easily cause numerical overflow, such as directly performing Softmax calculations. Calculations can instantly lead to gradient explosion and training crashes. To address this, the attention branch introduces a numerical hard truncation operation, forcibly truncating the attention score matrix to a second dynamic threshold interval. This effectively avoids gradient explosion while preserving the relative distribution of attention weights. At the same time, the dynamic threshold interval can adapt to the statistical distribution of different wafer batches, achieving deep synergy between frequency domain enhancement and explosion-proof attention.

[0059] Step A03: Fuse the frequency domain enhancement features with the spatial attention features to obtain the output features.

[0060] Understandably, since models like U-Net only perform limited fusion through skip connections, step S03, by fusing multi-scale features, allows the decoder to complementarily integrate shallow spatial details with deep semantic information, thereby enhancing the integrity of defect region boundaries and internal consistency. This resolves the contradiction that a single-level feature cannot simultaneously address defect localization accuracy and classification accuracy, as well as the semantic gap caused by the simple feature stacking of traditional decoders.

[0061] Step S04: Generate a wafer defect segmentation image containing preset defect types based on the output features.

[0062] It should be noted that the preset defect type is a predefined set of wafer surface defect categories, which may include central circle defects (regional anomalies located in the center of the wafer), edge defects (clumps or ring-shaped anomalies clustered in the edge region), random defects (irregularly distributed scattered noise), and scratch defects (thin, low-contrast linear edge anomalies). The wafer defect segmentation image is a pixel-level classification label map (i.e., a mask image) with the same spatial size as the input image of the wafer surface defects to be detected. Each pixel is assigned a category label, indicating whether the pixel belongs to the background or a certain type of defect. The image is presented in pseudo-color, with different colors corresponding to different defect types, thereby visualizing the precise outline and distribution of defects.

[0063] For example, to aid in understanding the technical concept or principles of this application, please refer to Figure 2 , Figure 2 The diagram provides an illustration of wafer surface defects, including central circle defects, scratch defects, edge defects, and random defects. Central circle defects appear as circular anomalies in the central region of the wafer, scratch defects appear as thin, elongated linear marks, edge defects are distributed in the outer periphery of the wafer, and random defects appear as irregular scattered points or clumps. These four types of defects cover the main abnormal morphologies commonly encountered in wafer manufacturing.

[0064] In one feasible implementation, step S04 includes steps A31 to A33: Step A31: Input the output features into the preset segmentation head, and perform spatial resolution upsampling operation on the output features based on the preset segmentation head to restore the spatial size of the output features to the same size as the image of the surface defects of the wafer to be detected, so as to obtain the restored feature map; It should be noted that the preset segmentation head is a neural network module deployed after the decoder, used to convert the feature map output by the preset decoder into pixel-level classification results. It can include upsampling layers (such as transposed convolutions or interpolation) and point-by-point classification layers (such as 1×1 convolutions). Spatial resolution upsampling is a computational process that enlarges the height and width dimensions of the feature map corresponding to the output features; for example, it enlarges the feature map of the original size... Figure 3 The 1 / 12th feature map is progressively enlarged to the original image size using bilinear interpolation or transposed convolution, ensuring that each spatial location corresponds to a pixel region in the original image. The spatial dimensions are the height and width values ​​(in pixels) of the feature map corresponding to the output features. The recovered feature map is obtained after upsampling, and its spatial dimensions are completely identical to the original input wafer image.

[0065] For example, the output features of the decoder are input into a preset segmentation head. Since the spatial size of the output features is usually smaller than the original input image (e.g., one-thirty-second of the original image), the segmentation head gradually enlarges the height and width of the feature map through upsampling methods such as bilinear interpolation or transposed convolution until its spatial size is completely consistent with the original image of the defect on the surface of the wafer to be detected, thus obtaining the recovered feature map.

[0066] Step A32: For any pixel in the recovered feature map, classify the pixel into categories and determine the target defect type corresponding to the pixel. It should be noted that the target defect type is the specific defect category to which any pixel belongs after it has been classified.

[0067] For example, an independent category classification operation is performed on each pixel in the restored feature map. Specifically, the preset segmentation head first maps the channel dimension of the restored feature map to the total number of preset defect categories (including background, central circle defect, edge defect, random defect, and scratch) through a pointwise convolutional layer, obtaining the original score of each pixel in each preset defect category. Subsequently, a normalized exponential function is applied to the original scores to obtain the probability distribution of each category. Finally, the category with the highest probability value is selected as the target defect type of the pixel. For example, for a wafer image containing a central circle defect, pixels located within the defect area have the highest probability in the "central circle defect" category after classification, so the target defect type of the pixel is determined to be "central circle defect"; while pixels located in the normal background area are determined to be "background".

[0068] Step A33: Output a wafer defect segmentation image based on the recovered feature map and the target defect type. In the wafer defect segmentation image, different colors represent different preset defect types, including the target defect type.

[0069] For example, in the output segmented image, the central circular defect region can be rendered as purple, the edge defect region as red, the scratch defect region as blue, the random defect region as green, and the background region as black. Then, the color values ​​of all pixels are combined into a pseudo-color image of the same size as the original wafer image, which is the final segmentation mask image. This image can simultaneously present the precise contours and spatial distribution of multiple defects in an intuitive visualization.

[0070] For example, to aid in understanding the technical concept or principles of this application, please refer to Figure 3 , Figure 3 A schematic diagram of wafer defect segmentation images is provided. Preset defect types include central circular defects, edge defects, scratch defects, and random defects. These defects overlap and intersect spatially. Decoupling and labeling the preset defect types allows different types to be marked with different colors in the wafer defect segmentation image, thus providing a more intuitive identification of defects. For example, a fine distinction can be made between annular edge defects distributed along the wafer periphery (assigned a red pseudo-color mask) and clustered edge defects concentrated in a certain area (assigned a green pseudo-color mask). This not only achieves accurate defect extraction but also further decouples them into subclasses based on their spatial distribution and morphological characteristics. This multi-dimensional color mapping capability provides a more granular basis for fault attribution in subsequent analysis of semiconductor processes.

[0071] Understandably, due to the subjective qualitative assessment of manual visual inspection and the coarse positioning of traditional machine vision, step S04 is performed to generate a pixel-level segmentation mask. Each pixel is classified as a central circle defect, edge defect, random defect, scratch, or background. This achieves the effect of outputting intuitive and visual results and quantifying the defect morphology, solving the problem that traditional detection methods can only determine whether a defect exists but cannot output pixel-level information on the defect type, location, and contour.

[0072] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 The wafer defect detection method also includes steps S11 to S13: Step S11: Obtain a preset wafer surface defect image set. For any preset wafer surface defect image in the preset wafer surface defect image set, use the preset wafer surface defect image and its corresponding image label as training samples. It should be noted that the preset wafer surface defect image set is a pre-constructed collection of wafer defect images used to train the feature extraction network. This set contains multiple wafer surface defect images, each with a corresponding ground truth label. The preset wafer surface defect image is a single wafer surface defect image from the preset wafer surface defect image set; it serves as the basic input unit during training. The image label is the ground truth annotation information corresponding one-to-one with the preset wafer surface defect image. In segmentation tasks, this label is typically a pixel-level segmentation mask, where each pixel is labeled as either background or a specific defect category (such as a scratch). Training samples consist of a single preset wafer surface defect image and its corresponding image label, forming a data pair.

[0073] In one feasible implementation, step S11, the step of acquiring a preset wafer surface defect image set, includes steps B01 to B04: Step B01: Obtain a preset basic defect image set, wherein the preset basic defect image set includes basic defect images of each preset defect type, and each preset defect type includes at least one of center circle defect, edge defect, random defect and scratch defect; It should be noted that the preset basic defect image set is a collection of raw wafer defect images pre-acquired, containing only a single defect type, with each image labeled with only one defect type. The basic defect image is a single raw wafer defect image from the preset basic defect image set. A center circle defect is a defect type located in the central region of the wafer, such as a circular or ring-shaped anomalous area. Edge defects are ring-shaped or cluster-like defects concentrated at the wafer edge. Random defects are defect types that are irregularly distributed on the wafer surface, such as discrete point-like or small-area anomalies. Scratch defects are defect types distributed in a linear pattern, usually caused by mechanical scratching, such as thin, weak edge marks.

[0074] Step B02: For any basic defect image in the preset basic defect image set, perform random horizontal flipping, random vertical flipping, and random rotation operations with preset probabilities on the basic defect image to obtain the initial transformed image. It should be noted that the preset probability is a predefined probability value for performing a certain random operation (such as flipping or rotating), for example, 70%. The random horizontal flip operation is an operation that mirrors the image horizontally along the vertical central axis with a preset probability. The random vertical flip operation is an operation that mirrors the image vertically along the horizontal central axis with a preset probability. The random rotation operation is an operation that randomly rotates the image by a specific angle (e.g., 90 degrees, 180 degrees, or 270 degrees) with a preset probability. The initial transformed image is the enhanced image obtained by sequentially performing the above random horizontal flip, random vertical flip, and random rotation operations on a single basic defect image; this image still retains the original defect type label.

[0075] Step B03: Merge the initial transformation images containing different preset defect types to obtain a preset wafer surface defect image, and record the preset defect types contained in the preset wafer surface defect image, wherein the preset wafer surface defect image contains at least one preset defect type. It should be noted that multiple initial transformed images, acquired separately from different preset defect types and after data augmentation, are synthesized into a single image according to preset spatial layout rules (such as random overlay or stitching), thereby generating a preset wafer surface defect image containing multiple defect types. Simultaneously, all defect type labels contained in this synthesized image are recorded (e.g., simultaneously containing a central circle defect and a scratch defect), ensuring that the image contains at least one defect type. For example, a flipped image of a central circle defect and a rotated image of a scratch defect are pixel-level overlaid to generate a synthesized image containing both central region anomalies and edge linear scratches, and its defect type set is labeled.

[0076] Step B04: Merge each preset wafer surface defect image into a preset wafer surface defect image set.

[0077] Step S12: Input the training samples into the preset pyramid network for forward propagation to obtain the hierarchical features corresponding to the output of each network layer of the preset pyramid network. The preset pyramid network includes shallow network layers and deep network layers. The weight parameters of the shallow network layers remain unchanged, while the weight parameters of the deep network layers are updated during the training process of the preset pyramid network. It's important to note that the Preset Pyramid Network is a neural network structure with multiple stages (layers). Each stage outputs a feature map with a different spatial resolution, used to extract multi-scale hierarchical features from the input image. It can be a hierarchical network such as PVTv2 or Swin Transformer (Shifted Window Transformer, which achieves linear complexity cross-window information interaction through a moving window attention mechanism). Forward propagation is the process where data starts from the input layer of the Preset Pyramid Network, is computed layer by layer, and passed to the output layer. During this process, the data sequentially passes through each layer of the network, generating feature maps for each layer. Network layers are different stages arranged sequentially in the Preset Pyramid Network. Each layer is responsible for processing features at a specific resolution; for example, the first layer processes high-resolution details, and the fourth layer processes low-resolution semantics. Hierarchical features are the feature maps output after forward propagation at each network layer. Features at different layers have different spatial resolutions and semantic abstraction levels; for example, shallow features have high resolution and rich edge information, while deep features have low resolution and strong semantic information. Shallow layers in a pre-defined pyramid network are located near the input and are used to extract high-resolution general features (such as edges and textures). Their weights are fixed and do not participate in updates. Deep layers in a pre-defined pyramid network are located near the output and are used to extract low-resolution task-relevant semantic features. Their weights participate in gradient backpropagation and are updated. Weights are the learning coefficients of the connections between layers in the neural network.

[0078] For example, the preset pyramid network can contain four stages, where stages 1 and 2 correspond to the shallow layers of the network, and stages 3 and 4 correspond to the deep layers. Stage 1 can process feature maps at the H / 4×W / 4 scale, extracting high-resolution details such as basic edges and corners; weights are frozen in this stage. Stage 2 can process feature maps at the H / 8×W / 8 scale, recognizing local texture patterns; weights are frozen in this stage to stabilize the output of lower-level features. Stage 3 can process feature maps at the H / 16×W / 16 scale, beginning to learn the unique semantics of wafer defects; weights are unfrozen in this stage to participate in training. Stage 4 can process feature maps at the H / 32×W / 3 scale, perceiving the global context and performing macroscopic localization of large-scale targets such as central circular defects and edge defects; weights are unfrozen in this stage to participate in training. The preset pyramid network ultimately outputs four hierarchical features at different resolutions.

[0079] Understandably, given the scarcity of samples and the complexity of background textures in wafer defect detection, fixing the weight parameters of the shallow layers of the network preserves the general high-resolution edge and texture extraction capabilities already learned by the pre-trained model. This avoids "catastrophic forgetting" during fine-tuning on small sample datasets, which could lead to mislearning background noise as defect features. At the same time, only updating the weight parameters of the deeper layers of the network allows the model to adaptively learn the mid-to-high-level semantic relationships and global context of wafer defects. This effectively balances generalization ability and task specificity, significantly alleviates overfitting, and improves the robustness and accuracy of defect segmentation.

[0080] Step S13: Calculate the network loss based on the hierarchical features and image labels corresponding to each network layer, and update the weight parameters of the deeper layers of the network in reverse based on the network loss until the preset pyramid network meets the preset training conditions, thus obtaining the preset feature extraction network.

[0081] It should be noted that network loss is a scalar value calculated based on the difference between the hierarchical features output by the network (or the prediction result obtained after subsequent processing) and the image label. It is used to measure the accuracy of the current network prediction; the smaller the loss, the closer the prediction is to the true label. Preset training conditions are pre-defined conditions for terminating training. For example, the network loss may fall below a certain threshold, performance on the validation set may no longer improve, or the preset maximum number of training epochs may be reached. Once these conditions are met, training stops, and the current network parameters are saved as the final preset feature extraction network.

[0082] For example, the hierarchical features output from each network layer are upsampled to the same spatial resolution as the image labels, and the cross-entropy loss between each layer and the pixel-level label is calculated. The total network loss is then obtained by weighted summation of the losses from each layer. The gradient of the loss with respect to the deep network weights is calculated using the backpropagation algorithm, and these deep weights are updated using an optimizer (such as Adam (Adaptive Moment Estimation, an optimization algorithm combining momentum and adaptive learning rate, dynamically adjusting the learning rate based on the first and second moments of the gradient)). The shallow network weights remain unchanged at their initial pre-training values. This forward and backward process is repeated until the total network loss no longer decreases for several consecutive epochs or reaches a preset maximum number of epochs. At this point, training is stopped, and the network parameters are saved, resulting in the trained preset feature extraction network.

[0083] In this embodiment, by constructing a training sample set containing images and corresponding labels, the problem of scarce labeled data and diverse defect types in wafer defect detection is solved, providing a supervisory signal for model learning. By adopting a pyramid network and freezing shallow weights while only updating deep weights, the problems of overfitting and forgetting pre-trained knowledge on small industrial datasets are solved, achieving the effect of retaining general edge extraction capabilities while adapting to specific defect semantics. The loss is calculated based on hierarchical features and labels, and deep parameters are updated in reverse, solving the problem of shallow features being disturbed by noise and deep features lacking discriminative power, thus achieving stable convergence of the feature extraction network.

[0084] For example, to aid in understanding the technical concept or principles of this application, please refer to Figure 5 , Figure 5 A schematic diagram of the network architecture for a wafer defect detection method is provided. First, the input image is processed by a PVTv2 backbone network to extract features X1, X2, X3, and X4 at each level. The weight parameters for stage 1 (high-resolution detail extraction) and stage 2 (local pattern recognition) are frozen, while the weight parameters for stage 3 (mid-level semantic encoding) and stage 4 (global context awareness) are trainable. Each level of feature is fed into an FES-FORMER (Functional Electrical Stimulation Transformer, a deep learning model based on the Transformer architecture for processing and analyzing signals or control tasks related to functional electrical stimulation) decoder. The features are then processed by a CL fusion module (the first fusion module) and a CC fusion module (the second fusion module), and then proceed in parallel to wavelet transform and attention branches. The CL fusion module may include two 2D convolutional layers, one layer normalization layer, and one linear rectified activation layer. The CC fusion module may include one 2D convolutional layer, one layer normalization layer, one linear rectified activation layer, and one residual direct connection channel. In the wavelet transform branch, the input features... The low-frequency component LL and the horizontal high-frequency components LH, vertical high-frequency components HL, and diagonal high-frequency components HH are decomposed by 2D discrete wavelet transform (DWT2d). After each component is hard-truncated in the interval [-3,3], the truncated high frequencies are added together and multiplied by a coefficient of 1.2 to obtain the high-frequency comprehensive feature HF. Then, the frequency domain feature (i.e., amplification feature) is obtained by combining the variance of the low-frequency component LL. The normalized frequency domain features (i.e., frequency domain enhancement features) are output after layer normalization. In the attention branch, an attention matrix is ​​generated by matrix multiplication with the QKV (three matrices in the attention mechanism, where Q represents Query, K represents Key, and V represents Value; attention weights are calculated by the dot product of Q and K, and then weighted summation is performed on V to obtain the output) vector. After hard truncation in the interval [-10, 10], the matrix is ​​combined with layer normalization and spatial features. After splicing and value fusion, the data is sent to the CC fusion module to finally obtain the output features. The output image is then obtained through resolution reconstruction.

[0085] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the wafer defect detection method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0086] This application also provides a wafer defect detection system; please refer to [reference needed]. Figure 6 The wafer defect detection system includes: Image acquisition module 10 is used to acquire images of surface defects on the wafer to be inspected; Feature extraction module 20 is used to extract features from the surface defect image of the wafer to be detected based on a preset feature extraction network and output the multi-scale features corresponding to the surface defect image of the wafer to be detected. The feature fusion module 30 is used to fuse multi-scale features based on a preset decoder to obtain fused features, and to perform frequency domain enhancement processing and explosion-proof attention processing in parallel based on the fused features to obtain output features; The defect determination module 40 is used to generate a wafer defect segmentation image containing preset defect types based on the output features.

[0087] The wafer defect detection system provided in this application, employing the wafer defect detection method described in the above embodiments, can solve the technical problem of poor wafer surface defect detection performance. Compared with the prior art, the beneficial effects of the wafer defect detection system provided in this application are the same as those of the wafer defect detection method provided in the above embodiments, and other technical features of the wafer defect detection system are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0088] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the wafer defect detection method in Embodiment 1 above.

[0089] The following is for reference. Figure 7The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, and PADs (Portable Application Description: Tablet computers), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0090] like Figure 7 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagrams show electronic devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0091] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0092] The electronic device provided in this application, employing the wafer defect detection method described in the above embodiments, can solve the technical problem of poor wafer surface defect detection performance. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the wafer defect detection method provided in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0093] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0094] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0095] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the wafer defect detection method in the above embodiments.

[0096] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0097] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.

[0098] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by an electronic device, the wafer defect detection device causes the following actions: acquiring an image of a surface defect on a wafer to be detected; extracting features from the image of the surface defect on the wafer to be detected based on a preset feature extraction network, and outputting multi-scale features corresponding to the image of the surface defect on the wafer to be detected; fusing the multi-scale features based on a preset decoder to obtain fused features; performing frequency domain enhancement processing and explosion-proof attention processing in parallel based on the fused features to obtain output features; and generating a wafer defect segmentation image containing a preset defect type based on the output features.

[0099] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0100] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0101] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0102] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described wafer defect detection method, thereby solving the technical problem of poor wafer surface defect detection performance. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the wafer defect detection method provided in the above embodiments, and will not be repeated here.

[0103] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the wafer defect detection method described above.

[0104] The computer program product provided in this application can solve the technical problem of poor wafer surface defect detection effect. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the wafer defect detection method provided in the above embodiments, and will not be repeated here.

[0105] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for detecting wafer defects, characterized in that, The wafer defect detection method includes: Acquire images of surface defects on the wafer to be inspected; Based on a preset feature extraction network, feature extraction is performed on the surface defect image of the wafer to be detected, and multi-scale features corresponding to the surface defect image of the wafer to be detected are output. The multi-scale features are fused based on a preset decoder to obtain fused features. Frequency domain enhancement processing and explosion-proof attention processing are performed in parallel based on the fused features to obtain output features. Generate a wafer defect segmentation image containing preset defect types based on the output features; The steps of fusing the multi-scale features based on a preset decoder to obtain fused features, and then performing frequency domain enhancement processing and explosion-proof attention processing in parallel based on the fused features to obtain output features include: The multi-scale features are input into a preset decoder, and the multi-scale features are fused based on the first fusion module and the second fusion module in the preset decoder to obtain fused features. The first fusion module is composed of a preset residual direct connection channel and a series channel connected in parallel. The series channel is composed of a first preset convolutional layer, a first preset layer normalization layer and a first preset linear rectified activation layer connected in series. The second fusion module is composed of a second preset convolutional layer, a second preset layer normalization layer and a second preset linear rectified activation layer stacked alternately. The fused features are input in parallel into the wavelet transform branch and the attention branch of the preset decoder. In the wavelet transform branch, the fused features are subjected to frequency domain enhancement processing to obtain frequency domain enhanced features. In the attention branch, the fused features are subjected to explosion-proof attention processing to obtain spatial attention features. The frequency domain enhancement feature is fused with the spatial attention feature to obtain the output feature; The step of performing frequency domain enhancement processing on the fused features in the wavelet transform branch to obtain frequency domain enhanced features includes: A two-dimensional discrete wavelet transform is performed on the fused features to obtain low-frequency components and high-frequency components in each preset orientation, wherein the preset orientation includes the horizontal direction, the vertical direction and the diagonal direction; A first dynamic threshold interval is generated based on the historical statistics of the fusion features. The low-frequency components are numerically hard-truncated through the first dynamic threshold interval to obtain the first truncated components. The high-frequency components are numerically hard-truncated through the first dynamic threshold interval to obtain the second truncated components. Each of the second truncated components is added element by element to generate a high-frequency comprehensive feature; Based on the edge response intensity of each high-frequency component to each preset defect type, dynamic weighting coefficients are assigned to the high-frequency integrated features, and the high-frequency integrated features are amplified based on the dynamic weighting coefficients to obtain amplified features. The amplification feature is added to the first truncated component to obtain the frequency domain enhancement feature; The step of performing explosion-proof attention processing on the fused features in the attention branch to obtain spatial attention features includes: Based on the fusion features, a query matrix, a key matrix, and a value matrix are generated. The attention score matrix is ​​obtained by multiplying the query matrix with the transpose of the key matrix. Based on a preset piecewise linear truncation function, each matrix element in the attention score matrix is ​​truncated to a second dynamic threshold interval to obtain an attention weight matrix, wherein the second dynamic threshold interval is determined by learnable parameters, and the learnable parameters are updated by the attention branch; The attention weight matrix is ​​multiplied by the value matrix to obtain the attention enhancement feature, and the attention enhancement feature is concatenated with the spatial feature of the preset decoder to obtain the spatial attention feature.

2. The wafer defect detection method as described in claim 1, characterized in that, The wafer defect detection method further includes: Obtain a preset set of wafer surface defect images. For any preset wafer surface defect image in the preset set of wafer surface defect images, use the preset wafer surface defect image and its corresponding image label as training samples. The training samples are input into a preset pyramid network for forward propagation to obtain the layer features corresponding to the output of each network layer of the preset pyramid network. The preset pyramid network includes shallow layers and deep layers. The weight parameters of the shallow layers remain unchanged, while the weight parameters of the deep layers are updated during the training process of the preset pyramid network. The network loss is calculated based on the hierarchical features corresponding to each network layer and the image label, and the weight parameters of the deeper layers of the network are updated in reverse based on the network loss until the preset pyramid network meets the preset training conditions, thus obtaining the preset feature extraction network.

3. The wafer defect detection method as described in claim 2, characterized in that, The step of acquiring a preset set of wafer surface defect images includes: Obtain a preset basic defect image set, wherein the preset basic defect image set includes basic defect images of each preset defect type, and each preset defect type includes at least one of center circle defect, edge defect, random defect and scratch defect; For any basic defect image in the preset basic defect image set, perform a random horizontal flip operation, a random vertical flip operation, and a random rotation operation with a preset probability on the basic defect image to obtain an initial transformed image. The initial transformation images containing different preset defect types are merged to obtain a preset wafer surface defect image, and the preset defect types contained in the preset wafer surface defect image are recorded, wherein the preset wafer surface defect image contains at least one preset defect type; The preset wafer surface defect images are merged into a preset wafer surface defect image set.

4. The wafer defect detection method as described in claim 1, characterized in that, The step of generating a wafer defect segmentation image containing a preset defect type based on the output features includes: The output features are input to a preset segmentation head, and the output features are upsampled in terms of spatial resolution based on the preset segmentation head to restore the spatial size of the output features to the same size as the image of the surface defect of the wafer to be detected, so as to obtain a restored feature map. For any pixel in the restored feature map, classify the pixel into categories and determine the target defect type corresponding to the pixel; Based on the recovered feature map and the target defect type, a wafer defect segmentation image is output, wherein different colors in the wafer defect segmentation image represent different preset defect types, and the preset defect types include the target defect type.

5. An electronic device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the wafer defect detection method as described in any one of claims 1 to 4.

6. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the wafer defect detection method as described in any one of claims 1 to 4.

7. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the wafer defect detection method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Visual identification method and device based on computer processing

    CN119180988A