Semiconductor defect detection method and system, electronic device, storage medium

CN122820692APending Publication Date: 2026-09-25NEXCHIP SEMICON CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611239718.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本发明提供一种半导体缺陷检测方法及系统、电子设备、存储介质,用以解决现有的半导体缺陷检测技术中的抽样复检方式效率和准确性不佳的缺陷,实现更准确、高效的晶圆缺陷识别

Benefits of technology

[0019]第五方面,本发明还提供一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现如上述任一种所述的半导体缺陷检测方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820692A_ABST
    Figure CN122820692A_ABST
Patent Text Reader

Abstract

The application provides a semiconductor defect detection method and system, electronic equipment and a storage medium, and belongs to the technical field of semiconductor defect detection. The method comprises the following steps: extracting defect features of a defect image, a scanning signal and layout design data corresponding to each potential defect respectively to obtain multi-modal features corresponding to each potential defect; performing cross-modal feature fusion on the multi-modal features to obtain a comprehensive defect feature vector corresponding to each potential defect; clustering the comprehensive defect feature vectors corresponding to all potential defects in wafer scanning data; sampling from the clustering results by using different sampling strategies to obtain a sampling list; and performing re-inspection according to the sampling list to obtain a wafer defect detection result. By performing multi-modal feature extraction and clustering on potential defects, the application designs a directional sampling strategy, can effectively cover various defect types, reduces the interference of false defects, and can improve defect detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of semiconductor defect detection technology, and in particular to a semiconductor defect detection method and system, electronic device, and storage medium. Background Technology

[0002] In the semiconductor manufacturing industry, wafer defect detection is a core component in ensuring chip yield and reliability. As integrated circuit process nodes continue to shrink and process complexity increases dramatically, the sensitivity to defects also grows exponentially. Even minute particle contamination, pattern anomalies, or material defects can lead to device failure. Therefore, efficient and accurate defect detection technology has become an indispensable support for advanced process development and mass production.

[0003] The current industry standard for defect detection is primarily based on a hybrid approach of "optical initial screening + electron beam re-inspection." Specifically, firstly, a high-speed optical scanning device or electron beam scanning system is used to perform a full-surface scan of the wafer, capturing the coordinates of a massive number of potential defects and their corresponding localized tiny images (patch images). Subsequently, hundreds of samples are randomly selected from tens of thousands of inspection results and re-inspected using a high-resolution scanning electron microscope (SEM) to confirm the authenticity of the defects and determine their physical type (such as bridging, residue, voids, particles, etc.).

[0004] However, with continuous advancements in process nodes, the limitations of this traditional sampling and re-inspection model have become increasingly apparent. First, random sampling is inherently prone to blindness; in the context of highly non-uniform defect distribution, it is extremely easy to miss critical but rare systemic defects or emerging defect patterns, leading to incomplete information acquisition. Second, the increasing complexity and concealment of defect patterns make traditional rule-based or simple statistical sampling criteria inadequate, failing to accurately pinpoint critical defects affecting yield. These shortcomings result in severe challenges for existing methods in terms of detection efficiency, cost control, and the depth of root cause analysis, thereby slowing down process debugging cycles and hindering yield ramp-up.

[0005] Therefore, the semiconductor industry urgently needs to develop more intelligent, targeted and adaptive defect detection strategies to break through the inherent bottlenecks of the random sampling framework and achieve more accurate and efficient defect identification. Summary of the Invention

[0006] This invention provides a semiconductor defect detection method and system, electronic device, and storage medium to address the shortcomings of existing semiconductor defect detection technologies, such as the poor efficiency and accuracy of sampling and re-inspection methods, and to achieve more accurate and efficient wafer defect identification.

[0007] In a first aspect, the present invention provides a semiconductor defect detection method, comprising: Extract the defect image and scan signal corresponding to each potential defect from the wafer scan data, and extract the layout design data corresponding to each potential defect from the wafer layout data; Defect features are extracted from the defect image, the scanning signal, and the layout design data respectively to obtain the multimodal features corresponding to each potential defect; Cross-modal feature fusion is performed on the multimodal features to obtain a comprehensive defect feature vector corresponding to each potential defect; Cluster the comprehensive defect feature vectors corresponding to all potential defects in the wafer scanning data, and identify false defects and real defects based on the clustering results; Different sampling strategies are used to sample from the false defects and the real defects respectively to obtain a sampling list; The wafer is re-inspected according to the sampling list to obtain the defect detection results.

[0008] According to a semiconductor defect detection method provided by the present invention, the step of extracting defect features from the defect image, the scanning signal, and the layout design data respectively to obtain multimodal features corresponding to each potential defect includes: The visual difference features between the defect image and the reference image are extracted using a visual difference feature extraction model. The location and geometric features of the potential defects are extracted from the scan signal using a scan feature extraction model. The design context features of the potential defects are extracted from the layout design data using a layout feature extraction model. The multimodal features corresponding to the potential defects include the visual difference features, the positional features, the geometric features, and the design context features.

[0009] According to a semiconductor defect detection method provided by the present invention, the visual differential feature extraction model includes a multi-scale feature extraction network, a feature alignment network, a differential fusion module, an attention focusing module, and a feature description-transformation module; The step of extracting visual difference features between the defect image and the reference image using a visual difference feature extraction model includes: The multi-scale feature extraction network is used to extract multi-scale features from the defect image and the reference image respectively, to obtain the multi-scale defect features of the defect image and the multi-scale reference features of the reference image. The feature alignment network is used to perform sub-pixel alignment of the defect features and the reference features at each scale to obtain reference features aligned with the defect features at each scale. The differential fusion module calculates the absolute difference between the defect feature at each scale and the reference feature aligned with the defect feature, and performs scale adjustment and multi-scale fusion on the absolute difference to obtain the fused differential feature. The attention focusing module enhances the fused differential features by channel attention and spatial attention, resulting in enhanced fused differential features. The enhanced fusion difference features are described and concatenated using the feature description-transformation module to obtain concatenated features; wherein the diversified feature description includes at least one of statistical moment features, quantile features, and texture features. The splicing features are subjected to dimensionality reduction and normalization to obtain the visual difference features between the defective image and the reference image.

[0010] According to a semiconductor defect detection method provided by the present invention, the feature alignment network includes a convolutional network and a deformable convolutional layer; The step of performing sub-pixel alignment of the defect features and the reference features at each scale using the feature alignment network to obtain reference features aligned with the defect features at each scale includes: The convolutional network is used to predict the dense offset field from the reference feature to the defect feature at each scale; The pixel offset between the reference feature and the defect feature is determined based on the dense offset field. Based on the pixel offset, the reference features are resampled using the deformable convolutional layer to obtain reference features aligned with the defect features at each scale.

[0011] According to a semiconductor defect detection method provided by the present invention, the loss function of the visual difference feature extraction model includes: A symmetric consistency loss function that measures the feature differences between two different reference images; A reconstruction loss function that measures the difference between the reference image and the reconstructed reference image; wherein the reconstructed reference image is obtained based on the visual difference features; A difference maximization loss function measures the feature differences between the defective image and the reference image; The total loss function of the visual difference feature extraction model is obtained by weighted summation of the symmetric consistency loss function, the reconstruction loss function, and the difference maximization loss function.

[0012] According to a semiconductor defect detection method provided by the present invention, the scanning signal includes the location, area, and size information of the potential defect; The step of extracting the location and geometric features of the potential defect from the scan signal using a scan feature extraction model includes: The location features of the potential defect are obtained by performing different types of coordinate transformations and feature extractions on the location coordinates using the scanning feature extraction model. The geometric features of the potential defect are obtained by performing logarithmic transformation and Gaussian mixture model on the area and size using the scanning feature extraction model.

[0013] According to a semiconductor defect detection method provided by the present invention, the step of extracting the design context features of the potential defects from the layout design data through a layout feature extraction model includes: The layout design data is embedded, encoded, concatenated, and extracted using a layout feature extraction model to obtain the design context features of the potential defects. The layout design data corresponding to the potential defects includes at least one of the following: layer information, local graphic density, adjacent structure information, functional area information, and design rule information.

[0014] According to a semiconductor defect detection method provided by the present invention, the step of clustering the comprehensive defect feature vectors corresponding to all potential defects in the wafer scanning data, and identifying false defects and real defects based on the clustering results, includes: The comprehensive defect feature vectors corresponding to all potential defects in the wafer scanning data are subjected to dimensionality reduction processing to obtain the comprehensive defect feature vectors of each potential defect in the low-dimensional space. A density-based clustering method is used to cluster the comprehensive defect feature vector in the low-dimensional space to obtain clustering results; wherein, the clustering results include at least one noise point and multiple clusters; Each potential defect corresponding to a noise point is treated as a false defect, and all noise points are combined to form a set of false defects. The potential defects corresponding to the remaining clusters are taken as real defects; wherein, the real defects include known defects and unknown defects.

[0015] According to a semiconductor defect detection method provided by the present invention, the step of sampling from the clustering results using different sampling strategies to obtain a sampling list includes: For each cluster corresponding to the known defects, sampling is carried out using equal sampling or weighted sampling strategies based on the number of samples in each cluster to obtain the sampled samples of the known defects. For each cluster corresponding to the unknown defect, an exploratory sampling strategy is used to sample and obtain a sample of the unknown defect. For the set of false defects, a random sampling strategy is used to obtain a sample of the false defects. A sampling list is constructed based on the sampling samples of the known defects, the sampling samples of the unknown defects, and the sampling samples of the false defects; wherein the sampling quantity of the known defects is greater than the sampling quantity of the false defects, and the sampling quantity of the false defects is greater than the sampling quantity of the unknown defects.

[0016] In a second aspect, the present invention also provides a semiconductor defect detection system, comprising: The data acquisition module is used to extract the defect image and scanning signal corresponding to each potential defect from the wafer scan data, and to extract the layout design data corresponding to each potential defect from the wafer layout data; The feature extraction module is used to extract the defect features from the defect image, the scanning signal, and the layout design data respectively, to obtain the multimodal features corresponding to each potential defect; The feature fusion module is used to perform cross-modal feature fusion on the multimodal features to obtain a comprehensive defect feature vector corresponding to each potential defect; The clustering module is used to cluster the comprehensive defect feature vectors corresponding to all potential defects in the wafer scanning data, and to identify false defects and real defects based on the clustering results. The sampling and re-inspection module is used to sample from the false defects and the real defects using different sampling strategies to obtain a sampling list; and to perform re-inspection based on the sampling list to obtain the defect detection results of the wafer.

[0017] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the semiconductor defect detection method as described above.

[0018] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the semiconductor defect detection method as described above.

[0019] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the semiconductor defect detection method as described above.

[0020] The beneficial effects of the technical solutions provided by some embodiments of the present invention include at least the following: This invention provides a semiconductor defect detection method and system, electronic device, and storage medium. It extracts multimodal features corresponding to each potential defect from wafer scan data and wafer layout data, improving the feature richness of potential defects and achieving accurate representation of their features. Furthermore, this invention clusters the comprehensive defect feature vectors of all potential defects to identify false and true defects. It also designs diverse sampling strategies to achieve a shift from "random" to "targeted" sampling, effectively covering various defect types, reducing interference from false defects, and effectively detecting emerging defects. Re-inspection based on this intelligent sampling result can significantly improve re-inspection efficiency, achieving accurate and complete defect type confirmation and providing clear and quantitative data support for subsequent process improvements.

[0021] This invention provides a semiconductor defect detection method and system, electronic device, and storage medium. Its unexpected technical advantages are as follows: This invention performs multimodal feature extraction and feature fusion on defect images, scan signals, and layout design data of potential defects. Then, utilizing the powerful adaptive and noise recognition capabilities of unsupervised clustering, it identifies false defects, overcoming the shortcomings of traditional methods that can only passively accept interference from false defects. This reduces the interference from false defects and automatically discovers patterns in the data, making it applicable to scenarios where defect types are unknown in the early stages of R&D, thus reducing the difficulty of defect identification. Furthermore, this invention employs different sampling strategies to sample from known defects, unknown defects, and false defects, comprehensively covering all key defect types in the wafer, significantly improving the efficiency of wafer defect improvement during the R&D process and shortening the R&D cycle. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0023] Figure 1 This is a schematic flowchart of a semiconductor defect detection method provided by the present invention; Figure 2 This is a schematic diagram of the structure of the visual difference feature extraction model provided by the present invention; Figure 3 This is a schematic diagram of the defect image encoder in the multi-scale feature extraction network provided by the present invention; Figure 4 This is a schematic diagram of the residual block structure in the defect image encoder provided by the present invention; Figure 5This is a schematic diagram of the feature alignment network provided by the present invention; Figure 6 This is a schematic diagram of the differential fusion module provided by the present invention; Figure 7 This is a schematic diagram of the attention focusing module provided by the present invention; Figure 8 This is a schematic diagram of the feature description-transformation module provided by the present invention; Figure 9 This is a schematic diagram of the sampling strategy provided by the present invention; Figure 10 This is a schematic diagram of the semiconductor defect detection system provided by the present invention; Figure 11 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0025] Example 1 Please see Figure 1 , Figure 1 One of the schematic flowcharts of a semiconductor defect detection method provided in an embodiment of the present invention includes: S101. Extract the defect image and scanning signal corresponding to each potential defect from the wafer scan data, and extract the layout design data corresponding to each potential defect from the wafer layout data. S102. Extract the defect features from the defect image, scan signal, and layout design data respectively to obtain the multimodal features corresponding to each potential defect; S103. Perform cross-modal feature fusion on the multimodal features to obtain the comprehensive defect feature vector corresponding to each potential defect; S104. Cluster the comprehensive defect feature vectors corresponding to all potential defects in the wafer scanning data, and identify false defects and real defects based on the clustering results. S105. Using different sampling strategies, samples are taken from false defects and real defects respectively to obtain a sampling list; S106. Conduct a re-inspection based on the sampling list to obtain the defect detection results of the wafer.

[0026] This invention extracts multimodal features corresponding to each potential defect from wafer scan data and wafer layout data, improving the feature richness of potential defects and achieving accurate expression of potential defect features. Based on this, the invention clusters the comprehensive defect feature vectors of all potential defects to identify false and real defects. Furthermore, it designs diverse sampling strategies to achieve a shift from "random" to "targeted" sampling, effectively covering various defect types, reducing interference from false defects, and effectively identifying emerging defects. Re-inspection based on this intelligent sampling result can significantly improve re-inspection efficiency, achieving accurate and complete defect type confirmation, and providing clear and quantitative data support for subsequent process improvements.

[0027] The semiconductor defect detection method provided by this invention has the following unexpected technical effects: This invention performs multimodal feature extraction and feature fusion on defect images, scan signals, and layout design data of potential defects. Then, it utilizes the powerful adaptive and noise recognition capabilities of unsupervised clustering to identify false defects, overcoming the shortcomings of traditional methods that can only passively accept interference from false defects. This reduces the interference from false defects and can automatically discover patterns in the data, making it applicable to scenarios where defect types are unknown in the early stages of R&D, thus reducing the difficulty of defect identification. Furthermore, this invention employs different sampling strategies to sample from various types of real and false defects, comprehensively covering all key defect types in the wafer, significantly improving the efficiency of wafer defect improvement during the R&D process and shortening the R&D cycle.

[0028] In step S101, wafer scan data is acquired in advance, and the defect image and scan signal corresponding to each potential defect are extracted from the wafer scan data. Based on the coordinates of each potential defect, the layout design data corresponding to each potential defect is extracted from the corresponding wafer layout data.

[0029] During or after wafer fabrication, optical or electron beam inspection methods are typically used to scan the entire wafer surface. Combined with a reference patch (a standard image without defects), the grayscale differences between the defect image and the reference image are used to initially identify potential defects on the wafer. After the inspection equipment completes the scan, it typically generates a scan result file containing rich information. This file is usually in a standard spreadsheet format, such as KLARF or other industry-standard formats. This scan result file serves as wafer scan data, including the coordinates of each potential defect and the corresponding small defect patch image. In addition to this basic defect location and image data, the wafer scan data also includes scan signals corresponding to potential defects, such as geometric features like size, area, and morphology, basic attributes like signal strength, spatial distribution information, and other metadata. These scan signals are numerical data derived from the scanning results of the inspection equipment and are used to discover and quantify defects.

[0030] Wafer layout data is design data used to describe the physical topology of integrated circuits. It is usually stored in GDS files, which use a hierarchical structure and geometric primitives to accurately represent all the mask patterns required in the chip manufacturing process.

[0031] Based on the coordinates of each potential defect obtained from wafer scanning data, this invention extracts the layout design data corresponding to each potential defect from the wafer layout data corresponding to that wafer. For example, the layer information, orientation information, geometric shape information, local pattern density, adjacent structure information, functional area information, design rule information, etc. of the potential defect.

[0032] The information includes: Layer information (e.g., the process layer where the potential defect falls, such as the active area layer (AA), polysilicon gate layer (POLY), metal interconnect layer (Metal1), etc.), and if the potential defect spans multiple layers, the adjacent layers that the potential defect may affect in the vertical direction can also be extracted through inter-layer relationships. Orientation information can be the orientation angle of the pattern containing the potential defect relative to the wafer coordinate system. Geometric morphology information can include the radius of curvature of the corners of the pattern containing the potential defect, the type of corner, and whether the potential defect is located at a line end. Local pattern density refers to the area proportion of the pattern on a specific process layer within a local window centered on the coordinates of the potential defect. Adjacent structure information refers to the pattern layout characteristics within a certain range around the defect, including the relationship between the potential defect and nearby patterns, such as the straight-line distance from the defect center to the edge of the nearest geometric pattern, the type of adjacent patterns, and whether the area around the potential defect is a dense wiring area, an isolated line area, or a blank area. Functional area information refers to the circuit function type corresponding to the location of the potential defect, such as the core logic area, SRAM cache area, analog / RF area, I / O interface area, etc. Design rule information refers to the relationship between the location of a defect and the design rule constraints of the corresponding process layer, such as minimum spacing rules, minimum width rules, enclosing / covering rules, area rules, etc. Based on the design rule information, it is possible to check whether the location of a potential defect violates or is close to violating the design rules, and it can also be used to distinguish between process problems and design problems.

[0033] This invention extracts multi-dimensional information such as defect images, scanning signals, and layout design data corresponding to each potential defect to form multimodal data, which serves as the basis for subsequent wafer defect monitoring. This achieves comprehensive integration of potential defect information, allowing for complementary advantages and corrections, reducing false detection and false negative rates, and improving the robustness and generalization ability of overall defect classification.

[0034] In step S102 above, multimodal feature extraction is performed on the multimodal data such as defect image, scanning signal and layout design data of each potential defect to obtain the multimodal features corresponding to each potential defect.

[0035] For example, for defective images, local detail features, such as edge features and texture features, can be quickly extracted using a convolutional neural network model. Furthermore, based on these local detail features, a Transformer model can be used to capture their global contextual information to obtain the corresponding defect features of the defective image. In this way, defect features spanning large areas, such as scratches, can be extracted more effectively.

[0036] For example, for scan signals, since they are numerical data related to potential defects, these scan signals can be normalized and spliced, and then feature extraction can be performed through a small neural network model to obtain the defect features corresponding to the scan signals.

[0037] For example, for layout design data, non-numerical data such as layer information and functional area information can be embedded and encoded, and then concatenated with numerical data in local graphic density, orientation information, and adjacent structure information to form a vectorized layout design data. Then, it is input into a small neural network model for feature extraction to obtain the defect features corresponding to the layout design data.

[0038] The multimodal features of a potential defect can be constructed by combining the defect features corresponding to the defect image, the defect features corresponding to the scan signal, and the defect features corresponding to the layout design data.

[0039] In some possible embodiments, defect features are extracted from defect images, scan signals, and layout design data respectively to obtain multimodal features corresponding to each potential defect, including: Visual difference features between the defect image and the reference image are extracted using a visual difference feature extraction model. The location and geometric features of potential defects are extracted from the scan signal using a scan feature extraction model. The layout feature extraction model is used to extract design context features of potential defects from layout design data. Among them, the multimodal features corresponding to potential defects include visual difference features, positional features, geometric features, and design context features.

[0040] Specifically, when scanning a wafer using inspection equipment, the scanned image is typically compared with a reference image to initially locate defects. Therefore, the defect image and its reference image can be used as paired images, input into a visual difference feature extraction model to extract visual difference features between the defect image and the reference image, ignoring identical backgrounds and highlighting differences. This invention also extracts numerical features such as the location, size, and area of ​​potential defects from the scanning signal to enrich the feature representation of potential defects. Furthermore, layout design data contains multi-dimensional quantitative parameters that describe the physical design environment, geometric layout characteristics, process sensitivity, and electrical effects of the defect location. These features transform abstract layout graphics into quantifiable numerical or categorical information, which can be used to establish a correlation model between defects and the design. Therefore, this invention extracts multi-level, multi-dimensional design context features from layout design data, such as basic geometric features like layer, linewidth, and spacing; statistical distribution features like graphic density, orientation distribution, and line end density; and electrical functional features like functional area type and current density.

[0041] The following sections will describe the implementation methods for feature extraction using the visual difference feature extraction model, the scanning feature extraction model, and the layout feature extraction model.

[0042] For example, the visual difference feature extraction model can be a pre-trained convolutional neural network model. For instance, the convolutional neural network model can be used to extract features from the defect image and the reference image respectively, and then the extracted image features can be differentially calculated to obtain visual difference features.

[0043] In some possible embodiments, the visual difference feature extraction model includes a multi-scale feature extraction network, a feature alignment network, a difference fusion module, an attention focusing module, and a feature description-transformation module connected in sequence. A multi-scale feature extraction network is used to extract multi-scale features from the defect image and the reference image respectively, resulting in multi-scale defect features of the defect image and multi-scale reference features of the reference image. The feature alignment network is used to perform sub-pixel alignment of the defect features and reference features at each scale to obtain the reference features aligned with the defect features at each scale. The differential fusion module is used to calculate the absolute difference between the defect features and the reference features aligned with the defect features at each scale, and to perform scale adjustment and multi-scale fusion on the absolute difference to obtain the fused differential features. The attention focusing module is used to enhance the channel attention and spatial attention of the fused difference features to obtain enhanced fused difference features; The feature description-transformation module is used to perform diversified feature description and feature concatenation on the enhanced fusion difference features to obtain a concatenated feature matrix; the concatenated feature matrix is ​​subjected to dimensionality reduction and normalization to obtain the visual difference features between the defective image and the reference image; wherein, the diversified feature description includes at least one of statistical moment features, quantile features and texture features.

[0044] Specifically, such as Figure 2 The diagram shown is a structural schematic of the visual difference feature extraction model provided in an embodiment of the present invention. The following is a description of the model in conjunction with... Figure 2 The multi-scale feature extraction network, feature alignment network, difference fusion module, attention focusing module, and feature description-transformation module of the visual difference feature extraction model of the present invention will be described separately.

[0045] exist Figure 2 In the visual difference feature extraction model shown, the multi-scale feature extraction network is used to extract multi-scale features of the defect image and the reference image respectively. For example, the defect image and the reference image can be extracted using a dual encoder with shared weights to ensure that the defect image and the reference image are understood in the same visual language system.

[0046] For example, a pre-trained convolutional neural network model, such as a ResNet34 model, can be loaded. The last layer of the ResNet34 model can be removed, and its convolutional layers can be retained as a defect image encoder for a visual difference feature extraction model, used to extract multi-scale features of defect images. The formula for layer forward propagation is:

[0047] in, For the first The output of the layer is also the current layer. Layer input, For its first The output features of the layer , They are the first Layer weight matrix and bias terms, For example, the ReLU function.

[0048] Extracting from three scales: shallow, intermediate, and deep ( For example, the characteristics of ) Figure 3 The diagram shows the structure of the defect image encoder, which includes an input layer, a preprocessing unit, a shallow feature extraction unit, a mid-level feature extraction unit, a deep feature extraction unit, and an output layer connected in sequence. The input layer is used to input the defect image. The preprocessing unit includes a 7×7 convolutional layer, a batch normalization layer (BatchNorm), a ReLU activation layer, and a max pooling layer connected in sequence. This preprocessing unit is mainly used for preliminary feature extraction and downsampling of the input defect image, facilitating further feature refinement. The shallow feature extraction unit includes three residual blocks connected in sequence for extracting shallow features. The mid-level feature extraction unit includes four residual blocks connected in sequence for extracting mid-level features. The deep feature extraction unit includes six residual blocks connected in sequence for extracting deep features. Each residual block includes two 3×7 convolutional layers, such as... Figure 4 The diagram shown is a schematic of the residual block. This represents an element-wise addition operation. The mapping layer is either an identity mapping or a projection mapping. If it is a projection mapping, a batch normalization layer is added after it.

[0049] For example, a convolutional neural network model with the same structure as the defect image encoder described above can be used as the reference image encoder for the visual difference feature extraction model to extract multi-scale features from the reference image. This reference image encoder shares weights with the defect image encoder, forming a dual-channel encoder that extracts multi-scale features from the defect image and the reference image respectively.

[0050] Then, the defect image and the reference image are input into the visual difference feature extraction model. The outputs of the intermediate layers of the defect image encoder and the reference image encoder of the visual difference feature extraction model are extracted as multi-scale features, resulting in two strictly corresponding feature pyramids. Figure 3 Taking the defect image encoder shown as an example, assuming that the reference image encoder has the same structure as the defect image encoder, the outputs of the middle three layers of the defect image encoder and the reference image encoder are taken respectively. Defect images can be obtained. Corresponding multi-scale defect features Reference image Corresponding multi-scale reference features :

[0051]

[0052] in, , The three dimensions correspond to shallow, medium, and deep semantic information, respectively, and represent features at three scales. (Taking defect images as an example.) For example, its characteristics at the shallow, medium, and deep scales can be represented as follows: Shallow features : The output size is [64, H / 4, W / 4], used to capture edge / texture features; Mid-layer features : Output dimensions [128, H / 8, W / 8], used to capture shape features; Deep features : The output size [256, H / 16, W / 16] is used to capture semantic concepts.

[0053] Where H and W are the input defect images, respectively. Height and width, for reference image Its height and width should be the same as the defect image; For the softmax function, , The first l ( l The weight matrix and bias terms of the features of layers 1, 2, and 3.

[0054] exist Figure 2 In the visual difference feature extraction model shown, the feature alignment network is used to perform feature alignment on each scale in the feature space. The following defect characteristics Reference features Perform differentiable, non-rigid fine alignment to achieve sub-pixel alignment in the feature space.

[0055] For example, the cross-attention mechanism of Transformer can be used to allow defective features to actively search for pixel information that needs to be aligned in reference features, achieving implicit alignment. For instance, the U-Net-style Transformer model UFormer is an end-to-end feature point matching model based on Transformer and U-Net structures. It can establish feature associations between defects and reference features through cross-attention during the encoding stage, fuse global and local attention information at each scale to achieve fine semantic alignment, and finally refine the coarse matching to the sub-pixel level through upsampling and expectation calculation.

[0056] In some possible embodiments, the feature alignment network includes a convolutional network and deformable convolutional layers; A convolutional network is used to predict the dense offset field from the reference feature to the defect feature at each scale, and to determine the pixel offset between the reference feature and the defect feature based on the dense offset field. Deformable convolutional layers are used to deformably resample reference features based on pixel offsets to obtain reference features aligned with defect features at each scale.

[0057] Specifically, such as Figure 5 The diagram shows a schematic of the feature alignment network provided in an embodiment of the present invention. The feature alignment network includes a feature splicing layer, a convolutional network, and a deformable convolutional layer. The feature splicing layer is used to perform a splicing operation on the input reference features and defect features before feeding them into the convolutional network. The convolutional network includes a 3×3 convolutional layer, a batch normalization layer, a ReLU activation layer, and a 1×1 convolutional layer connected in sequence. The deformable convolutional layer is used to determine the sampling position based on the pixel offset and to perform sampling using bilinear interpolation to achieve sub-pixel alignment in the feature space and output the reference features aligned with the defect features at each scale.

[0058] like Figure 5 As shown, the expression for predicting the dense offset field from the reference feature to the defect feature at each scale using a convolutional network can be:

[0059] in, Represents a dense offset field. This represents the splicing operation of the feature splicing layer. Representing scale respectively Reference features and defect features are listed below. Representing a convolutional network, its expression is:

[0060] in, These represent 3×3 convolution operations, batch normalization operations, and... Activation function, 1×1 convolution operation. X represents the input of the convolutional network. In this embodiment, X = .

[0061] The dense migration field obtained from the above formula It can extract the corresponding two-dimensional offset for each target location on the reference feature. Based on this two-dimensional offset, the reference features are processed through deformable convolutional layers. Deformable resampling is performed to obtain the defect features at each scale. Alignment reference features .

[0062] For reference features The expression for standard convolutional sampling is:

[0063] in, Sampling points on the reference feature, The position of the convolution kernel sampling grid R The weight, This is the sampling result of standard convolution.

[0064] For reference features The expression for deformable convolution sampling is:

[0065] in, For the aligned reference feature at the current target position The value, R represents the sampling points on the sampling object (i.e., the reference feature), and R is the sampling grid of the convolution kernel. For the sampling points in R relative to the target position A fixed offset; Position in R The corresponding weights.

[0066] The comparison shows that, compared to standard convolution, deformable convolutional layers increase the learnable two-dimensional offset. It can perform flexible sampling with offset, improving alignment accuracy.

[0067] Furthermore, to achieve sub-pixel precision alignment, bilinear interpolation is used during deformable convolution to achieve fractional-pixel precision shifting, ensuring gradient propagation:

[0068] in, These are non-integer sampling points. x and y coordinates For reference features Non-integer sampling points The sampling results For reference features Integer sampling points on Eigenvalues ​​at; These are the floor down and floor up operators, respectively.

[0069] Through the above deformable sampling, sub-pixel alignment of the feature space from the reference feature to the defect feature can be achieved, resulting in the reference feature aligned with the defect feature at each scale. Aligning reference features with defect features can avoid introducing noise when transforming defect features, thus preventing the loss or obscuring of true defects and preserving the original defect information completely. This ensures that downstream feature clustering tasks are based on true defect features.

[0070] This invention uses a feature alignment network to perform deformable alignment of defective images and reference images in the feature space, achieving sub-pixel precision feature alignment and reducing the impact of noise.

[0071] exist Figure 2 In the visual difference feature extraction model shown, the difference fusion module is used to perform difference calculation and multi-scale fusion on the defect features and reference features aligned by the feature alignment network to obtain fused difference features.

[0072] like Figure 6 The diagram shown illustrates the structure of a differential fusion module provided in an embodiment of the present invention. This differential fusion module may include an absolute difference calculation layer, an upsampling layer, and a feature concatenation layer connected sequentially. The absolute difference calculation layer is used to perform shallow feature absolute difference calculation, mid-level feature absolute difference calculation, and deep feature absolute difference calculation, respectively. For each layer, the corresponding layer (scale) is calculated. The absolute difference between the defect features and the reference features aligned with the defect features is calculated; the upsampling layer is used to upsample the shallow and deep absolute differences obtained from the absolute difference calculation layer to the intermediate scale to unify the size; the feature stitching layer is used to fuse the sampling results obtained from the upsampling layer with the absolute differences at the intermediate scale to obtain the fused difference features.

[0073] Specifically, each scale is calculated separately using an absolute difference layer. Calculate the absolute difference:

[0074] in, For scale Below and defect characteristics Alignment reference features, For a very small positive number, for example, we can take .

[0075] The absolute differences at each scale are upsampled to an intermediate scale m using an upsampling layer, still using three scales. For example, the intermediate scale m is... Scale:

[0076] in, To upsample the absolute difference of shallow features to the sampling results at an intermediate scale m, To upsample the absolute difference of deep features to the sampling results at an intermediate scale m, The feature map size is at the intermediate scale m. This represents an upsampling operation based on bilinear interpolation.

[0077] Finally, multi-scale difference fusion is performed through a feature concatenation layer to obtain fused difference features. :

[0078] in, The concatenation operation of the representative feature concatenation layer fuses the differential features. The number of channels C= , These are fusion difference features Height and width.

[0079] exist Figure 2 In the visual difference feature extraction model shown, the attention focusing module is used to enhance the fused difference features through channel attention mechanism and spatial attention mechanism to obtain enhanced fused difference features.

[0080] like Figure 7The diagram shows a schematic of the attention focusing module provided in an embodiment of the present invention. This attention focusing module includes a channel attention module and a spatial attention module. The channel attention module enhances the fused differential features through a channel attention mechanism. It includes a globally average pooling layer and two fully connected layers connected in sequence, each followed by an activation function layer. Finally, it performs recalibration through a channel-wise multiplication operation. The spatial attention module further enhances the features enhanced by the channel attention module through the channel attention mechanism. It includes an average pooling layer, a max pooling layer, a feature concatenation layer, a 7×7 convolutional layer, and an activation layer. Finally, it performs spatial broadcasting through an element-wise multiplication operation. Figure 7 middle, This is a channel-by-channel multiplication operation. This represents element-wise multiplication.

[0081] Specifically, the goal of the channel attention module is to generate a one-dimensional weight vector to measure the importance of different feature channels in the fused differential features. The feature enhancement process using the channel attention module can include three steps: compression, activation, and recalibration. a. Compression: To aggregate spatial information, a global average pooling layer (GAP) is used to compress the two-dimensional feature map on each channel c into a scalar:

[0082] in, To The global feature of the c-th channel after compression, where c = 1, 2, ..., C, and C is... The total number of channels; Represents the global average pooling operation of the global average pooling layer; Representative fusion difference features In channel c, position The feature value at point c. Global features of each channel c. The complete channel descriptor vector Z is formed. .

[0083] b. Incentives: To capture the dependencies between channels, two fully connected layers are used to learn the non-linear interactions between channels, generating weights for each channel. The first fully connected layer is used for dimensionality reduction, and the second fully connected layer is used to restore dimensionality. Finally, an activation function is used to normalize the weight values ​​to the [0, 1] interval.

[0084] in, This is the final channel attention weight vector. , These represent the weights and biases of the first fully connected layer, respectively. , These represent the weights and biases of the second fully connected layer, respectively. All are activation functions, such as It can be the Sigmoid activation function. It can be the ReLU activation function.

[0085] c. Recalibration: Apply the learned channel attention weights s to the original fused differential features. The above involves performing channel-by-channel weighting to obtain the channel-enhanced feature map. :

[0086] in, This is a channel-by-channel multiplication operation.

[0087] In obtaining channel-enhanced feature maps Next, a two-dimensional spatial weight map is generated through the spatial attention module to indicate which spatial locations (pixel regions) in the feature map should be given priority. The feature enhancement process of the spatial attention module can include several steps: channel aggregation, spatial weight generation, and spatial broadcasting. d. Channel aggregation: Feature maps obtained from the channel attention module are processed through average pooling and max pooling layers in the channel dimension, respectively. Compression is performed to generate two single-channel feature maps, representing the average pooling feature and the max pooling feature of the channel dimension, respectively:

[0088] in, These represent the average pooling operation of the average pooling layer and the max pooling operation of the max pooling layer, respectively. These represent the average pooling characteristic and the max pooling characteristic, respectively. D c ( c ,:,:) represents a feature map of channel enhancement. Features in channel c dimension.

[0089] The generated average pooling and max pooling feature maps are concatenated to form a two-channel feature map:

[0090] in, This is a channel aggregation feature map. This represents the splicing operation of the feature splicing layer.

[0091] e. Spatial weight generation: Aggregate feature maps of channels using a standard 7×7 convolutional layer. The process is performed to generate the final spatial attention weight map:

[0092] in, For activation function, This represents a 7×7 convolution operation. The calculation process can be expanded as follows:

[0093] in, represent In position eigenvalues ​​at that location The convolution kernel for the e-th input channel at the offset The weights on the surface, e=1,2, Channel aggregation feature map In the e-th channel, at position ( ) eigenvalues ​​on ).

[0094] f. Spatial Broadcasting: This generates a spatial attention weight map. Feature maps applied to channel enhancement The feature map is obtained by applying both channel and spatial attention mechanisms. :

[0095] in, This represents element-wise multiplication. In channel c, position eigenvalues ​​at , Feature maps for channel enhancement In channel c, position Eigenvalues ​​at; Spatial attention weight map In position The weight value at that location.

[0096] The feature map enhanced by both channel and spatial attention mechanisms This refers to the enhanced fusion difference feature ultimately obtained by the attention-focusing module.

[0097] This invention achieves multi-level difference and dual attention fusion through a difference fusion module and an attention focusing module, effectively integrating multi-scale difference information and enhancing the expressive power of key difference information.

[0098] exist Figure 2In the visual difference feature extraction model shown, the feature description-transformation module is used to perform diversified feature descriptions and feature splicing, dimensionality reduction, normalization and other processing on the enhanced fusion difference features output by the attention focusing module to obtain the visual difference features between the defective image and the reference image. These visual difference features are the final output of the visual difference feature extraction model.

[0099] like Figure 8 The diagram shows the structure of the feature description-transformation module provided by the present invention. The feature description-transformation module includes a feature description layer, a feature concatenation layer, a feature mapping layer, and a normalization layer. The feature description layer is used to perform diversified feature descriptions on the input enhanced fusion difference features through generalized pooling operations, quantile feature calculation, texture feature calculation, etc. The feature concatenation layer is used to concatenate the output results of the feature description layer. The feature mapping layer is used to perform dimensionality reduction processing on the concatenated feature matrix obtained by the feature concatenation layer, that is, to map from a high-dimensional space to a low-dimensional space.

[0100] Specifically, first, the enhanced fusion difference features... Diverse feature descriptions can be implemented, such as capturing rich statistical information through generalized pooling operations to obtain statistical moment features, and calculating quantile features, texture features, etc. These diverse feature descriptions can capture multi-dimensional information, enrich feature representation, and have complementary robustness to different types of noise and interference.

[0101] Among them, statistical moment features can be calculated by computing enhanced fusion difference features. The mean, variance, skewness, kurtosis, and other indicators are used to measure the performance of each channel c.

[0102] ①Mean (First-order moment):

[0103] in, for The eigenvalue at position (i,j) in channel c.

[0104] ② Variance (Second-order central moment):

[0105] ③ Skewness (Third-order normalized moments):

[0106] ④ Kurtosis (Fourth-order normalized moments):

[0107] in,i =1,2,…, , j =1,2,…, , , These are fusion difference features Height and width To fuse differential features In channel c, position The eigenvalue at that location.

[0108] Quantile features can be used to calculate enhanced fusion difference features. Q quantile To measure, for example, quartiles .

[0109] Texture features can be enhanced by fusing differential features. It is measured by metrics such as gray-level co-occurrence matrix, contrast, and energy. The expression for the gray-level co-occurrence matrix with a distance of 1 and an orientation of 0° is:

[0110] in, The gray-level co-occurrence matrix of channel c In position The value, at the same time, It also represents grayscale levels; , indicating that the statistics satisfy and All pixels under this condition Quantity, For counting symbols, For logical AND, The step size for discretizing continuous features, To enhance the fusion difference features In channel c, at position ( eigenvalues ​​at that location This is the rounding function.

[0111] Texture contrast is used to measure local grayscale changes. The expression for contrast is:

[0112] in, To enhance the fusion difference features In terms of contrast in channel c, =1,2,...,L, where L is the total number of gray levels after discretization.

[0113] Texture energy measures the uniformity of a texture, and its expression is:

[0114] in, To enhance the fusion difference features Texture energy in channel c.

[0115] Then, the diverse feature descriptions of channel c, such as statistical moment features, quantile features, and texture features, are concatenated through a feature concatenation layer to obtain the feature vector for each channel c. It is understandable that the above diverse feature descriptions can be set or selected according to actual requirements, and multiple feature descriptions can be selected and combined; this is not limited here. For example, The following features can be used to obtain the splicing results:

[0116] in, This refers to the splicing operation of the feature splicing layer.

[0117] Based on the feature vectors of all channels A spliced ​​feature matrix can be constructed:

[0118] Where K is The dimension of c = 1, 2, ..., C, where C is The number of channels.

[0119] Finally, the concatenated feature matrix is ​​reduced in dimensionality by a feature mapping layer, and the reduced features are normalized by a normalization layer to obtain the visual difference features between the defective image and the reference image.

[0120] For example, taking the concatenated feature matrix V as an example, dimensionality reduction can be performed using PCA: 1) Calculate the covariance matrix of V. :

[0121] in, This represents the mean eigenvector of each channel.

[0122] 2) For the covariance matrix Perform eigenvalue decomposition:

[0123]

[0124]

[0125] in, It is an orthogonal matrix. It is a diagonal matrix. for The eigenvalues ​​on the diagonal, Eigenvalues The corresponding eigenvectors. 3) Principal component selection: Select the eigenvectors corresponding to the k largest eigenvalues ​​to form the projection matrix. :

[0126] 4) Perform projection (dimensionality reduction):

[0127] in, eigenvectors of channel c The result of dimensionality reduction is c=1,2,…,C.

[0128] Feature matrix after PCA dimensionality reduction :

[0129] For example, when the number of channels C is large, an autoencoder can be used for dimensionality reduction, or, based on PCA dimensionality reduction, an autoencoder can be further used for nonlinear compression. An autoencoder generally includes an encoder. and decoder Two parts, a compact representation of the data through self-supervised learning. For example, regarding... Taking nonlinear compression as an example, the main process is as follows: 1) Encoder compression process:

[0130] in, , , encoders The weights, biases, and activation functions, for The compression result.

[0131] 2) Decoder reconstruction process:

[0132] in, , , Decoders The weights, biases, and activation functions, Based on compression results The reconstructed vector.

[0133] 3) Optimization objective (minimize reconstruction error):

[0134] The reconstruction error is solved by the solution. Minimal compression result As a feature of the autoencoder after compression.

[0135] Finally, the reduced features are subjected to L2 normalization:

[0136] in, This is the normalized result for channel c. It is the L2 norm. It can be or .

[0137] Visual difference features between the defect image and the reference image obtained after normalization processing It can be represented as:

[0138] This invention captures deep semantic features and explicit statistical feature distributions through a feature description-transformation module, which can improve the discriminative power of features.

[0139] This invention constructs a visual difference feature extraction model that includes a multi-scale feature extraction network, a feature alignment network, a difference fusion module, an attention focusing module, and a feature description-transformation module. This model extracts visual difference features between defective images and reference images, and can effectively measure the deep-level differences between defective images and reference images.

[0140] It is understood that the visual difference feature extraction model of the present invention can be trained independently or jointly trained with the scanning feature extraction model and the map feature extraction model.

[0141] In some possible embodiments, the loss function of the visual difference feature extraction model includes: A symmetric consistency loss function that measures the feature differences between two different reference images; A reconstruction loss function that measures the difference between a reference image and a reconstructed reference image; wherein the reconstructed reference image is obtained based on visual difference features. The difference maximization loss function measures the feature differences between the defective image and the reference image; The total loss function of the visual difference feature extraction model is obtained by weighted summation of the symmetric consistency loss function, the reconstruction loss function, and the difference maximization loss function.

[0142] Specifically, the symmetric consistency loss function The expression is:

[0143] in, These represent two different reference images, such as images of the same defect-free wafer region taken at different angles, under different lighting conditions, or at different times. The complete mapping function representing the visual difference feature extraction model. Represents the L1 norm.

[0144] Theoretically, the visual difference between two flawless reference images should be close to 0, according to the symmetric consistency loss function. This ensures that the model does not fabricate differences between defect-free image pairs, enhancing the model's discriminative ability and ensuring that it only produces non-zero features when there are genuine defects.

[0145] Reconstruction loss function The expression is:

[0146] in, These represent the input visual difference feature extraction models, respectively. Reference image, defect image, i.e. visual difference features , The function representing the decoder is used to extract the model based on visual difference features. The output visual difference features are used to reconstruct the reference image. This ensures the visual difference feature extraction model... It can preserve the complete feature information of the image.

[0147] Difference Maximization Loss Function The expression is:

[0148] Where τ is a temperature parameter that controls the sensitivity to differences; the smaller τ is, the more sensitive the device is to differences. This is the normalization term. The difference maximization loss function maximizes the difference features between the defective image and the reference image, allowing the model to learn to highlight the difference features and improve the discriminative power of the features.

[0149] Combining the above loss functions, the total loss function of the visual difference feature extraction model is... for:

[0150] in, For L2 regularization terms, , , , These are the weighting coefficients for each loss function.

[0151] This invention achieves a visual difference feature extraction model through a total loss function that integrates the symmetric consistency loss function, the reconstruction loss function, and the difference maximization loss function. Multi-objective joint optimization can yield high-quality visual difference features.

[0152] The above describes the feature extraction process of the visual difference feature extraction model. The following describes the feature extraction process of the scanning feature extraction model and the map feature extraction model.

[0153] In embodiments of the present invention, features such as position, area, and size in the scanning signal are extracted using a scanning feature extraction model.

[0154] For example, the scanning feature extraction model can be a shallow fully connected neural network model, such as a symmetric autoencoder containing an encoder and a decoder. The encoder compresses the input data into a low-dimensional feature vector through three fully connected layers (such as Linear with ReLU activation), and the decoder reconstructs the feature vector back to the original input dimension through three symmetric fully connected layers.

[0155] In some possible embodiments, the scanning signal includes information on the location, area, and size of the potential defect; The location and geometric features of potential defects are extracted from the scan signal using a scan feature extraction model, including: Different types of coordinate transformations and feature extractions are performed on the location of potential defects to obtain the location features of the potential defects; Logarithmic transformation and Gaussian mixture modeling are applied to the area and size of the potential defect to obtain its geometric features.

[0156] Specifically, the scan signal typically records the physical coordinates of potential defects across the entire wafer, thus allowing direct acquisition of wafer-level absolute coordinates. These wafer-level absolute coordinates help identify systemic process issues, such as whether defects are consistently concentrated at the wafer edges. Coordinate transformation is then performed on the potential defect's coordinates to extract its sub-pixel-level relative coordinates within the die. Normalization is then applied to map the coordinate values ​​to a uniform range of 0-1, eliminating the dimensional influence of different die sizes and pinpointing the defect's precise location. Polar coordinates can also be extracted through polar coordinate transformation to capture the radial distribution of defects.

[0157] Then, the absolute coordinates of the defect at the wafer level, the relative coordinates within the die, and the polar coordinates are concatenated into a complete positional feature vector, which serves as the positional feature of the potential defect. This positional feature of the potential defect is represented by multiple coordinate features; different coordinate features can reflect different defect characteristics. For example, the defect characteristics corresponding to the absolute coordinates at the wafer level and the relative coordinates within the die are different. Defects generated in different process stages have significantly different spatial distribution patterns. For instance, defects caused by mechanical scratches usually appear as a long line on the entire wafer, representing a global defect feature; while some defects in the finishing process often recur at specific locations within each die, and their identification mainly relies on the relative coordinates within the die. This invention, by integrating multiple coordinate features, can effectively distinguish different causes, which is beneficial for improving the accuracy of defect identification and process monitoring.

[0158] For the area and size data of potential defects, a logarithmic transformation is performed to lengthen the tail, so that the size of all defects falls within a relatively uniform range. The expression is as follows: log_size = self.log_transform(size).unsqueeze(-1) Where `size` represents the area or size of the potential defect (e.g., maximum length), `log_size` represents the result of the logarithmic transformation, and `self.log_transform(`)` represents the logarithmic transformation result. ) is a user-defined logarithmic transformation function, such as f (size) = log(size + ), It is a very small positive number; unsqueeze(-1) means inserting a new dimension after the last dimension.

[0159] Since the area / size of defects is not randomly and uniformly distributed, but often clusters into several typical clusters, such as small defects G1, medium defects G2, and large defects G3, this invention, based on the logarithmic transformation results, automatically learns the mean, variance, and weight ratio of these distributions through a Gaussian mixture model. The core idea is to represent the distribution of these potential defects using the superposition of multiple Gaussian distributions, and to convert the area / size data of each defect into its probability density on these Gaussian components G1, G2, and G3. For example, the mean, variance, and weight ratio are used as learnable parameters for each Gaussian component of the Gaussian mixture model. For each specific defect area / size value, the probability of it belonging to each Gaussian component G1, G2, and G3 is calculated, obtaining the Gaussian encoded vector corresponding to the area / size of each potential defect. The Gaussian encoded vectors corresponding to the area and size of each potential defect are concatenated as the geometric features of the potential defect. Alternatively, a small fully connected network can be used to extract higher-order geometric features from the concatenated Gaussian encoded vectors.

[0160] The present invention extracts design context features of potential defects from the layout design data obtained in step S101 using a layout feature extraction model.

[0161] In some possible embodiments, design context features of potential defects are extracted from layout design data using a layout feature extraction model, including: By using a layout feature extraction model to embed, encode, concatenate, and extract features from layout design data, design context features of potential defects are obtained.

[0162] Specifically, the layout design data corresponding to potential defects can include one or more of the following: layer information, local graphic density, proximity structure information, functional area information, and design rule information. The layout feature extraction model can also be a small, fully connected neural network. For example, the selected layer information, local graphic density, proximity structure information, functional area information, and design rule information can be embedded and encoded separately, then concatenated to obtain a merged GDS feature vector. This vector is then input into the layout feature extraction model for feature extraction, yielding deeper design context features.

[0163] This invention integrates multi-dimensional information such as visual difference features, positional features, geometric features, and design context features corresponding to potential defects to form multimodal data, thereby achieving comprehensive information integration of potential defects. This information can be used for complementary advantages and correction, reducing false detection rate and false negative rate, and improving the robustness and generalization ability of overall defect identification.

[0164] In step S103 above, feature fusion is performed on the multimodal features corresponding to each potential defect to obtain the comprehensive defect feature vector corresponding to each potential defect.

[0165] For example, the multimodal features corresponding to each potential defect can be normalized, and then the features can be directly concatenated to form a comprehensive defect feature vector corresponding to each potential defect.

[0166] For example, the Transformer model can be used to perform cross-modal deep fusion of the multimodal features corresponding to each potential defect, and the correlation between the defect features of different modalities can be learned through the attention mechanism to obtain a unified and information-rich comprehensive defect feature vector. This makes the characterization of potential defects more comprehensive and accurate, thereby effectively distinguishing visually similar but different defects and obtaining a more discriminative defect feature representation.

[0167] In step S104 above, clustering is performed on the comprehensive defect feature vectors corresponding to all potential defects, and a clustering algorithm with the ability to automatically determine the number of clusters and noise points is used to distinguish between real defects and false defects.

[0168] In some possible embodiments, clustering is performed on the comprehensive defect feature vectors corresponding to all potential defects in the wafer scan data, and false defects and real defects are identified based on the clustering results, including: The dimensionality reduction of the comprehensive defect feature vectors corresponding to all potential defects in the wafer scanning data is performed to obtain the comprehensive defect feature vectors of each potential defect in the low-dimensional space. A density-based clustering method is used to cluster the comprehensive defect feature vector in low-dimensional space to obtain clustering results; wherein the clustering results include at least one noise point and multiple clusters; Each potential defect corresponding to a noise point is treated as a false defect, and all noise points are combined to form a set of false defects. The potential defects corresponding to the remaining clusters are treated as actual defects.

[0169] Specifically, the UMAP (Uniform Manifold Approximation and Projection) method can be used to reduce the dimensionality of the composite defect feature vectors corresponding to each potential defect, obtaining the composite defect feature vectors of each potential defect in a low-dimensional space (such as 2D or 3D). The main steps include: finding the nearest composite defect feature vector based on the Euclidean distance between each composite defect feature vector. We construct a weighted graph in a high-dimensional space from two neighboring graphs; we initialize a graph in a low-dimensional space and design an optimization objective to optimize its layout so that it is as similar as possible to the structure of the weighted graph in the high-dimensional space, thus obtaining a comprehensive defect feature vector in the low-dimensional space. The optimization objective can be to minimize the cross-entropy between the two graphs.

[0170] Then, density-based spatial clustering algorithms are used to cluster the comprehensive defect feature vectors in low-dimensional space. For example, the HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) algorithm can be used, which can automatically determine the number of clusters and identify points that do not belong to any dense region as noise.

[0171] To distinguish between real and false defects, this invention defines clustering labels for each cluster and noise point obtained from clustering. For dense, stable clusters, the clustering label is ≥0, representing known or unknown (fuzzy / emerging) repetitive defect patterns, which are all identified as real defects. For sparse, scattered noise points, the clustering label is -1, which are identified as false defects. In this invention, these are redefined as process background noise that needs to be monitored.

[0172] For example, given 10,000 potential defects (ea), clustering reveals 5 clusters with 2200 ea of ​​noise. These 5 clusters are labeled with a value greater than 0, representing true defects. The cluster label for the set of 4200 ea of ​​noise is defined as -1, representing false defects. True defects include both known and unknown defects, which can be distinguished by feature comparison with a reference image. Clustering details are shown in Table 1. Table 1. Examples of clustering details for real defects

[0173] This invention employs a density-based clustering method to cluster comprehensive defect feature vectors. Leveraging the powerful adaptive and noise recognition capabilities of unsupervised clustering, it can identify false defects based on the marked noise points in the clustering results. This overcomes the shortcomings of traditional methods that can only passively accept the interference of false defects, reduces the interference of false defects, and can automatically discover patterns from the data. It is applicable to scenarios where the defect type is unknown in the early stages of R&D, reducing the difficulty of defect identification.

[0174] In step S104 above, a variety of sampling strategies are used to sample from the clustering results to obtain a sampling list.

[0175] For example, random sampling can be performed from the clusters corresponding to each real defect to form a sampling list.

[0176] In some possible embodiments, different sampling strategies are used to sample from the clustering results to obtain a sampling list, including: For each cluster corresponding to a known defect, sampling is carried out using equal sampling or weighted sampling strategies based on the number of samples in each cluster to obtain a sample of the known defects. For each cluster corresponding to the unknown defect, an exploratory sampling strategy is used to sample and obtain a sample of the unknown defect. For the set of false defects, a random sampling strategy is used to obtain a sample of false defects. A sampling list is constructed based on the sampling samples of known defects, unknown defects, and false defects; wherein the sampling quantity of known defects is greater than the sampling quantity of false defects, and the sampling quantity of false defects is greater than the sampling quantity of unknown defects.

[0177] Specifically, let the clustering result be T non-noise clusters Cluster1, Cluster2, ..., Cluster3. T Given a set of pseudo-defects called Noise, then these T clusters are Cluster1, Cluster2, ..., Cluster3. T These are all clusters corresponding to real defects, including clusters corresponding to known defects and clusters corresponding to unknown defects. For example... Figure 9 The diagram shown is a schematic representation of a sampling strategy provided in an embodiment of the present invention. Figure 9 As shown, for known defects, equal sampling or weighted sampling strategies are used for key sampling; for false defects, the sampling quantity is monitored and a small number of samples are taken; for unknown defects, exploratory sampling is used and a very small number of samples are taken.

[0178] For example, an equal-sampling strategy can be used to draw a fixed number of samples from each cluster corresponding to a known defect, i.e., for each cluster corresponding to a known defect... t Extract from 10 samples, of which For the preset sample size, |Cluster t | for cluster t The number of samples in the sample.

[0179] For example, a weighted sampling strategy can also be used, allocating sampling quotas proportionally based on the number of samples in each cluster or the clustering stability index generated by the HDBSCAN algorithm. Let the total number of samples be... total Then, from the clusters corresponding to the known defects... t Sampling quantity for:

[0180] Where t=1,2,…, , This represents the number of clusters with known defects.

[0181] For example, other specific sampling strategies can also be used, such as prioritizing the most dispersed points in the feature space within a single cluster to ensure that subtle morphological variations within that type are covered.

[0182] For the set of pseudo-defects (Noise), a random sampling strategy can be used, such as randomly selecting a fixed number of samples. noise A few points (e.g., 5-10). However, it is necessary to monitor the sampling quantity and only a small number of samples should be taken. The sampling quantity of known defects should be much larger than the sampling quantity of false defects.

[0183] For unknown defects, which may include fuzzy defects or emerging defects, a very small number of samples can be taken by exploring sampling strategies.

[0184] Based on the sampling samples corresponding to the known defects, unknown defects, and pseudo defects, a sampling list is constructed.

[0185] This invention employs different sampling strategies to perform intelligent sampling and re-inspection from clustering results, which can comprehensively cover all key defect types, significantly improve the efficiency of defect improvement during the R&D process, and shorten the R&D cycle.

[0186] In step S106 above, a re-inspection is performed based on the sampling list obtained in step S105.

[0187] For example, a scanning electron microscope (SEM) can be used for re-examination to confirm the existence and authenticity of defects, thus obtaining the final defect detection result.

[0188] The final defect detection results can also be fed back to the visual difference feature extraction model, scan feature extraction model, and map feature extraction model to optimize the accuracy of model feature extraction. For example, new training samples can be created based on the final defect detection results to participate in model training and optimization, gradually improving the defect detection effect.

[0189] Overall, this invention optimizes the decision-making process between wafer scanning and re-inspection. It is equivalent to inserting a software-defined intelligent middleware into the key link between wafer scanning and re-inspection, which revitalizes dormant wafer scanning data assets. The SEM machine is almost entirely used to analyze real and meaningful defects, releasing the maximum value of scanning electron microscope in defect detection. It is expected to gradually evolve into a standardized defect analysis platform in the semiconductor manufacturing field and become an indispensable tool for advanced process development.

[0190] Example 2 Please see Figure 10 , Figure 10 A schematic diagram of a semiconductor defect detection system provided as an embodiment of the present invention is shown. The system includes: The data acquisition module 1010 is used to extract the defect image and scanning signal corresponding to each potential defect from the wafer scan data, and to extract the layout design data corresponding to each potential defect from the wafer layout data. The feature extraction module 1020 is used to extract defect features from defect images, scanning signals and layout design data respectively, and obtain multimodal features corresponding to each potential defect; The feature fusion module 1030 is used to perform cross-modal feature fusion on multimodal features to obtain a comprehensive defect feature vector corresponding to each potential defect; The clustering module 1040 is used to cluster the comprehensive defect feature vectors corresponding to all potential defects in the wafer scanning data, and to identify false defects and real defects based on the clustering results. The sampling and re-inspection module 1050 is used to sample from false defects and real defects using different sampling strategies to obtain a sampling list; and to re-inspect according to the sampling list to obtain the defect detection results of the wafer.

[0191] The semiconductor defect detection system and the semiconductor defect detection method described above can be referred to each other.

[0192] Example 3 Figure 11 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 11 As shown, the electronic device may include a processor 1110, a communications interface 1120, a memory 1130, and a communication bus 1140, wherein the processor 1110, the communications interface 1120, and the memory 1130 communicate with each other via the communication bus 1140. The processor 1110 can call logical instructions in the memory 1130 to execute the semiconductor defect detection method provided in the above-described method embodiments.

[0193] Furthermore, the logical instructions in the aforementioned memory 1130 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0194] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the semiconductor defect detection method provided in the above-described method embodiments.

[0195] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the semiconductor defect detection method provided in the above-described method embodiments.

[0196] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0197] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0198] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A semiconductor defect detection method, characterized in that, include: Extract the defect image and scan signal corresponding to each potential defect from the wafer scan data, and extract the layout design data corresponding to each potential defect from the wafer layout data; Defect features are extracted from the defect image, the scanning signal, and the layout design data respectively to obtain the multimodal features corresponding to each potential defect; Cross-modal feature fusion is performed on the multimodal features to obtain a comprehensive defect feature vector corresponding to each potential defect; Cluster the comprehensive defect feature vectors corresponding to all potential defects in the wafer scanning data, and identify false defects and real defects based on the clustering results; Different sampling strategies are used to sample from the false defects and the real defects respectively to obtain a sampling list; The wafer is re-inspected according to the sampling list to obtain the defect detection results.

2. The semiconductor defect detection method according to claim 1, characterized in that, The step of extracting defect features from the defect image, the scan signal, and the layout design data to obtain multimodal features corresponding to each potential defect includes: The visual difference features between the defect image and the reference image are extracted using a visual difference feature extraction model. The location and geometric features of the potential defects are extracted from the scan signal using a scan feature extraction model. The design context features of the potential defects are extracted from the layout design data using a layout feature extraction model. The multimodal features corresponding to the potential defects include the visual difference features, the positional features, the geometric features, and the design context features.

3. The semiconductor defect detection method according to claim 2, characterized in that, The visual difference feature extraction model includes a multi-scale feature extraction network, a feature alignment network, a difference fusion module, an attention focusing module, and a feature description-transformation module connected in sequence. The multi-scale feature extraction network is used to extract multi-scale features from the defect image and the reference image respectively, to obtain multi-scale defect features of the defect image and multi-scale reference features of the reference image. The feature alignment network is used to perform sub-pixel alignment of the defect features and reference features at each scale to obtain reference features aligned with the defect features at each scale. The differential fusion module is used to calculate the absolute difference between the defect feature and the reference feature aligned with the defect feature at each scale, and to perform scale adjustment and multi-scale fusion on the absolute difference to obtain the fused differential feature. The attention focusing module is used to perform channel attention enhancement and spatial attention enhancement on the fused differential features to obtain enhanced fused differential features; The feature description-transformation module is used to perform diversified feature description and feature splicing on the enhanced fusion difference features to obtain a spliced ​​feature matrix; and to perform dimensionality reduction and normalization processing on the spliced ​​feature matrix to obtain the visual difference features between the defective image and the reference image; wherein, the diversified feature description includes at least one of statistical moment features, quantile features and texture features.

4. The semiconductor defect detection method according to claim 3, characterized in that, The feature alignment network includes a convolutional network and deformable convolutional layers; The convolutional network is used to predict the dense offset field from the reference feature to the defect feature at each scale, and to determine the pixel offset between the reference feature and the defect feature based on the dense offset field. The deformable convolutional layer is used to deformably resample the reference features based on the pixel offset to obtain reference features aligned with the defect features at each scale.

5. The semiconductor defect detection method according to claim 4, characterized in that, The loss function of the visual difference feature extraction model includes: A symmetric consistency loss function that measures the feature differences between two different reference images; A reconstruction loss function that measures the difference between the reference image and the reconstructed reference image; wherein the reconstructed reference image is obtained based on the visual difference features; A difference maximization loss function measures the feature differences between the defective image and the reference image; The total loss function of the visual difference feature extraction model is obtained by weighted summation of the symmetric consistency loss function, the reconstruction loss function, and the difference maximization loss function.

6. The semiconductor defect detection method according to claim 2, characterized in that, The scanning signal includes information on the location, area, and size of the potential defect; The step of extracting the location and geometric features of the potential defect from the scan signal using a scan feature extraction model includes: Different types of coordinate transformations and feature extractions are performed on the location of the potential defect to obtain the location features of the potential defect; The area and size of the potential defect are modeled using a logarithmic transformation and a Gaussian mixture model to obtain the geometric features of the potential defect.

7. The semiconductor defect detection method according to claim 2, characterized in that, The step of extracting the design context features of the potential defects from the layout design data using the layout feature extraction model includes: The layout design data is embedded, encoded, concatenated, and extracted using a layout feature extraction model to obtain the design context features of the potential defects. The layout design data corresponding to the potential defects includes at least one of the following: layer information, local graphic density, adjacent structure information, functional area information, and design rule information.

8. The semiconductor defect detection method according to claim 1, characterized in that, The step of clustering the comprehensive defect feature vectors corresponding to all potential defects in the wafer scanning data, and identifying false defects and real defects based on the clustering results, includes: The comprehensive defect feature vectors corresponding to all potential defects in the wafer scanning data are subjected to dimensionality reduction processing to obtain the comprehensive defect feature vectors of each potential defect in the low-dimensional space. A density-based clustering method is used to cluster the comprehensive defect feature vector in the low-dimensional space to obtain clustering results; wherein, the clustering results include at least one noise point and multiple clusters; Each potential defect corresponding to a noise point is treated as a false defect, and all noise points are combined to form a set of false defects. The potential defects corresponding to the remaining clusters are taken as real defects; wherein, the real defects include known defects and unknown defects.

9. The semiconductor defect detection method according to claim 8, characterized in that, The sampling list is obtained by sampling from the clustering results using different sampling strategies, including: For each cluster corresponding to the known defects, sampling is carried out using equal sampling or weighted sampling strategies based on the number of samples in each cluster to obtain the sampled samples of the known defects. For each cluster corresponding to the unknown defect, an exploratory sampling strategy is used to sample and obtain a sample of the unknown defect. For the set of false defects, a random sampling strategy is used to obtain a sample of the false defects. A sampling list is constructed based on the sampling samples of the known defects, the sampling samples of the unknown defects, and the sampling samples of the false defects; wherein the sampling quantity of the known defects is greater than the sampling quantity of the false defects, and the sampling quantity of the false defects is greater than the sampling quantity of the unknown defects.

10. A semiconductor defect detection system, characterized in that, include: The data acquisition module is used to extract the defect image and scanning signal corresponding to each potential defect from the wafer scan data, and to extract the layout design data corresponding to each potential defect from the wafer layout data; The feature extraction module is used to extract the defect features from the defect image, the scanning signal, and the layout design data respectively, to obtain the multimodal features corresponding to each potential defect; The feature fusion module is used to perform cross-modal feature fusion on the multimodal features to obtain a comprehensive defect feature vector corresponding to each potential defect; The clustering module is used to cluster the comprehensive defect feature vectors corresponding to all potential defects in the wafer scanning data, and to identify false defects and real defects based on the clustering results. The sampling and re-inspection module is used to sample from the false defects and the real defects using different sampling strategies to obtain a sampling list; and to perform re-inspection based on the sampling list to obtain the defect detection results of the wafer.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the semiconductor defect detection method as described in any one of claims 1 to 9.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the semiconductor defect detection method as described in any one of claims 1 to 9.