Semiconductor chip surface defect detection method based on deep adversarial supervision

By employing a deep adversarial supervision approach and utilizing generators and multi-scale supervision strategies, the problems of speed and accuracy in semiconductor chip surface defect detection are solved, enabling efficient and accurate detection of minute defects in complex backgrounds.

CN121032983APending Publication Date: 2025-11-28SHANGHAI IND U TECH RES INST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511174333.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify subtle defects in complex backgrounds while maintaining speed in semiconductor chip surface defect detection, and deep learning methods suffer from insufficient supervision and inaccurate feature representation.

Method used

A deep adversarial supervision-based approach is adopted, which uses a generator consisting of an encoder and a decoder, combined with multi-scale supervision and a multi-layer discriminator for adversarial training, to achieve multi-scale feature extraction and defect segmentation of chip surface images. A phased training strategy is used to optimize the model parameters.

Benefits of technology

It significantly improves the accuracy and quality of defect segmentation, enhances the model's ability to identify subtle defects and improves boundary accuracy, and increases detection stability and robustness in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121032983A_ABST
    Figure CN121032983A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of defect detection, in particular to a semiconductor chip surface defect detection method based on deep adversarial supervision, which comprises the following steps: acquiring a chip surface image, and inputting a pre-constructed defect detection model to obtain a defect segmentation result of the chip surface image; the defect detection model comprises an encoder, a decoder and a discriminator; the encoder is used for extracting multi-scale features of the chip surface image; the decoder is used for recovering the spatial resolution of the image layer by layer based on the multi-scale features and outputting a defect segmentation result of a corresponding scale; the discriminator is used for discriminating the difference between the defect segmentation result output by the decoder at the corresponding scale and the real label image at the corresponding scale; the method comprises the following steps: updating parameters of an encoder and a decoder by calculating semantic segmentation loss; and then adversarial training is carried out in combination with the discrimination result of the discriminator, and the parameters of the encoder, the decoder and the discriminator are jointly optimized, so that iterative updating of the parameters of the defect detection model is realized. The defect detection precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of defect detection, and in particular to a semiconductor chip surface defect detection method based on deep adversarial supervision. BACKGROUND

[0002] With the continuous advancement of semiconductor manufacturing process nodes to 7nm, 5nm and more advanced processes, chip surface defect detection has become a key link to ensure wafer yield and product performance. Any tiny scratch, pattern distortion, particle contamination and other defects on the chip surface can directly affect the performance, power consumption and yield of the chip. Therefore, how to achieve high-precision and real-time detection of semiconductor chip surface defects has become a key link to ensure product quality and improve production capacity.

[0003] Currently, there are three main technical paths for semiconductor chip surface defect detection methods. The first type is based on traditional image processing methods, which usually use edge detection, threshold segmentation, template matching and other computer vision techniques to complete defect positioning and discrimination through manually set rules. The second type is based on machine learning methods, which use support vector machines (SVM), random forests and other traditional machine learning algorithms to identify and classify defects after manually extracting image features, in order to improve detection performance under standard process conditions. The third type is based on deep learning methods. With the great potential of convolutional neural networks (CNN) in image segmentation and object detection in recent years, wafer defect detection methods based on deep convolutional neural networks, surface defect segmentation algorithms based on residual networks, and defect detection frameworks enhanced by attention mechanisms to enhance feature expression have been developed. This type of method can automatically extract multi-level features from raw images through end-to-end training, which improves the accuracy of defect recognition and the precision of segmentation to some extent.

[0004] However, although existing methods have improved defect detection performance to some extent, with the increasing refinement of chip structures and the complexity of defect types, existing technologies still have limitations in practical applications. On the one hand, traditional image processing and machine learning-based methods still rely on human involvement, making it difficult to capture diverse and subtle defect patterns, and the detection accuracy will decrease significantly under changing process conditions or complex backgrounds. On the other hand, although deep learning-based detection methods have end-to-end feature learning capabilities, they are still prone to misclassification, missed detection and inaccurate boundary positioning due to feature expression confusion, boundary detail loss and insufficient supervision in complex background and subtle defect scenarios. In addition, the high computational overhead of high-resolution wafer images required by deep learning further restricts the feasibility of real-time detection. In summary, existing technologies cannot accurately identify subtle defects in complex backgrounds while ensuring the speed of semiconductor chip surface detection. SUMMARY

[0005] This application provides a semiconductor chip surface defect detection method based on deep adversarial supervision, which can accurately identify subtle defects in complex backgrounds while ensuring the detection speed of semiconductor chip surfaces. The technical solution provided in this application is as follows: In a first aspect, this application provides a method for detecting surface defects in semiconductor chips based on deep adversarial supervision, the method comprising: Acquire images of the chip surface; The chip surface image is input into a pre-constructed defect detection model to obtain defect segmentation results of the chip surface image; the defect detection model includes an encoder, a decoder, and at least one discriminator; the encoder is used to extract multi-scale features of the chip surface image; the decoder is used to recover the spatial resolution of the image layer by layer based on the multi-scale features and output defect segmentation results at the corresponding scale; the discriminator is used to determine the difference between the defect segmentation results output by the decoder at the corresponding scale and the real label image at the corresponding scale; Based on the chip surface image and the corresponding real label image, the parameters of the encoder and the decoder are first updated by calculating the semantic segmentation loss; then, adversarial training is performed in combination with the discrimination result of the discriminator to jointly optimize the parameters of the encoder, the decoder and the discriminator, thereby realizing the iterative update of the defect detection model parameters.

[0006] In one specific implementation, the encoder and the decoder constitute the generator in the defect detection model; the generator uses the U-Net architecture as its basic framework; the encoder includes several convolutional layers, which process the chip surface image through convolution operations and output feature maps of corresponding scales, and gradually reduce the size of the feature maps through downsampling; the decoder includes several decoding layers, which receive feature maps output from the corresponding convolutional layers of the encoder, and restore the spatial resolution of the image layer by layer through upsampling operations, and output defect segmentation results corresponding to the resolution of that layer.

[0007] In one specific implementation, the discriminator corresponds one-to-one with the decoding layer in the decoder. Each discriminator includes two convolutional structures connected in series and a nonlinear mapping structure. The nonlinear mapping structure is connected to the output of the second convolutional structure. The convolutional structure includes a 3×3 convolutional kernel, a normalization layer, and a LeakyReLU activation layer, used to extract the feature differences between the input defect segmentation result and the real label image. The nonlinear mapping structure includes a Sigmoid activation layer, used to map the features output by the convolutional structure to represent the probability that the input belongs to the real label.

[0008] In a specific feasible implementation, a multi-scale supervision strategy is adopted, which involves training by introducing real labels of corresponding resolutions at different scales; the multi-scale supervision strategy is divided into three categories: high-resolution supervision, medium-resolution supervision, and low-resolution supervision. The high-resolution supervision is performed at the output of the last layer of the decoder using real labels with the same resolution as the chip surface image. The intermediate-resolution supervision is performed in the middle layer of the decoder, using the corresponding downsampled real labels for feature maps that have been downsampled once or multiple times. The low-resolution supervision is performed at the deep output of the decoder using large-scale downsampled label images.

[0009] In one specific implementation, a channel attention mechanism, a spatial attention mechanism, and a pyramid pooling module are introduced into the encoder; The channel attention mechanism is used to perform global average pooling and global max pooling on the channel dimension of the input feature map to obtain channel description vectors respectively. The channel description vectors are then processed by a multilayer perceptron and a channel weight vector is generated by the Sigmoid activation function. The spatial attention mechanism is used to perform global average pooling and global max pooling on the input feature map in the channel dimension to obtain two spatial feature maps. The spatial feature maps are concatenated and then generated into a spatial weight map through a convolution operation. The pyramid pooling module performs spatial downsampling on the input feature map through multi-scale pooling and then upsampling it to the same size. The multi-scale pooling result is then concatenated with the original input feature map in the channel dimension to form a fused feature.

[0010] In one specific implementation, the iterative update of the defect detection model parameters adopts a staged training strategy, and uses a generator loss function. With discriminator loss function To optimize the objective; The generator's loss function The definition is as follows: ; in, For semantic segmentation loss, For feature matching loss, To combat the losses; ; in, For the number of discriminators, For the first The segmentation result or feature map obtained from the output of the layer decoding layer. Indicates the relationship with the first The discriminator output corresponding to the layer decoding layer; This represents the binary cross-entropy loss; The loss function of the discriminator The definition is as follows: ; in, In order to be with the first The corresponding scale of the real label output by the layer decoding layer.

[0011] In one specific implementation, the iterative update of the defect detection model parameters includes three stages: The first stage is the generator pre-training stage, which utilizes semantic segmentation loss. The parameters of the encoder and decoder in the generator are trained; The second stage involves introducing a discriminator for adversarial training, based on the discriminator loss. The discriminator is trained, and semantic segmentation loss is also incorporated. Combating losses and feature matching loss Train the generator; The third stage is the joint optimization stage, which involves jointly optimizing all parameters of the generator and discriminator, using semantic segmentation loss. Combating losses and feature matching loss To achieve the training objective, the model parameters are iterated continuously until the predetermined training termination condition is met.

[0012] Secondly, this application provides a semiconductor chip surface defect detection system based on deep adversarial supervision, employing the following technical solution: A semiconductor chip surface defect detection system based on deep adversarial supervision, comprising: Image acquisition module, used to acquire images of the chip surface; A defect segmentation module is used to input the chip surface image into a pre-constructed defect detection model to obtain defect segmentation results of the chip surface image; the defect detection model includes an encoder, a decoder, and at least one discriminator; the encoder is used to extract multi-scale features of the chip surface image; the decoder is used to recover the spatial resolution of the image layer by layer based on the multi-scale features and output the defect segmentation results at the corresponding scale; the discriminator is used to determine the difference between the defect segmentation results output by the decoder at the corresponding scale and the real label image at the corresponding scale. The parameter update module is used to update the parameters of the encoder and the decoder by calculating the semantic segmentation loss based on the chip surface image and the corresponding real label image; then, adversarial training is performed in combination with the discrimination result of the discriminator to jointly optimize the parameters of the encoder, the decoder and the discriminator, thereby realizing the iterative update of the defect detection model parameters.

[0013] Thirdly, this application provides an electronic device, the device including a processor and a memory; the memory stores a program, the program being loaded and executed by the processor to implement a semiconductor chip surface defect detection method based on deep adversarial supervision as described in the first aspect.

[0014] Fourthly, this application provides a computer-readable storage medium storing a program that, when executed by a processor, is used to implement a semiconductor chip surface defect detection method based on deep adversarial supervision as described in the first aspect.

[0015] In summary, the beneficial effects of this application include at least the following: (1) By introducing a deep supervision mechanism, the multi-layer discriminator is used to conduct multi-scale and multi-level adversarial training on the segmentation results output by the generator, enabling the model to capture more delicate and richer defect feature expressions. This mechanism not only strengthens the generator's ability to identify subtle defects in complex backgrounds, but also promotes the consistency between generated features and real features through feature matching loss, thereby significantly improving the overall accuracy and quality of defect segmentation.

[0016] (2) A pixel-level semantic segmentation method is adopted to achieve accurate boundary recognition of defects on the chip surface. Through a multi-scale supervision strategy, the model refines the defects layer by layer at different spatial resolution levels, so that defects of different sizes can be accurately identified and located. In particular, it effectively solves the problems of small defects and blurred edges, ensuring the detail integrity and accuracy of the segmentation results.

[0017] (3) In order to improve the adaptability of the model under diverse process conditions and complex backgrounds, the present invention integrates a multi-scale feature fusion module, integrates contextual information of different scales through pyramid pooling and other technologies, and combines rich training signals provided by deep supervision mechanism, so that the model can take into account defect features of different sizes and shapes, enhance the generalization ability to unknown scenes and samples, and effectively improve the stability and robustness of detection.

[0018] By introducing a generative adversarial network (GAN) framework, a multi-layered discriminator is designed to distinguish between segmentation results at different scales output by the decoder, forming a deep supervision mechanism that enhances the ability to capture subtle defect features and improves boundary accuracy. Simultaneously, a staged training strategy is adopted: the generator is first pre-trained using semantic segmentation loss, and then adversarial training is performed in conjunction with the discriminator, achieving synergistic optimization between the generator and discriminator, effectively improving the model's robustness and generalization ability. This approach not only overcomes the limitations of traditional methods but also solves the problems of insufficient supervision and imprecise feature representation in deep learning methods, thus achieving more accurate and stable defect detection while ensuring computational efficiency.

[0019] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, the preferred embodiments of this application are described in detail below with reference to the accompanying drawings. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating the semiconductor chip surface defect detection method based on deep adversarial supervision in the embodiments of this application.

[0021] Figure 2 This is a logic block diagram of the semiconductor chip surface defect detection method based on deep adversarial supervision in the embodiments of this application.

[0022] Figure 3 This is a structural block diagram of the semiconductor chip surface defect detection system based on deep adversarial supervision in the embodiments of this application.

[0023] Figure 4 This is a block diagram of an electronic device for detecting surface defects in semiconductor chips based on deep adversarial supervision, as described in this application. Detailed Implementation

[0024] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate this application, but are not intended to limit the scope of this application.

[0025] Optionally, this application uses the semiconductor chip surface defect detection method based on deep adversarial supervision provided in various embodiments as an example for description in electronic devices. The electronic device is a terminal or server. The terminal can be a computer, tablet computer, etc. This embodiment does not limit the type of electronic device.

[0026] Reference Figure 1 This is a flowchart illustrating a semiconductor chip surface defect detection method based on deep adversarial supervision according to an embodiment of this application. The method includes at least the following steps: Step S101: Obtain an image of the chip surface.

[0027] In step S101, the surface of the semiconductor chip to be inspected is acquired by an imaging device to obtain a high-resolution image of the chip surface.

[0028] Optionally, the imaging device may include an optical microscope, an electron microscope, or other image acquisition device suitable for capturing the microstructure and defect features of a chip. This application does not limit the specific type of imaging device.

[0029] During implementation, the acquired images should possess sufficient clarity and detail to ensure that the subsequent defect detection model can accurately extract multi-scale feature information of surface defects. During acquisition, appropriate adjustments can be made to ambient lighting, shooting angle, and imaging parameters to ensure the stability and consistency of image quality, providing a reliable data foundation for subsequent defect segmentation and recognition.

[0030] Step S102: Input the chip surface image into the pre-constructed defect detection model to obtain the defect segmentation result of the chip surface image; the defect detection model includes an encoder, a decoder and at least one discriminator; the encoder is used to extract multi-scale features of the chip surface image; the decoder is used to recover the spatial resolution of the image layer by layer based on the multi-scale features and output the defect segmentation result at the corresponding scale; the discriminator is used to distinguish the difference between the defect segmentation result output by the decoder at the corresponding scale and the real label image at the corresponding scale.

[0031] In step S102, the encoder and decoder together constitute the generator in the defect detection model, and the generator adopts the U-Net architecture as the basic framework.

[0032] Alternatively, other semantic segmentation networks such as PSPNet, DeepLab, or Transformer architectures can be used as the basic framework. This application does not impose any restrictions on the network architecture of the generator.

[0033] The encoder includes several convolutional layers. The convolutional layers process the chip surface image through convolution operations and output feature maps of the corresponding scale. The feature map size is gradually reduced by downsampling to increase the receptive field. In addition, residual connections are preferably used in the encoder to ensure deep gradient propagation and information preservation.

[0034] In implementation, to further enhance the encoder's ability to extract features from chip surface defects and improve the model's sensitivity to defects of different scales and spatial locations, a feature enhancement module was introduced into the encoder. This module mainly improves the accuracy and robustness of defect detection by integrating attention mechanisms and multi-scale feature fusion techniques, thereby effectively weighting the input features and enriching contextual information.

[0035] Specifically, firstly, the attention mechanism helps the model automatically focus on more discriminative defect features by assigning different weights to different channels and spatial locations in the feature map. This includes channel attention mechanism and spatial attention mechanism.

[0036] The channel attention mechanism globally summarizes information from the channel dimensions of the input feature map. It calculates the average and maximum values ​​for each channel using global average pooling and global max pooling, respectively, resulting in two channel description vectors. These vectors are then subjected to nonlinear transformations via a multilayer perceptron and finally normalized to channel weight vectors using the sigmoid function. The mathematical expression for the channel attention mechanism is: ; in, Represents the input feature map, and These represent the global average pooling and max pooling operations on the feature map, respectively. It is a multilayer perceptron containing two fully connected layers, responsible for learning the nonlinear relationships between channels. This is the Sigmoid activation function.

[0037] The spatial attention mechanism obtains two spatial feature maps by performing global average pooling and max pooling on the input feature map along the channel dimension. These two maps are then concatenated and mapped to spatial attention weights via convolution, helping the model highlight defective regions at key spatial locations. The mathematical expression for the spatial attention mechanism is: ; in, This indicates that two feature maps are concatenated along the channel dimension. This indicates a convolution operation.

[0038] Secondly, to further integrate contextual information at different scales, the encoder also employs a pyramid pooling module. This module performs spatial compression on the feature maps through multi-scale pooling, extracting multi-level semantic information. Finally, these multi-scale features are concatenated and fused with the original features, enhancing the adaptability to diverse defect scales. The mathematical expression for the pyramid pooling module is: ; in, This indicates that pooling operations of different sizes spatially downsample the input feature map and then upsample it to the same size. This means concatenating all pooling results and the original feature map along the channel dimension. Through the integration of the above feature enhancement modules, the encoder can not only capture the key information of defects more accurately, but also take into account the representation of defects in different spatial locations and sizes, thus providing a solid foundation for the subsequent decoder to generate a fine defect segmentation map, significantly improving the overall detection performance and generalization ability of the model.

[0039] The decoder comprises several decoding layers. Each decoding layer receives feature maps from the corresponding convolutional layers of the encoder and recovers the spatial resolution of the image layer by layer through upsampling operations, outputting a defect segmentation result corresponding to the resolution of that layer. During the upsampling process, skip connections are used to fuse feature information at different scales to preserve edge and detail information.

[0040] The discriminator and the decoding layer in the decoder are in one-to-one correspondence. The output of each decoding layer serves as both the input to the next decoding layer and the input to the corresponding discriminator. Each discriminator consists of two concatenated convolutional structures and a nonlinear mapping structure. The nonlinear mapping structure is connected to the output of the second convolutional structure. The convolutional structure includes a 3×3 convolutional kernel, a normalization layer, and a LeakyReLU activation layer, used to extract the feature differences between the input defect segmentation result and the real label image. The nonlinear mapping structure includes a Sigmoid activation layer, used to map the features output by the convolutional structure to represent the probability that the input belongs to the real label.

[0041] Alternatively, PatchGAN or a multi-scale discriminator can be selected instead of the discriminator in the embodiments of this application. PatchGAN can improve the discrimination efficiency, and the multi-scale discriminator can perform discrimination at different resolutions.

[0042] Furthermore, in the deep supervised generative adversarial framework, the decoder progressively restores the spatial resolution of the image through multiple layers, outputting defect segmentation results at different scales. Due to the diverse sizes and shapes of defects on the chip surface, single-scale supervision is insufficient to comprehensively capture all defect features, easily leading to inadequate learning of fine-grained and macroscopic structural information. Therefore, a multi-scale supervision strategy is adopted, introducing ground truth labels of corresponding resolutions at different scales for training. This fully utilizes multi-layered feature information, improving the model's defect detection accuracy and generalization ability. Specifically, the multi-scale supervision strategy is divided into three categories: high-resolution supervision, medium-resolution supervision, and low-resolution supervision. High-resolution supervision uses ground truth labels of the same resolution as the original chip surface image at the last layer output of the decoder. This layer provides rich fine-grained supervision signals, primarily ensuring the model's accurate restoration of defect boundaries and details. Medium-resolution supervision, in the middle layer of the decoder, uses corresponding downsampled ground truth labels for feature maps that have undergone one or more downsampling iterations. This supervision layer balances spatial resolution and semantic information, helping the model learn details and overall structure at a medium scale. Low-resolution supervision uses coarser, larger-scale downsampled label images at deeper decoder outputs. The focus of this stage is to guide the model to learn the global distribution and macroscopic structural features of defects, avoiding the neglect of overall morphological information. Through this multi-scale label and corresponding decoder output layer-by-layer supervision, the model can obtain more comprehensive and hierarchical training signals, improving the richness and accuracy of feature representation. Multi-scale supervision not only helps solve the problems of blurred details and unclear boundaries but also effectively alleviates false positives or false negatives caused by insufficient information in single-scale supervision.

[0043] Step S103: Based on the chip surface image and the corresponding real label image, first update the parameters of the encoder and decoder by calculating the semantic segmentation loss; then combine the discrimination results of the discriminator to perform adversarial training, jointly optimize the parameters of the encoder, decoder and discriminator, and realize the iterative update of the defect detection model parameters.

[0044] In step S103, based on the obtained chip surface image and corresponding pixel-level real labels The parameters of the pre-built defect detection model are iteratively updated. The parameter iterative update employs a staged training strategy and uses the generator loss function. With discriminator loss function To optimize the objective.

[0045] Specifically, the generator's loss function The definition is as follows: ; in, For semantic segmentation loss, in this embodiment of the application, the sum or weighted combination of Dice Loss and cross-entropy loss is used to measure the difference between the generator output and the real label at the pixel level. It is a feature matching loss used to measure the difference in intermediate features between the generator output and the real sample in the discriminator or other predetermined feature space; To counteract the loss, a measure is taken of the generator's ability to cause the discriminator to misclassify the generated result as a real sample. In the examples, the following can be used: ; in, For the number of discriminators, For the first The segmentation result or feature map obtained from the output of the layer decoding layer. Indicates the relationship with the first The discriminator output corresponding to the layer decoding layer; This represents the binary cross-entropy loss.

[0046] loss function of discriminator The definition is as follows: ; in, In order to be with the first The corresponding scale of the real labels output by the layer decoding layer. The first term indicates that the discriminator should classify a non-real sample as having a loss target value of 0 when receiving the generator output; the second term indicates that the discriminator should classify a real sample as having a loss target value of 1 when receiving the real label.

[0047] Based on the above loss function, the parameter iterative update includes three stages: The first stage is the generator pre-training stage. This stage only utilizes semantic segmentation loss. The encoder and decoder parameters in the generator are trained to enable the generator to initially learn the feature representation and segmentation capabilities of chip surface defects. This stage does not involve the discriminator, ensuring that the generator completes the basic segmentation task under single supervision.

[0048] The second stage involves introducing a discriminator for adversarial training. In this stage, both the discriminator and generator are trained simultaneously. The discriminator calculates its loss function based on the difference between the defect segmentation results output by the generator and the true labels. To improve its own discrimination ability; the generator combines semantic segmentation loss. Combating losses and feature matching loss The model optimizes its output to deceive the discriminator, making the segmentation results closer to the true labels. Through alternating training of the discriminator and generator, the model continuously improves the generated segmentation results.

[0049] The third stage is the joint optimization stage. After the training in the first two stages, the parameters of the generator and discriminator reach a certain level. In this stage, all parameters of the generator and discriminator are jointly optimized. The model performance is further refined by targeting semantic segmentation loss, adversarial loss, and discriminator loss, so as to improve the accuracy of defect segmentation and the ability to express details, until the predetermined training termination condition is reached.

[0050] In implementation, a progressive training strategy is adopted to optimize the network parameters of the defect detection model, which effectively promotes the coordinated learning of the generator and discriminator. By gradually introducing the discriminator and adversarial loss in stages, the generator first ensures stable feature representation capabilities on the basic semantic segmentation task. Subsequently, adversarial training is introduced to enhance the model's detail capture and realistic generation. Finally, joint optimization achieves an overall improvement in model performance. This strategy avoids the training instability that may occur when training the generator and discriminator simultaneously, and is beneficial to the convergence of the training process and performance optimization.

[0051] Alternatively, this application may also employ other training strategies to optimize the network parameters of the defect detection model according to actual needs, such as adopting an adaptive weight adjustment strategy or applying a course learning strategy. This application does not impose any restrictions on the specific training strategy.

[0052] In summary, combining Figure 2 This application, based on a generative adversarial network (GAN) framework, introduces a multi-layer discriminator into a semantic segmentation model, constructing a defect detection model comprising an encoder, decoder, and discriminator. The encoder extracts multi-scale features from the chip surface image, the decoder recovers the image spatial resolution layer by layer and outputs the corresponding scale of defect segmentation results, and the discriminator distinguishes the segmentation results at each scale from the ground truth labels, achieving deep supervision and adversarial training. By pre-training the encoder and decoder parameters using semantic segmentation loss, and then jointly optimizing the generator and discriminator parameters based on the discriminator's discrimination results, iterative updates of the model are achieved, effectively improving the model's ability to identify complex and subtle defects. By introducing a GAN framework, a multi-layer discriminator is designed to discriminate the segmentation results at different scales output by the decoder, forming a deep supervision mechanism that enhances the ability to capture subtle defect features and improves boundary accuracy. Simultaneously, a staged training strategy is adopted: the generator is first pre-trained using semantic segmentation loss, and then adversarial training is performed in conjunction with the discriminator, achieving collaborative optimization of the generator and discriminator, effectively improving the model's robustness and generalization ability. This approach not only overcomes the limitations of traditional methods but also solves the problems of insufficient supervision and imprecise feature representation in deep learning methods, thereby achieving more accurate and stable defect detection while ensuring computational efficiency.

[0053] Figure 3This is a structural block diagram of a semiconductor chip surface defect detection system based on deep adversarial supervision according to an embodiment of this application. The system includes at least the following modules: Image acquisition module, used to acquire images of the chip surface; The defect segmentation module is used to input a chip surface image into a pre-built defect detection model to obtain the defect segmentation result of the chip surface image. The defect detection model includes an encoder, a decoder, and at least one discriminator. The encoder is used to extract multi-scale features of the chip surface image. The decoder is used to recover the spatial resolution of the image layer by layer based on the multi-scale features and output the defect segmentation result at the corresponding scale. The discriminator is used to distinguish the difference between the defect segmentation result output by the decoder at the corresponding scale and the real label image at the corresponding scale. The parameter update module is used to update the parameters of the encoder and decoder by calculating the semantic segmentation loss based on the chip surface image and the corresponding real label image. Then, adversarial training is performed by combining the discrimination results of the discriminator to jointly optimize the parameters of the encoder, decoder and discriminator, so as to realize the iterative update of the defect detection model parameters.

[0054] For relevant details, please refer to the above method implementation examples.

[0055] Figure 4 This is a block diagram of an electronic device provided in one embodiment of this application. The device includes at least a processor 401 and a memory 402.

[0056] Processor 401 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 401 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0057] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 are used to store at least one instruction, which is executed by the processor 401 to implement the semiconductor chip surface defect detection method based on deep adversarial supervision provided in the method embodiments of this application.

[0058] In some embodiments, the electronic device may also optionally include: a peripheral device interface and at least one peripheral device. The processor 401, memory 402, and peripheral device interface can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface via a bus, signal line, or circuit board. Indicatively, peripheral devices include, but are not limited to: radio frequency circuits, touch displays, audio circuits, and power supplies.

[0059] Of course, electronic devices may also include fewer or more components, and this embodiment does not limit this.

[0060] Optionally, this application also provides a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the semiconductor chip surface defect detection method based on deep adversarial supervision in the above-described method embodiments.

[0061] Optionally, this application also provides a computer product including a computer-readable storage medium storing a program, which is loaded and executed by a processor to implement the semiconductor chip surface defect detection method based on deep adversarial supervision described in the above method embodiments.

[0062] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0063] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for detecting surface defects in semiconductor chips based on deep adversarial supervision, characterized in that, The method includes: Acquire images of the chip surface; The chip surface image is input into a pre-constructed defect detection model to obtain defect segmentation results of the chip surface image; the defect detection model includes an encoder, a decoder, and at least one discriminator; the encoder is used to extract multi-scale features of the chip surface image; the decoder is used to recover the spatial resolution of the image layer by layer based on the multi-scale features and output defect segmentation results at the corresponding scale; the discriminator is used to determine the difference between the defect segmentation results output by the decoder at the corresponding scale and the real label image at the corresponding scale; Based on the chip surface image and the corresponding real label image, the parameters of the encoder and the decoder are first updated by calculating the semantic segmentation loss; then, adversarial training is performed in combination with the discrimination result of the discriminator to jointly optimize the parameters of the encoder, the decoder and the discriminator, thereby realizing the iterative update of the defect detection model parameters.

2. The semiconductor chip surface defect detection method based on deep adversarial supervision according to claim 1, characterized in that, The encoder and the decoder constitute the generator in the defect detection model; the generator adopts the U-Net architecture as the basic framework; the encoder includes several convolutional layers, which process the chip surface image through convolution operations and output feature maps of corresponding scales, and gradually reduce the size of the feature maps through downsampling; the decoder includes several decoding layers, which receive feature maps output from the corresponding convolutional layers of the encoder, and restore the spatial resolution of the image layer by layer through upsampling operations, and output the defect segmentation result corresponding to the resolution of that layer.

3. The semiconductor chip surface defect detection method based on deep adversarial supervision according to claim 2, characterized in that, The discriminator corresponds one-to-one with the decoding layer in the decoder. Each discriminator includes two convolutional structures connected in series and a nonlinear mapping structure. The nonlinear mapping structure is connected to the output of the second convolutional structure. The convolutional structure includes a 3×3 convolutional kernel, a normalization layer, and a LeakyReLU activation layer, used to extract the feature differences between the input defect segmentation result and the real label image. The nonlinear mapping structure includes a Sigmoid activation layer, used to map the features output by the convolutional structure to represent the probability that the input belongs to the real label.

4. The semiconductor chip surface defect detection method based on deep adversarial supervision according to claim 2, characterized in that, A multi-scale supervision strategy is adopted, which involves training by introducing real labels of corresponding resolutions at different scales; the multi-scale supervision strategy is divided into three categories: high-resolution supervision, medium-resolution supervision, and low-resolution supervision. The high-resolution supervision is performed at the output of the last layer of the decoder using real labels with the same resolution as the chip surface image. The intermediate-resolution supervision is performed in the middle layer of the decoder, using the corresponding downsampled real labels for feature maps that have been downsampled once or multiple times. The low-resolution supervision is performed at the deep output of the decoder using large-scale downsampled label images.

5. The semiconductor chip surface defect detection method based on deep adversarial supervision according to claim 2, characterized in that, The encoder incorporates a channel attention mechanism, a spatial attention mechanism, and a pyramid pooling module. The channel attention mechanism is used to perform global average pooling and global max pooling on the channel dimension of the input feature map to obtain channel description vectors respectively. The channel description vectors are then processed by a multilayer perceptron and a channel weight vector is generated by the Sigmoid activation function. The spatial attention mechanism is used to perform global average pooling and global max pooling on the input feature map in the channel dimension to obtain two spatial feature maps. The spatial feature maps are concatenated and then generated into a spatial weight map through a convolution operation. The pyramid pooling module performs spatial downsampling on the input feature map through multi-scale pooling and then upsampling it to the same size. The multi-scale pooling result is then concatenated with the original input feature map in the channel dimension to form a fused feature.

6. The semiconductor chip surface defect detection method based on deep adversarial supervision according to claim 2, characterized in that, The iterative update of the defect detection model parameters adopts a phased training strategy, and uses the generator loss function. With discriminator loss function To optimize the objective; The generator's loss function The definition is as follows: ; in, For semantic segmentation loss, For feature matching loss, To combat the losses; ; in, For the number of discriminators, For the first The segmentation result or feature map obtained from the output of the layer decoding layer. Indicates the relationship with the first The discriminator output corresponding to the layer decoding layer; This represents the binary cross-entropy loss; The loss function of the discriminator The definition is as follows: ; in, In order to be with the first The corresponding scale of the real label output by the layer decoding layer.

7. The semiconductor chip surface defect detection method based on deep adversarial supervision according to claim 6, characterized in that, The iterative update of the defect detection model parameters includes three stages: The first stage is the generator pre-training stage, which utilizes semantic segmentation loss. The parameters of the encoder and decoder in the generator are trained; The second stage involves introducing a discriminator for adversarial training, based on the discriminator loss. The discriminator is trained, and semantic segmentation loss is also incorporated. Combating losses and feature matching loss Train the generator; The third stage is the joint optimization stage, which involves jointly optimizing all parameters of the generator and discriminator, using semantic segmentation loss. Combating losses and feature matching loss To achieve the training objective, the model parameters are iterated continuously until the predetermined training termination condition is met.

8. A semiconductor chip surface defect detection system based on deep adversarial supervision, characterized in that, include: Image acquisition module, used to acquire images of the chip surface; The defect segmentation module is used to input the chip surface image into a pre-built defect detection model to obtain the defect segmentation result of the chip surface image; The defect detection model includes an encoder, a decoder, and at least one discriminator; the encoder is used to extract multi-scale features from the chip surface image. The decoder is used to recover the spatial resolution of the image layer by layer based on the multi-scale features and output the defect segmentation result at the corresponding scale; the discriminator is used to determine the difference between the defect segmentation result output by the decoder at the corresponding scale and the real label image at the corresponding scale. The parameter update module is used to update the parameters of the encoder and the decoder by calculating the semantic segmentation loss based on the chip surface image and the corresponding real label image; then, adversarial training is performed in combination with the discrimination result of the discriminator to jointly optimize the parameters of the encoder, the decoder and the discriminator, thereby realizing the iterative update of the defect detection model parameters.

9. An electronic device, characterized in that, The device includes a processor and a memory; the memory stores a program, which is loaded and executed by the processor to implement a semiconductor chip surface defect detection method based on deep adversarial supervision as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a program that, when executed by a processor, is used to implement a semiconductor chip surface defect detection method based on deep adversarial supervision as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Road defect target detection method based on domain adaptation

    CN121582690A