Wafer defect detection method, wafer defect detection device, wafer defect detection equipment and storage medium

By cutting the wafer image and multi-level feature extraction and fusion, the insufficient sensitivity and false detection of micro defect detection in traditional methods are solved, and efficient defect detection is achieved.

CN120259243AActive Publication Date: 2025-07-04HEFEI ZHE TOWER TECH CO LTD +1

Patent Information

Application Number
CN202510343172.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-04
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

Traditional wafer defect detection methods have low sensitivity to micro defects in complex scenarios and are prone to false detection due to hierarchical information fragmentation.

Method used

The wafer image is cut into multiple sub-maps, and the size is uniformized. Multi-level feature maps are extracted through a pre-trained encoder, and multi-level feature extraction is performed using global and local decoders, and predicted defect heat maps are generated by a hybrid decoder.

Benefits of technology

High sensitivity detection for small defects is achieved, error detection caused by hierarchical information fragmentation is avoided, and detection accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259243A_ABST
    Figure CN120259243A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of semiconductor detection, and discloses a wafer defect detection method, device and equipment and a storage medium. The method comprises the following steps: cutting a wafer image to be detected into a plurality of sub-images; performing size unification processing on the sub-graphs, and extracting multi-level feature graphs of the sub-graphs through a pre-trained encoder; sampling feature maps of the same size from the feature maps of different hierarchies, stacking the feature maps, and inputting the stacked images into a global decoder and a local decoder for multi-level feature extraction; inputting the features extracted by the global decoder and the local decoder into a hybrid decoder for feature fusion, and generating a prediction defect heat map; and determining the wafer defect position based on the predicted defect heat map. By means of the mode, multi-level feature distillation from macroscopic global features to microscopic detail signals is achieved based on the multi-level distillation model, sensitivity to tiny defects is reserved, and the false detection problem caused by hierarchical information splitting of a traditional model is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semiconductor detection, and particularly to a wafer defect detection method, device, equipment and storage medium. Background Art

[0002] In the field of industrial inspection, the identification of tiny defects in complex scenarios has always been the core difficulty in technology research. Traditional methods are often limited by problems such as insufficient feature decoupling ability and weak dynamic adaptability, and there are problems of low sensitivity to tiny defects and misdetection caused by hierarchical information fragmentation.

[0003] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main object of the present invention is to provide a wafer defect detection method, device, equipment and storage medium, aiming to solve the technical problems that traditional methods are often limited by problems such as insufficient feature decoupling ability and weak dynamic adaptability, and there are problems of low sensitivity to tiny defects and misdetection caused by hierarchical information fragmentation.

[0005] To achieve the above object, the present invention provides a wafer defect detection method, and the wafer defect detection method includes the following steps:

[0006] Cut the wafer image to be detected into multiple sub-images;

[0007] Perform size unification processing on the sub-images, and extract multi-level feature maps of the sub-images through a pre-trained encoder;

[0008] Sample feature maps of the same size from feature maps of different levels and stack them, and input the stacked image into a global decoder and a local decoder for multi-level feature extraction;

[0009] Input the features extracted by the global decoder and the local decoder into a hybrid decoder for feature fusion to generate a predicted defect heat map;

[0010] Determine the wafer defect position based on the predicted defect heat map.

[0011] In some embodiments, the cutting the wafer image to be detected into multiple sub-images includes:

[0012] Select a size with a length and width exceeding the original chip area by at least 4 pixels as the cutting size according to the chip arrangement;

[0013] Cut the wafer image to be detected into multiple sub-images according to the cutting size, wherein each sub-image contains a complete chip area.

[0014] In some embodiments, the selected size for the size unification process is 512×512, the pre-trained encoder is a wide_resnet_101 model, and its parameters are kept frozen during training. The extracted feature map levels include the 1st, 3rd, and 5th layers.

[0015] In some embodiments, there are 2 decoders in the global decoder and the local decoder respectively. The global decoder uses a variational autoencoder (VAE) trained on a large dataset, and the local decoder uses a VAE trained on a small semantic defect dataset. The final reconstructed size is the size of the original image. During the training process of the global decoder and the local decoder, the reconstruction losses at multiple levels are calculated respectively, and the model parameters are updated through backpropagation.

[0016] In some embodiments, the calculation formula for the reconstruction loss is where i represents the reconstruction level, Flat represents the operation of flattening a high-dimensional vector, d represents the decoder, and e represents the encoder.

[0017] In some embodiments, the hybrid decoder adopts a channel attention mechanism to adjust the weights between feature maps at different levels, so as to adjust the proportion of the global decoder and the local decoder in the reconstructed image, and calculate the total reconstruction loss for backpropagation to update the model.

[0018] In some embodiments, the calculation formula for the total reconstruction loss is where L i is the encoding-decoder loss at the corresponding level, and a i is the adjustment coefficient.

[0019] In addition, to achieve the above object, the present invention also proposes a wafer defect detection device, which includes:

[0020] An image division module, configured to cut the wafer image to be detected into multiple sub-images;

[0021] An image processing module, configured to perform size unification processing on the sub-images, and extract multi-level feature maps of the sub-images through a pre-trained encoder;

[0022] A feature extraction module, configured to sample feature maps of the same size from feature maps at different levels and stack them, and input the stacked image into the global decoder and the local decoder for multi-level feature extraction;

[0023] A feature fusion module, configured to input the features extracted by the global decoder and the local decoder into the hybrid decoder for feature fusion to generate a predicted defect heat map;

[0024] A defect detection module, configured to determine the wafer defect location based on the predicted defect heat map.

[0025] In addition, to achieve the above object, the present invention also provides a wafer defect detection device, which includes: a memory, a processor, and a wafer defect detection program stored on the memory and executable on the processor, where the wafer defect detection program is configured to implement the steps of the wafer defect detection method as described above.

[0026] In addition, to achieve the above object, the present invention also provides a storage medium, on which a wafer defect detection program is stored, and when the wafer defect detection program is executed by a processor, it implements the steps of the wafer defect detection method as described above.

[0027] In the present invention, the wafer image to be detected is cut into multiple sub-images; the sub-images are subjected to size unification processing, and multi-level feature maps of the sub-images are extracted through a pre-trained encoder; feature maps of the same size are sampled from feature maps of different levels and stacked, and the stacked image is input into a global decoder and a local decoder for multi-level feature extraction; the features extracted by the global decoder and the local decoder are input into a hybrid decoder for feature fusion to generate a predicted defect heat map; the wafer defect location is determined based on the predicted defect heat map. In this way, based on the multi-level distillation model, multi-level feature distillation from macroscopic global features to microscopic detail signals is realized, which not only retains the sensitivity to tiny defects but also avoids the misdetection problem caused by the split of hierarchical information in traditional models. Description of the Drawings

[0028] Figure 1 It is a schematic flowchart of the first embodiment of the wafer defect detection method of the present invention;

[0029] Figure 2 It is a schematic diagram of the basic architecture of the multi-level distillation model in the wafer defect detection method of the present invention;

[0030] Figure 3 It is a schematic diagram of the global decoder and the local decoder in the wafer defect detection method of the present invention;

[0031] Figure 4 It is a schematic diagram of the abnormal prediction output in the wafer defect detection method of the present invention;

[0032] Figure 5 It is a structural block diagram of the first embodiment of the wafer defect detection device of the present invention.

[0033] The realization, functional features, and advantages of the object of the present invention will be further described in conjunction with the embodiments and with reference to the drawings. Detailed Embodiments

[0034] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0035] An embodiment of the present invention provides a wafer defect detection method. Referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of a wafer defect detection method of the present invention.

[0036] In this embodiment, the wafer defect detection method includes the following steps:

[0037] Step S10: Cut the wafer image to be detected into multiple sub-images.

[0038] In this embodiment, the execution subject of this embodiment is a wafer defect detection device. Among them, the wafer defect detection device has functions such as data processing, data communication, and program operation. The wafer defect detection device can be a computer terminal device or other network devices. Of course, it can also be other devices with similar functions. This embodiment does not limit this.

[0039] It should be noted that in the field of industrial inspection, the identification of tiny defects in complex scenarios has always been the core difficulty in technology research. Traditional methods are often limited by problems such as insufficient feature decoupling ability and weak dynamic adaptability, and there are problems of low sensitivity to tiny defects and false detection caused by the disconnection of hierarchical information.

[0040] To solve the above technical problems, in this embodiment, the wafer image to be detected is cut into multiple sub-images; the sub-images are subjected to size unification processing, and multi-level feature maps of the sub-images are extracted through a pre-trained encoder; feature maps of the same size are sampled from feature maps of different levels and stacked, and the stacked image is input into a global decoder and a local decoder for multi-level feature extraction; the features extracted by the global decoder and the local decoder are input into a hybrid decoder for feature fusion to generate a predicted defect heat map; the wafer defect position is determined based on the predicted defect heat map. Through the above method, multi-level feature distillation from macroscopic global features to microscopic detail signals is realized based on a multi-level distillation model, which not only retains the sensitivity to tiny defects but also avoids the problem of false detection caused by the disconnection of hierarchical information in traditional models. Specifically, it can be implemented in the following way.

[0041] In a specific implementation, in this embodiment, the wafer image to be detected needs to be cut into multiple sub-images. Specifically, the size of the cut is selected according to the chip arrangement, with the length and width exceeding the original chip area by at least 4 pixels. The wafer image to be detected is cut into multiple sub-images according to the cut size. During the cutting process, it is necessary to ensure that each sub-image contains a complete chip area. The chip arrangement, that is, the Die distribution. Dies are usually arranged in a regular grid pattern on the wafer, and this arrangement is called an array. Common array forms include squares, hexagons, or other geometric shapes to maximize the utilization efficiency of the wafer.

[0042] It should also be noted that to accurately divide the wafer according to the Die distribution, it is only necessary to ensure that each image has the same size, and at the same time, one image almost contains a complete Die. Only the known normal images in this wafer are used in the training process, without additional annotation.

[0043] Step S20: Perform size unification processing on the sub-images, and extract multi-level feature maps of the sub-images through a pre-trained encoder.

[0044] In a specific implementation, after cutting multiple sub-images, size unification processing needs to be performed on the sub-images. Specifically, the size selected for size unification processing in this embodiment is 512×512. The pre-trained encoder is the wide_resnet_101 model, and its parameters remain frozen during the training process. The extracted feature layers include the 1st, 3rd, and 5th layers. It should be understood that the above size and encoder model can also be adjusted according to actual needs, and this embodiment does not limit this.

[0045] Step S30: Sample feature maps of the same size from feature maps of different levels and stack them, and input the stacked image into a global decoder and a local decoder for multi-level feature extraction.

[0046] In a specific implementation, after obtaining feature maps of different levels, in this embodiment, feature maps of the same size are sampled from feature maps of different levels and stacked respectively. For example, the extracted feature layers in this embodiment are the 1st, 3rd, and 5th layers respectively. The feature maps of the 3rd and 5th layers are upsampled to the same size as the 1st layer feature and stacked in the channel dimension. Then, the stacked image is input into the global decoder and the local decoder for multi-level feature extraction. The process of feature extraction is essentially to input multi-level feature maps into the global decoder and the local decoder for global feature and local feature screening.

[0047] It should be noted that there are 2 decoders in the global decoder and the local decoder respectively. The global decoder uses the variational autoencoder (VAE) trained on a large dataset, and the local decoder uses the VAE trained on a small semantic defect dataset. For example, the decoder of the VAE trained on ImageNet is used as the global feature decoder, and the decoder of the VAE trained on the texture type data in MVTEC is used as the local feature decoder. The final reconstruction size is the same as the original image size. During the training process of the global decoder and the local decoder, the reconstruction losses at multiple levels are calculated respectively, and the model parameters are updated through backpropagation. Among them, the calculation formula of the reconstruction loss is

[0048]

[0049] where i represents the reconstruction level, Flat represents the operation of flattening a high-dimensional vector, d represents the decoder, and e represents the encoder. Among them, the basic architecture of the multi-level distillation model adopted in this embodiment can be referred to Figure 2 as shown Figure 2Water in the middle represents the input data stream (such as a wafer image), which is the original input of the model. Dataset represents the labeled or unlabeled training / testing dataset, which is used for model training and validation. E1, E2, and E3 are the three-level encoder hierarchies respectively, extracting the abstract features of the input data layer by layer. E1 is the shallow encoder, capturing high-frequency details (such as edges and textures); E2 is the middle encoder, extracting medium-grained semantic features; E3 is the deep encoder, learning the global structure and context information. D local-1 and D local-2 are the local decoders respectively, focusing on reconstructing local detailed features, reading local features from LDB1 / LDB2, and gradually restoring high-frequency information (such as tiny defects). D global-2 is the global decoder, reconstructing the overall structure (such as the wafer layout) based on the global features of E3. Dmixture-3 is the hybrid decoder, dynamically integrating local and global features through the channel attention mechanism and outputting the final reconstruction result. Filter1local and Filter2 global represent the local decoder and the global decoder respectively, extracting specific-level local features and global features from LDB1 and LDB2 respectively. Reconstructed Water is the finally reconstructed output data (such as the denoised wafer image), and the anomaly detection result is generated by comparing it with the original image. The specific workflow is as follows: Encoding stage: The input data passes through the E1-E3 encoders in sequence, extracting multi-level features and storing them in the LDB series databases. Feature fusion: By merging features across levels (such as combining E2 features with E1 features), the semantic coherence is enhanced. In the decoding stage, the local decoders (D local-1 / 2) read features from LDB1 / LDB2 and restore the detailed information; the global decoder (D global-2) reconstructs the overall structure based on the E3 features; the hybrid decoder (D mixture-3) integrates local and global features and outputs a high-fidelity reconstruction result. Anomaly detection: By comparing the reconstructed data with the original image, the defective area is located (such as heat map difference analysis). Further, in this embodiment, the structures of the global decoder and the local decoder are as Figure 3 shown.

[0050] Step S40: Input the features extracted by the global decoder and the local decoder into the hybrid decoder for feature fusion to generate a predicted defect heat map.

[0051] It should be noted that considering that the sizes of the reconstructed features of the global decoder and the local decoder are already the initial sizes, the hybrid decoder is adopted to adjust the proportions of the global and local decoders in the final reconstructed image. The specific implementation method is to stack the feature maps and adjust the weights between different layers through the channel attention mechanism and finally output the reconstructed image, calculate the total reconstruction loss and perform backpropagation to update the model.

[0052] Specifically, the hybrid decoder adopts a channel attention mechanism to adjust the weights between feature maps of different levels, so as to adjust the proportion of the global decoder and the local decoder in the reconstructed image, and calculates the total reconstruction loss for backpropagation to update the model. The calculation formula of the total reconstruction loss is as follows:

[0053]

[0054] Among them, L i is the loss of the encoding-decoder corresponding to the corresponding level, and a i is the adjustment coefficient. In specific applications, the total reconstruction loss is calculated and the model is updated by backpropagation. The final loss calculation weight is 0.6 for the hybrid layer loss and 0.2 for the rest.

[0055] Finally, all network parameters are frozen, weights are loaded for prediction, and a predicted defect heat map is obtained by comparing with the original image. All network parameters are frozen during the inference process. During the inference, the image to be inferred only performs reconstruction in the last layer in the decoder, and the high-frequency information is replaced by the stored values in the high-frequency information library. After the reconstruction is completed, the difference is taken with the original image to obtain the anomaly detection map.

[0056] Step S50: Determine the wafer defect position based on the predicted defect heat map.

[0057] It should be understood that the wafer defect can be obtained according to the predicted defect heat map, and the wafer defect position can be accurately located. Refer to Figure 4 as shown, Figure 4 In it, Input Image represents the input image, such as the original wafer image, GroundTruth is the true label corresponding to the defect feature, Segmentation is the segmentation result, and the defect detection result output by the model. The wafer defect can be accurately obtained through the color difference.

[0058] In this embodiment, the wafer image to be detected is cut into multiple sub-images; the sub-images are subjected to size unification processing, and multi-level feature maps of the sub-images are extracted by a pre-trained encoder; feature maps of the same size are sampled from feature maps of different levels and stacked, and the stacked image is input into the global decoder and the local decoder for multi-level feature extraction; the features extracted by the global decoder and the local decoder are input into the hybrid decoder for feature fusion to generate a predicted defect heat map; the wafer defect position is determined based on the predicted defect heat map. In the above manner, multi-level feature distillation from macroscopic global features to microscopic detail signals is realized based on the multi-level distillation model, which not only retains the sensitivity to tiny defects but also avoids the misdetection problem caused by the hierarchical information fragmentation of the traditional model.

[0059] In addition, an embodiment of the present invention further provides a storage medium, on which a wafer defect detection program is stored. When the wafer defect detection program is executed by a processor, the steps of the wafer defect detection method described above are implemented.

[0060] Referring to Figure 5 , Figure 5 is a structural block diagram of the first embodiment of the wafer defect detection device of the present invention.

[0061] As Figure 5 shown, the wafer defect detection device proposed in the embodiment of the present invention includes:

[0062] An image division module 10, configured to cut the wafer image to be detected into multiple sub-images;

[0063] An image processing module 20, configured to perform size unification processing on the sub-images, and extract multi-level feature maps of the sub-images through a pre-trained encoder;

[0064] A feature extraction module 30, configured to sample feature maps of the same size from feature maps of different levels and stack them, and input the stacked image into a global decoder and a local decoder for multi-level feature extraction;

[0065] A feature fusion module 40, configured to input the features extracted by the global decoder and the local decoder into a hybrid decoder for feature fusion to generate a predicted defect heat map;

[0066] A defect detection module 50, configured to determine the wafer defect position based on the predicted defect heat map.

[0067] In this embodiment, the wafer image to be detected is cut into multiple sub-images; size unification processing is performed on the sub-images, and multi-level feature maps of the sub-images are extracted through a pre-trained encoder; feature maps of the same size are sampled from feature maps of different levels and stacked, and the stacked image is input into a global decoder and a local decoder for multi-level feature extraction; the features extracted by the global decoder and the local decoder are input into a hybrid decoder for feature fusion to generate a predicted defect heat map; the wafer defect position is determined based on the predicted defect heat map. In the above manner, multi-level feature distillation from macroscopic global features to microscopic detail signals is achieved based on a multi-level distillation model, which not only retains the sensitivity to tiny defects but also avoids the false detection problem caused by the disconnection of hierarchical information in traditional models.

[0068] In some embodiments, the image division module 10 is configured to select a size with a length and width exceeding the original chip area by at least 4 pixels as the cutting size according to the chip arrangement manner;

[0069] According to the cutting size, the wafer image to be detected is cut into multiple sub-images, where each sub-image contains a complete chip area.

[0070] In some embodiments, the selected size for the size unification process is 512×512, the pre-trained encoder is the wide_resnet_101 model, and its parameters are kept frozen during training. The extracted feature map levels include the 1st, 3rd, and 5th layers.

[0071] In some embodiments, there are 2 decoders in the global decoder and the local decoder respectively. The global decoder uses the variational autoencoder (VAE) trained on a large dataset, and the local decoder uses the VAE trained on a small semantic defect dataset. The final reconstructed size is the size of the original image. During the training process of the global decoder and the local decoder, the reconstruction losses at multiple levels are calculated respectively, and the model parameters are updated through backpropagation.

[0072] In some embodiments, the calculation formula for the reconstruction loss is where i represents the reconstruction level, Flat represents the operation of flattening the high-dimensional vector, d represents the decoder, and e represents the encoder.

[0073] In some embodiments, the hybrid decoder adopts a channel attention mechanism to adjust the weights between feature maps at different levels, so as to adjust the proportion of the global decoder and the local decoder in the reconstructed image, and calculate the total reconstruction loss for backpropagation to update the model.

[0074] In some embodiments, the calculation formula for the total reconstruction loss is where L i is the encoding-decoder loss corresponding to the corresponding level, and a i is the adjustment coefficient.

[0075] The embodiment of the present application also provides a wafer defect detection device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus. The memory is used to store the wafer defect detection program; the processor is used to implement the above-mentioned wafer defect detection method when executing the program stored on the memory.

[0076] The communication bus mentioned in the above wafer defect detection device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0077] The communication interface is used for communication between the above-mentioned wafer defect detection device and other devices.

[0078] The memory may include a random access memory (English: Random Access Memory, abbreviated as: RAM), or may also include a non-volatile memory (English: Non-Volatile Memory, abbreviated as: NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0079] The above-mentioned processor may be a general-purpose processor, including a central processing unit (English: Central Processing Unit, abbreviated as: CPU), a network processor (English: Network Processor, abbreviated as: NP), etc.; it may also be a digital signal processor (English: Digital Signal Processing, abbreviated as: DSP), an application-specific integrated circuit (English: Application Specific Integrated Circuit, abbreviated as: ASIC), a field-programmable gate array (English: Field-Programmable Gate Array, abbreviated as: FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0080] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

[0081] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0082] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the relevant part of the method embodiment for the related content.

[0083] The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

[0084] It should be understood that the above is only an example for illustration and does not constitute any limitation to the technical solutions of the present invention. In specific applications, those skilled in the art can set according to needs, and the present invention does not make any restrictions in this regard.

[0085] It should be noted that the above-described work process is only illustrative and does not constitute a limitation to the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no restrictions are made here.

[0086] In addition, for the technical details not described in detail in this embodiment, reference can be made to the wafer defect detection method provided in any embodiment of the present invention, and details will not be repeated here.

[0087] In addition, it should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.

[0088] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0089] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a read-only memory (ROM) / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0090] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the description of the present invention's specification and drawings, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

[0091] It can be understood that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention. For the explanations, examples and beneficial effects of related content, reference can be made to the corresponding parts in the above methods.

Claims

1. A wafer defect detection method, characterized in that, The described wafer defect detection method includes: Cutting the wafer image to be detected into multiple sub-images; Performing size unification processing on the sub-images, and extracting multi-level feature maps of the sub-images through a pre-trained encoder; Sampling feature maps of the same size from feature maps of different levels and stacking them, and inputting the stacked image into a global decoder and a local decoder for multi-level feature extraction; Inputting the features extracted by the global decoder and the local decoder into a hybrid decoder for feature fusion to generate a predicted defect heat map; Determining the wafer defect position based on the predicted defect heat map.

2. The wafer defect detection method according to claim 1, wherein The step of cutting the wafer image to be detected into multiple sub-images includes: Selecting a size with a length and width exceeding the original chip area by at least 4 pixels as the cutting size according to the chip arrangement; Cutting the wafer image to be detected into multiple sub-images according to the cutting size, where each sub-image contains a complete chip area.

3. The wafer defect detection method according to claim 1, wherein The size selected for the size unification processing is 512×512, the pre-trained encoder is the wide_resnet_101 model, and its parameters are kept frozen during training. The extracted feature map levels include the 1st, 3rd, and 5th layers.

4. The wafer defect detection method according to claim 1, characterized in that There are 2 decoders respectively in the global decoder and the local decoder. The global decoder uses a variational autoencoder (VAE) trained on a large dataset, and the local decoder uses a VAE trained on a small semantic defect dataset. The final reconstruction size is the original image size. During the training process of the global decoder and the local decoder, the reconstruction losses of multiple levels are calculated respectively, and the model parameters are updated through backpropagation.

5. The wafer defect detection method according to claim 4, characterized in that, The calculation formula of the reconstruction loss is where i represents the reconstruction level, Flat represents the operation of flattening the high-dimensional vector, d represents the decoder, and e represents the encoder.

6. The wafer defect detection method according to claim 1, wherein The hybrid decoder adopts a channel attention mechanism to adjust the weights between feature maps of different levels, so as to adjust the proportion of the global decoder and the local decoder in the reconstructed image, and calculates the total reconstruction loss for backpropagation to update the model.

7. The wafer defect detection method according to claim 6, characterized in that The calculation formula of the total reconstruction loss is where L i is the loss of the hierarchical encoding decoder corresponding to it, and a i is the adjustment coefficient.

8. A wafer defect detection device, characterized in that, The described wafer defect detection device includes: An image division module for cutting the wafer image to be detected into multiple sub-images; An image processing module for performing size unification processing on the sub-images and extracting multi-level feature maps of the sub-images through a pre-trained encoder; A feature extraction module for sampling feature maps of the same size from feature maps of different levels and stacking them, and inputting the stacked image into a global decoder and a local decoder for multi-level feature extraction; A feature fusion module for inputting the features extracted by the global decoder and the local decoder into a hybrid decoder for feature fusion to generate a predicted defect heat map; A defect detection module for determining the wafer defect position based on the predicted defect heat map.

9. A wafer defect detection device, characterized in that, The described wafer defect detection device includes: a memory, a processor, and a wafer defect detection program stored on the memory and executable on the processor. The wafer defect detection program is configured to implement the steps of the wafer defect detection method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, A wafer defect detection program is stored on the storage medium. When the wafer defect detection program is executed by a processor, it implements the steps of the wafer defect detection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Defect detection method based on joint optimization and mixed attention feature fusion

    CN115294038A

  • Surface defect detection method and system based on Swin Transform

    CN116703885A

  • Lightweight defect detection method based on mixed multi-scale knowledge distillation

    CN118154607A

  • Remote sensing image segmentation method based on dual-branch multi-scale feature fusion

    CN118314353A

  • Image anomaly detection method fusing local-global feature reconstruction

    CN119251629A

Cited By

  • Silicon wafer edge defect detection method and device under special chamfering process

    CN120765638A

  • A method and apparatus for detecting edge defects in silicon wafers under a special chamfering process

    CN120765638B