Fourier transform-based machine learning for defect inspection of semiconductor samples
By adopting a computerized system based on machine learning in the semiconductor manufacturing process, the difference image between the test image and the reference image is processed, and multiple machine learning models are used for Fourier transform and defect detection, the problem of insufficient defect detection accuracy and efficiency in the prior art is solved, and high-precision and efficient defect detection are achieved.
Patent Information
- Application Number
- CN202411869503.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-18
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art lacks the accuracy and efficiency of sample defect detection in the semiconductor manufacturing process, especially in the manufacturing of high-density and high-performance equipment, and it is difficult to effectively monitor and classify small defects.
Using a computerized system based on machine learning, the difference image between the test image and the reference image is obtained through the processing circuit system, and Fourier transform and defect detection are used to generate an output image representing the defect map.
The accuracy and efficiency of defect detection are improved, the residual noise in the output image is significantly reduced, the detection sensitivity is enhanced, and the tiny defects in semiconductor samples can be effectively identified.
Smart Images

Figure CN120182173A_ABST
Abstract
Description
Technical Field
[0001] The subject matter of the present disclosure generally relates to the field of inspecting semiconductor samples, and more particularly to machine learning-based defect detection of samples. Background Art
[0002] Current demands for high density and high performance associated with the very large scale integration of fabricated devices require sub-micron features, and require increased transistor and circuit speeds, as well as enhanced reliability. With the progress of semiconductor processes, pattern dimensions such as line widths and other types of critical dimensions are continuously shrinking. Such demands require the formation of device features with high precision and high uniformity, which in turn requires careful monitoring of the manufacturing process, including automated inspection of the device while it is still in the form of a semiconductor wafer.
[0003] Run-time inspection can generally adopt a two-stage procedure. For example, a sample is inspected, and then the sampling locations of potential defects are reviewed. Inspection generally involves generating certain outputs (e.g., images, signals, etc.) of the sample by directing light or electrons onto the wafer and detecting the light or electrons from the wafer. During the first stage, the surface of the sample is inspected at high speed and relatively low resolution. Defect detection is typically performed by applying a defect detection algorithm to the inspection output. A defect map is generated to show the suspicious locations on the sample with a high probability of defects. During the second stage, at least some of the suspicious locations are analyzed more thoroughly at a relatively high resolution to determine different parameters of the defects, such as category, thickness, roughness, size, etc.
[0004] During or after manufacturing the sample to be inspected, inspection can be performed by using non-destructive inspection tools. As non-limiting examples, various non-destructive inspection tools include scanning electron microscopes, atomic force microscopes, optical inspection tools, etc.
[0005] The inspection process can include multiple inspection steps. The manufacturing process of semiconductor devices can include various procedures, such as etching, deposition, planarization, growth such as epitaxial growth, implantation, etc. For example, after certain process procedures and / or after manufacturing certain layers, etc., the inspection step can be performed multiple times. Additionally or alternatively, for example, for different wafer locations, or for the same wafer location with different inspection settings, each inspection step can be repeated multiple times.
[0006] During various steps in semiconductor manufacturing, the inspection process is used to detect and classify defects on the sample, and to perform metrology-related operations. The effectiveness of inspection can be improved by automating processes such as, for example, defect detection, automatic defect classification (ADC), automatic defect review (ADR), image segmentation, automatic metrology-related operations, etc.
[0007] An automated inspection system ensures that the manufactured components meet the expected quality standards and provides useful information on the adjustments that may be required to the manufacturing tools, equipment, and / or compositions based on the type of identified defects.
[0008] In some cases, machine learning techniques can be used to assist the automated inspection process to facilitate the improvement of the yield. For example, supervised machine learning can be used to achieve an accurate and efficient solution for automating a specific inspection application based on fully annotated training images. SUMMARY OF THE INVENTION
[0009] According to certain aspects of the subject matter of the present disclosure, a computerized system for performing runtime defect inspection on a semiconductor sample is provided. The system includes processing circuitry configured to: obtain an input image indicative of a difference between an inspection image of the sample and a corresponding reference image; and use a trained machine learning (ML) system to process the input image to generate an output image representing a defect map indicative of the distribution of defect of interest (DOI) candidates in the input image, wherein the ML system includes a plurality of ML models that are operatively connected to each other and are pre-trained together to perform defect detection on the input image based on a Fourier transform, and wherein the output image can be used for further defect inspection.
[0010] In addition to the above features, the system according to this aspect of the subject matter of the present disclosure can include one or more of the following features (i) to (viii) in any desired combination or arrangement that is technically possible:
[0011] (i). The ML system can include a first ML model that is operatively connected in parallel to a second ML model and a third ML model. The first ML model is configured to generate one or more feature maps representing the input image, and the second ML model and the third ML model are configured to separately process the one or more feature maps and generate the real and imaginary parts of a Fourier image corresponding to the input image, respectively.
[0012] (ii). The ML system further includes an analysis model configured to perform an inverse Fourier transform (IFT) on the Fourier image to reconstruct the output image.
[0013] (iii). The first ML model is an autoencoder, and the second ML model and the third ML model are convolutional neural networks (CNNs).
[0014] (iv). The ML system includes a first ML model that is operatively connected to a second model. The first ML model is configured to generate one or more feature maps representing an input image, and the second model is configured to process the one or more feature maps and reconstruct an output image.
[0015] (v). The ML system is trained to: detect the presence of a DOI in a training image based on the frequency response of the Fourier image corresponding to the training image and the ground truth frequency response of the DOI; and identify the location of the DOI in the training image based on the training image and the reconstructed image of the training image.
[0016] (vi). The processing circuitry is further configured to: obtain a second input image indicative of the difference between a test image and a second reference image; use the trained ML system to process the second input image to obtain a second output image; and determine the presence of a DOI candidate based on the output image and the second output image.
[0017] (vii). The ML system includes a classifier. The processing circuitry is further configured to: use the classifier to process the output image to provide a classification score indicative of the confidence level of the presence of a DOI in the output image; and determine the presence of a DOI based on the output image and the classification score.
[0018] (viii). The input image is a difference image obtained by comparing the test image with a reference image. The output image suppresses residual noise relative to the input image, which improves detection sensitivity when used for further defect inspection.
[0019] According to other aspects of the subject matter of the present disclosure, there is provided a computerized method for performing runtime defect inspection on a semiconductor sample, the method comprising: obtaining an input image indicative of the difference between a test image of the sample and a corresponding reference image; and using a trained machine learning (ML) system to process the input image to generate an output image representing a defect map indicative of the distribution of defects of interest (DOI) candidates in the input image, wherein the ML system includes a plurality of ML models that are operatively connected to each other and are pre-trained together to perform defect detection on the input image based on a Fourier transform, and wherein the output image is usable for further defect inspection.
[0020] With necessary modifications, these aspects of the subject matter of the present disclosure may include any desired combination or arrangement technically possible, including one or more of the features (i) to (viii) listed above with respect to the system.
[0021] According to other aspects of the subject matter of the present disclosure, a computerized method of training a machine learning (ML) system that can be used for defect inspection of semiconductor samples is provided. The method includes: obtaining a training set that includes a first subset of training images and a second subset of training images. Each training image in the first subset of training images includes a DOI, and the second subset of training images does not have any DOI. Each training image indicates the difference between the inspection image of the sample and the corresponding reference image and is associated with the ground truth (GT) frequency response of each training image in the Fourier domain; for each given training image in the training set, using the ML system to process the given training image to generate a Fourier image of the given training image and the frequency response of the Fourier image; performing an inverse Fourier transform (IFT) process on the Fourier image to obtain a reconstructed image; and using a loss function having two components to optimize the ML system: a first component based on the frequency response associated with the given training image and the GT frequency response, and a second component based on the reconstructed image and the given training image.
[0022] According to other aspects of the subject matter of the present disclosure, a computerized method of training a machine learning (ML) system that can be used for defect inspection of semiconductor samples is provided. The method includes: obtaining a training set that includes a first subset of training images and a second subset of training images. Each training image in the first subset of training images includes a DOI, and the second subset of training images does not have any DOI. Each training image indicates the difference between the inspection image of the sample and the corresponding reference image and is associated with the ground truth (GT) frequency response of each training image in the Fourier domain; for each given training image in the training set, using the ML system to process the given training image to obtain a reconstructed image; processing the reconstructed image to generate a Fourier image of the given training image and the frequency response of the Fourier image; and using a loss function having two components to optimize the ML system: a first component based on the frequency response associated with the given training image and the GT frequency response, and a second component based on the reconstructed image and the given training image.
[0023] In addition to the above features, these aspects of the subject matter of the present disclosure may include any desired combination or arrangement technically possible, including one or more of the features (ix) to (xvi) listed below:
[0024] (ix). The first subset includes at least one training image synthetically generated by implanting a DOI into a defect-free difference image.
[0025] (x). The ML system includes: a first ML model configured to generate, for each given training image, one or more feature maps representing the given training image; a second ML model and a third ML model configured to separately process the one or more feature maps and generate, respectively, the real part and the imaginary part of a Fourier image corresponding to the given training image; and an analysis model configured to perform IFT based on the Fourier image to obtain a reconstructed image.
[0026] (xi). Optimization using the first loss component enables the second ML model and the third ML model to learn to generate Fourier images having a frequency response close to the GT frequency response, which in turn causes the first ML model to learn to extract one or more feature maps that are more relevant to the presence of DOI while suppressing noise in the given training image.
[0027] (xii). The second loss component is a reconstruction loss representing the difference between the reconstructed image and the defect information of the given training image, and optimization using the second loss component enables the ML system to identify the location where DOI is present in the given training image.
[0028] (xiii). Depending on the type of DOI, the ground truth frequency response is one of the following: Gaussian response, incremented response, cylindrical response, and rectangular response.
[0029] (xiv). The analysis model is configured to perform IFT to generate a composite image including the real part representing the reconstructed image and the imaginary part. The analysis model is further configured to control the weight assigned to the imaginary part so as to introduce uncertainty about the presence of DOI in the reconstructed image.
[0030] (xv). The method further includes: using the ML system to process the reconstructed image to provide a classification score indicating the confidence level of the presence of DOI in the reconstructed image; and using a classification loss function based on the training image and the classification score to optimize the ML system.
[0031] (xvi). The ML system includes: a first ML model configured to generate one or more feature maps representing a given training image; a second model configured to generate a reconstructed image based on the one or more feature maps; and an analysis model configured to perform a Fourier transform based on the reconstructed image.
[0032] According to other aspects of the subject matter of the present disclosure, there is provided a non-transitory computer-readable medium including instructions that, when executed by a computer, cause the computer to perform the method steps of any one of the above methods. Description of the Drawings
[0033] To understand the present disclosure and to see how it may be carried out in practice, embodiments will now be described by way of non-limiting examples only with reference to the accompanying drawings, in which:
[0034] Figure 1 A generalized block diagram of an inspection system in accordance with certain embodiments of the subject matter of the present disclosure is illustrated.
[0035] Figure 2 A generalized flowchart of training a machine learning system that can be used to perform defect inspection on semiconductor samples in accordance with certain embodiments of the subject matter of the present disclosure is illustrated.
[0036] Figure 3 A generalized flowchart of an adaptive process of block 203 in an alternative ML system implementation in accordance with certain embodiments of the subject matter of the present disclosure is illustrated.
[0037] Figure 4 A generalized flowchart of performing runtime defect inspection on semiconductor samples using a trained ML system in accordance with certain embodiments of the subject matter of the present disclosure is illustrated.
[0038] Figure 5 A generalized flowchart of an additional process of performing runtime defect inspection using a second input image in accordance with certain embodiments of the subject matter of the present disclosure is shown.
[0039] Figure 6 A schematic diagram of a training process of an ML system in accordance with certain embodiments of the subject matter of the present disclosure is shown.
[0040] Figure 7 A schematic diagram of a training process of an ML system with an alternative implementation in accordance with certain embodiments of the subject matter of the present disclosure is shown. Detailed Description
[0041] The processes of semiconductor manufacturing typically require multiple successive processing steps and / or layers, and each of the processing steps and / or layers may introduce errors that can lead to yield loss. Examples of various processing steps may include lithography, etching, deposition, planarization, growth (such as epitaxial growth), and implantation, among others. Various inspection operations such as defect-related inspections (e.g., defect detection, defect review, and defect classification, etc.) and / or metrology-related inspections (e.g., critical dimension (CD) measurement, etc.) may be performed at different processing steps / layers during the manufacturing process to monitor and control the process. For example, the inspection operations may be performed multiple times, such as after certain processing steps and / or after manufacturing certain layers, etc.
[0042] As described above, defect inspection can generally adopt a two-stage procedure, for example, inspection of a sample followed by review of the sampled locations for potential defects. In the first stage, the surface of the sample is inspected at high speed and relatively low resolution. Defect detection is typically performed by applying a defect detection algorithm to the inspection output. Various detection algorithms, such as die-to-die (D2D), die-to-history (D2H), die-to-database (D2DB), etc., can be used to detect defects on the sample.
[0043] As an example, a typical die-to-reference detection algorithm, such as die-to-die (D2D), is often used in some cases. In D2D, an inspection image of a target die is captured. In order to detect defects in the inspection image, one or more reference images are captured from one or more reference dies (e.g., one or more adjacent dies) of the target die. The inspection image and the reference image are aligned and compared to each other. One or more difference images (and / or derivatives of one or more difference images, such as grade images) may be generated based on the difference between the pixel values of the inspection image and the pixel values derived from the one or more reference images. A detection threshold is then applied to the difference map, and a defect map indicating defect candidates in the target die is created.
[0044] Typically, the different changes between the inspection image and the reference image may be caused by the physical effects of the manufacturing process and / or the inspection process of the sample, resulting in false alarms and random noise in the difference image and / or defect map. Examples of such changes include process changes and color changes, etc. Process changes refer to changes caused by changes in the manufacturing process of the sample. As an example, the manufacturing process may cause slight shifts / scaling / distortion of certain structures / patterns, which results in pattern changes between different images (such as the inspection image and the reference image). As another example, the manufacturing process may cause thickness changes in the sample, which affects the reflectivity and thus affects the grayscale of the resulting inspection image. For example, die-to-die material thickness changes may result in different reflectivities between the two in the die, which results in different background grayscale values for the images of the two dies. Color changes may be caused by process changes and / or (multiple) inspection tools used to inspect the sample. As an example, changes and / or calibration of the inspection tool, such as different settings of the inspection tool (e.g., optical mode, detector, etc.), may result in grayscale differences of different inspection images. Color changes may occur within a single image (e.g., due to layer thickness changes) or between the inspection image and the reference image.
[0045] In some cases, image preprocessing techniques can be applied to the inspection and reference images to reduce the impact of these variations. However, these variations cannot usually be completely eliminated, leaving residual variations and noise in the difference image and / or defect map, which affects the defect detection sensitivity and thus reduces the detection performance.
[0046] As semiconductor manufacturing processes continue to advance, semiconductor devices are developed into increasingly complex structures, and the feature sizes of semiconductor devices are gradually reduced, making it more difficult for conventional inspection methods to meet inspection performance.
[0047] Accordingly, certain embodiments of the subject matter of the present disclosure propose using a machine learning-based defect inspection system that does not have one or more of the disadvantages described above. The present disclosure proposes using a trained machine learning (ML) system to process an input image (such as a difference image) and generate an output image representing a defect map that indicates the distribution of defect of interest (DOI) candidates in the input image. The proposed runtime inspection system can generate an output image with significantly reduced residual noise and variations, thereby improving defect detection sensitivity. The ML model is constructed and trained in a specific manner, which will be described in detail below.
[0048] Based on this, note that Figure 1 , the figure illustrates a functional block diagram of an inspection system according to certain embodiments of the subject matter of the present disclosure.
[0049] In Figure 1 The inspection system 100 shown can be used to inspect semiconductor samples (e.g., wafers, dies, and / or portions thereof) as part of a sample manufacturing process. As described above, the inspection mentioned herein can be interpreted to cover any type of operation involving defect inspection / detection, defect classification, segmentation, and / or metrology operations, etc. regarding the sample. The system 100 includes one or more inspection tools 120 configured to scan the sample and capture an image of the sample to be further processed for various inspection applications.
[0050] The term "(s) inspection tool" as used herein should be broadly interpreted to cover any tool that can be used in an inspection-related process. As a non-limiting example, inspection-related processes include scanning (in single or multiple scans), imaging, sampling, reviewing, measuring, classifying, and / or other processes performed on the sample or a portion of the sample. Without limiting the scope of the present disclosure in any way, it should also be noted that the inspection tool 120 can be implemented as various types of inspection machines, such as optical inspection machines, electron beam inspection machines (e.g., scanning electron microscope (SEM), atomic force microscope (AFM), or transmission electron microscope (TEM), etc.).
[0051] One or more inspection tools 120 may include one or more inspection tools and / or one or more review tools. In some cases, at least one of the inspection tools 120 may be an inspection tool configured to scan a sample (e.g., an entire wafer, an entire die, or a portion of the sample) to capture inspection images (typically at a relatively high speed and / or low resolution) to detect potential defects (i.e., defect candidates). During inspection, during exposure, the wafer may be moved in a stepwise manner relative to the detector of the inspection tool (or the wafer and the tool may be moved relative to each other in opposite directions), and the wafer may be scanned step by step by the inspection tool along a strip of the wafer, where the inspection tool images a part / portion of the sample (within the strip) at a time. As an example, the inspection tool may be an optical inspection tool. At each step, light may be detected from a rectangular portion of the wafer, and this detected light may be converted into a plurality of intensity values at a plurality of points in that portion, thereby forming an image corresponding to a part / portion of the wafer. For example, in optical inspection, an array of parallel laser beams may scan the surface of the wafer along the strip. The strips are laid in parallel rows / columns that abut each other to build an image of the surface of the wafer strip by strip. For example, the tool may scan the wafer along the strips from top to bottom, then switch to the next strip and scan the wafer from bottom to top, and so on, until the entire wafer is scanned and the inspection image of the wafer is collected.
[0052] In some cases, at least one of the inspection tools 120 may be a review tool configured to capture review images of at least some of the defect candidates detected by the inspection tool to determine whether the defect candidates are indeed defects of interest (DOI). Such review tools are typically configured to inspect fragments of the sample one by one (usually at a relatively low speed and / or high resolution). As an example, the review tool may be an electron beam tool, such as a scanning electron microscope (SEM), etc. An SEM is an electron microscope that produces an image of a sample by scanning the sample with a focused electron beam. The electrons interact with the atoms in the sample, thereby generating various signals containing information about the surface topography and / or composition of the sample. The SEM is capable of accurately inspecting and measuring features during the manufacture of semiconductor wafers.
[0053] The inspection tool and the review tool can be different tools located at the same or different positions, or a single tool operating in two different modes. In some cases, the same inspection tool can provide low-resolution image data and high-resolution image data. The resulting image data (low-resolution image data and / or high-resolution image data) can be transmitted directly or via one or more intermediate systems to system 101. The present disclosure is not limited to any particular type of inspection tool and / or the resolution of the image data generated by the inspection tool. In some cases, at least one of the inspection tools 120 has metrology capabilities and can be configured to capture an image and perform metrology operations on the captured image. Such an inspection tool is also referred to as a metrology tool.
[0054] According to certain embodiments of the subject matter of the present disclosure, inspection system 100 includes a computer-based system 101 that is operatively connected to inspection tool 120 and is capable of performing automatic defect detection on a semiconductor sample at runtime based on runtime images obtained during sample fabrication. System 101 is also referred to as a defect detection system.
[0055] System 101 includes processing circuitry 102 that is operatively connected to a hardware-based I / O interface 126 and is configured to provide the processing required to operate the system, as further detailed Figures 2 to 5 below. Processing circuitry 102 can include one or more processors (not shown separately) and one or more memories (not shown separately). One or more processors of processing circuitry 102 can be configured to execute a number of functional modules individually or in any suitable combination according to computer-readable instructions implemented on a non-transitory computer-readable memory included in the processing circuitry. Such functional modules are hereinafter referred to as being included in the processing circuitry.
[0056] One or more processors mentioned herein can represent one or more general-purpose processing devices, such as a microprocessor, a central processing unit, etc. More particularly, a given processor can be one of the following: a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. One or more processors can also be one or more dedicated processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. One or more processors are configured to execute the instructions for performing the operations and steps described herein.
[0057] The memories mentioned in this document may include one or more of the following: internal memories such as processor registers and caches, main memories such as read-only memories (ROMs), flash memories, and dynamic random access memories (DRAMs) such as synchronous DRAMs (SDRAMs) or Rambus DRAMs (RDRAMs).
[0058] According to certain embodiments of the subject matter of this disclosure, system 101 may be a runtime defect detection system configured to perform defect detection operations using a trained machine learning (ML) system based on runtime images obtained during sample fabrication. In such a case, one or more functional modules included in processing circuitry 102 of system 101 may include an ML system 106 pre-trained for defect detection during a training / setup phase and include an optional defect inspection module 108.
[0059] Specifically, processing circuitry 102 may be configured to: obtain an input image indicating a difference between an inspection image of a sample (acquired at runtime by an inspection tool) and a corresponding reference image via an I / O interface 126; and provide the input image as an input to a machine learning system (such as ML system 106) for processing. ML system 106 may generate an output image representing a defect map indicating the distribution of defect of interest (DOI) candidates in the input image. ML system 106 has been pre-trained using a training set during a setup / training phase. Specifically, the ML system includes a plurality of ML models operatively connected to each other and pre-trained together to perform defect detection on the input image based on a Fourier transform. In some cases, the optional defect inspection module 108 may be configured to perform further defect inspections on the output image, such as further damage filtering, defect review, defect classification, etc.
[0060] In such a case, ML system 106 and defect inspection module 108 may be regarded as part of a defect inspection recipe that can be used to perform runtime defect inspection operations on the acquired runtime images. System 101 may be regarded as a runtime defect inspection system capable of performing runtime defect-related operations using the defect inspection recipe. Details of the runtime inspection process are described below with reference to Figures 4 to 5 Describe the details of the runtime inspection process.
[0061] In some embodiments, system 101 may be configured as a training system capable of training an ML model using a specific training set during a training / setup phase. In such a case, one or more functional modules included in the processing circuitry 102 of system 101 may include a training module 104 and an ML system 106 to be trained. Specifically, the training module 104 may be configured to obtain a training set that includes a first subset of training images and a second subset of training images, where each training image in the first subset of training images includes a DOI and the second subset of training images does not have any DOI. Each training image indicates the difference between a test image of a sample and a corresponding reference image and is associated with a ground truth (GT) frequency response of each training image in the Fourier domain.
[0062] The training module 104 may be configured to use the training set to train the ML system 106. Specifically, for each training image in the training set, the ML system may be configured to process the training image to generate a transformed Fourier image of the training image and a frequency response of the transformed Fourier image. An inverse Fourier transform (IFT) process may be performed on the transformed Fourier image to obtain a reconstructed image. A loss function having two components may be used to optimize the ML system: a first component based on the frequency response associated with the training image and the GT frequency response, and a second component based on the reconstructed image and the training image.
[0063] As described above, the ML system can be used for runtime defect detection after being trained. Details of the training process are described below with reference to Figures 2 to 3 and Figures 6 to 7 Describe the details of the training process.
[0064] Reference will be made to Figures 2 to 5 Further details of the operation of systems 100 and 101, the processing circuitry 102, and the functional modules therein will be described.
[0065] According to certain embodiments, the ML system 106 may include a plurality of ML models operatively connected to each other. The ML models mentioned herein may be implemented as various types of machine learning models. As an example, the ML model may be implemented as one of the following: a support vector machine (SVM), a neural network, a Bayesian network, a transformer, and / or a collection / combination thereof. The learning algorithm used by the ML model may be any one of the following: supervised learning, unsupervised learning, self-supervised learning, or semi-supervised learning, etc. The subject matter of the present disclosure is not limited to a specific type of ML model or a specific type of learning algorithm used by the ML model.
[0066] In some embodiments, the ML model may be implemented as a deep neural network (DNN). The DNN may include multiple layers organized according to a corresponding DNN architecture. As a non-limiting example, the layers of the DNN may be organized according to the architecture of a convolutional neural network (CNN), a recurrent neural network, a recursive neural network, an autoencoder, a generative adversarial network (GAN), etc. Optionally, at least some of the layers may be organized into multiple DNN sub-networks. Each layer of the DNN may include multiple basic computing units (CEs), which are commonly referred to in the art as dimensions, neurons, or nodes.
[0067] The weighted values and / or thresholds associated with the CEs of the deep neural network and the connections of the deep neural network may be initially selected before training and may be further iteratively adjusted or modified during training to achieve an optimal set of weighted values and / or thresholds in the trained DNN. After each iteration, the difference between the actual output produced by the DNN module and the target output associated with the corresponding training data set may be determined. The difference may be referred to as an error value. Training may be determined to be complete when the loss / cost function indicating the error value is less than a predetermined value or when a limited change in performance between iterations is achieved. The input data set used to adjust the weights / thresholds of the deep neural network is referred to as the training set.
[0068] Note that the teachings of the subject matter of the present disclosure are not limited by the specific architecture of the ML model or DNN as described above.
[0069] It should be noted that although some embodiments of the present disclosure relate to a processing circuitry 102 configured to perform the operations described above, the functionality / operations of the foregoing functional modules may be performed in various ways by one or more processors in the processing circuitry 102. As an example, the operations of each functional module may be performed by a specific processor or a combination of processors. Thus, the operations of the various functional modules may be performed by the corresponding processors (or combinations of processors) in the processing circuitry 102, such as processing input images and performing defect inspections, etc., while optionally, these operations may be performed by the same processor. The present disclosure should not be construed as being limited to a single processor that always performs all operations.
[0070] In some cases, in addition to the system 101, the inspection system 100 may include one or more inspection modules, such as a defect detection module, a damage filtering module, an automated defect review module (ADR), an automated defect classification module (ADC), a metrology operation module, and / or other inspection modules that can be used to inspect semiconductor samples. One or more inspection modules may be implemented as stand-alone computers, or the functionality (or at least a portion of the functionality) of one or more inspection modules may be integrated with the inspection tool 120. In some cases, the output of the system 101 (e.g., a trained ML system, a generated output image, and / or a defect inspection result) may be provided to one or more inspection modules (such as ADR, ADC, etc.) for further processing.
[0071] According to certain embodiments, the system 100 may include a storage unit 122. The storage unit 122 may be configured to store any data required by the operating system 101, such as data related to the input and output of the system 101, and intermediate processing results generated by the system 101. As an example, the storage unit 122 may be configured to store images of samples and / or derivatives of the images generated by the inspection tool 120, such as the runtime input images and training sets as described above. Thus, these input images may be retrieved from the storage unit 122 and provided to the processing circuitry 102 for further processing. The output of the system 101 (such as a trained ML system, a generated output image, and / or a defect inspection result) may be sent to the storage unit 122 for storage.
[0072] In some embodiments, the system 100 may optionally include a computer-based graphical user interface (GUI) 124, which is configured to enable user-specified input related to the system 101. For example, a visual representation of the sample, including an image of the sample, etc., may be presented to the user (e.g., via a display forming part of the GUI 124). Options for defining certain operation parameters may be provided to the user through the GUI. The user may also view operation results or intermediate processing results on the GUI, such as generated output images and / or defect inspection results, etc.
[0073] In some cases, system 101 may be further configured to send the operation result to inspection tool 120 via I / O interface 126 for further processing. In some cases, system 101 may be further configured to send the result to storage unit 122 and / or an external system (e.g., a yield management system (YMS) of a manufacturing plant (semiconductor foundry)). In the context of semiconductor manufacturing, a yield management system (YMS) is a data management, analysis, and tool system that, especially during manufacturing ramp-up, collects data from a semiconductor foundry and helps engineers find ways to improve yield. YMS helps semiconductor manufacturers and semiconductor foundries manage a large amount of production analysis with fewer engineers. These systems analyze yield data and generate reports. Integrated device manufacturers (IDMs), semiconductor foundries, fabless semiconductor companies, and outsourced semiconductor assembly and test (OSAT) can use YMS.
[0074] Those skilled in the art will readily understand that the teachings of the subject matter of the present disclosure are not limited by Figure 1 the systems shown. Figure 1 Each system component and module in may be constituted by any combination of software, hardware, and / or firmware. Correspondingly, software, hardware, and / or firmware are executed on one or more appropriate devices, and the one or more devices perform the functions defined and explained herein. The equivalent and / or modified functionality described for each system component and module may be combined or divided in another way. Thus, in some embodiments of the subject matter of the present disclosure, the system may include fewer, more, modified, and / or different components, modules, and functions than Figure 1 those shown.
[0075] Figure 1 Each component in may represent multiple specific components that are adapted to operate independently and / or collaboratively to process various data and electrical inputs and to implement operations related to a computerized inspection system. In some cases, multiple instances of a component may be utilized for reasons of performance, redundancy, and / or availability. Similarly, in some cases, multiple instances of a component may be utilized for reasons of functionality or application. For example, different parts of a specific functionality may be placed in different instances of a component.
[0076] It should be noted that the inspection system shown in Figure 1 can be implemented in a distributed computing environment, where Figure 1One or more of the foregoing components and functional modules shown may be distributed across several local and / or remote devices. By way of example, inspection tool 120 and system 101 may be located at the same entity (hosted by the same device in some cases), or distributed across different entities. As another example, as described above, in some cases, system 101 may be configured as a training system for training an ML model, while in some other cases, system 101 may be configured as a runtime defect detection system that uses the trained ML model. The training system and the runtime detection system may be located at the same entity (hosted by the same device in some cases), or distributed across different entities, depending on the specific system configuration and implementation requirements.
[0077] In some examples, certain components utilize cloud implementations, such as implemented in a private or public cloud. In cases where the various components of the inspection system are not all located at one location or one physical entity, communication between the various components of the inspection system may be achieved via any signaling system or communication components, modules, protocols, software languages, and drive signals, and the communication may appropriately be wired and / or wireless.
[0078] It should also be noted that in some embodiments, at least some of inspection tool 120, storage unit 122, and / or GUI 124 may be external to inspection system 100 and operate in data communication with systems 100 and 101 via I / O interface 126. As described above, system 101 may be implemented as a (multiple) standalone computer for use in conjunction with inspection tools and / or additional inspection modules. Alternatively, the corresponding functions of system 101 may be at least partially integrated with one or more inspection tools 120, thereby enhancing and augmenting the functionality of inspection tool 120 in inspection-related processes.
[0079] Although not necessarily so, the operational procedures of systems 101 and 100 may correspond to some or all of the stages of the method described with respect to Figures 2 to 5 Similarly, the methods described with respect to Figures 2 to 5 and their possible implementations may be implemented by systems 101 and 100. Therefore, note that, with necessary modifications, the embodiments discussed with respect to the methods described with respect to Figures 2 to 5 may also be implemented as various embodiments of systems 101 and 100, and vice versa.
[0080] Reference Figure 2 , a generalized flowchart of training a machine learning system that can be used for defect inspection of semiconductor samples is illustrated in accordance with certain embodiments of the subject matter of the present disclosure.
[0081] A training set can be obtained (202) (e.g., by the training module 104 in the processing circuitry 102). The training set includes a first subset of training images and a second subset of training images, where the training images in the first subset each include a defect of interest (DOI), and the second subset of training images does not have any DOI. The training images including DOI can also be referred to as defect training images, while the training images without any DOI can also be referred to as defect-free training images, i.e., clean images without defect features. In some cases, the training images can be obtained as image tiles of a predefined size from patches of inspection images of samples acquired by an inspection tool.
[0082] Each training image in the training set can indicate the difference between an inspection image of a sample and a corresponding reference image (the reference image corresponding to the inspection image in the sense of capturing a similar region containing a pattern similar to the pattern of the inspection image). As an example, the training image can be a difference image or a derivative of a difference image generated by a comparison of the pixel values of the inspection image and the reference image of the inspection image. For example, the difference image can be generated by subtracting the reference image from the inspection image. In some cases, a rank image as a derivative of the difference image can be generated by applying a predefined difference normalization factor to the difference image. The difference normalization factor can be determined based on the behavior of a normal population of pixel values and can be used to normalize the pixel values of the difference image. As an example, the rank of a pixel can be calculated as the ratio between the corresponding pixel value of the difference image and the predefined difference normalization factor. The difference image or the rank image (or any further derivatives thereof) can be used as the training image.
[0083] In some embodiments, some of the training images in the first subset and / or the second subset can be synthetic images generated by image simulation, compared with real images generated by actual image acquisition by an inspection tool (e.g., real difference images obtained from a comparison of an actually acquired inspection image and a reference image). As an example, the first subset of training images and / or the second subset of training images can in some cases include only real images, or only synthetic images, or a combination of both types of images in any possible proportion. For example, the first subset can include at least one defective training image synthetically generated by implanting a DOI into a difference image, where the difference image can be a real or synthetic defect-free image. Similarly, the second subset can include one or more real images and / or synthetic defect-free images. The present disclosure is not limited to the type or number of training images and / or the specific manner of acquiring / generating them.
[0084] According to some embodiments, each training image in the training set is associated with a ground truth (GT) frequency response of each training image in the Fourier domain. The Fourier transform (FT) generally refers to a transform used to convert a signal / function into a form that describes the frequencies present in the original signal / function (such as decomposing the signal into a sum of sine waves). The output of the transform is a complex function of frequency, including the real and imaginary parts of this function. Thus, the FT is sometimes referred to as the frequency-domain representation of the original function. When used in image processing, the Fourier transform can transform an image from the spatial domain into the frequency domain (also referred to herein as the Fourier domain) and provide information about the frequency content of the image. The transformed image (also referred to herein as a Fourier image or a transformed Fourier image) can be regarded as a complex image including a real part and an imaginary part, and the real and imaginary parts together represent the energy distribution of the image in the frequency domain.
[0085] The FT can be performed to transform the training image into the Fourier domain. As an example, if the transformed Fourier image is represented by the following complex function: F = Re + j*Im, the frequency response (also referred to as the power spectrum) of the transformed image can be represented by, for example, |F| or |F| which represents a quantitative measure of the magnitude of the Fourier image. 2 In the case where the training image includes a DOI, the frequency response of the corresponding Fourier image of the training image can be represented by the shape of various distributions depending on the type of the DOI.
[0086] For example, if the DOI in the spatial domain can be represented by a Gaussian distribution, the corresponding frequency response in the frequency domain of the training image can be represented as a Gaussian response, as illustrated by 614 in Figure 6 Thus, the training image is associated with the frequency response of the training image in the Fourier domain (as the ground truth (GT) (also referred to as the ground truth feature response)). In the case where the training image does not contain any DOI, the frequency response of the corresponding Fourier image of the training image can be represented as the power spectrum of random noise.
[0087] In some embodiments, regardless of whether the defective training images in the first subset are real or synthetic, the defective training images in the first subset consist of a single DOI. This may be necessary because an image with a single DOI typically has a specific feature response of the corresponding Fourier image in the Fourier domain. For example, the Gaussian response as illustrated by Figure 6 corresponds to a training image consisting of a single DOI. In the case where there are multiple DOIs in the training image, the spectrum of the feature response is not deterministic and thus may be difficult to use to identify the presence of the DOI.
[0088] It should be noted that there may be other types of feature responses depending on the specific type of the DOI associated with different types of DOIs, such as, for example, incremental response, cylindrical response, rectangular response, etc.
[0089] Once the training set is obtained, for each given training image in the training set, the ML system can process (204) the given training image to generate a Fourier image of the training image and the frequency response of the Fourier image. The ML system can include multiple ML models that are operatively connected to each other and are trained together to perform defect detection on a given input image based on the Fourier transform.
[0090] The ML system can be constructed in different ways. According to some embodiments, the ML system includes a first ML model that is operatively connected in parallel to a second ML model and a third ML model. For a given training image, the first ML model can be configured to generate one or more feature maps (e.g., tensors) representing the training image. To mimic the effect of the Fourier transform, the second ML model and the third ML model are used in parallel to separately process the one or more feature maps and generate the real part and the imaginary part of the Fourier image in the Fourier domain corresponding to the training image, respectively.
[0091] An exemplary implementation of the above ML system is illustrated in Figure 6 block 600 of Figure 6 FIG. 6 shows a schematic diagram of a training process of an ML system according to some embodiments of the subject matter of the present disclosure. As shown, a training image 602 is fed into the ML system 600. In this example, the training image 602 is a synthetic defective image generated by implanting a DOI at a specific location in a defect-free difference image. As shown, the training image 602 is quite noisy, having random noise and residual patterns. The ML system 600 includes a first ML model 604 that is operatively connected in parallel to a second ML model 606 and a third ML model 608. As an example, the first ML model 604 can be implemented as an autoencoder (also written as auto - encoder or AE).
[0092] An autoencoder is a neural network that is typically used for data reproduction by learning efficient data encoding and reconstructing the input of the data encoding (e.g., minimizing the difference between the input and the output). An autoencoder typically has an input layer, an output layer, and one or more hidden layers connecting them. Generally, an autoencoder can be regarded as including two components: an encoder and a decoder. The autoencoder learns to encode the data from the input layer into a short code (i.e., the encoder), and then decodes the code into an output that closely matches the original data (i.e., the decoder component). The output of the encoder is called the code, the latent variable, or the latent representation in the latent space representing the input image. The code can pass through the hidden layers in the decoder and can be reconstructed as an output image corresponding to the input image in the output layer.
[0093] When processing an input image through the layers in an autoencoder, certain layers of the encoder and decoder can provide layer outputs in the form of, for example, feature maps. For example, in the case of an autoencoder with a convolutional encoder and decoder, the output of each layer can be represented as a 2D output feature map. The size of the feature map gradually decreases in the encoder and gradually increases in the decoder. For example, an output feature map can be generated by convolving each filter of a particular layer across the width and height of the input feature map and by producing a two-dimensional activation map that gives the response of the filter at each spatial position.
[0094] Stacking the activation maps of all filters along the depth dimension forms the complete output feature map of a particular layer, and the complete output feature map can be represented as a 3D output feature map with multiple channels, where each channel corresponds to the activation map of a given filter. The 3D output feature map can also be regarded as multiple output feature maps corresponding to multiple channels / filters. The output feature maps of each layer in the encoder and decoder can be used to represent the features learned by the autoencoder.
[0095] In some embodiments, one or more output feature maps (also referred to herein as feature maps) from a given decoder layer of the first ML model 604 (e.g., an autoencoder) can be extracted and provided to the second ML model 606 and the third ML model 608 for processing. To achieve the effect of Fourier transform and obtain the output of the transformed Fourier image through machine learning, the second ML model and the third ML model are designed to be connected in parallel to the first ML model and are configured to process one or more feature maps separately. The two ML models can respectively generate the real part and the imaginary part of the Fourier image in the Fourier domain corresponding to the training images.
[0096] As an example, the second ML model and the third ML model can be implemented as CNNs. For example, the input of the 3D feature map can be processed through the deconvolution layer of the second ML model or the third ML model to generate a 2D feature map. The 2D feature map generated by the second ML model 606 can represent the real part Re of the Fourier image, and the 2D feature map generated by the third ML model 608 can represent the imaginary part Im of the Fourier image. The composite Fourier image 610 can be represented in the following form: F = Re + j*Im.
[0097] As described above, the frequency response 612 of the Fourier image 610 can be obtained, for example, by calculating F (such as |F| or |F| 2 ) based on the Fourier image 610 to obtain the absolute value 611, and the absolute value represents the magnitude of the Fourier image.
[0098] Continue Figure 2The process can perform an inverse Fourier transform (IFT) (206) on the Fourier image to obtain a reconstructed image. IFT refers to the inverse process of the Fourier transform, which converts a signal / function from the frequency domain to the original domain. It is used to recover the original signal from the spectrum of the Fourier image. In some embodiments, an analytical model can be used to perform IFT on the Fourier image to reconstruct the original image in the spatial domain. Although the analytical model is a mathematical model rather than a learning-based model, in some cases the analytical model can be considered to be included in the ML system. As Figure 6 shown, IFT 616 is applied to the Fourier image 610. Since IFT is a complex mathematical operation, a composite image in the spatial domain is generated, including a real part 618 and an imaginary part 620. The real part 618 represents the reconstructed image, while the imaginary part 620 represents the phase of the reconstructed image.
[0099] A loss function with two components can be used (e.g., by the training module 104) to optimize (208) the ML system: a first component based on the frequency response associated with a given training image and the GT frequency response, and a second component based on the reconstructed image and the given training image.
[0100] As Figure 6 illustrated, the frequency response 612 of the Fourier image 610 can be evaluated relative to the GT frequency response 614 based on the first component of the loss function (also referred to as the frequency loss, such as Figure 6 the frequency loss 622 shown). The ML system 600 can be optimized to reduce or minimize the frequency loss between the frequency response 612 and the GT frequency response 614, such that for a given training image, the ML system can learn to generate a GT response as close as possible to the GT response.
[0101] For example, for a training image with a DOI (original or implanted), the GT frequency response of the training image can be a Gaussian response as illustrated in 614. The output frequency response 612 generated by the ML system can be compared with the GT frequency response 614, and the difference between them represented by the frequency loss 622 can be minimized to optimize the parameters of the ML system. In the case where the training image has no DOI, the GT frequency response of the training image can be a random noise spectrum. As an example, when the training image is an image tile I j M*N, for example, the jth tile of size M*N in a relatively large difference image, the frequency loss can be expressed as follows (F represents the GT frequency response):
[0102]
[0103] Using the first loss component can present the ML system, especially Figure 6In the second and third ML models in the examples, to learn to generate a transformed Fourier image having a frequency response close to the GT frequency response (such as a Gaussian response). This in turn can guide the first ML model to learn to extract feature maps more relevant to the DOI while suppressing any noise / artifacts in the training images.
[0104] In particular, in a first ML model such as an autoencoder, the encoder of the network receives an input and produces a low-dimensional representation of the input. The intermediate layer between the encoder and the decoder (i.e., the latent space layer) is regarded as a bottleneck layer. This network architecture is designed to determine which aspects of the observed data (e.g., the feature maps representing the input) are relevant information to be passed and which aspects can be discarded. The autoencoder thus acts as a low-pass filter, which can learn to pass features relevant to the DOI while discarding / suppressing features related to noise / artifacts in the input image when trained in the ML system of the present disclosure (e.g., by imposing the GT frequency response as described above).
[0105] As described above, imposing the GT frequency response can only help detect the presence of the DOI in the input image. It cannot identify the exact location of the DOI within the input image because as long as there is one DOI in the image, the frequency response in the Fourier domain remains the same. Therefore, a second loss component is needed to identify the exact location of any DOI in the image.
[0106] As described above, the IFT 616 produces a composite image having a real part 618 and an imaginary part 620, where the real part 618 represents the reconstructed image. The reconstructed image is a reconstructed difference image corresponding to the input training image representing the detected defect information. In some cases, the IFT is configured to control the weight assigned to the imaginary part so as to implant a certain level of uncertainty about the presence of the DOI into the reconstructed image. As an example, the tunable controllable weight is used to "smear" / spread the energy in the reconstructed image, thereby allowing more than one DOI to be detected (in some cases, no DOI may be detected).
[0107] The reconstructed image can be evaluated based on a second component of the loss function (also referred to as the reconstruction loss, such as Figure 6 the reconstruction loss 624 shown) with respect to the input training image. The reconstruction loss represents the difference between the reconstructed image and the actual defect information in the training image. The reconstruction loss based on the difference between the reconstructed image and the defect information in the training image can be used to optimize the ML system.
[0108] For example, in the case where the training image includes a DOI, the defect information in the training image includes, for example, the location of the DOI in the training image. The location of the DOI can be regarded as the ground truth location of the DOI associated with the training image. As Figure 6Illustrated, the training image 602 is a synthetic defective image with an implanted DOI at a specific location in the image. The defect information 626 for the given training image includes the ground truth location 628 of the implanted DOI. In the case where the training image is a defect-free image without any DOI, the defect information 626 will be represented as an image with random noise in the training image. In either case, the ML system 600 can be optimized to reduce or minimize the reconstruction loss between the reconstructed image 618 of the training image and the defect information 626, such that for a given training image, the ML system can learn to generate a reconstructed image that is as close as possible to the GT defect information. This enables the ML system to identify the exact location of any DOI in the input image.
[0109] Using a loss function that includes the two components described above enables the ML system to learn not only to detect the presence of the DOI, but also to identify the exact location of the DOI in the input image. In particular, the reconstructed image represents a much clearer difference image compared to the input difference image, where residual noise and patterns are removed and there is a DOI, as Figure 6 illustrated in the reconstructed image 618 in
[0110] In some embodiments, optionally, the ML system can be used to further process the reconstructed image (210) to provide a classification score indicating the confidence level of the presence of a DOI in the reconstructed image. A classification loss function based on the training image and the classification score can be used to further optimize the ML system. As Figure 6 illustrated, in this case, the ML system 600 can further include a classifier 630. The reconstructed image 618 can be fed into the classifier 630 for processing, and the classifier can output a classification score 632 that indicates how confident the (multiple) DOI(s) presented in the reconstructed image are real defects. As an example, the classification score can be represented in the form of a probability score.
[0111] During training, the classifier as part of the ML system can be optimized using a classification loss function based on the defect information (as the ground truth) of the training image and the classification score. In some cases, an overall loss function that includes a frequency loss, a reconstruction loss, and a classification loss can be used to overall optimize the ML system that includes all the learning components (such as the first ML model, the second ML model, the third ML model, and the classifier) of the ML system.
[0112] According to certain other embodiments, the ML system can be constructed in an alternative manner. The ML system can include a first ML model operatively connected to a second ML model. For a given training image, the first ML model can be configured to generate one or more feature maps representing the training image. The second model can be configured to generate a reconstructed image based on the one or more feature maps. The ML system can further include an analysis model configured to perform a Fourier transform based on the reconstructed image.
[0113] In this alternative implementation, box 203 can be modified accordingly. Figure 2 of FIG. Figure 3 FIG. shows a generalized flowchart of the adaptive process of box 203 (i.e., box 203') in an alternative implementation according to certain embodiments of the subject matter of the present disclosure. Specifically, for each given training image in the training set, the ML system can process (304) the given training image to generate a reconstructed image of the given training image. Specifically, as described above, the first ML model in the ML system can be configured to generate one or more feature maps representing the training image. The second model can be configured to generate a reconstructed image based on the one or more feature maps. The reconstructed image can be Fourier-transformed (306) (e.g., by the analysis model) to obtain a Fourier image and the frequency response of the Fourier image. The processes of boxes 304 and 306 constitute box 203', which replaces Figure 2 box 203 in the flow of FIG. to be able to accommodate the alternative implementation.
[0114] Figure 7 FIG. shows an example of the ML system in this alternative implementation as described below. Specifically, Figure 7 FIG. shows a schematic diagram of the training process of an ML system including a first ML model and a second ML model according to certain embodiments of the subject matter of the present disclosure.
[0115] As shown, the training image 702 is fed into the ML system 700. The training image 702 is the same as the training image 602, which is a synthetic defective image generated by implanting a DOI at a specific location in a defect-free difference image. The ML system 700 includes a first ML model 704 operatively connected to a second ML model 706. Similar to the ML model 604, the first ML model 704 can be implemented as an autoencoder, as described above.
[0116] As an example, one or more feature maps from a given decoder layer of the first ML model 704 (e.g., an autoencoder) can be extracted and provided to the second ML model 706 for processing. In some cases, the second ML model can be implemented as a CNN. For example, a 3D feature map (or multi-channel feature map) output by the first ML model can be fed as input to the second ML model. The second ML model can process the feature map and generate a reconstructed image 718 corresponding to the training image 702. The reconstructed image is a reconstructed difference image representing the detected defect information.
[0117] Next, the reconstructed image 718 can be processed in two paths. As Figure 7 shown, the reconstructed image can be processed by a Fourier transform 710 (e.g., by an analysis model) to obtain a Fourier image and the frequency response 712 of the Fourier image. The frequency response 712 of the Fourier image can be evaluated based on a first component of the loss function (such as Figure 7 the frequency loss 722 shown) with respect to the GT frequency response 714. The ML system 700 can be optimized to reduce or minimize the frequency loss between the frequency response 712 and the GT frequency response 714, such that for a given training image, the ML system can learn to generate a reconstructed image whose Fourier transform image has a frequency response as close as possible to the GT response.
[0118] On the other hand, the reconstructed image 718 can be evaluated based on a second component of the loss function (such as Figure 7 the reconstruction loss 724 shown) with respect to the input training image 702. The reconstruction loss represents the difference between the reconstructed image and the actual defect information in the training image. As described above, in the case where the training image includes a DOI, the defect information in the training image includes, for example, the location of the DOI in the training image, which can be considered the ground truth location of the DOI associated with the training image.
[0119] As Figure 7 illustrated, the training image 702 is a synthetic defective image with an implanted DOI at a specific location in the image. The defect information 726 for a given training image 702 includes the ground truth location 728 of the implanted DOI. The ML system 700 can be optimized to reduce or minimize the reconstruction loss between the reconstructed image 718 of the training image 702 and the defect information 726, such that for a given input image, the ML system can learn to generate a reconstructed image as close as possible to the GT defect information. This enables the ML system to identify the exact location of any DOI in the input image.
[0120] Similarly, as described above with reference to the reference box 210, in an alternative implementation, an ML system can also be used to further process the reconstructed image to provide a classification score indicating the confidence level of the presence of a DOI in the reconstructed image. A classification loss function based on the training images and the classification score can be used to further optimize the ML system.
[0121] Turning now to Figure 4 , a generalized flowchart of performing in - line defect inspection of a semiconductor sample using a trained ML system in accordance with certain embodiments of the subject matter of the present disclosure is illustrated.
[0122] As described above, semiconductor samples are typically made of multiple layers. During the manufacturing process of the sample, for example, after a processing step of a particular layer, the inspection process of the sample can be performed multiple times. In some cases, a sampling set of processing steps can be selected for in - line inspection based on the known effects of multiple processing steps on device characteristics or yield. Images of the sample or a portion of the sample can be acquired at the sampling set of processing steps for inspection.
[0123] For illustrative purposes only, certain embodiments described below are described with respect to an image of a given processing step / layer relative to a sampling set of processing steps. Those skilled in the art will readily understand that the teachings of the subject matter of the present disclosure can be performed after any layer and / or processing step of the sample, such as a process of machine - learning - based inspection. The present disclosure should not be limited to the number of layers included in the sample and / or the specific layer(s) to be inspected.
[0124] As Figure 4 illustrated, an input image indicating the difference between a test image of the sample and a corresponding reference image can be obtained (402) (e.g., by the inspection tool 120) during in - line inspection of the sample. A semiconductor sample can refer herein to a semiconductor wafer, die, or a portion of a semiconductor sample, which is manufactured and inspected in a semiconductor foundry during the manufacturing process of the semiconductor sample. A test image of the sample can refer to an image that captures at least a portion of the sample to be inspected by an inspection tool. As an example, the test image can capture a target region of interest or a target structure (e.g., a structural feature or pattern on the semiconductor sample) for inspecting the semiconductor sample.
[0125] For each inspection image, one or more reference images can be used for defect detection. A reference image refers to a nominal / defect-free image that does not have defective features or has a high probability of not including any defective features, such that the reference image can be used as a reference for the corresponding inspection image for defect inspection. The reference images can be obtained in various ways, and the number of reference images used herein and the way of obtaining such images should not be construed as limiting the present disclosure in any way. In some cases, one or more reference images can be captured from one or more reference dies (e.g., adjacent dies to the inspection die) of the same sample or different samples. In some cases, reference images can be synthetically generated through image simulation. As an example, a simulated image can be generated based on design data (e.g., CAD data) of a die or a part of a die.
[0126] Various inspection tools can be used to acquire inspection images and / or reference images of inspection images. For example, the image can be an electron beam (e-beam) image acquired by an electron beam tool or can be an optical image acquired in real-time by an optical inspection tool during in-line inspection of a semiconductor sample.
[0127] The in-run input image as described in reference box 402 can be a difference image or a derivative of a difference image, which is generated by a comparison between the pixel values of the inspection image and the reference image of the inspection image. A difference image can be generated in a manner similar to the manner described above. For example, a difference image can be generated by subtracting the reference image from the inspection image. In some cases, a rank image as a derivative of the difference image can be generated by applying a predefined difference normalization factor to the difference image. The difference normalization factor can be determined based on the behavior of a normal population of pixel values and can be used to normalize the pixel values of the difference image. As an example, the rank of a pixel can be calculated as the ratio between the corresponding pixel value of the difference image and the predefined difference normalization factor. The difference image or the rank image (or any further derivatives thereof) can be used as the input image. Figure 2 The in-run input image can be processed by a trained ML system (e.g., by ML system 106) to generate (404) an output image representing a defect map that indicates the distribution of defect of interest (DOI) candidates in the input image. The ML system includes a plurality of ML models that are operatively connected to each other and are pre-trained together to perform defect detection on the input image based on Fourier transform. As described in the above references
[0128] and Figures 2 to 3 and Figures 6 to 7 the ML system is pre-trained during the training / setup phase.
[0129] Specifically, the ML system can be constructed in different ways. In some embodiments, as Figure 6Illustratively, the ML system may include a first ML model that is operatively connected in parallel to a second ML model and a third ML model. After being trained, the first ML model may be configured to generate one or more feature maps representing an input image. The second ML model and the third ML model may be configured to separately process the one or more feature maps and generate the real and imaginary parts of a Fourier image corresponding to the input image, respectively. The ML system may further include an analysis model configured to perform an inverse Fourier transform (IFT) on the Fourier image to reconstruct an output image.
[0130] The ML system is trained in such a way that the second ML model and the third ML model (via a frequency loss) learn to generate a transformed Fourier image having a frequency response close to the GT frequency response, which in turn results in a first ML model that acts like a low-pass filter to learn to extract feature maps that are more relevant to the DOI in a given input image while suppressing noise in the image. This enables the ML system to detect the presence of the DOI in a runtime input image once it is trained. Additionally, during training, the ML system also learns (via a reconstruction loss) to identify the location of the DOI in the image. This enables the ML system to identify the exact location of the DOI in a runtime input image.
[0131] In some other embodiments, as Figure 7 Illustratively, the ML system may include a first ML model that is operatively connected to a second model. After being trained, the first ML model may be configured to generate one or more feature maps representing an input image. The second model may be configured to process the one or more feature maps and reconstruct an output image.
[0132] The ML system is trained in such a way that the second ML model (via a reconstruction loss) learns to generate a reconstructed image presenting the exact location of the DOI, and when transformed into a Fourier image, the image (via a frequency loss) has a frequency response close to the GT frequency response.
[0133] The ML system constructed and trained as described above (in any embodiment) is capable of not only detecting the presence of the DOI but also identifying the exact location of the DOI. The output image reconstructed by the ML system has reduced / suppressed residual noise relative to the input image, which significantly improves the detection sensitivity when used for defect detection.
[0134] In some cases, more than one reference image may be acquired, and the Figure 4 described process may be repeated multiple times in order to obtain multiple output images (e.g., defect maps). Figure 5 Illustrates a generalized flowchart of an additional process for runtime defect inspection using a second input image in accordance with certain embodiments of the subject matter of the present disclosure.
[0135] A second input image indicating the difference between the inspection image and the second reference image can be obtained (502). This is the case where two reference images are obtained for the inspection image. Figure 4 The processing of can be understood as using the first difference image generated by comparing the inspection image with the first reference image as the input image. In addition, the second input image can be obtained in a manner similar to that described in reference box 402 as the second difference image generated based on the difference between the inspection image and the second reference image.
[0136] The trained ML system can be used to process (504) the second input image in a manner similar to that described in reference box 404 to obtain a second output image. The second output image represents a second defect map indicating the distribution of defects of interest (DOI) candidates in the input image. The presence of DOI candidates can be determined (506) based on the output image (i.e., the first output image) and the second output image. As an example, the two output images can be combined (e.g., by any kind of multiplication or averaging technique) into a combined output image, which can be used to determine the presence of DOI. In this way, the true DOI is selected as those candidates whose presence is indicated / determined by the two output images (e.g., where both images show a strong signal indication of the presence of DOI). Using multiple input images with a trained ML system can further eliminate residual variations and noise and enhance detection sensitivity.
[0137] Optionally, the ML system may further include a classifier. As Figure 6 illustrated, the classifier can be used to process the output image to provide a classification score indicating the confidence level of the presence of DOI in the output image. In this case, the classification score can be used together with the output image to determine the presence of DOI. As an example, only candidates with a classification score higher than a certain threshold will be determined as DOI. It should be noted that the classifier can also be included in Figure 7 the implementation manner of.
[0138] The ML models used in the above embodiments can be implemented with various types of learning models. As an example, the first ML model can be implemented as an autoencoder. The second ML model and / or the third ML model can be implemented as a convolutional neural network (CNN).
[0139] The output image generated by the ML system can be used for further defect inspection (e.g., by the defect inspection module 108). Such defect inspection can refer to one or more of the following operations: defect detection, defect review, and defect classification.
[0140] As an example, additional filtering techniques can be further used to select a list of DOI candidates from the output image. As another example, DOI candidates can be provided to a defect review tool (such as ADR). The review tool is configured to: capture a review image (usually with a higher resolution) at the location of the corresponding DOI candidate; and review the review image to determine whether the candidate is indeed a DOI. In some cases, as an addition or alternative to a defect review (DR) tool, a defect classification tool (such as ADC) is used. As an example, the classification tool can provide category data that provides information on whether each candidate is a DOI, and for candidates classified as DOIs, the classification tool also provides the category or type of the DOI.
[0141] Note that the examples shown in this disclosure (such as exemplary ML models and ML systems, loss functions, defect inspection applications, etc.) are illustrated for exemplary purposes and should not be considered as limiting this disclosure in any way. As an addition or alternative to the above, other suitable examples / implementations can be used.
[0142] One of the advantages of certain embodiments of the subject matter of the present disclosure as described herein is that an ML system capable of performing defect inspection operations based on Fourier transform is provided. The proposed ML system is a learning-based system specifically constructed based on the effects of Fourier transform using multiple ML models, particularly the frequency response representation of DOIs in the Fourier domain.
[0143] The ML system can not only detect the presence of DOIs in the input image but also identify the exact location of the DOIs via various implementations. The output image reconstructed by the ML system has reduced / suppressed residual noise and variations relative to the input image, which significantly improves the detection sensitivity when used for defect detection.
[0144] Other advantages of certain embodiments of the subject matter of the present disclosure as described herein are that the ML system can be used to process multiple input difference images generated by comparing a test image with multiple reference images. Multiple output images (e.g., multiple defect maps) can be combined and used to determine the presence of DOIs. In this way, the presence of true DOIs being selected as candidates is indicated by these candidates in multiple output images. Using multiple input images with a trained ML system can further eliminate residual variations and noise and enhance the detection sensitivity.
[0145] It should be understood that the present disclosure is not limited to the details set forth in the specification included herein or shown in the drawings in the application of the present disclosure.
[0146] In this detailed description, numerous specific details are set forth to provide a thorough understanding of the present disclosure. However, those skilled in the art will understand that the subject matter of the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the subject matter of the present disclosure.
[0147] Unless otherwise specifically stated, as will be apparent from this discussion, it should be understood that throughout the specification, discussions using terms such as "obtaining," "examining," "processing," "providing," "training," "using," "generating," "executing," "optimizing," "reconstructing," "detecting," "identifying," "determining," "comparing," "controlling," etc. refer to the (multiple) actions and / or (multiple) processes of a computer that manipulate data and / or transform data into other data, where the data is represented as physical quantities (such as electrical quantities) and / or the data represents physical objects. The term "computer" should be broadly interpreted to cover any type of hardware-based electronic device having data processing capabilities. By way of non-limiting example, computers include inspection systems, defect detection systems, and their corresponding parts disclosed in the present application.
[0148] As used herein, the terms "non-transitory memory" and "non-transitory storage medium" should be broadly interpreted to cover any volatile or non-volatile computer memory suitable for the subject matter of the present disclosure. The terms should be considered to include a single medium or multiple media that store one or more instruction sets (e.g., a centralized or distributed database, and / or associated caches and servers). The terms should also be understood to include any medium that is capable of storing an instruction set executable by a computer or encoding such an instruction set and causing the computer to perform any one or more of the methods of the present disclosure. Thus, the terms should be considered to include, but not be limited to, read-only memory ("ROM"), random access memory ("RAM"), disk storage media, optical storage media, flash memory devices, etc.
[0149] As used in this specification, the term "sample" should be broadly interpreted to cover any type of physical object or substrate used in the manufacture of semiconductor integrated circuits, magnetic heads, flat panel displays, and other semiconductor finished products, including wafers, masks, reticles, and other structures and their combinations and / or parts. Samples are also referred to herein as semiconductor samples and may be produced by manufacturing equipment that performs the corresponding manufacturing processes.
[0150] As used in this specification, the term "inspection" should be broadly interpreted to cover any type of operation involving defect detection, defect review, and / or various types of defect classification, segmentation, and / or metrology operations during and / or after a sample manufacturing process. Inspection is performed by using non-destructive inspection tools during or after manufacturing the sample to be inspected. As a non-limiting example, an inspection process may include runtime scans (either single scan or multiple scans), imaging, sampling, detection, review, measurement, classification, and / or other operations on the sample or a portion of the sample using the same or different inspection tools. Similarly, inspection may be performed before manufacturing the sample to be inspected, and the inspection may include, for example, generating (multiple) inspection recipes and / or other setup operations. Note that, unless otherwise specifically stated, the term "inspection" or derivatives of "inspection" used in this specification are not limited to the resolution or size of the inspection area. As a non-limiting example, various non-destructive inspection tools include scanning electron microscopes (SEM), atomic force microscopes (AFM), optical inspection tools, etc.
[0151] As used in this specification, the term "metrology operation" should be broadly interpreted to cover any metrology operation procedure for extracting metrology information related to one or more structural elements on a semiconductor sample. In some embodiments, a metrology operation may include measurement operations, such as critical dimension (CD) measurements performed on certain structural elements on the sample, including but not limited to the following: dimensions (e.g., line width, line pitch, contact diameter, size of an element, edge roughness, gray level statistics, etc.); shape of an element; distance within or between elements; associated angles; overlap information associated with elements corresponding to different design levels, etc. For example, measurement results such as measured images are analyzed by adopting image processing techniques. Note that, unless otherwise specifically stated, the term "metrology" or derivatives of "metrology" used in this specification are not limited to measurement techniques, measurement resolution, or the size of the inspection area.
[0152] As used in this specification, the term "defect" should be broadly interpreted to cover any type of abnormality or undesired feature / functionality formed on a sample. In some cases, a defect may be a defect of interest (DOI), where a DOI is a real defect that has some impact on the functionality of the manufactured device, and thus the customer is interested in detecting the defect. For example, any "killer" defect that may cause a yield loss may be represented as a DOI. In some other cases, a defect may be a negligible impairment (also referred to as a "false alarm" defect) because the defect has no impact on the functionality of the completed device and does not affect the yield.
[0153] As used in this specification, the term "defect candidate" should be broadly interpreted to cover suspected defect locations on a sample that has been detected as having a relatively high probability of being a defect of interest (DOI). Thus, when being inspected / tested, a defect candidate can actually be a DOI, or in some other cases, it can be a damage as described above, or it can be any noise caused by different variations (e.g., process variations, color variations, mechanical and electrical variations, etc.) during inspection.
[0154] As used in this specification, the term "design data" should be broadly interpreted to cover any data that indicates the hierarchical physical design (layout) of a sample. Design data can be provided by the corresponding designer, and / or can be derived from the physical design (e.g., through complex simulations, simple geometric and Boolean operations, etc.). Design data can be provided in different formats, for example, as non-limiting examples, the GDSII format, the OASIS format, etc. Design data can be presented in vector format, grayscale intensity image format, or other formats.
[0155] As used in this specification, the term "(a)n image" or "image data" should be broadly interpreted to cover any original image / frame of a sample captured by an inspection tool during a manufacturing process, derivatives of the captured image / frame obtained by various preprocessing stages, and / or computer-generated synthetic images (in some cases based on design data). Depending on the specific scanning method (e.g., one-dimensional scanning such as line scanning, two-dimensional scanning in the x and y directions, or point scanning at a specific point, etc.), image data can be represented in different formats, for example, as grayscale profiles, two-dimensional images, or discrete pixels, etc. It should be noted that in some cases, in addition to the image (e.g., the captured image, the processed image, etc.), the image data mentioned herein can also include digital data associated with the image (e.g., metadata, manual attributes, etc.). It should also be noted that the image or image data can include data related to the processing step / layer of interest or multiple processing steps / layers of the sample.
[0156] It should be understood that, unless otherwise specifically stated, certain features of the subject matter of the present disclosure described in the context of separate embodiments can also be provided in combination in a single embodiment. Conversely, the various features of the subject matter of the present disclosure described in the context of a single embodiment can also be provided separately or in any suitable sub-combination. In this detailed description, numerous specific details are set forth in order to provide a thorough understanding of the methods and apparatuses.
[0157] It should also be understood that the systems according to the present disclosure can be implemented, at least in part, on a suitably programmed computer. Similarly, the present disclosure contemplates a computer-readable computer program for performing the methods of the present disclosure. The present disclosure further contemplates a non-transitory computer-readable memory for performing the methods of the present disclosure, the non-transitory computer-readable memory tangibly embodying a program of instructions executable by a computer.
[0158] The present disclosure can have other embodiments and can be practiced and carried out in various ways. Accordingly, it should be understood that the language and terminology used herein are for the purpose of description and should not be regarded as limiting. Thus, those skilled in the art will understand that the concepts upon which this disclosure is based can readily be utilized as a basis for designing other structures, methods, and systems for several purposes that implement the subject matter of this disclosure.
[0159] Those skilled in the art will readily understand that various modifications and changes can be made to the embodiments of the present disclosure described above without departing from the scope of the present disclosure as defined by the appended claims.
Claims
1. A computerized system for performing runtime defect inspection on a semiconductor sample, the system comprising processing circuitry configured to: obtaining an input image indicating a difference between a test image of the sample and a corresponding reference image; and The input image is processed using a trained machine learning (ML) system to generate an output image representing a defect map indicating a distribution of defect of interest (DOI) candidates in the input image, wherein the ML system includes a plurality of ML models that are operatively connected to each other and are together pre-trained to perform defect detection on the input image based on Fourier transform, and wherein the output image can be used for further defect inspection.
2. A computerized system according to claim 1, wherein the ML system includes a first ML model, the first ML model is operatively connected in parallel to a second ML model and a third ML model, wherein the first ML model is configured to generate one or more feature maps representing the input image, and the second ML model and the third ML model are configured to separately process the one or more feature maps and respectively generate a real part and an imaginary part of a Fourier image corresponding to the input image.
3. The computerized system of claim 2, wherein the ML system further comprises an analysis model configured to perform an inverse Fourier transform (IFT) on the Fourier image to reconstruct the output image.
4. The computerized system of claim 2, wherein the first ML model is an autoencoder, and the second and third ML models are convolutional neural networks (CNNs).
5. The computerized system of claim 1 , wherein the ML system comprises a first ML model operatively connected to a second model, wherein the first ML model is configured to generate one or more feature maps representing the input image, and the second model is configured to process the one or more feature maps and reconstruct the output image.
6. The computerized system of claim 1 , wherein the ML system is trained to: detect the presence of a DOI in a training image based on a frequency response of a Fourier image corresponding to the training image and a ground truth frequency response of the DOI; and identify a location of the DOI in the training image based on the training image and a reconstructed image of the training image.
7. The computerized system of claim 1 , wherein the processing circuitry is further configured to: obtaining a second input image indicating a difference between the test image and a second reference image; processing the second input image using the trained ML system to obtain a second output image; and The presence of a DOI candidate is determined based on the output image and the second output image.
8. The computerized system of claim 1 , wherein the ML system includes a classifier, and the processing circuit system is further configured to: process the output image using the classifier to provide a classification score indicating a confidence level of the presence of a DOI in the output image; and determine the presence of a DOI based on the output image and the classification score.
9. The computerized system of claim 1 , wherein the input image is a difference image obtained by comparing the inspection image with the reference image, and wherein the output image suppresses residual noise relative to the input image, which improves detection sensitivity when used for further defect inspection.
10. A computerized method for training a machine learning (ML) system for defect inspection of semiconductor samples, the method comprising: Obtaining a training set, the training set comprising a first subset of training images and a second subset of training images, the first subset of training images each comprising a DOI, the second subset of training images being free of any DOI, each training image indicating a difference between a test image of a sample and a corresponding reference image, and being associated with a ground truth (GT) frequency response of each training image in a Fourier domain; For each given training image in the training set, processing the given training image using the ML system to generate a Fourier image of the given training image and a frequency response of the Fourier image; Performing inverse Fourier transform (IFT) processing on the Fourier image to obtain a reconstructed image; as well as The ML system is optimized using a loss function having two components: a first component based on the frequency response associated with the given training image and a GT frequency response, and a second component based on the reconstructed image and the given training image.
11. The computerized method of claim 10, wherein the first subset comprises at least one training image synthetically generated by implanting the DOI into a defect-free difference image.
12. The computerized method of claim 10, wherein the ML system comprises: a first ML model configured to generate, for each given training image, one or more feature maps representing the given training image; a second ML model and a third ML model, the second ML model and the third ML model being configured to individually process the one or more feature maps and respectively generate a real part and an imaginary part of a Fourier image corresponding to the given training image; as well as An analysis model is configured to perform the IFT based on the Fourier image to obtain the reconstructed image.
13. A computerized method according to claim 12, wherein the optimization using the first loss component causes the second ML model and the third ML model to learn to generate the Fourier image having a frequency response close to the GT frequency response, which in turn causes the first ML model to learn to extract the one or more feature maps that are more relevant to the presence of DOI while suppressing noise in the given training image.
14. The computerized method of claim 12, wherein the second loss component is a reconstruction loss representing the difference between the defect information of the reconstructed image and the given training image, and the optimization using the second loss component enables the ML system to identify locations where DOIs exist in the given training image.
15. The computerized method of claim 10, wherein the ground truth frequency response is one of: a Gaussian response, a delta response, a cylindrical response, and a rectangular response, depending on the type of DOI.
16. The computerized method of claim 12, wherein the analysis model is configured to perform the IFT to generate a composite image comprising a real part representing the reconstructed image, and an imaginary part, and wherein the analysis model is further configured to control a weight assigned to the imaginary part so as to introduce uncertainty in the presence of the DOI in the reconstructed image.
17. The computerized method of claim 10, further comprising: processing the reconstructed image using the ML system to provide a classification score indicating a confidence level of the presence of the DOI in the reconstructed image; and optimizing the ML system using a classification loss function based on the training images and the classification scores.
18. A computerized method for training a machine learning (ML) system for defect inspection of semiconductor samples, the method comprising: Obtaining a training set, the training set comprising a first subset of training images and a second subset of training images, the first subset of training images each comprising a DOI, the second subset of training images being free of any DOI, each training image indicating a difference between a test image of a sample and a corresponding reference image, and being associated with a ground truth (GT) frequency response of each training image in a Fourier domain; For each given training image in the training set, use the ML system to process the given training image to obtain a reconstructed image; processing the reconstructed image to generate a Fourier image of the given training image and a frequency response of the Fourier image; as well as The ML system is optimized using a loss function having two components: a first component based on the frequency response associated with the given training image and a GT frequency response, and a second component based on the reconstructed image and the given training image.
19. The computerized method of claim 18, wherein the ML system comprises: a first ML model configured to generate one or more feature maps representing the given training image; a second model configured to generate a reconstructed image based on the one or more feature maps; and an analysis model configured to perform a Fourier transform based on the reconstructed image.
20. A non-transitory computer-readable storage medium tangibly embodying a program of instructions that, when executed by a computer, cause the computer to perform a method of defect inspection for a semiconductor sample, the method comprising: obtaining an input image indicating a difference between a test image of the sample and a corresponding reference image; as well as The input image is processed using a trained machine learning (ML) system to generate an output image representing a defect map indicating a distribution of defect of interest (DOI) candidates in the input image, wherein the ML system includes a plurality of ML models that are operatively connected to each other and are together pre-trained to perform defect detection on the input image based on Fourier transform, and wherein the output image can be used for further defect inspection.