Training machine learning model to predict image representative of defect on substrate
By adding defects to the trained image and using neural network to train the model, the problem of low defect capture rate in substrate defect recognition is solved, and more efficient defect recognition and more accurate image quality improvement is achieved.
Patent Information
- Application Number
- CN202380074430.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-23
- Filing Date
- 2023-09-21
- Publication Date
- 2025-08-08
AI Technical Summary
When training machine learning models, it is difficult to effectively distinguish defects and interference on the substrate, resulting in low defect capture rate, and traditional methods do not fully utilize training data, resulting in inaccurate defect identification.
By adding defects to the trained image and training the model using neural networks, the loss function is calculated to modify the network configuration, improve the defect capture rate, enhance the defect signal, and select image pairs with high defect scores for training, and optimize the model with characterization feature extraction filters and classifier feature maps.
The capture rate of substrate defects is improved, ensuring that the model can accurately identify defects and reduce the recognition of false defects, and improving image quality and training efficiency.
Smart Images

Figure CN120457385A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. application 63 / 418,578, filed on October 23, 2022, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The disclosure herein relates to semiconductor manufacturing, and more particularly to inspecting semiconductor substrates. Background Art
[0004] A lithography apparatus is a machine that applies a desired pattern to a target portion of a substrate. Lithography apparatuses are used, for example, in the manufacture of integrated circuits (ICs). For example, the IC chip in a smartphone can be as small as a human thumb and may include over 2 billion transistors. Manufacturing ICs is a complex and time-consuming process, in which circuit elements are arranged in different layers and involve hundreds of individual steps. Even a mistake in one step can cause problems in the final IC and potentially lead to device failure. The presence of defects can compromise high process yields and high wafer production.
[0005] Metrology processes are used at various steps in the patterning process to monitor and / or control the process. For example, the metrology process is used to measure one or more characteristics of the substrate, such as the relative position (e.g., registration, overlay, alignment, etc.) or size (e.g., line width, critical dimension (CD), thickness, etc.) of features formed on the substrate during the patterning process or random variations, so that, for example, the performance of the patterning process can be determined based on the one or more characteristics. If one or more characteristics are unacceptable (e.g., outside a predetermined range of characteristic(s),) one or more variables of the patterning process can be designed or changed, for example, based on the measurement of the one or more characteristics, so that the substrate produced by the patterning process has acceptable characteristics(s).
[0006] Wafer inspection is a process used to find defects on a wafer. Wafer inspection tools can be used to perform wafer inspection. During the inspection process, the wafer inspection tool takes a picture of the die. The inspection tool then takes a picture of another die and compares them. If there is a change, it is usually a defect. The inspection tool can find defects and can also detect false defects, often called "noises." At more advanced nodes, noises and defects appear to be clustered together on the image, making it difficult to distinguish the difference between the two. Detecting noisy defects often requires high-quality or high-resolution images to find the defect of interest. However, capturing high-resolution images (hence the term "slow scan images") consumes a lot of time.
[0007] Low-quality or low-resolution images have much faster image capture times than high-quality images (hence the term "fast scan images"), but may not be helpful in identifying defects and interferences due to poor quality. Machine learning (ML) models offer a solution to improve image quality from low quality to high quality with an acceptable defect capture rate (defect to interference ratio). ML models may require low-resolution and high-resolution image pairs as training data to convert low-resolution images to high-resolution images. Summary of the Invention
[0008] In some aspects, the technology described herein relates to a non-transitory computer-readable medium having instructions that, when executed by a computer, cause the computer to perform a method for training a machine learning model to generate an image representing defects on a substrate. The method includes: inputting a first image and a reference image representing images captured using different image capture conditions into a neural network, the first image and the reference image indicating defects on a substrate patterned using a target layout; generating a predicted image using the neural network in response to the first image; calculating a loss function indicating a difference between a defect distribution in the predicted image and a defect distribution in the reference image; and modifying the neural network based on the loss function.
[0009] In some aspects, the technology described herein relates to a non-transitory computer-readable medium having instructions that, when executed by a computer, cause the computer to perform a method for training a machine learning model to generate an image representing a defect on a substrate. The method includes: inputting a first image and a reference image representing images captured using different image capture conditions into a neural network, the first image and the reference image representing a defect on a substrate patterned using a target layout; generating a predicted image using the neural network in response to the first image; calculating a loss function indicating a difference between a first set of feature vectors of the predicted image and a second set of feature vectors of the reference image; and modifying the neural network based on the loss function.
[0010] In some aspects, the technology described herein relates to a non-transitory computer-readable medium having instructions that, when executed by a computer, cause the computer to perform a method for training a machine learning model to generate an image indicating a defect on a substrate. The method includes acquiring a first image and a reference image captured using different image capture conditions, the first image and the reference image representing a defect on a substrate patterned using a target layout; adding the defect to the first image and the reference image to generate an updated first image and an updated reference image; and training a neural network using the updated first image and the updated reference image to convert the updated first image into a predicted image using the updated reference image, wherein the predicted image represents the defect on the substrate and corresponds to the image capture conditions of the reference image.
[0011] In some aspects, the technology described herein relates to a non-transitory computer-readable medium having instructions that, when executed by a computer, cause the computer to perform a method for training a machine learning model to generate an image indicative of a defect on a substrate. The method includes acquiring a first image and a reference image captured using different image capture conditions, the first image and the reference image representing a defect on a substrate patterned using a target layout; modifying a region of the reference image representing the defect to generate an updated reference image; and training a neural network to convert the first image into a predicted image using the updated reference image, wherein the predicted image represents the defect on the substrate and corresponds to the image capture conditions of the reference image.
[0012] In some aspects, the technology described herein relates to a non-transitory computer-readable medium having instructions that, when executed by a computer, cause the computer to perform a method for training a machine learning model to generate an image indicative of a defect on a substrate. The method includes acquiring a plurality of image pairs, wherein each image pair includes a first image and a reference image captured using different image capture conditions, the first image and the reference image representing a defect on a substrate patterned using a target layout; determining a defect detection probability for the reference image of the image pair; selecting a subset of the image pairs based on the defect detection probability; and training a neural network using the subset of image pairs to convert the first image of the image pair into a predicted image using the reference image of the image pair of the subset of image pairs, wherein the predicted image represents a defect on the substrate and corresponds to the image capture conditions of the reference image.
[0013] In some aspects, the technology described herein relates to a method for training a machine learning model to generate an image representing defects on a substrate. The method includes: inputting a first image and a reference image representing images captured using different image capture conditions into a neural network, the first image and the reference image indicating defects on a substrate patterned using a target layout; generating a predicted image using the neural network in response to the first image; calculating a loss function indicating a difference between a defect distribution in the predicted image and a defect distribution in the reference image; and modifying the neural network based on the loss function.
[0014] In some aspects, the technology described herein relates to a method for training a machine learning model to generate an image representing a defect on a substrate. The method includes: inputting a first image and a reference image representing images captured using different image capture conditions into a neural network, the first image and the reference image representing a defect on a substrate patterned using a target layout; generating a predicted image using the neural network in response to the first image; calculating a loss function indicating a difference between a first set of feature vectors of the predicted image and a second set of feature vectors of the reference image; and modifying the neural network based on the loss function.
[0015] In some aspects, the technology described herein relates to an apparatus for training a machine learning model to generate an image representing defects on a substrate. The apparatus includes: a memory storing an instruction set; and a processor configured to execute the instruction set so that the apparatus performs the following method: inputting a first image and a reference image representing images captured using different image capture conditions into a neural network, the first image and the reference image indicating defects on a substrate patterned using a target layout; generating a predicted image using the neural network in response to the first image; calculating a loss function indicating a difference between a defect distribution in the predicted image and a defect distribution in the reference image; and modifying the neural network based on the loss function.
[0016] In some aspects, the technology described herein relates to an apparatus for training a machine learning model to generate an image representing a defect on a substrate. The apparatus includes: a memory storing an instruction set; and a processor configured to execute the instruction set so that the apparatus performs the following method: inputting a first image and a reference image representing images captured using different image capture conditions into a neural network, the first image and the reference image representing a defect on a substrate patterned using a target layout; generating a predicted image using the neural network in response to the first image; calculating a loss function indicating a difference between a first set of feature vectors of the predicted image and a second set of feature vectors of the reference image; and modifying the neural network based on the loss function. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0018] Figure 1 is a schematic diagram illustrating an exemplary electron beam inspection (EBI) system according to an embodiment.
[0019] Figure 2 is a schematic diagram of an exemplary electron beam tool according to an embodiment.
[0020] Figure 3 Depicted is a schematic representation of overall lithography showing the collaboration between three technologies to optimize semiconductor manufacturing, according to an embodiment.
[0021] Figure 4 is a block diagram of a system for enhancing defects in an image for training a prediction model to convert a fast-scanning image to a slow-scanning image, consistent with various embodiments.
[0022] Figure 5 is a flow chart of a method for enhancing defects in an image for training a prediction model to convert a fast-scan image to a slow-scan image, consistent with various embodiments.
[0023] Figure 6 is a block diagram of a system for enhancing defects in an image for training a prediction model to convert a fast-scanning image to a slow-scanning image, consistent with various embodiments.
[0024] Figure 7 is a flow chart of a method for enhancing defects in an image for training a prediction model to convert a fast-scan image to a slow-scan image, consistent with various embodiments.
[0025] Figure 8 is a block diagram of a system for training a prediction model to convert a fast-scan image to a slow-scan image based on defect distribution in the image of a substrate, consistent with various embodiments.
[0026] Figure 9 is a flow chart of a method for training a prediction model to convert a fast-scan image to a slow-scan image based on defect distribution in the image of a substrate, consistent with various embodiments.
[0027] Figure 10 is a block diagram of a system for training a prediction model to convert a fast-scan image to a slow-scan image based on a classifier feature map associated with an image of a substrate, consistent with various embodiments.
[0028] Figure 11 is a flow chart of a method for training a prediction model to convert a fast-scan image to a slow-scan image based on a classifier feature map associated with an image of a substrate, consistent with various embodiments.
[0029] Figure 12 is a block diagram of a system for selecting images for training a prediction model to convert fast-scan images to slow-scan images, consistent with various embodiments.
[0030] Figure 13 is a flow chart of a method for selecting images for training a prediction model to convert fast-scan images to slow-scan images, consistent with various embodiments.
[0031] Figure 14 is a block diagram of an example computer system according to an embodiment.
[0032] The embodiments will now be described in detail with reference to the accompanying drawings, which are provided as illustrative examples to enable those skilled in the art to practice the embodiments. It is worth noting that the following figures and examples are not intended to limit the scope to a single embodiment, but rather that other embodiments are possible by exchanging some or all of the elements described or shown. Whenever convenient, the same reference numerals will be used throughout the figures to refer to the same or similar parts. Where certain elements of these embodiments can be partially or fully implemented using known components, only the portions of such known components necessary to understand the embodiment will be described, and detailed descriptions of the remaining portions of such known components will be omitted to avoid obscuring the description of the embodiment. In this specification, embodiments showing a single component should not be considered limiting; rather, unless otherwise expressly stated herein, the scope is intended to encompass other embodiments including multiple identical components, and vice versa. In addition, unless expressly provided, the applicant does not intend to assign unusual or special meanings to any term in the specification or claims. Furthermore, the scope encompasses current and future known equivalents of the components described herein. DETAILED DESCRIPTION
[0033] A lithographic apparatus is a machine that applies a desired pattern to a target portion of a substrate. The process of transferring the desired pattern to the substrate is called a patterning process. The patterning process may include a patterning step for transferring the pattern from a patterning device (such as a mask) to the substrate. Various changes (e.g., changes in the patterning process or the lithographic apparatus) may limit the implementation of lithography for high-volume semiconductor manufacturing (HVM). High-resolution images of the substrate (e.g., images with a resolution above a specified threshold), such as images acquired using a scanning electron microscope (SEM), can be inspected to identify any defects in the patterning process.
[0034] Conventional technologies employ various computational methods to acquire high-resolution (HR) images of defects on substrates. For example, a machine learning (ML) model is used to generate HR images of defects on substrates based on low-resolution (LR) images of the defects (e.g., images with a resolution below a specified threshold) (e.g., acquired using a SEM). For example, the low-resolution images are captured under high-speed beam scanning conditions, while the corresponding HR images are captured at a much slower speed. The ML model is trained using LR and HR image pairs of the defect region to predict the HR image of the defect region based on the LR images. However, the number of available image pairs of the defect region is typically limited and insufficient for training the ML model. In some cases, the image quality is poor and the defects are too weak to be useful in training the ML model. In some cases, conventional training methods do not consider the defect scores associated with the defects when training the ML model, and therefore, some defects are not captured or some false defects are captured, resulting in a reduced defect capture rate. In some cases, conventional training methods do not consider the characterizing features associated with the defects or interferences when training the ML model, and therefore, some defects are not captured or some interferences are captured as defects, resulting in a reduced defect capture rate. Characteristic features can be features of defects or interferences extracted using a characteristic feature extraction filter. For example, some characteristic feature extraction filters (such as low-pass filters or wavelet filters) capture low-frequency and high-frequency properties of defects or interferences. For example, in some cases, due to insufficient training data coverage, the training process may not pay enough attention to cases where defects and interferences are difficult to distinguish when training the ML model. For example, if image pairs are randomly selected, the selected image pairs may over-include interferences and defects that are easy to distinguish, resulting in the ML model performing poorly in distinguishing interferences from defects.
[0035] Embodiments are disclosed for training a prediction model (e.g., an ML model) to generate high-resolution images of defects on a substrate, thereby improving the capture rate of defects. In one embodiment, the number of training images with defects is increased by adding one or more defects to the training images (e.g., by a user) and training the prediction model with the updated images. For example, a fast scan image acquired using fast scan image capture conditions (e.g., an LR image of a defect area on a substrate) and a slow scan image acquired using slow scan image capture conditions (e.g., a corresponding HR image of the defect area) are modified by adding defects to the images, and the prediction model is trained with multiple such updated image pairs to convert the fast scan image into a slow scan image.
[0036] In one embodiment, the problem of weak defect signals can be overcome by enhancing the defect signals in the image. For example, the contrast of the defect area can be enhanced in the slow scan image, and a prediction model can be trained with multiple such updated image pairs to convert the fast scan image to the slow scan image.
[0037] In one embodiment, the defect capture rate (e.g., the ratio of the number of actual or "golden" defects captured to the total number of defects captured) can be improved by training a prediction model based on a defect distribution associated with a training image pair. For example, a defect score indicating the probability that a defect exists in a portion of the image is determined for each of the fast scan and slow scan images, and a loss function of the prediction model is customized to include the difference between the defect scores of the fast scan image and the slow scan image, and the prediction model can be trained based on the customized loss function to convert the fast scan image to the slow scan image.
[0038] In one embodiment, the defect capture rate can be improved by training a prediction model based on a classifier feature map associated with a training image pair. For example, a set of characterization feature vectors representing characteristic features of a defect or interference is extracted from a slow scan image and a fast scan image using a characterization feature extraction filter (e.g., a wavelet filter, a low-pass filter, etc.). The loss function of the prediction model is customized to include the difference between the classifier feature maps (e.g., images generated based on the characterization feature vectors) of the fast scan image and the slow scan image, and the prediction model can be trained based on the customized loss function to convert the fast scan image into a slow scan image.
[0039] In one embodiment, defect capture rates can also be improved by selecting training image pairs based on defect scores associated with actual defects and distractions. For example, defect candidates with defect scores below a first threshold score and defect candidates with defect scores above a second threshold score can be more easily classified as distractions and defects, respectively, compared to defect candidates within a "target" range between the first and second threshold scores. By selecting image pairs with defect scores within the target range and training the prediction model with these selected image pairs, the prediction model can be configured to convert fast-scanning images into slow-scanning images without missing any actual defects and ignoring distractions, thereby improving capture rates.
[0040] Now refer to Figure 1 , which illustrates an exemplary electron beam inspection (EBI) system 100 consistent with embodiments of the present disclosure. Figure 1 As shown, the EBI system 100 includes a main chamber 110, a load lock chamber 120, an electron beam tool 140, and an equipment front end module (EFEM) 130. The electron beam tool 140 is located within the main chamber 110. The exemplary EBI system 100 can be a single beam or multi-beam system. Although the description and drawings are directed to electron beams, it should be understood that the embodiments are not intended to limit the present disclosure to specific charged particles.
[0041] EFEM 130 includes a first load port 130a and a second load port 130b. EFEM 130 may include additional load port(s). First load port 130a and second load port 130b receive front-opening pods (FOUPs) containing wafers (e.g., semiconductor wafers or wafers made of other materials) or samples to be inspected (wafers and samples are collectively referred to as "wafers" hereinafter). One or more robotic arms (not shown) in EFEM 130 transport the wafers to load lock chamber 120.
[0042] The load lock chamber 120 is connected to a load / lock vacuum pump system (not shown), which removes gas molecules from the load lock chamber 120 to reach a first pressure below atmospheric pressure. After reaching the first pressure, one or more robotic arms (not shown) transport the wafer from the load lock chamber 120 to the main chamber 110. The main chamber 110 is connected to a main chamber vacuum pump system (not shown), which removes gas molecules from the main chamber 110 to reach a second pressure below the first pressure. After reaching the second pressure, the wafer is inspected by an electron beam tool 140. In one embodiment, the electron beam tool 140 may include a single beam inspection tool.
[0043] The controller 150 may be electrically connected to the electron beam tool 140 and may also be electrically connected to other components. The controller 150 may be a computer configured to perform various controls of the EBI system 100. The controller 150 may also include processing circuitry configured to perform various signal and image processing functions. Although the controller 150 is Figure 1 1. The controller 150 is shown as being located outside of the structure including the main chamber 110, the load lock chamber 120, and the EFEM 130, but it is understood that the controller 150 may be part of the structure.
[0044] Figure 2 A schematic diagram of an exemplary imaging system 200 is illustrated, in accordance with an embodiment of the present disclosure. Figure 2 The electron beam tool 140 can be configured for the EBI system 100. The electron beam tool 140 can be a single beam device or a multi-beam device. Figure 2As shown, the electron beam tool 140 includes a motorized sample stage 201 and a wafer holder 202, which is supported by the motorized sample stage 202 to hold a wafer 203 to be inspected. The electron beam tool 140 also includes an objective lens assembly 204, an electron detector 206 (which includes electron sensor surfaces 206a and 206b), an objective aperture 208, a converging lens 210, a beam limiting aperture 212, a gun aperture 214, an anode 216, and a cathode 218. In one embodiment, the objective lens assembly 204 can include a modified swinging objective delay immersion lens (SORIL) including a pole piece 204a, a control electrode 204b, a deflector 204c, and an excitation coil 204d. The electron beam tool 140 can additionally include an energy dispersive X-ray spectrometer (EDS) detector (not shown) to characterize the material on the wafer 203.
[0045] A primary electron beam 220 is emitted from cathode 218 by applying a voltage between anode 216 and cathode 218. Primary electron beam 220 passes through gun aperture 214 and beam-limiting aperture 212, both of which determine the size of the electron beam entering converging lens 210, which is located below beam-limiting aperture 212. Converging lens 210 focuses primary electron beam 220 before entering objective aperture 208, thereby setting its size before entering objective lens assembly 204. Deflector 204c deflects primary electron beam 220 to facilitate beam scanning across the wafer. For example, during scanning, deflector 204c can be controlled to sequentially deflect primary electron beam 220 to different locations on the top surface of wafer 203 at different times, thereby providing data for image reconstruction of different portions of wafer 203. Furthermore, deflector 204c can be controlled to deflect primary electron beam 220 to different sides of wafer 203 at specific locations and at different times, thereby providing data for stereoscopic image reconstruction of the wafer structure at that location. Furthermore, in one embodiment, the anode 216 and the cathode 218 may be configured to generate multiple primary electron beams 220, and the electron beam tool 140 may include multiple deflectors 204c to simultaneously project the multiple primary electron beams 220 onto different portions / sides of the wafer, thereby providing data for image reconstruction of different portions of the wafer 203.
[0046] The excitation coil 204d and the pole piece 204a generate a magnetic field that originates at one end of the pole piece 204a and terminates at the other end of the pole piece 204a. A portion of the wafer 203 scanned by the primary electron beam 220 can be immersed in the magnetic field and can become charged, which in turn generates an electric field. The electric field reduces the energy of the primary electron beam 220 that strikes the wafer 203 near its surface before colliding with the wafer 203. The control electrode 204b, electrically isolated from the pole piece 204a, controls the electric field on the wafer 203 to prevent micro-arching of the wafer 203 and ensure proper beam focusing.
[0047] Upon receiving the primary electron beam 220, a secondary electron beam 222 may be emitted from a portion of the wafer 203. The secondary electron beam 222 may form a beam spot on the sensor surfaces 206a and 206b of the electron detector 206. The electron detector 206 may generate a signal (e.g., voltage, current, etc.) representing the intensity of the beam spot and provide the signal to the image processing system 250. The intensity of the secondary electron beam 222 and the resulting beam spot may vary depending on the external or internal structure of the wafer 203. Furthermore, as described above, the primary electron beam 220 may be projected onto different locations on the top surface of the wafer or onto different sides of the wafer at a specific location to generate secondary electron beams 222 (and the resulting beam spots) of varying intensities. Thus, by mapping the intensity of the beam spot with the location of the wafer 203, the processing system may reconstruct an image reflecting the internal or surface structure of the wafer 203.
[0048] The imaging system 200 can be used to inspect a wafer 203 on a sample stage 201 and includes the electron beam tool 140, as described above. The imaging system 200 can also include an image processing system 250, which includes an image acquisition device 260, a storage device 270, and a controller 150. The image acquisition device 260 can include one or more processors. For example, the image acquisition device 260 can include a computer, a server, a mainframe, a terminal, a personal computer, any type of mobile computing device, or a combination thereof. The image acquisition device 260 can be connected to the detector 206 of the electron beam tool 140 via a medium such as an electrical conductor, a fiber optic cable, a portable storage medium, infrared (IR), Bluetooth, the internet, a wireless network, wireless radio, or a combination thereof. The image acquisition device 260 can receive signals from the detector 206 and construct an image. Thus, the image acquisition device 260 can acquire an image of the wafer 203. The image acquisition device 260 can also perform various post-processing functions, such as generating a profile, overlaying indicators on the acquired image, and the like. The image acquisition device 260 can also be configured to adjust the brightness and contrast of the acquired image. The storage device 270 may be a storage medium such as a hard disk, a cloud storage device, a random access memory (RAM), or other types of computer-readable memory. The storage device 270 may be coupled to the image acquirer 260 and may be used to store scanned raw image data as a raw image and to store post-processed images. The image acquirer 260 and the storage device 270 may be connected to the controller 150. In one embodiment, the image acquirer 260, the storage device 270, and the controller 150 may be integrated together as a control unit.
[0049] In one embodiment, the image acquirer 260 may acquire one or more images of the sample based on the imaging signal received from the detector 206. The imaging signal may correspond to a scanning operation for performing charged particle imaging. The acquired image may be a single image including multiple imaging regions. The single image may be stored in the storage device 270. The single image may be an original image that may be divided into multiple regions. Each region may include an imaging region that includes a feature of the wafer 203.
[0050] Figure 3 A schematic representation of the overall lithography is depicted, which shows the collaboration between the three technologies to optimize semiconductor manufacturing. Generally, the patterning process in the lithography apparatus LA is one of the most critical steps in the process, requiring high precision in determining the size and placement of structures on the substrate W ( Figure 1 ). To ensure this high level of accuracy, the three systems (in this example) can be combined in a so-called “holistic” control environment, such as Figure 3 Schematically shown. One of these systems is a lithography apparatus LA, which is (virtually) connected to a metrology apparatus (e.g., a metrology tool) MT (a second system) and a computer system CL (a third system). The "holistic" environment can be configured to optimize the collaboration between these three systems to enhance the overall process window and provide a tight control loop to ensure that the patterning performed by the lithography apparatus LA remains within the process window. The process window defines the range of process parameters (e.g., dose, focus, overlay) within which a particular manufacturing process will produce a defined result (e.g., a functional semiconductor device)—typically, the process parameters in the lithography process or patterning process are allowed to vary within this range.
[0051] The computer system CL can use (a portion of) the design layout to be patterned to predict which resolution enhancement technology to use and perform computational lithography simulations and calculations to determine which mask layout and lithography apparatus settings achieve the maximum overall process window (in Figure 2 Typically, the resolution enhancement technique is arranged to match the patterning possibilities of the lithographic apparatus LA. The computer system CL may also be used to detect the current operating position of the lithographic apparatus LA within the process window (e.g. using input from the metrology tool MT) to, for example, predict whether defects may exist due to suboptimal processes (e.g. Figure 2 (as indicated by the arrow pointing to “0” in the second scale SC2).
[0052] The metrology device (tool) MT can provide input to the computer system CL to enable accurate simulation and prediction, and can provide feedback to the lithographic apparatus LA to identify possible drift, for example in the calibration state of the lithographic apparatus LA (e.g. Figure 3(as indicated by multiple arrows in the third scale SC3).
[0053] The following paragraphs describe a system and method for training a predictive model (e.g., an ML model) to convert a low-resolution image of a defect on a substrate into a high-resolution image that can be used to improve the defect capture rate. Note that the predictive model discussed below can be implemented as an ML model (e.g., a neural network), a non-machine learning model, a physical model, a statistical model, an analytical model, a rule-based model, or any other empirical model. The training image pair input to the predictive model can include a first image of the substrate captured under a first image capture condition and a second image of the substrate captured under a second image capture condition. The second image can be used as a reference or true image when training the predictive model. In one embodiment, the first image is a low-resolution image of an area of the substrate captured using a fast scan mode of the inspection system (hence referred to as a "fast scan image"), and the second image / true / reference image is a corresponding high-resolution image of an area of the substrate captured using a slow scan mode of the inspection system (hence referred to as a "slow scan image"). Typically, an inspection system (e.g., Figure 1-Figure 3 The fast scan mode of an inspection system (e.g., an inspection system) captures images of a substrate faster than the slow scan mode, and the resolution of the fast scan image is generally lower than the resolution of the slow scan image. The training image pair, or at least one of the fast scan image or the slow scan image, may be obtained using a SEM or other imaging system (e.g., at least a reference image). Figure 1-Figure 3 The first image and the second image may be acquired using an inspection system described herein or other methods such as simulation. The following paragraphs use fast scan and slow scan images as examples of the first image and the second image, respectively, but the first image and the second image are not limited to fast scan images and slow scan images. The first image and the second image may also be acquired using other image capture conditions. For example, the first image may be a simulation image, and the second image is a high-resolution version of the simulation image. In one embodiment, the resolution of the low-resolution (LR) image is lower than a specified resolution threshold, and the resolution of the high-resolution image (HR) is higher than the specified resolution threshold.
[0054] Figure 4 is a block diagram of an exemplary system 400 for enhancing defects in an image for training a prediction model to convert a fast-scanning image to a slow-scanning image, consistent with various embodiments. Figure 5 is a flow chart of an exemplary method 500 for enhancing defects in an image for training a prediction model to convert a fast-scan image to a slow-scan image, consistent with various embodiments.
[0055] At process P505, an image pair 401 is acquired having a fast scan image 402 of an area of the substrate and a corresponding slow scan image 404 of the area of the substrate. The fast scan image 402 and the slow scan image 404 may or may not indicate any defects on the substrate. Figure 4 In the example of , image pair 401 does not indicate any defects on the substrate.
[0056] At process P510, one or more defects are added to image pair 401. In one embodiment, adding defects to an image includes editing a portion of the image to add a marker representing the defect, or matching a portion of any reference image of the substrate that indicates a defect on the substrate. For example, fast scan image 402 is edited to add defect 406, thereby generating updated fast scan image 403, and slow scan image 404 is edited to add defect 408, thereby generating updated slow scan image 405. Defects may be added to the image by a user or in other ways. In one embodiment, a statistical analysis may be performed on the defects in the actual SEM image of the substrate to determine various attributes, such as shape, size, intensity, signal value (e.g., pixel value at the location of the defect in the image), etc. The artificial defect may be added to image pair 401 such that one or more of the attributes of the artificial defect match attributes determined based on the statistical analysis. In one embodiment, the attributes of the artificial defect may be randomly selected from the attributes of the actual defect determined based on the statistical analysis.
[0057] In process P515, image generator 450 is trained using updated image pair 407 to generate a predicted slow scan image 415a from updated fast scan image 403 using updated slow scan image 405 as a true image or reference image. Image generator 450 can be implemented as a prediction model. Image generator 450 generates a predicted slow scan image 415a corresponding to updated slow scan image 405. Image generator 450 calculates an image reconstruction loss 420, which is determined as the difference between predicted slow scan image 415a and a reference image (such as updated slow scan image 405). Image reconstruction loss 420 can be calculated as the difference between the pixel value of each pixel of predicted slow scan image 415a and updated slow scan image 405. The configuration of image generator 450 can be updated to reduce image reconstruction loss 420. For example, updating the image generator 450 includes updating the configuration of the neural network (e.g., weights, biases, or other parameters) based on the image reconstruction loss 420. For example, the connection weights can be adjusted to reconcile the difference between the neural network's prediction (e.g., the predicted slow scan image 415a) and the reference feedback (e.g., the updated slow scan image 405). In another use case, one or more neurons (or nodes) of the neural network may need to have their corresponding errors sent back through the neural network to facilitate the updating process (e.g., backpropagation of the error). For example, the update to the connection weights can reflect the magnitude of the error (e.g., the loss function) propagated back after the forward pass is completed. In this way, for example, the image generator 450 can be trained to generate better predictions (e.g., SEM images of the substrate).
[0058] In one embodiment, training the image generator 450 is an iterative process, wherein each iteration includes generating a predicted image (e.g., predicted slow scan image 415a), calculating a loss function (e.g., image reconstruction loss 420), determining whether the loss function is minimized, and updating the configuration of the image generator 450 to reduce the loss function. Iterations may be performed until a specified condition is met (e.g., a predetermined number of iterations, until the loss function is minimized, or another condition). After training is complete, the image generator 450 is considered trained and can be used to predict a slow scan image for a fast scan image of a defective region of any given substrate.
[0059] Figure 6 is a block diagram of an exemplary system 600 for enhancing defects in an image for training a prediction model to convert a fast-scanning image to a slow-scanning image, consistent with various embodiments. Figure 7is a flow chart of an exemplary method 700 for enhancing defects in an image for training a prediction model to convert a fast-scan image to a slow-scan image, consistent with various embodiments.
[0060] At process P705, an image pair 601 is acquired that includes a fast scan image 602 of a region of a substrate and a corresponding slow scan image 604 of the region of the substrate. Fast scan image 602 and slow scan image 604 indicate defects 606 on the substrate. In one embodiment, even in slow scan image 604, the defect signal may be very weak and may not be useful in training a prediction model. A prediction model trained using such images may predict that the slow scan image may not indicate a defect at all or may indicate an inaccurate defect.
[0061] At process P710, the region of slow-scan image 604 indicating defect 406 is modified to enhance the defect signal. For example, the contrast of slow-scan image 604 is adjusted (e.g., enhanced) in the region of defect 606 to improve the defect signal, thereby generating an updated slow-scan image 605 indicating defect 608.
[0062] At process P715, image generator 650 is trained to generate a predicted slow scan image 615a from fast scan image 602 using updated slow scan image 605 as a true image or reference image. Image generator 650 can be implemented as a prediction model. Image generator 650 generates a predicted slow scan image 615a corresponding to updated slow scan image 605. Image generator 650 calculates an image reconstruction loss 620, which is determined as the difference between predicted slow scan image 615a and a reference image (such as updated slow scan image 605). Image reconstruction loss 620 can be calculated as the difference between the pixel value of each pixel of predicted slow scan image 615a and updated slow scan image 605. In one embodiment, the loss function can include any of the loss functions described below. The configuration of image generator 650 can be updated to reduce image reconstruction loss 620. For example, updating the image generator 650 includes updating the configuration of the neural network (e.g., weights, biases, or other parameters) based on the image reconstruction loss 620. For example, the connection weights can be adjusted to reconcile the difference between the neural network's prediction (e.g., the predicted slow scan image 615a) and the reference feedback (e.g., the updated slow scan image 605). In another use case, one or more neurons (or nodes) of the neural network may need to have their corresponding errors sent back through the neural network to facilitate the updating process (e.g., backpropagation of the error). For example, the update to the connection weights can reflect the magnitude of the error (e.g., the loss function) propagated back after the forward pass is completed. In this way, for example, the image generator 650 can be trained to generate better predictions (e.g., SEM images of the substrate).
[0063] In one embodiment, training the image generator 650 is an iterative process, wherein each iteration includes generating a predicted image (e.g., predicted slow scan image 615a), calculating a loss function (e.g., image reconstruction loss 620), determining whether the loss function is minimized, and updating the configuration of the image generator 650 to reduce the loss function. Iterations may be performed until a specified condition is met (e.g., a predetermined number of iterations, until the loss function is minimized, or another condition). After training is complete, the image generator 650 is considered trained and can be used to predict a slow scan image for a fast scan image of a defective region of any given substrate.
[0064] Figure 8 is a block diagram of an exemplary system 800 for training a predictive model to convert a fast-scan image to a slow-scan image based on defect distribution in the image of a substrate, consistent with various embodiments. Figure 9is a flow chart of an exemplary method 900 for training a prediction model to convert a fast-scan image to a slow-scan image based on defect distribution in the image of a substrate, consistent with various embodiments.
[0065] At process P905, an image pair 801 having a fast scan image 802 and a corresponding slow scan image 804 of an area of a substrate is input to an image generator 850. Fast scan image 802 may indicate a defect on the substrate as defect 806, and slow scan image 804 may indicate the same as defect 808. Image generator 850 may be implemented as a predictive model.
[0066] At process P910 , the image generator generates a predicted slow scan image 815 a from the fast scan image 802 using the slow scan image 804 as a true image or reference image.
[0067] At process P915, image generator 850 calculates a loss function that indicates the difference between the defect distribution in the predicted slow-scan image 815a and the defect distribution in a reference image (such as slow-scan image 804). In one embodiment, the defect distribution in the image is represented using a defect score map that includes a plurality of defect scores. Each defect score can indicate the probability that a defect exists in a portion of the image (such as a pixel of the image). Defect score component 1025 can be configured to calculate the defect scores in various ways. For example, defect score component 825 can be configured to compare the image of a first die with a reference image of another die (e.g., a die known to be defect-free) and, if a difference exists, deem the image to include a defect. Defect score component 825 can be configured to assign a score that indicates the magnitude of the difference. For example, defect score component 825 can compare each pixel of the image of the first die with the corresponding pixel at the same location in a reference image of another die (e.g., a second reference image of a second die) and, if a difference exists between the pixel values, deem the pixel at that location in the image to be defective. The defect score component 825 may also compare the image of the first die with reference images of other dies (e.g., a third image of the third die, a fourth image of the fourth die, and so on). The probability of a defect being present in all reference images may be low. Therefore, if there is a similar difference between a pixel of the first image and a corresponding pixel in any reference image, the first image may have a defect at that location of the pixel. The defect score component 825 may determine a defect score for the pixel based on the difference (e.g., by normalizing the differences across multiple comparisons). In one embodiment, the defect score component 825 may also consider differences associated with one or more neighboring pixels of the pixel (e.g., differences between the neighboring pixels of the pixel in the first image and corresponding pixels in the reference images that are located at the same location as the neighboring pixels) when determining the defect score for the pixel. For example, when determining the defect score for the pixel, the defect score component 825 may aggregate the differences associated with the neighboring pixels with the differences associated with the pixel. In one embodiment, portions of the image (e.g., pixels) with a defect score above a specified threshold may be considered indicative of a defect.
[0068] The predicted slow scan image 815a can be input to a defect score component 825, which generates a predicted defect score map 832 whose defect scores indicate the probability of a defect being present in the predicted slow scan image 815a. Similarly, the defect score component 825 can generate a reference defect score map 831 that indicates the probability of a defect being present in the slow scan image 804. The image generator 850 calculates a defect-based loss 830 as the difference between the defect scores between the two images. For example, the defect-based loss can be expressed as:
[0069] lossdefect =dsm_weightt*|dsm_pred-dsm_gt| x
[0070] Equation (1)
[0071] Where dsm_weight is the weight associated with the defect distribution, dsm_pred is the defect score associated with the predicted slow-scan image 815a, dsm_gt is the defect score associated with the slow-scan image 804, and "x" is the order or degree (e.g., 2).
[0072] In one embodiment, calculating the loss function may further include calculating an image reconstruction loss 820, which is determined as the difference between the predicted slow scan image 815a and the slow scan image 804. The image reconstruction loss 820 may be calculated as the difference between the pixel values of each pixel of the predicted slow scan image 815a and the slow scan image 804. For example, the image reconstruction loss 820 may be expressed as:
[0073] loss reconstruction =|img_pred-img_gt| x
[0074] Equation (2)
[0075] Where img_pred is the pixel value associated with the predicted slow-scan image 815a, img_gt is the pixel value associated with the slow-scan image 804, and "x" is the order or degree (e.g., 2).
[0076] The image generator 850 may calculate the loss function as a function of both the defect-based loss 830 and the image reconstruction loss 820, which may be expressed as:
[0077] Loss = loss reconstruction +loss defect
[0078] Equation (3)
[0079] At process P920, the image generator 850 may be modified based on a loss function (e.g., Equation (3)). For example, the configuration of the image generator 850 may be updated to reduce the loss function (e.g., Equation (3)). In one embodiment, updating the image generator 850 includes updating the configuration of the neural network (e.g., weights, biases, or other parameters) based on the loss function. For example, the connection weights may be adjusted to reconcile the difference between the prediction of the neural network (e.g., the predicted slow scan image 815a) and the reference feedback (e.g., the slow scan image 804). In another use case, one or more neurons (or nodes) of the neural network may need to have their corresponding errors sent back through the neural network to facilitate the update process (e.g., back propagation of the error). For example, the update to the connection weights may reflect the magnitude of the error (e.g., the loss function) propagated back after the forward pass is completed. In this manner, for example, the image generator 850 may be trained to generate better predictions (e.g., SEM images of the substrate).
[0080] In one embodiment, training the image generator 850 is an iterative process, wherein each iteration includes generating a predicted image (e.g., predicted slow scan image 815a), calculating a loss function (e.g., Equation (3)), determining whether the loss function is minimized, and updating the configuration of the image generator 850 to reduce the loss function. Iterations may be performed until a specified condition is met (e.g., a predetermined number of iterations, until the loss function is minimized, or another condition). After training is complete, the image generator 850 is considered trained and can be used to predict a slow scan image for a fast scan image of a defective region of any given substrate.
[0081] In one embodiment, by training the image generator 850 based on a defect distribution (e.g., a defect score map), the image generator 850 is trained to predict images having a defect score map similar to a true image, which minimizes errors such as predicting non-defective areas as defects or missing any defective areas, thereby improving the defect capture rate.
[0082] Figure 10 is a block diagram of an exemplary system 1000 for training a prediction model to convert a fast-scan image to a slow-scan image based on a classifier feature map associated with an image of a substrate, consistent with various embodiments. Figure 11 is a flow chart of an exemplary method 1100 for training a prediction model to convert a fast-scan image to a slow-scan image based on a classifier feature map associated with an image of a substrate, consistent with various embodiments.
[0083] At process P1105, an image pair 1001 having a fast scan image 1002 and a corresponding slow scan image 1004 of an area of a substrate is input to an image generator 1050. Fast scan image 1002 may indicate a defect on the substrate as defect 1006, and slow scan image 1004 may indicate the same as defect 1008. Image generator 1050 may be implemented as a predictive model.
[0084] In process P1110 , the image generator generates a predicted slow scan image 1015 a from the fast scan image 1002 using the slow scan image 1004 as a true image or reference image.
[0085] In process P1115, the image generator 1050 calculates a loss function that indicates the difference between a first set of characterizing feature vectors associated with the predicted slow scan image 1015a and a second set of characterizing feature vectors associated with a reference image (such as the slow scan image 1004). In one embodiment, the characterizing feature vector represents a characteristic of the image. For example, the characterizing feature vector can be used to represent the characteristics of defects and interference (e.g., false defects) in the image. The characterizing feature vector includes a set of numbers (e.g., pixel values) that represent the characteristics of a pixel, which can be generated using a feature extraction filter (e.g., a wavelet filter, a low-pass image filter, etc.). For example, when a low-pass image filter is applied to the image, a characterizing feature vector is generated that indicates the low-frequency characteristics of the pixel, while when a wavelet filter is applied to the image, a characterizing feature vector is generated that indicates the high-frequency characteristics of the pixel. A classifier feature map can be generated as the image based on the pixel values in the characterizing feature vector. Different classifier feature maps can be generated using different characterizing feature extraction filters, and each classifier feature map indicates a specific characteristic of the image.
[0086] The predicted slow scan image 1015a can be input to a characterization feature vector generation component 1035, which generates a predicted classifier feature map 1037 having a first set of characterization feature vectors associated with the predicted slow scan image 1015a. Similarly, the characterization feature vector generation component 1035 can generate a reference classifier feature map 1036 having a second set of characterization feature vectors associated with the slow scan image 1004. The characterization feature vector generation component 1035 can be configured to generate a plurality of classifier feature maps (CFMs) for each image by applying various feature extraction filters. The image generator 1050 calculates a CFM-based loss 1040 as the difference between the characterization feature vectors of the two images. For example, the CFM-based loss can be expressed as:
[0087]
[0088] Where CFM_weight is the weight associated with the CFM component of the loss function, w i is the weight of the i-th CFM, CFM_pred is the CFM associated with the predicted slow-scan image 1015a, CFM_gt is the CF associated with the slow-scan image 1004, and "x" is the order or degree (eg, 2).
[0089] The image generator 1050 can also calculate a defect-based loss 1030 associated with the image. For example, the defect-based loss 1030 can be calculated as the difference in defect scores between the two images using defect score maps 1031 and 1032, which are calculated using a defect score component 1025 (e.g., similar to Figure 8 The defect score component 825) generates, as at least with reference to the above Figure 8 and Figure 9 As stated.
[0090] In one embodiment, calculating the loss function may further include calculating an image reconstruction loss 1020, which is determined as the pixel-by-pixel difference between the predicted slow scan image 1015a and the slow scan image 1004 (e.g., at least the reference pixel). Figure 8 and Figure 9 described above).
[0091] The image generator 1050 may calculate a loss function based on one or more of the CFM-based loss 1040, the defect-based loss 1030, or the image reconstruction loss 1020, which may be expressed as:
[0092] Loss = loss reconstruction +loss defect +loss CFM
[0093] Equation (5)
[0094] At process P1120, the image generator 1050 may be modified based on a loss function (e.g., Equation (5)). For example, the configuration of the image generator 1050 may be updated to reduce the loss function. In one embodiment, updating the image generator 1050 includes updating the configuration of the neural network (e.g., weights, biases, or other parameters) based on the loss function. For example, the connection weights may be adjusted to reconcile the difference between the prediction of the neural network (e.g., the predicted slow scan image 1015a) and the reference feedback (e.g., the slow scan image 1004). In another use case, one or more neurons (or nodes) of the neural network may need to have their corresponding errors sent back to them through the neural network to facilitate the update process (e.g., back propagation of the error). For example, the update to the connection weights may reflect the magnitude of the error (e.g., the loss function) propagated back after the forward pass is completed. In this manner, for example, the image generator 1050 may be trained to generate better predictions (e.g., SEM images of the substrate).
[0095] In one embodiment, training the image generator 1050 is an iterative process, wherein each iteration includes generating a predicted image (e.g., predicted slow scan image 1015a), calculating a loss function (e.g., Equation (5)), determining whether the loss function is minimized, and updating the configuration of the image generator 1050 to reduce the loss function. Iterations may be performed until a specified condition is met (e.g., a predetermined number of iterations, until the loss function is minimized, or another condition). After training is complete, the image generator 1050 is considered trained and can be used to predict a slow scan image for a fast scan image of a defective region of any given substrate.
[0096] In one embodiment, by training the image generator 1050 based on the CFM, the image generator 1050 is trained to predict images with classifier feature maps similar to real images, which minimizes errors such as predicting defects as noise or vice versa, thereby improving the capture rate of defects.
[0097] Figure 12 is a block diagram of an exemplary system 1200 for selecting images for training a prediction model to convert fast-scan images to slow-scan images, consistent with various embodiments. Figure 13 is a flow chart of an exemplary method 1300 for selecting images for training a prediction model to convert fast-scan images to slow-scan images, consistent with various embodiments.
[0098] At process P1305, a set of image pairs 1205 is acquired, wherein an image pair in the set of image pairs includes a fast scan image 1202 and a corresponding slow scan image 1204 of an area of a substrate. The fast scan image 1202 and the slow scan image 1204 may or may not indicate any defects on the substrate. Figure 12 In the example, at least some of the image pairs 1205 indicate one or more defects on the substrate.
[0099] At process P1310, image pair 1205 is input to defect score component 1225 to generate a defect score map 1210 for the slow scan image in image pair 1205. For example, defect score component 1225 generates defect score map 1210a for slow scan image 1204 in the first image pair 1205. As described above, the defect score map includes a plurality of defect scores (e.g., one score for each pixel of the image), and the defect score indicates the probability that a defect exists in the corresponding pixel. Any portion of the image with a defect score above a specified threshold can be identified as a defect candidate. Candidate defects can be golden defects (e.g., actual defects) or noise (e.g., false defects).
[0100] At process P1315, defect score map 1210 can be input to image selector 1230, which is configured to identify defect score maps in defect score map 1210 whose defect scores are within a target range. As described above, defect candidates with defect scores below a first threshold score or above a second threshold score can be easily classified as interference or defects, respectively. However, for defect candidates with defect scores within the "target" range between the first threshold score and the second threshold score, the probability of defect detection (i.e., the probability of distinguishing between gold defects and interference defects) is very low. Therefore, image selector 1230 is configured to identify a subset of defect score maps 1210 whose defect scores are within the target range, such as defect score map 1207. For example, image selector 1230 can select the following defect score maps in defect score map 1210 in which defect candidates classified as gold defects or interference are associated with defect scores within the target range. Image selector 1230 also identifies the slow scan image associated with defect score map 1207.
[0101] At process P1320 , image selector 1230 may select a subset of image pairs 1205 , such as image pair 1215 , that have slow-scan images associated with defect score map 1207 (eg, selected at process P1315 ).
[0102] At process P1325, the selected image pair 1215 is input to image generator 1250 to train image generator 1250 to generate predicted slow scan images from fast scan images. Image generator 1250 can be implemented as a prediction model. For example, image generator 1250 can generate predicted slow scan image 1215a from fast scan image 1217 using corresponding slow scan image 1219 as a true image or reference image. Image generator 1250 generates predicted slow scan image 1215a corresponding to slow scan image 1219. Image generator 1250 calculates an image reconstruction loss 1220, which is determined as the difference between predicted slow scan image 1215a and a reference image, such as slow scan image 1219. Image reconstruction loss 1220 can be calculated as the difference between the pixel values of each pixel in predicted slow scan image 1215a and slow scan image 1219. The configuration of the image generator 1250 can be updated to reduce the image reconstruction loss 1220. For example, updating the image generator 1250 includes updating the configuration of the neural network (e.g., weights, biases, or other parameters) based on the image reconstruction loss 1220. For example, the connection weights can be adjusted to reconcile the difference between the prediction of the neural network (e.g., the predicted slow scan image 1215a) and the reference feedback (e.g., the slow scan image 1219). In another use case, one or more neurons (or nodes) of the neural network may need to have their corresponding errors sent back through the neural network to facilitate the update process (e.g., back propagation of the error). For example, the update to the connection weights can reflect the magnitude of the error (e.g., the loss function) propagated back after the forward pass is completed. In this way, for example, the image generator 1250 can be trained to generate better predictions (e.g., SEM images of the substrate).
[0103] In one embodiment, training the image generator 1250 is an iterative process, wherein each iteration includes generating a predicted image (e.g., predicted slow scan image 1215a), calculating a loss function (e.g., image reconstruction loss 1220), determining whether the loss function is minimized, and updating the configuration of the image generator 1250 to reduce the loss function. Iterations may be performed until a specified condition is met (e.g., a predetermined number of iterations, until the loss function is minimized, or another condition). After training is complete, the image generator 1250 is considered trained and can be used to predict a slow scan image for a fast scan image of a defective region of any given substrate.
[0104] In one embodiment, by selecting image pairs whose defect scores are within a target range and using the selected image pairs to train a prediction model, the accuracy of the prediction model in distinguishing defects from interference is achieved or improved (for example, including defect candidates with scores within the target range), thereby converting a fast-scanning image into a slow-scanning image with an improved defect capture rate.
[0105] Note that the image pair 801 for training the image generator 850 or the image pair 1001 for training the image generator 1050 may be obtained by at least one of the following: (a) adding defects to the fast scan image and the corresponding slow scan image, such as at least reference Figure 4 and Figure 5 (b) modifying a portion of the slow scan image, such as enhancing the contrast of the defective area, such as by at least referring to Figure 6 and Figure 7 or (c) selecting an image pair from a plurality of image pairs based on a defect detection probability, such as at least referring to Figure 12 and Figure 13 As stated.
[0106] In one embodiment, any of the trained image generators described above can be used to predict a slow scan image from a fast scan image indicating a defect on any given substrate. For example, a slow scan image of a defect region on a substrate or any other LR image of a defect region on a substrate can be input to a trained image generator. The image generator is executed to predict a slow scan image or HR image of the defect region on the substrate (e.g., an image having a resolution greater than the resolution of the input image).
[0107] The predicted slow-scan image can be used for various purposes. For example, after inspecting the predicted slow-scan image for defects, the patterning process or lithography apparatus (e.g., one or more parameters of the patterning process or lithography apparatus) can be optimized or adjusted to minimize defects when patterning a target layout on a substrate. The optimized patterning process is then executed to print a pattern corresponding to the target layout on the substrate.
[0108] Figure 14 1 is a block diagram illustrating a computer system 1400 that can help implement the various methods and systems disclosed herein. The computer system 1400 can be used to implement any entity, component, module, or service depicted in the examples of the accompanying figures (as well as any other entity, component, module, or service described in this specification). The computer system 1400 can be programmed to execute computer program instructions to perform the functions, methods, processes, or services described herein (e.g., of any entity, component, or module). The computer system 1400 can be programmed to execute computer program instructions through at least one of software, hardware, or firmware.
[0109] Computer system 1400 includes a bus 1402 or other communication mechanism for communicating information, and a processor 1404 (or multiple processors 1404 and 1405) coupled to bus 1402 for processing information. Computer system 1400 also includes a main memory 1406, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 1402 for storing information and instructions to be executed by processor 1404. Main memory 1406 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by processor 1404. Computer system 1400 also includes a read-only memory (ROM) 1408 or other static storage device coupled to bus 1402 for storing static information and instructions for processor 1404. A storage device 1410, such as a magnetic disk or optical disk, is provided and coupled to bus 1402 for storing information and instructions.
[0110] The computer system 1400 may be coupled via bus 1402 to a display 1412 (such as a cathode ray tube (CRT) or a flat-panel or touch-panel display) for displaying information to a computer user. An input device 1414 (including alphanumeric and other keys) is coupled to bus 1402 for communicating information and command selections to processor 1404. Another type of user input device is a cursor control 1416, such as a mouse, trackball, or cursor direction keys, which is used to communicate direction information and command selections to processor 1404 and to control cursor movement on display 1412. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), which allows the device to specify a position in a plane. A touch-panel (screen) display may also be used as an input device.
[0111] According to one embodiment, portions of one or more methods as described herein may be performed by computer system 1400 in response to processor 1404 executing one or more sequences of one or more instructions contained in main memory 1406. Such instructions may be read into main memory 1406 from another computer-readable medium, such as storage device 1410. Execution of the sequences of instructions contained in main memory 1406 causes processor 1404 to perform the processing steps described herein. One or more processors in a multi-processing arrangement may also be employed to execute the sequences of instructions contained in main memory 1406. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions. Thus, the description herein is not limited to any specific combination of hardware circuitry and software.
[0112] As used herein, the term "computer-readable medium" refers to any medium that participates in providing instructions to processor 1404 for execution. Such media can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 1410. Volatile media include dynamic memory, such as main memory 1406. Transmission media include coaxial cables, copper wire, and optical fibers, including the wires that make up bus 1402. Transmission media can also take the form of sound waves or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs, any other optical media, punch cards, paper tape, any other physical media with a pattern of holes, RAM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves described below, or any other medium that a computer can read.
[0113] Various forms of computer-readable media may be involved in carrying one or more sequences of instructions to processor 1404 for execution. For example, the instructions may initially be stored on a disk of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 1400 can receive the data on the telephone line and use an infrared transmitter to convert the data to an infrared signal. An infrared detector coupled to bus 1402 can receive the data carried in the infrared signal and place the data on bus 1402. Bus 1402 transfers the data to main memory 1406, from which processor 1404 retrieves and executes the instructions. The instructions received by main memory 1406 may optionally be stored on storage device 1410 before or after execution by processor 1404.
[0114] The computer system 1400 also preferably includes a communication interface 1418 coupled to the bus 1402. The communication interface 1418 provides a two-way data communication coupling with a network link 1420 connected to a local network 1422. For example, the communication interface 1418 can be an integrated services digital network (ISDN) card or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, the communication interface 1418 can be a local area network (LAN) card to provide a data communication connection to a compatible LAN. A wireless link can also be implemented. In any such implementation, the communication interface 1418 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.
[0115] Network link 1420 typically provides data communication through one or more networks to other data devices. For example, network link 1420 can provide a connection to a host computer 1424 or to data equipment operated by an Internet Service Provider (ISP) 1426 through a local network 1422. ISP 1426, in turn, provides data communication services through a global packet data communication network, now commonly referred to as the "Internet" 1428. Both local network 1422 and Internet 1428 use electrical, electromagnetic, or optical signals that carry digital data streams. The signals through the various networks, as well as the signals on network link 1420 and through communication interface 1418, are exemplary forms of carrier waves transporting the information, and these signals carry the digital data to and from computer system 1400.
[0116] Computer system 1400 can send messages and receive data, including program code, through the network(s), network link 1420, and communication interface 1418. In the Internet example, server 1430 can send the requested code for an application through Internet 1428, ISP 1426, local area network 1422, and communication interface 1418. For example, one such downloaded application can provide illumination optimization according to an embodiment. The received code can be executed by processor 1404 upon receipt or stored in storage device 1410 or other non-volatile memory for later execution. In this manner, computer system 1400 can obtain application code in the form of a carrier wave.
[0117] While the concepts disclosed herein may be used for imaging on substrates such as silicon wafers, it should be understood that the disclosed concepts may be used with any type of lithographic imaging system, such as systems for imaging on substrates other than silicon wafers.
[0118] As used herein, the terms "perform optimization" and "optimize" refer to or indicate adjusting a patterning device (e.g., a lithographic device), a patterning process, or the like so that the result and / or process has more desirable characteristics, such as greater accuracy in projection of the design pattern on the substrate, a larger process window, or the like. Thus, as used herein, the terms "perform optimization" and "optimize" refer to or indicate a process that identifies one or more values of one or more parameters that provide an improvement, such as a local optimum, of at least one relevant metric compared to an initial set of one or more values of the one or more parameters. "Optimal value" and other related terms should be interpreted accordingly. In one embodiment, the optimization steps may be applied iteratively to provide further improvements in one or more metrics.
[0119] Aspects of the present invention may be implemented in any convenient form. For example, an embodiment may be implemented by one or more appropriate computer programs, which may be carried on an appropriate carrier medium, which may be a tangible carrier medium (e.g., a disk) or an intangible carrier medium (e.g., a communication signal). Embodiments of the present invention may be implemented using a suitable device, which may specifically take the form of a programmable computer running a computer program, which is arranged to implement the methods described herein. Therefore, embodiments of the present disclosure may be implemented in hardware, firmware, software, or any combination thereof. Embodiments of the present disclosure may also be implemented as instructions stored on a machine-readable medium, which may be read and executed by one or more processors. A machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium may include a read-only memory (ROM); a random access memory (RAM); a disk storage medium; an optical storage medium; a flash memory device; an electrical, optical, acoustic, or other form of propagation signal (e.g., a carrier wave, an infrared signal, a digital signal, etc.), etc. In addition, firmware, software, routines, instructions may be described herein as performing certain actions. However, it should be understood that such descriptions are for convenience only and that such actions are actually caused by computing devices, processors, controllers or other devices executing firmware, software, routines, instructions or the like.
[0120] In the block diagrams, the components shown are depicted as discrete functional blocks, but embodiments are not limited to systems in which the functionality described herein is organized as shown. The functionality provided by each component may be provided by software or hardware modules organized differently than presently depicted, for example, such software or hardware may be mixed, connected, replicated, decomposed, distributed (e.g., within a data center or geographically), or otherwise organized differently. The functionality described herein may be provided by one or more processors of one or more computers executing code stored on a tangible, non-transitory, machine-readable medium. In some cases, a third-party content delivery network may host some or all of the information delivered over the network, in which case, to the extent that information (e.g., content) is supplied or otherwise provided, that information may be provided by sending instructions to retrieve that information from the content delivery network.
[0121] Unless expressly stated otherwise, it will be apparent from the discussion that discussions in this specification using terms such as "process," "calculate," "compute," "determine," etc., refer to actions or processes of a specific device, such as a special purpose computer or similar special purpose electronic processing / computing device.
[0122] The reader should understand that this application describes several inventions. Rather than being separated into multiple independent patent applications, these inventions have been grouped into one document because their related subject matter helps save costs during the application process. However, the unique advantages and aspects of these inventions should not be lumped together. In some cases, embodiments address all of the deficiencies described herein, but it should be understood that the invention is independently useful and that some embodiments address only a subset of these problems or provide other unmentioned benefits that would be clear to a person skilled in the art reading this disclosure. Due to cost limitations, some of the inventions disclosed herein are not currently claimed and may be claimed in a later application, such as a continuation application or by amending an existing claim. Likewise, due to length limitations, neither the Abstract nor the Summary of the Invention sections of this document should be considered to be a comprehensive listing of all such inventions or all aspects of such inventions.
[0123] It should be understood that the description and drawings are not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present invention as defined by the appended claims.
[0124] In view of this description, modifications and alternative embodiments of various aspects of the present invention will be clear to those skilled in the art. Therefore, this description and the drawings are to be interpreted as illustrative only and are intended to teach those skilled in the art the general manner of implementing the invention. It should be understood that the forms of the invention shown and described herein are to be regarded as examples of embodiments. Elements and materials may be substituted for those shown and described herein, parts and processes may be reversed or omitted, certain features may be used independently, and embodiments or features of embodiments may be combined, all of which will be clear to those skilled in the art. Changes may be made to the elements described herein without departing from the spirit and scope of the invention as described in the following claims. The headings used herein are for organizational purposes only and are not meant to limit the scope of the description.
[0125] As used herein, unless otherwise specifically stated, the term "or" includes all possible combinations unless impractical. For example, if a component is specified to include either A or B, then, unless otherwise specifically stated or impractical, the component may include A, or B, or A and B. As a second example, if a component is specified to include either A, B, or C, then, unless otherwise specifically stated or impractical, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C. Expressions such as "at least one of..." do not necessarily modify the entirety of the following list, nor do they necessarily modify each member of the list, so "at least one of A, B, and C" should be understood to include only one of A, only one of B, only one of C, or any combination of A, B, and C. The phrases "one of A and B" or "any of A and B" should be interpreted in the broadest sense to include either one of A or one of B.
[0126] Examples are provided under the following terms:
[0127] 1. A non-transitory computer-readable medium having instructions, which, when executed by a computer system, cause the computer system to at least perform a method for training a machine learning model to generate an image representing defects on a substrate, the method comprising:
[0128] inputting a first image and a reference image into a neural network, the first image and the reference image representing images captured using different image capture conditions, the first image and the reference image indicating defects on a substrate patterned using a target layout;
[0129] generating a predicted image in response to the first image using the neural network;
[0130] calculating a loss function indicating a difference between a defect distribution in the predicted image and a defect distribution in the reference image; and
[0131] The neural network is modified based on the loss function.
[0132] 2. The computer-readable medium of clause 1, wherein calculating the loss function comprises:
[0133] determining the defect distribution in the predicted image as a predicted defect score map, in which a defect score indicates a probability of a defect being present in a portion of the predicted image;
[0134] determining the defect distribution in the reference image as a reference defect score map, in which a defect score indicates a probability of a defect being present in a portion of the reference image; and
[0135] A difference between the predicted defect score map and the reference defect score map is calculated.
[0136] 3. The computer-readable medium of clause 2, wherein the defect score satisfying a threshold score indicates that a defect is at a location on the substrate corresponding to the portion of the reference image.
[0137] 4. The computer-readable medium of any one of clauses 1 to 3, wherein calculating the loss function further comprises calculating a difference between a first set of eigenvectors of the predicted image and a second set of eigenvectors of the reference image.
[0138] 5. The computer-readable medium of clause 4, wherein calculating the difference comprises:
[0139] applying a feature extraction filter to the reference image to obtain the first set of feature vectors as a reference classifier feature map, the reference classifier feature map representing characteristics of defects or interference in the reference image;
[0140] applying the feature extraction filter to the predicted image to obtain the second set of feature vectors as a predicted classifier feature map, the predicted classifier feature map representing features of defects or interference in the predicted image; and
[0141] The difference between the reference classifier feature map and the predicted classifier feature map is calculated.
[0142] 6. The computer-readable medium of clause 1, wherein calculating the loss function further comprises calculating a pixel-by-pixel difference between the predicted image and the reference image.
[0143] 7. The computer-readable medium of clause 1 , wherein modifying the neural network based on the loss function comprises modifying parameters of the neural network until the loss function is minimized.
[0144] 8. The computer-readable medium of clause 1, wherein the first image corresponds to a fast-scan image capture condition, and the predicted image and the reference image correspond to a slow-scan image capture condition.
[0145] 9. The computer-readable medium of clause 8, wherein the first image has a lower resolution than the reference image.
[0146] 10. The computer-readable medium of clause 1, wherein inputting the first image and the reference image comprises adding defects to the first image and the reference image.
[0147] 11. The computer-readable medium of clause 10, wherein adding a defect to the first image and the reference image comprises editing a portion of the first image and the reference image to match a portion of a designated image that indicates a defect on the substrate.
[0148] 12. A computer-readable medium according to claim 1, wherein inputting the first image and the reference image includes: selecting a first image pair among the multiple image pairs based on a defect detection probability of a reference image among the multiple image pairs, wherein the first image pair includes the first image and the reference image.
[0149] 13. A computer-readable medium according to clause 12, wherein selecting the first image pairs comprises selecting some of the image pairs in which the reference image is associated with a defect score map having defect scores for defects and interferences within a first range.
[0150] 14. The computer-readable medium of clause 1, further comprising:
[0151] inputting a specified image of a specified substrate captured under rapid scan image capture conditions into the neural network; and
[0152] The neural network is executed to generate a specified predicted image based on the specified image, the specified predicted image representing a defect on the specified substrate and corresponding to a slow scan image capture condition.
[0153] 15. The computer-readable medium of clause 14, further comprising adjusting parameters of at least one of a patterning process or a lithographic apparatus based on the specified predicted image to minimize defects in patterning a target layout on the specified substrate.
[0154] 16. The computer-readable medium of clause 15, further comprising: performing the patterning process via the photolithography apparatus to print a pattern corresponding to the target layout on the substrate.
[0155] 17. A non-transitory computer-readable medium having instructions thereon, which, when executed by a computer, cause the computer to perform a method for training a machine learning model to generate an image representative of defects on a substrate, the method comprising:
[0156] inputting a first image and a reference image into a neural network, the first image and the reference image representing images captured using different image capture conditions, the first image and the reference image representing defects on a substrate patterned using a target layout;
[0157] generating a predicted image in response to the first image using the neural network;
[0158] calculating a loss function indicating a difference between a first set of feature vectors of the predicted image and a second set of feature vectors of the reference image; and
[0159] The neural network is modified based on the loss function.
[0160] 18. The computer-readable medium of clause 17, wherein calculating the loss function comprises:
[0161] applying a feature extraction filter to the reference image to obtain the first set of feature vectors as a reference classifier feature map, the reference classifier feature map representing characteristics of defects or interference in the reference image;
[0162] applying a feature extraction filter to the predicted image to obtain the second set of feature vectors as a predicted classifier feature map, the predicted classifier feature map representing features of defects or interference in the predicted image; and
[0163] The difference between the predicted classifier feature map and the reference classifier feature map is calculated.
[0164] 19. The computer-readable medium of clause 17, wherein calculating the loss function further comprises calculating a difference between a defect distribution in the predicted image and a defect distribution in the reference image.
[0165] 20. The computer-readable medium of clause 19, wherein calculating the difference comprises:
[0166] determining the defect distribution in the predicted image as a predicted defect score map, in which a defect score indicates a probability of a defect being present in a portion of the predicted image;
[0167] determining the defect distribution in the reference image as a reference defect score map, in which a defect score indicates a probability of a defect being present in a portion of the reference image; and
[0168] A difference between the predicted defect score map and the reference defect score map is calculated.
[0169] 21. The computer-readable medium of clause 20, wherein the defect score satisfying a threshold score indicates that a defect is at a location on the substrate corresponding to the portion of the reference image.
[0170] 22. The computer-readable medium of clause 17, wherein calculating the loss function further comprises calculating a pixel-by-pixel difference between the predicted image and the reference image.
[0171] 23. The computer-readable medium of clause 17, wherein modifying the neural network based on the loss function comprises modifying parameters of the neural network until the loss function is minimized.
[0172] 24. The computer-readable medium of clause 17, wherein the first image corresponds to a first image capture condition, and the reference image and the predicted image correspond to a second image capture condition.
[0173] 25. The computer-readable medium of clause 24, wherein the first image has a lower resolution than the reference image.
[0174] 26. The computer-readable medium of clause 17, wherein inputting the first image and the reference image comprises adding defects to the first image and the reference image.
[0175] 27. The computer-readable medium of clause 26, wherein adding a defect to the first image and the reference image comprises editing a portion of the first image and the reference image to match a portion of a designated image that indicates a defect on the substrate.
[0176] 28. A computer-readable medium according to clause 17, wherein inputting the first image and the reference image includes: selecting a first image pair among the multiple image pairs based on a defect detection probability of a reference image among the multiple image pairs, wherein the first image pair includes the first image and the reference image.
[0177] 29. The computer-readable medium of clause 28, wherein selecting the first image pair comprises:
[0178] Some of the image pairs are selected in which a reference image is associated with a defect score map having defect scores for defects and disturbances within a first range.
[0179] 30. The computer-readable medium of clause 17, further comprising:
[0180] inputting a specified image of a specified substrate captured under rapid scan image capture conditions into the neural network; and
[0181] The neural network is executed to generate a specified predicted image based on the specified image, the specified predicted image representing a defect on the specified substrate and corresponding to a slow scan image capture condition.
[0182] 31. A non-transitory computer-readable medium having instructions that, when executed by a computer, cause the computer to perform a method for training a machine learning model to generate an image indicative of defects on a substrate, the method comprising:
[0183] acquiring a first image and a reference image captured using different image capture conditions, the first image and the reference image representing defects on a substrate patterned using a target layout;
[0184] adding defects to the first image and the reference image to generate an updated first image and an updated reference image; and
[0185] A neural network is trained using the updated first image and the updated reference image to convert the updated first image into a predicted image using the updated reference image, wherein the predicted image represents defects on the substrate and corresponds to image capture conditions of the reference image.
[0186] 32. The computer-readable medium of clause 31, wherein adding a defect to an image comprises editing a portion of the image to match a portion of the reference image or the first image that indicates the defect on the substrate.
[0187] 33. The computer-readable medium of clause 31 , wherein training the neural network is an iterative process, and each iteration comprises:
[0188] determining a difference between the predicted image and the updated reference image;
[0189] determining whether the difference has decreased; and
[0190] In response to a determination that the difference has not decreased, parameters of the neural network are modified and iterations are repeated.
[0191] 34. The computer-readable medium of clause 31 , wherein the first image corresponds to a first image capture condition and the reference image corresponds to a second image capture condition.
[0192] 35. The computer-readable medium of clause 34, wherein the first image has a lower resolution than the reference image.
[0193] 36. The computer-readable medium of clause 31, further comprising:
[0194] inputting a specified image of a specified substrate captured under rapid scan image capture conditions into the neural network; and
[0195] The neural network is executed to generate a specified predicted image based on the specified image, the specified predicted image representing a defect on the specified substrate and corresponding to a slow scan image capture condition.
[0196] 37. A non-transitory computer-readable medium having instructions that, when executed by a computer, cause the computer to perform a method for training a machine learning model to generate an image indicative of defects on a substrate, the method comprising:
[0197] acquiring a first image and a reference image captured using different image capture conditions, the first image and the reference image representing defects on a substrate patterned using a target layout;
[0198] modifying the region of the reference image representing the defect to generate an updated reference image; and
[0199] A neural network is trained to convert the first image into a predicted image using the updated reference image, wherein the predicted image represents defects on the substrate and corresponds to image capture conditions of the reference image.
[0200] 38. The computer-readable medium of clause 37, wherein modifying the region of the reference image comprises enhancing contrast of the region of the reference image.
[0201] 39. The computer-readable medium of clause 37, wherein training the neural network is an iterative process, and each iteration comprises:
[0202] determining a difference between the predicted image and the updated reference image;
[0203] determining whether the difference has decreased; and
[0204] In response to a determination that the difference has not decreased, parameters of the neural network are modified and iterations are repeated.
[0205] 40. The computer-readable medium of clause 37, wherein the first image corresponds to a first image capture condition and the reference image corresponds to a second image capture condition.
[0206] 41. The computer-readable medium of clause 40, wherein the first image has a lower resolution than the reference image.
[0207] 42. The computer-readable medium of clause 37, further comprising:
[0208] inputting a specified image of a specified substrate captured under rapid scan image capture conditions into the neural network; and
[0209] The neural network is executed to generate a specified predicted image based on the specified image, the specified predicted image representing a defect on the specified substrate and corresponding to a slow scan image capture condition.
[0210] 43. A non-transitory computer-readable medium having instructions that, when executed by a computer, cause the computer to perform a method for training a machine learning model to generate an image indicative of defects on a substrate, the method comprising:
[0211] acquiring a plurality of image pairs, wherein each image pair includes a first image and a reference image captured using different image capture conditions, the first image and the reference image representing defects on a substrate patterned using a target layout;
[0212] determining a defect detection probability for a reference image of the image pair;
[0213] selecting the subset of image pairs based on the defect detection probability; and
[0214] A neural network is trained using a subset of image pairs to convert the first image of an image pair into a predicted image using the reference image of the image pair of the subset of image pairs, wherein the predicted image represents a defect on the substrate and corresponds to image capture conditions of the reference image.
[0215] 44. The computer-readable medium of clause 43, wherein determining the defect detection probability for the reference image comprises:
[0216] For each reference image of the image pair,
[0217] determining a defect score map indicating a defect score for each pixel of the reference image, wherein the defect score indicates a probability of a defect being present in the corresponding pixel;
[0218] Obtaining a defect score of a portion of the reference image classified as a defect; and obtaining a defect score of a portion of the reference image classified as noise.
[0219] 45. A computer-readable medium according to clause 43, wherein selecting the subset of image pairs comprises: selecting some of the image pairs in which a reference image is associated with a defect score map having defect scores for defects and interferences within a first range.
[0220] 46. The computer-readable medium of clause 45, wherein the first range represents a range of defect scores within which the probability of defect detection of determining a defect from a disturbance is below a specified threshold.
[0221] 47. The computer-readable medium of clause 43, wherein training the neural network is an iterative process, and each iteration comprises:
[0222] determining a difference between the predicted image and the reference image;
[0223] determining whether the difference has decreased; and
[0224] In response to a determination that the difference has not decreased, parameters of the neural network are modified and iterations are repeated.
[0225] 48. The computer-readable medium of clause 43, wherein the first image corresponds to a first image capture condition and the reference image corresponds to a second image capture condition.
[0226] 49. The computer-readable medium of clause 48, wherein the first image has a lower resolution than the reference image.
[0227] 50. The computer-readable medium of clause 43, further comprising:
[0228] inputting a specified image of a specified substrate captured under first image capturing conditions into the neural network; and
[0229] The neural network is executed to generate a specified predicted image based on the specified image, the specified predicted image representing a defect on the specified substrate and corresponding to a second image capture condition.
[0230] 51. A method for training a machine learning model to generate an image representing defects on a substrate, the method comprising:
[0231] inputting a first image and a reference image into a neural network, the first image and the reference image representing images captured using different image capture conditions, the first image and the reference image indicating defects on a substrate patterned using a target layout;
[0232] generating a predicted image in response to the first image using the neural network;
[0233] calculating a loss function indicating a difference between a defect distribution in the predicted image and a defect distribution in the reference image; and
[0234] The neural network is modified based on the loss function.
[0235] 52. A method for training a machine learning model to generate an image representing defects on a substrate, the method comprising:
[0236] inputting a first image and a reference image representing images captured using different image capture conditions into a neural network, the first image and the reference image representing defects on a substrate patterned using a target layout;
[0237] generating a predicted image in response to the first image using the neural network;
[0238] calculating a loss function indicating a difference between a first set of feature vectors of the predicted image and a second set of feature vectors of the reference image; and
[0239] The neural network is modified based on the loss function.
[0240] 53. An apparatus for training a machine learning model to generate an image representing defects on a substrate, the apparatus comprising:
[0241] a memory storing an instruction set; and
[0242] A processor is configured to execute the instruction set so that the apparatus performs the following method:
[0243] inputting a first image and a reference image into a neural network, the first image and the reference image representing images captured using different image capture conditions, the first image and the reference image indicating defects on a substrate patterned using a target layout;
[0244] generating a predicted image in response to the first image using the neural network;
[0245] calculating a loss function indicating a difference between a defect distribution in the predicted image and a defect distribution in the reference image; and
[0246] The neural network is modified based on the loss function.
[0247] 54. An apparatus for training a machine learning model to generate an image representing defects on a substrate, the apparatus comprising:
[0248] a memory storing an instruction set; and
[0249] A processor is configured to execute the instruction set so that the apparatus performs the following method:
[0250] inputting a first image and a reference image into a neural network, the first image and the reference image representing images captured using different image capture conditions, the first image and the reference image representing defects on a substrate patterned using a target layout;
[0251] generating a predicted image in response to the first image using the neural network;
[0252] calculating a loss function indicating a difference between a first set of feature vectors of the predicted image and a second set of feature vectors of the reference image; and
[0253] The neural network is modified based on the loss function.
[0254] The description herein is intended to be illustrative and not restrictive. Therefore, it will be apparent to those skilled in the art that modifications may be made as described above without departing from the scope of the claims set forth below.
Claims
1. A non-transitory computer-readable medium having instructions that, when executed by a computer system, cause the computer system to at least: inputting a first image and a reference image into a neural network, the first image and the reference image representing images captured using different image capture conditions, the first image and the reference image indicating defects on a substrate patterned using a target layout; generating a predicted image in response to the first image using the neural network; calculating a loss function indicating a difference between a defect distribution in the predicted image and a defect distribution in the reference image; and The neural network is modified based on the loss function.
2. The computer-readable medium of claim 1 , wherein the instructions configured to cause the computer system to calculate the loss function are further configured to cause the computer system to: determining the defect distribution in the predicted image as a predicted defect score map, in which a defect score indicates a probability of a defect being present in a portion of the predicted image; determining the defect distribution in the reference image as a reference defect score map, in which a defect score indicates a probability of a defect being present in a portion of the reference image; and A difference between the predicted defect score map and the reference defect score map is calculated. 3 . The computer-readable medium of claim 2 , wherein the defect score meeting a threshold score indicates that a defect is at a location on the substrate corresponding to the portion of the reference image.
4. The computer-readable medium of claim 3 , wherein the instructions configured to cause the computer system to calculate the loss function are further configured to cause the computer system to: calculate a difference between a first set of feature vectors of the predicted image and a second set of feature vectors of the reference image.
5. The computer-readable medium of claim 4, wherein the instructions configured to cause the computer system to calculate the difference are further configured to cause the computer system to: applying a feature extraction filter to the reference image to obtain the first set of feature vectors as a reference classifier feature map, the reference classifier feature map representing characteristics of defects or interference in the reference image; applying the feature extraction filter to the predicted image to obtain the second set of feature vectors as a predicted classifier feature map, the predicted classifier feature map representing features of defects or interference in the predicted image; and The difference between the reference classifier feature map and the predicted classifier feature map is calculated.
6. The computer-readable medium of claim 1, wherein the instructions configured to cause the computer system to calculate the loss function are further configured to cause the computer system to: calculate a pixel-by-pixel difference between the predicted image and the reference image.
7. The computer-readable medium of claim 1, wherein the first image corresponds to a fast-scan image capture condition of an image capture device, and the predicted image and the reference image correspond to slow-scan image capture conditions of the image capture device.
8. The computer-readable medium of claim 7, wherein the first image has a lower resolution than the reference image.
9. The computer-readable medium of claim 1, wherein the instructions configured to cause the computer system to input the first image and the reference image are further configured to cause the computer system to: add defects to the first image and the reference image.
10. The computer-readable medium of claim 9, wherein the instructions configured to cause the computer system to add a defect to the first image and the reference image are further configured to cause the computer system to: edit a portion of the first image and the reference image to match a portion of a specified image that indicates a defect on the substrate.
11. The computer-readable medium of claim 1 , wherein the instructions configured to cause the computer system to input the first image and the reference image are further configured to cause the computer system to: select a first image pair from among the plurality of image pairs based on a defect detection probability of a reference image from among the plurality of image pairs, wherein the first image pair includes the first image and the reference image.
12. The computer-readable medium of claim 11 , wherein the instructions configured to cause the computer system to select the first image pair are further configured to cause the computer system to: select some of the image pairs in which a reference image is associated with a defect score map having defect scores for defects and interferences within a first range.
13. The computer-readable medium of claim 1 , wherein the instructions are further configured to cause the computer system to: inputting a specified image of a specified substrate captured under rapid scan image capture conditions into the neural network; and The neural network is executed to generate a specified predicted image based on the specified image, the specified predicted image representing a defect on the specified substrate and corresponding to a slow scan image capture condition.
14. The computer-readable medium of claim 1, wherein the instructions are further configured to modify an area of the reference image representing a defect to generate an updated reference image to be used for training the neural network.
15. The computer-readable medium of claim 14, wherein said modifying said region comprises: The contrast of the region of the reference image is enhanced.