Apparatus and method for generating a noise removal model

By converting design patterns into simulated images using a GAN to train a noise removal model, the method enhances pattern coverage and accuracy, addressing inefficiencies in existing noise removal models for patterned substrates.

JP7714632B2Active Publication Date: 2025-07-29ASML NETHERLANDS BV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023501901
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-14
Filing Date
2021-06-24
Publication Date
2025-07-29
Estimated Expiration
2041-06-24

AI Technical Summary

Technical Problem

Existing noise removal models for patterned substrates require extensive training data from SEM images, leading to inefficient pattern coverage and frequent retraining due to limited image capture, which is time-consuming and resource-intensive.

Method used

A method involving the conversion of design patterns into simulated images using a generative adversarial network (GAN) to train a noise removal model, allowing for extensive pattern coverage during offline training, reducing the need for real-time retraining.

Benefits of technology

Significantly improves the effectiveness and accuracy of noise removal by covering a larger number of patterns during training, minimizing the need for retraining, and reducing scan and training times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007714632000007
    Figure 0007714632000007
  • Figure 0007714632000008
    Figure 0007714632000008
  • Figure 0007714632000009
    Figure 0007714632000009
Patent Text Reader

Abstract

Described herein is a method for training a denoising model. The method includes obtaining a first set of simulated images based on a design pattern. The simulated images may be clean images, and noise may be added to these simulated images to generate noisy simulated images. The clean and noisy simulated images are used as training data for generating the denoising model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications

[0001] This application claims the benefit of U.S. Patent Application No. 63 / 051,500, filed Jul. 14, 2020, which is hereby incorporated by reference in its entirety.

[0002]

[0002] The description herein generally relates to the processing of images acquired by inspection tools or measurement tools, and more particularly, to the removal of noise from images by using machine learning.

Background Art

[0003]

[0003] A lithographic projection apparatus can be used, for example, in the manufacture of integrated circuits (ICs). In such a case, a patterning device (e.g., a mask) can contain, or provide, a pattern (a “design layout”) corresponding to an individual layer of the IC, and this pattern can be transferred onto a target portion (e.g., including one or more dies) on a substrate (e.g., a silicon wafer) coated with a layer of radiation-sensitive material (a “resist”), by, for example, irradiating the target portion through the pattern on the patterning device. In general, a single substrate contains a plurality of adjacent target portions (one target portion at a time) onto which the pattern is successively transferred by the lithographic projection apparatus. In one type of lithographic projection apparatus, the pattern over the entire patterning device is transferred onto one target portion at a time, and such an apparatus is generally called a stepper. In an alternative apparatus, generally called a step-and-scan apparatus, the projection beam scans the patterning device in a given reference direction (the “scan” direction) while synchronously moving the substrate parallel or anti-parallel to this reference direction. Different parts of the pattern on the patterning device are progressively transferred onto one target portion. In general, since the lithographic projection apparatus has a reduction ratio M (e.g., 4), the speed F at which the substrate is moved is the speed at which the projection beam scans the patterning device × 1 / M. Further information regarding lithographic devices as described herein can be learned, for example, from U.S. Patent No. 6,046,792, which is incorporated herein by reference.

[0004]

[0004] Before transferring a pattern from a patterning device to a substrate, the substrate may undergo various procedures such as priming, resist coating, and soft baking. After exposure, the substrate may undergo other procedures ( "post-exposure procedures") such as post-bake (PEB), development, hard bake, and measurement / inspection of the transferred pattern. This number of procedures is used as a basis for creating individual layers of a device, such as an IC. The substrate may then undergo various processes such as etching, ion implantation (doping), metallization, oxidation, chemical mechanical polishing, etc. (all intended to finish the individual layers of the device). If several layers are required for the device, the entire procedure or a variant thereof is repeated for each layer. Finally, there are devices on each target portion on the substrate. These devices are then separated from each other by techniques such as dicing or sawing, as a result of which it is possible to attach the individual devices onto a carrier, connect them to pins, etc.

[0005]

[0005] Thus, manufacturing devices such as semiconductor devices generally involve processing a substrate (e.g., a semiconductor wafer) using a number of fabrication processes to form various features and multiple layers of the device. Such layers and features are generally fabricated and processed using, for example, deposition, lithography, etching, chemical mechanical polishing, and ion implantation. Multiple devices may be fabricated on multiple dies on a substrate and then separated into individual devices. This device manufacturing process can be regarded as a patterning process. The patterning process includes patterning steps such as using a patterning device in a lithographic apparatus to transfer a pattern on the patterning device to a substrate using light and / or nanoimprint lithography, and generally (but optionally) includes one or more related pattern processing steps such as resist development by a development device, baking of the substrate using a baking tool, and etching using the pattern by an etching device.

Summary of the Invention

[0006] According to one embodiment, a method for training an image denoising model for processing an image is provided. The method includes converting a design pattern into a first set of simulated images and training a denoising model based on the first set of simulated images and image noise. After training, the denoising model is operable to remove noise from an input image to generate a denoised image.

[0007] In one embodiment, a system is provided that includes electron beam optics configured to capture an image of a patterned substrate and one or more processors configured to generate a de-noised image of the input image. The one or more processors are configured to execute a trained model configured to generate a simulated image from a design pattern on the substrate. In one embodiment, the one or more processors are configured to execute the de-noising model using the captured image as input to generate a de-noised image of the patterned substrate.

[0008]

[0008] In one embodiment, one or more non-transitory computer-readable media for storing a noise removal model are provided. In one embodiment, the one or more non-transitory computer-readable media are configured to generate a noise removal image by the stored noise removal model. In particular, the one or more non-transitory computer-readable media store instructions that provide the noise removal model when executed by one or more processors. In one embodiment, the noise removal model obtains a first set of simulated images based on a design pattern (e.g., by converting a GDS pattern into a simulated image using a trained GAN), provides the first set of simulated images as input to a basic noise removal model to obtain a second set of simulated images, where the second set of simulated images is a noise removal image associated with the design pattern, and uses a reference noise removal image as feedback to update one or more configurations of the reference noise removal model, where the one or more configurations are updated based on a comparison between the reference noise removal image and the second set of simulated images, and is generated by executing instructions for updating.

[0009]

[0009] In some embodiments, the GAN is trained to convert a GDS pattern image into a clean simulated SEM image. First, noise features are extracted from the scanned SEM image, and then noise is added to these clean images to generate a noisy simulated image. The clean simulated image and the noisy image are combined with the scanned SEM image and used to train the noise removal model. The noise removal model can be further fine-tuned using the captured SEM image. When trained, the noise removal model is operable to remove noise from the input SEM image to generate a noise removal image.

[0010]

[0010] According to an embodiment of the present disclosure, the noise removal model is trained by using a simulated image converted from a design pattern through a generator model as described above. Training data including such a simulated image can collectively cover significantly and sufficiently more patterns than an image captured by an SEM. As a result of the improved pattern coverage, training can advantageously significantly improve the effectiveness and accuracy of the noise removal model. The need for retraining can be significantly reduced or even eliminated.

[0011]

[0011] The above aspects as well as other aspects and features will become apparent to those skilled in the art by considering the following description of specific embodiments in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0012] [Figure 1]

[0012] A block diagram of various subsystems of a lithography system according to one embodiment is shown. [Figure 2]

[0013] A method for training a noise removal model. In one embodiment, the noise removal model is trained to convert a design pattern. [Figure 3]

[0014] A flowchart of a variant of a method for training a noise removal model according to one embodiment. [Figure 4]

[0015] An example of training a model according to one embodiment is shown. [Figure 5]

[0016] An example of obtaining a first set of simulated SEM images by using the trained model of FIG. 4 according to one embodiment is shown. [Figure 6]

[0017] An example of adding noise to the simulated SEM image of FIG. 5 according to one embodiment is shown. [Figure 7]

[0018] An example of training a noise removal model according to an embodiment is shown. [Figure 8]

[0019] An example of a trained noise removal model used to generate a noise-removed input SEM image according to an embodiment is shown. [Figure 9]

[0020] An embodiment of a scanning electron microscope (SEM) is schematically depicted according to an embodiment. [Figure 10]

[0021] An embodiment of an electron beam inspection apparatus is schematically shown according to an embodiment. [Figure 11]

[0022] A block diagram of an exemplary computer system according to an embodiment is shown. [Figure 12]

[0023] A schematic diagram of a lithographic projection apparatus according to an embodiment is shown. [Figure 13]

[0024] A schematic diagram of another lithographic projection apparatus according to an embodiment is shown. [Figure 14]

[0025] A more detailed view of the apparatus of FIG. 12 according to an embodiment is shown. [Figure 15]

[0026] A more detailed view of the source collector module SO of the apparatuses of FIGS. 13 and 14 according to an embodiment is shown.

DETAILED DESCRIPTION OF THE INVENTION

[0013]

[0027] Before describing the embodiments in detail, it is beneficial to present an exemplary environment in which the embodiments can be implemented.

[0014]

[0028] Although this book may specifically refer to the manufacture of ICs, it should be explicitly understood that the description in this specification has many other possible applications. For example, the description in this specification can be used in the manufacture of integrated optical systems, guidance and detection patterns for magnetic domain memories, liquid crystal display panels, thin film magnetic heads, etc. Those skilled in the art will understand that in the context of such alternative uses, the use of the terms "reticle", "wafer", or "die" in this book should be considered interchangeable with the more general terms "mask", "substrate", and "target portion", respectively.

[0015]

[0029] In this document, the terms "radiation" and "beam" can be used to encompass all types of electromagnetic radiation, including ultraviolet radiation (e.g., having wavelengths of 365, 248, 193, 157, or 126 nm) and EUV (extreme ultraviolet radiation, e.g., having wavelengths in the range of about 5 - 100 nm).

[0016]

[0030] A patterning device may include or form one or more design layouts. The design layouts can be generated using a CAD (computer-aided design) program, and this process is often referred to as EDA (electronic design automation). Most CAD programs follow a set of predefined design rules to fabricate a functional design layout / patterning device. These rules are set by processing and design constraints. For example, the design rules define the spacing tolerances between devices (gates, capacitors, etc.) or interconnect lines so as to ensure that devices or lines do not interact with each other in an undesirable manner. One or more of the design rule constraints may be referred to as "critical dimension" (CD). The critical dimension of a device may be defined as the minimum width of a line or hole, or the minimum space between two lines or two holes. Thus, the CD determines the overall size and density of the designed device. Naturally, one of the goals of device manufacturing is to faithfully reproduce the original design intent on the substrate (via the patterning device).

[0017]

[0031] The pattern layout design may include, by way of example, the application of resolution enhancement techniques such as optical proximity correction (OPC). OPC addresses the fact that the final size and placement of the image of the design layout projected onto the substrate may not be the same as, or may not depend only on, the size and placement of the design layout on the patterning device. Note that the terms “mask,” “reticle,” and “patterning device” are used interchangeably herein. Also, in the context of RET, a design layout can be used instead of necessarily using a physical patterning device to represent the physical patterning device, and those skilled in the art will recognize that the terms “mask,” “patterning device,” and “design layout” can be used interchangeably. When the size of features present in some design layouts is small and the density is high, the position of a particular edge of a given feature is affected to some extent by the presence or absence of other adjacent features. These proximity effects result from minute amounts of radiation coupled from one feature to another or from non-geometric optical effects such as diffraction and interference. Similarly, proximity effects can generally result from diffusion and other chemical effects during post-exposure bake (PEB), resist development, and etching following lithography.

[0018]

[0032] To increase the likelihood that the projected image of a design layout complies with the requirements of a given target circuit design, advanced numerical models, design layout correction, or pre-distortion can be used to predict and compensate for proximity effects. The paper "Full-Chip Lithography Simulation and Design Analysis-How OPC Is Changing IC Design", C. Spence, Proc. SPIE, Vol. 5751, pp1-14 (2005) provides an overview of current "model-based" optical proximity effect correction processes. In typical high-end designs, some modification is applied to almost all features of the design layout to achieve a high fidelity of the projected image to the target design. These modifications can include shifts or biases in edge position or line width, as well as the application of "assist" features intended to support the projection of other features.

[0019]

[0033] An assist feature can be regarded as the difference between a feature on the patterning device and a feature in the design layout. The terms "primary feature" and "assist feature" are not meant to imply that a particular feature on the patterning device must be labeled as either one.

[0020]

[0034] As used herein, the term "mask" or "patterning device" can be broadly interpreted to refer to a general patterning device that can be used to impart a patterned cross-section corresponding to a pattern to be generated in a target portion of a substrate to an incoming radiation beam, and the term "light valve" may also be used in this context. In addition to conventional masks (transmission or reflection, binary, phase shift, hybrid, etc.), examples of other such patterning devices include the following: - Programmable mirror array. An example of such a device is a matrix-addressable surface having a viscoelastic control layer and a reflective surface. The basic principle behind such an apparatus is that, for example, the addressed areas of the reflective surface reflect the incident radiation as diffracted radiation, while the non-addressed areas reflect the incident radiation as non-diffracted radiation. An appropriate filter can be used to remove the aforementioned non-diffracted radiation from the reflected beam, leaving only the diffracted radiation, and in this way, the beam is patterned according to the addressing pattern of the matrix-addressable surface. The required matrix addressing can be implemented using appropriate electronic means. - Programmable LCD array. An example of such a structure is given by U.S. Patent No. 5,229,872 incorporated herein.

[0021]

[0035] As a simple introduction, FIG. 1 shows an exemplary lithographic projection apparatus 10A. The main components can be a deep ultraviolet excimer laser source, or another type of source such as an extreme ultraviolet (EUV) source, a radiation source 12A (as considered herein, the lithographic projection apparatus itself need not have a radiation source), and an illumination optical system that defines partial coherence (represented as sigma), and may include optical systems 14A, 16Aa, and 16Ab that shape the radiation from the source 12A, a patterning device 18A, and a transmissive optical system 16Ac that projects an image of the patterning device pattern onto a substrate plane 22A. An adjustable filter or aperture 20A in the pupil plane of the projection optical system may limit the range of beam angles impinging on the substrate plane 22A, where the maximum possible angle defines the numerical aperture NA = n sin(Θmax) of the projection optical system, where n is the refractive index of the medium between the substrate and the last element of the projection optical system, and Θmax is the maximum angle of the beam emerging from the projection optical system that can still impinge on the substrate plane 22A.

[0022]

[0036] In a lithographic projection apparatus, a source provides illumination (i.e., radiation) to a patterning device, and a projection optical system guides and shapes the illumination via the patterning device onto a substrate. The projection optical system may include at least some of components 14A, 16Aa, 16Ab, and 16Ac. A spatial image (AI) is the radiation intensity distribution at the substrate level. The resist layer on the substrate is exposed, and the spatial image is transferred to the resist layer as a potential “resist image” (RI). The resist image (RI) can be defined as the spatial distribution of the solubility of the resist in the resist layer. A resist model can be used to calculate the resist image from the spatial image, an example of which can be found in U.S. Patent Application Publication No. 2009 / 0157360, the disclosure of which is hereby incorporated by reference in its entirety. The resist model relates only to the characteristics of the resist layer (e.g., the effects of chemical processes occurring during exposure, PEB, and development). The optical characteristics of the lithographic projection apparatus (e.g., the characteristics of the source, patterning device, and projection optical system) determine the spatial image. Since the patterning device used in a lithographic projection apparatus can be changed, it may be desirable to decouple the optical characteristics of the patterning device from the optical characteristics of the remainder of the lithographic projection apparatus, including at least the source and projection optical system.

[0023]

[0037] Although this book may refer to the use of lithographic apparatuses in the manufacture of ICs, it should be understood that the lithographic apparatuses described herein may have other applications, such as, for example, integrated optical systems, guidance and detection patterns for magnetic domain memories, liquid crystal display panels (LCDs), the manufacture of thin film magnetic heads, etc. Those skilled in the art will understand that, in the context of such alternative applications, the use of the terms "wafer" or "die" herein may be considered synonymous with the more general terms "substrate" or "target portion", respectively. The substrates referred to herein may be processed, before or after exposure, by, for example, a track (e.g., a tool that typically applies a resist layer to a substrate and develops the exposed resist), or a measurement tool or an inspection tool. Where applicable, the disclosures herein may be applied to such and other substrate processing tools. Further, a substrate may be processed multiple times, for example, to create a multilayer IC, and the term "substrate" as used herein may also refer to a substrate that already includes multiple processed layers.

[0024]

[0038] As used herein, the terms "radiation" and "beam" encompass all types of electromagnetic radiation, including ultraviolet (UV) radiation (e.g., having a wavelength of 365, 248, 193, 157, or 126 nm) and extreme ultraviolet (EUV) radiation (e.g., having a wavelength in the range of 5 - 20 nm), as well as particle beams such as ion beams or electron beams.

[0025]

[0039] Existing training methods for noise removal models require a large number of images of patterned substrates (e.g., SEM images) as training data. In such training methods, the pattern coverage of the design layout is not limited to the patterns in the SEM images. In one embodiment, pattern coverage refers to the number of unique patterns within a design layout. Typically, a design layout can have hundreds of millions to billions of patterns and millions of unique patterns. Measuring millions of patterns on a patterned substrate for training purposes requires a significant amount of measurement time and computing resources, so it is not practical. Therefore, for example, training data including SEM images is usually much less than the data sufficient to train a machine learning model. Thus, there may be a need to retrain a trained model in real time using new patterns.

[0026]

[0040] The method of the present disclosure has several advantages. For example, the pattern coverage of the design layout can be significantly increased during offline training. Only a limited number of SEM images (e.g., 10 - 20 actual SEM images) can be used for training and verification purposes. After training, the trained model can be used to remove noise from captured measurement images (e.g., SEM images) during execution. Since a relatively large number of patterns are covered during training, the amount of retraining of the existing model is significantly less than that of the existing model. Fine-tuning of the model can be quickly achieved, for example, by acquiring 10 - 20 actual SEM images. Therefore, compared with the existing model, a significant amount of machine scan time and online model training time can be reduced. For example, in this method, the scan time can be limited to 20 SEM images as opposed to thousands of SEM images, and the online model training time can be limited to about 0.5 hours compared to 4 - 8 hours.

[0027]

[0041] Figure 2 is an exemplary method 200 for training a noise removal model according to an embodiment of the present disclosure. In one embodiment, for training purposes, another model converts a design pattern (e.g., a design layout in GDS file data) into a clean simulated SEM image. Further, noise can be added to these clean images to generate noisy simulated images. The clean simulated images and the noisy images are combined with scanned SEM images and used to train the noise removal model. In one embodiment, the method includes processes P201 and P203, which are discussed in detail below.

[0028]

[0042] Process P201 includes converting a design pattern into a first set of simulated images 201, such as simulated SEM images. In one embodiment, the design pattern DP is in the graphic data signal (GDS) file format. For example, a design layout including millions of design patterns is represented as a GDS data file.

[0029]

[0043] In one embodiment, obtaining the first set of simulated images 201 includes using the design pattern DP as an input to execute a trained model MD1 to generate the simulated images 201. In one embodiment, the trained model MD1 is trained based on the design pattern DP and a captured image of the patterned substrate, and each captured image is associated with a design pattern. In one embodiment, the captured image is an SEM image obtained via a scanning electron microscope (SEM) (e.g., FIGS. 10 - 11).

[0030]

[0044] In one embodiment, the trained model MD1 can be any model such as a machine learning model that can be trained using an existing training method that uses training data as discussed herein. For example, the trained model MD1 can be a convolutional neural network (CNN) or a deep convolutional neural network (DCNN). The present disclosure is not limited to a particular training method or a particular neural network. By way of example, the model MD1 can be a first deep learning model (e.g., DCNN) trained using a training method such as a generative adversarial network (GAN), in which the design pattern DP and the SEM image are used as training data. In this example, the trained model MD1 (e.g., DCNN) is referred to as a generative model configured to generate a simulated SEM image from a given design pattern, e.g., a GDS pattern.

[0031]

[0045] In one embodiment, an adversarial generative network (GAN) includes two deep learning models that are trained together in particular adversarially, namely, a generative model (e.g., a CNN or DCNN) and a discriminator model (e.g., another CNN or another DCNN). The generator model can receive a design pattern DP and a captured image (e.g., an SEM image) as inputs and output a simulated image (e.g., a simulated SEM image). In one embodiment, the output simulated image can be labeled as a fake image or a real image. In one example, a fake image is an image of a particular class that has not actually existed so far (e.g., a noise-removed image of an SEM image). On the other hand, a real image used as a reference (or ground truth) is a previously existing image (e.g., an SEM of a printed circuit board) that can be used during the training of the generator model and the discriminator model. The purpose of training is to train the generator model to generate fake images that closely resemble real images. For example, the features of the fake image match at least 95% of the features of the real image (e.g., a noise-removed SEM image). As a result, the trained generator model can generate realistic simulation images with a high level of accuracy.

[0032]

[0046] In one embodiment, the generator model (G) can be a convolutional neural network. The generator model (G) receives a design pattern (z) as an input and generates an image. In one embodiment, the image may be referred to as a fake image or a simulated image. The fake image can be expressed as X fake = G(z). The generator model (G) can be associated with a first cost function. The first cost function enables adjustment of the parameters of the generator model such that the cost function is improved (e.g., maximized or minimized). In one embodiment, the first cost function includes a first log-likelihood term that determines the probability that the simulated image is a fake image given an input vector.

[0033]

[0047] An example of the first cost function can be expressed by Equation 1 below. L S = E[log P(S = fake|X fake )]…(1)

[0034]

[0048] In the above Equation 1, the log-likelihood of the conditional probability is calculated. In this equation, S refers to the source assignment as fake by the discriminator model, and X fake is the output of the generator model, that is, the fake image. Therefore, in one embodiment, the training method minimizes the first cost function (L). As a result, the generator model generates a fake image (i.e., a simulated image) such that the conditional probability that the discriminator model recognizes the fake image as fake is low. In other words, the generator model gradually generates a more realistic image or pattern.

[0035]

[0049] FIG. 4 shows an exemplary process of training a model (e.g., MD1) as considered in process P201. In one embodiment, the trained generator model MD1 is trained using an adversarial generation network. The training based on the adversarial generation network includes two deep learning models GM1 and DM1 that are jointly trained such that the generator model GM1 gradually generates more accurate and robust results.

[0036]

[0050] In one embodiment, the generator model GM1 can receive, for example, design patterns DP1 and DP2 and SEM images SEM1 and SEM2 corresponding to the design patterns DP1 and DP2 as inputs. The generator model GM1 outputs simulated images such as images DN1 and DN2.

[0037]

[0051] The simulated image of GM1 is received by a discriminator model DM1, which is another CNN. The discriminator model DM1 also receives real images SEM1 and SEM2 (or a set of real patterns) in the form of pixelated images. Based on the real images, the discriminator model DM1 determines whether the simulated image is a fake (e.g., label L1) or real (e.g., label L2) and assigns a label accordingly. If the discriminator model DM1 classifies the simulated image as fake, the parameters (e.g., bias and weights) of GM1 and DM1 are modified based on a cost function (e.g., the first cost function above). The models GM1 and DM1 are repeatedly modified until the discriminator DM1 consistently classifies the simulated images generated by GM1 as real. In other words, the generator model GM1 is configured to generate realistic SEM images for any input design pattern.

[0038]

[0052] FIG. 5 shows an example of obtaining a first set of simulated SEM images using the trained generator model MD1 of FIG. 4 according to one embodiment. In one embodiment, any design patterns DP11, DP12,... DPn can be input to the trained generator model MD1 to generate simulated SEM images S1, S2,... Sn, respectively. Although the present disclosure is not limited thereto, these simulated SEM images S1, S2,... Sn can be clean images that contain no typical noise or very low noise.

[0039]

[0053] In some embodiments, image noise can be added to generate the first set of simulated images 201 in FIG. 2. FIG. 6 shows an example of adding noise to simulated SEM images S1, S2,... Sn to generate a first set of simulated images S1 * , S2 * ,... Sn * for training a noise removal model, where the noise removal model is another machine learning model configured to remove noise from the input image. The training of the noise removal model is discussed below with respect to process P203.

[0040]

[0054] Process P203 includes training a noise removal model and uses the first set of simulated images 201 and image noise as training data. In one embodiment, additionally, the captured image 205 of the patterned substrate may be included in the training data. For example, the captured image 205 may be an SEM image of the patterned substrate.

[0041]

[0055] In one embodiment, the noise removal model MD2 is a second machine learning model. For example, the second machine learning model may be a second CNN or a second DCNN. The present disclosure is not limited to a specific deep learning training method or machine learning training method. In one embodiment, the training is an iterative process that is performed until the second set of simulated images fall within a specified threshold of ground truth, such as the first set of simulated images 201 (e.g., the simulated SEM images S1, S2,... Sn in FIG. 5) or a reference image, before adding noise. In one embodiment, the training of the noise removal model includes using the first set of simulated images, the image noise, and the captured image as training data.

[0042]

[0056] In one embodiment, the image noise is extracted from the captured image of the patterned substrate, such as the captured SEM image. For example, a noise filter may be applied to extract noise from the SEM image captured by the SEM tool. The extracted noise can be represented as image noise. In one embodiment, the image noise is Gaussian noise, white noise, or salt-and-pepper noise characterized by user-specified parameters. In one embodiment, the image noise includes pixels whose intensity values are statistically independent of each other. In one embodiment, Gaussian noise can be generated by changing the parameters of the Gaussian distribution function.

[0043]

[0057] In one embodiment, referring to FIG. 6, image noise such as Gaussian noise is added to, for example, the simulated images S1, S2,... Sn, S1 * , S2 * ,..., Sn* Images with noise such as can be generated. The noisy images are the first set of simulated images 201 used to train the noise removal model.

[0044]

[0058] In the conventional method, the noise removal model is trained by using the captured SEM image as the training image. Due to the imaging throughput of the SEM system, and thus the amount of images captured by the SEM, the training images can collectively cover only a relatively small number of patterns. As a result, the trained noise removal model is not useful for removing the noise of input images that may have diverse patterns. Unfortunately, the trained noise removal model needs to be retrained so that it can process images with new patterns. According to an embodiment of the present disclosure, the noise removal model is trained by using simulated images converted from design patterns through a generator model as described above. The training data including the simulated images can collectively cover significantly and sufficiently more patterns than the images captured by the SEM. As a result of the improved pattern coverage, training can advantageously significantly improve the effectiveness and accuracy of the noise removal model. The need for retraining can be significantly reduced or even eliminated.

[0045]

[0059] FIG. 7 shows an exemplary process for training a noise removal model (e.g., MD2) according to an embodiment of the present disclosure.

[0046]

[0060] The first simulated image 201, image noise, and / or reference image REF as training data for training the noise removal model MD2. In one embodiment, additionally, the captured image 710 (e.g., SEM image) of the patterned substrate may be included in the training data. In one embodiment, the number of captured images 710 may be relatively less than the number of simulated images 201. In one embodiment, the captured image 710 can be used to update the trained noise removal model MD2.

[0047]

[0061] In one embodiment, the model MD2 can receive, for example, the noisy images S1 * , S2 * , …, Sn * (e.g., generated using the model MD1) and image noise as inputs. The model MD2 outputs noise-removed images such as DN11, D12, … DNn. During the training process, one or more model parameters of the model MD2 (e.g., the weights and biases of different layers of the DCNN) can be modified until convergence is achieved or the noise-removed image falls within a specified threshold of the reference image REF. In one embodiment, convergence is achieved when a change in the model parameter values does not result in a significant improvement in the model output compared to the previous model output.

[0048]

[0062] In one embodiment, the method 200 further includes acquiring, via a measurement tool, an SEM image of the patterned substrate and using the SEM image as an input image to execute the trained noise removal model MD2 to generate a noise-removed SEM image.

[0049]

[0063] In one embodiment, the second machine learning model MD2 can also be trained using the GAN training method as described above, using the inputs considered in the process P203. For example, the first simulated image 201, image noise, and reference image REF are used as training data. In one embodiment, the reference image REF can be the first simulated image 201.

[0050]

[0064] Figure 8 shows an exemplary process of generating a noise-removed SEM image using the trained noise removal model MD2. Exemplary SEM images 801 and 802 of the patterned substrate are captured via an SEM tool. Note that SEM images 801 and 802 have complex patterns that are very different from the patterns used in the images in the training data. Since the trained model MD2 is trained based on simulated images related to the design pattern, advantageously, it can cover a large number of patterns. Therefore, the trained model MD2 can generate very accurate noise-removed images 811 and 812 of SEM images 801 and 802, respectively. The results in Figure 8 show that the trained model MD2 can handle new patterns without additional training. In one embodiment, in order to further improve the quality of the noise-removed image, further fine-tuning of the noise removal model MD2 can be performed using the newly captured SEM image.

[0051]

[0065] Figure 3 is a flowchart of another exemplary method 300 for generating a noise removal model according to an embodiment of the present disclosure. Method 300 includes processes P301, P303, and P305 discussed below.

[0052]

[0066] Process P301 includes obtaining a first set of simulated images 301 based on a design pattern. In one embodiment, each image in the first set of simulated images 301 is a combination of a simulated SEM image and image noise (e.g., see S1 * in Figure 6, S2 * ,... Sn * ).

[0053]

[0067] In one embodiment, obtaining a simulated SEM image includes using a design pattern as an input to execute a trained model to generate a simulated SEM image. For example, as discussed with respect to FIG. 5, the trained model MD1 is executed. In one embodiment, the trained model (e.g., MD1) is trained based on a design pattern and a captured image of a patterned substrate, and each captured image is associated with a design pattern. In one embodiment, the captured image is an SEM image obtained via a scanning electron microscope (SEM). In one embodiment, the image noise is noise extracted from the captured image of the patterned substrate. In one embodiment, the image noise is Gaussian noise, white noise, or salt-and-pepper noise characterized by user-specified parameters.

[0054]

[0068] In one embodiment, the trained model (e.g., MD1) is a first machine learning model. In one embodiment, the first machine learning model is a CNN or DCNN trained using an adversarial generative network. In one embodiment, the trained model MD1 is a generative model configured to generate a simulated SEM image for a given design pattern. For example, the trained model MD1 as discussed with respect to FIGS. 5 and 6. In one embodiment, the reference noise-removed image is a simulated SEM image associated with a design pattern. For example, the reference images can be S1, S2, … Sn generated by MD1 in FIG. 5.

[0055]

[0069] Process P303 includes providing the first set of simulated images 301 as input to the basic noise removal model BM1 to obtain an initial second set of simulated images, where the initial second set of simulated images is a noise removal image associated with a design pattern. In one embodiment, the basic model can be an untrained model or a trained model that needs to be fine-tuned. In one embodiment, the captured image of the patterned substrate can be used to train or fine-tune the noise removal model. Process P305 includes using the reference noise removal image as feedback to update one or more configurations of the basic noise removal model BM1. The one or more configurations are updated based on a comparison between the reference noise removal image and the second set of simulated images. For example, updating one or more configurations includes modifying the model parameters of the basic model. At the end of the training process, the basic model with updated configurations becomes the noise removal model. Such a noise removal model can generate a second set of simulated images using, for example, an SEM image as input.

[0056]

[0070] In an embodiment, the noise removal model is a second machine learning model. In an embodiment, the second deep learning model is trained using a deep learning method or a machine learning method. In one embodiment, the noise removal model can also be trained using a training method for adversarial generative networks. In one embodiment, the noise removal model is a convolutional neural network or other machine learning model. In one embodiment, the noise removal model is MD2 as discussed with respect to FIG. 7.

[0057]

[0071] As discussed herein, examples of noise removal models are machine learning models. Both unsupervised and supervised machine learning models can be used to generate noise removal images from input noisy images such as SEMs of patterned substrates. Without limiting the scope of the present invention, the application of supervised machine learning algorithms is described below.

[0058]

[0072] Supervised learning is a machine learning task of inferring a function from labeled training data. The training data includes a set of training examples. In supervised learning, each example is a pair having an input object (typically a vector) and a desired output value (also called a monitoring signal). A supervised learning algorithm analyzes the training data and generates a hypothesized function that can be used to map new examples. The optimal scenario would be one in which the algorithm is able to correctly determine the class labels of unseen instances. For this, it is necessary that the learning algorithm generalizes from the training data in a "reasonable way" to unseen situations.

[0059]

[0073] x i is the feature vector of the i-th example, and y i is its label (i.e., class), given a set of N training examples of the form {(x1, y1), (x2, y2), …, (x N , y N )}, the learning algorithm seeks a function g: X → Y, where X is the input space and Y is the output space. A feature vector is an n-dimensional vector of numerical features that represents an object. Many algorithms in machine learning require such a representation because the numerical representation of an object facilitates processing and statistical analysis. When representing an image, the feature values may correspond to the pixels of the image, and when representing text, the feature values may probably correspond to term occurrence frequencies. The vector space associated with these vectors is often called the feature space. The function g is usually an element of a space of possible functions G, called the hypothesis space. g is defined to return the y-value that gives the highest score:

Number

Number

[0060]

[0074] Both G and F can be in any space of functions, but many learning algorithms have g being a probabilistic model where g(x) = P(y|x) or f being a joint probability model where f(x,y) = P(x,y). For example, Naive Bayes and linear discriminant analysis are joint probability models, and logistic regression is a conditional probability model.

[0061]

[0075] There are two basic approaches regarding the selection of f or g, empirical risk minimization and structural risk minimization. Empirical risk minimization seeks the function that best fits the training data. Structural risk minimization includes a penalty function that controls the bias / variance trade-off.

[0062]

[0076] In both cases, it is assumed that the training set has samples of independent and identically distributed pairs (x i , y i ). To measure how well a function fits the training data, a loss function L:

Number

Number

Number

[0063]

[0077] The risk R(g) of the function g is defined as the expected loss of g. This is

Number

[0064]

[0078] Exemplary models of supervised learning include decision trees, ensemble methods (bagging, boosting, random forest), k-NN, linear regression, naive Bayes, neural networks, logistic regression, perceptron, support vector machines (SVM), relevance vector machines (RVM), and deep learning.

[0065]

[0079] SVM is an example of a supervised learning model that can be used to analyze data, recognize patterns, and perform classification and regression analysis. Given a set of training examples, each marked as belonging to one of two categories, the SVM training algorithm constructs a model that assigns new examples to one or the other category, making it a non-probabilistic binary linear classifier. The SVM model is a representation of the examples as points in space, mapped such that examples from separate categories are divided by the widest possible clear gap. New examples are then mapped into the same space and predicted to belong to a category based on which side of the gap they fall on.

[0066]

[0080] In addition to performing linear classification, SVM can efficiently perform non-linear classification using the so-called kernel method, which can implicitly map their inputs into a high-dimensional feature space.

[0067]

[0081] The kernel method requires a user-specified kernel, i.e., a similarity function on pairs of data points in the raw representation. The name "kernel method" comes from the use of a kernel function, which makes it possible to operate by simply computing the inner product between the images of all pairs of data in the feature space, rather than ever computing the coordinates of the data in that space. This operation is often computationally cheaper than an explicit computation of the coordinates. This technique is called the "kernel trick".

[0068]

[0082] The effectiveness of SVM depends on the choice of kernel, the parameters of the kernel, and the soft margin parameter C. A common choice is the Gaussian kernel with a single parameter γ. The best combination of C and γ is often selected by grid search (also known as "parameter sweep") using a sequence of exponentially increasing C and γ (e.g., C ∈ {2 -5 、2 -4 、…、2 15 、2 16}; γ ∈ {2 -15 、2 -14 、…、2 4 、2 5}).

[0069]

[0083] Grid search is an exhaustive search within a manually specified subset of the hyperparameter space of a learning algorithm. The grid search algorithm is generally guided by some performance metric, measured by cross-validation on the training set or evaluation on a provided validation set.

[0070]

[0084] Each combination of parameter selections may be checked using cross-validation, and the parameters with the best cross-validation accuracy are selected.

[0071]

[0085] Cross-validation, sometimes called rotation estimation, is a model validation technique for assessing how the results of a statistical analysis will generalize to an independent data set. It is mainly used in situations where the goal is prediction, and one wants to estimate how accurately a predictive model will perform in practice. In a prediction problem, a model is usually given a data set of known data on which training is to be performed (the training data set), and a data set of unknown data (or out-of-sample data) on which the model is to be tested (the test data set). The goal of cross-validation is to define a data set (i.e., the validation data set) for "testing" the model in the training stage in order to limit problems such as overfitting and to provide insights into how the model will generalize to an independent data set (i.e., an unknown data set from an actual problem, for example). One round of cross-validation involves partitioning the data samples into complementary subsets, performing an analysis on one subset (referred to as the training set), and validating the analysis on the other subset (referred to as the validation set or test set). To reduce variability, multiple rounds of cross-validation are performed using different partitions, and the validation results across those rounds are averaged.

[0072]

[0086] Then, for testing and for use in classifying new data, a final model that can be used is trained on the entire training set using the selected parameters.

[0073]

[0087] Another example of supervised learning is regression. Regression infers the relationship between a dependent variable and one or more independent variables from a set of values of the dependent variable and the corresponding values of the independent variables. Regression can estimate the conditional expected value of the dependent variable given the independent variables. The inferred relationship is sometimes called a regression function. The inferred relationship can be probabilistic.

[0074]

[0088] In one embodiment, a system is provided that can generate a denoised image using model MD2 after the system captures an image of a patterned substrate. In one embodiment, the system can be, for example, the SEM tool of FIG. 9 or the inspection tool of FIG. 10 configured to include, for example, models MD1 and / or MD2 discussed herein. For example, the metrology tool includes an electron beam generator for capturing an image of a patterned substrate and one or more processors including models MD1 and MD2. The one or more processors are configured to execute a trained model configured to generate a simulated image based on a design pattern used to pattern the substrate, and to execute a denoising model using the captured image and the simulated image as inputs to generate a denoised image of the patterned substrate. As described above, the denoising model (e.g., MD2) is a convolutional neural network.

[0075]

[0089] Further, in one embodiment, the one or more processors are further configured to update the denoising model based on the captured image of the patterned substrate. In one embodiment, updating the denoising model includes using the captured image to execute the denoising model to generate a denoised image and updating one or more parameters of the denoising model based on a comparison of the denoised image and a reference denoised image.

[0076]

[0090] The present disclosure is not limited to applications using noise-removed images. In the semiconductor industry, noise-removed images can be used, for example, for inspection and measurement. In one embodiment, a noise-removed image can be used to determine hot spots on a patterned substrate. The hot spots are determined based on the absolute CD values measured from the noise-removed image. Alternatively, the hot spots can be determined based on a set of predetermined rules, such as rules used in a design rule check system, including but not limited to line end pullback, corner rounding, proximity of adjacent features, pattern necking or pinching, and other measurement criteria of pattern deformation relative to the desired pattern.

[0077]

[0091] In one embodiment, a noise-removed image can be used to improve the patterning process. For example, a noise-removed image can be used in a simulation of the patterning process to predict, for example, contours, CDs, edge placement (e.g., edge placement errors), etc. in a resist and / or an etched image. The purpose of the simulation is to accurately predict, for example, the edge placement of the printed pattern, and / or the aerial image intensity gradient, and / or the CD, etc. These values can be compared with the intended design, for example, to correct the patterning process or identify locations where defects are predicted to occur. The intended design is generally defined as a pre-OPC design layout and can be provided in a standard digital file format such as GDSII or OASIS or other file formats.

[0078]

[0092] In some embodiments, the inspection or measurement device can be a scanning electron microscope (SEM) that obtains an image of a structure (e.g., a part or all of a device structure) exposed or transferred onto a substrate. FIG. 9 depicts an embodiment of an SEM tool. The primary electron beam EBP emitted from the electron source ESO is focused by the condenser lens CL and then passes through the beam deflectors EBD1, the E×B deflector EBD2, and the objective lens OL to irradiate the substrate PSub on the substrate table ST at a focus.

[0079]

[0093] When the substrate PSub is irradiated with the electron beam EBP, secondary electrons are generated from the substrate PSub. The secondary electrons are deflected by the E×B deflector EBD2 and detected by the secondary electron detector SED. For example, by detecting electrons generated from the sample in synchronization with a two-dimensional scan of the electron beam by the beam deflector EBD1 or an iterative scan of the electron beam EBP by the beam deflector EBD1 in the other of the X or Y direction together with a continuous movement of the substrate PSub by the substrate table ST in the X or Y direction, a two-dimensional electron beam image can be obtained.

[0080]

[0094] The signal detected by the secondary electron detector SED is converted into a digital signal by an analog / digital (A / D) converter ADC, and the digital signal is transmitted to the image processing system IPU. In one embodiment, the image processing system IPU may have a memory MEM for storing all or part of the digital image for processing by the processing unit PU. The processing unit PU (e.g., specially designed hardware or a combination of hardware and software) is configured to convert or process the digital image into a data set representing the digital image. Further, the image processing system IPU may have a storage medium STOR configured to store the digital image and the corresponding data set in a reference database. The display device DIS may be connected to the image processing system IPU, and as a result, the operator can perform the necessary operations of the device with the help of a graphical user interface.

[0081]

[0095] As described above, SEM images can be processed to extract contours that depict the edges of objects representing the device structure within the image. These contours are quantified using measurement criteria such as CD. Thus, typically, images of device structures are compared and quantified using simple measurement criteria such as edge-to-edge distance (CD) or simple pixel differences between images. A typical contour model for detecting the edges of objects within an image to measure CD uses image gradients. Indeed, these models rely on strong image gradients. However, in practice, images are usually noisy and have discontinuous boundaries. Techniques such as smoothing, adaptive thresholding, edge detection, erosion, and dilation can be used to process the results of the image gradient contour model to handle noisy and discontinuous images, but these techniques ultimately lead to low-resolution quantification of high-resolution images. Thus, in most cases, mathematical operations on images of device structures to reduce noise and automate edge detection lead to a loss of image resolution, thereby leading to a loss of information. Therefore, the result is a low-resolution quantification that is a simple representation of a complex and high-resolution structure.

[0082]

[0096] Therefore, for example, regardless of whether the structure is within a potential resist image, within a developed resist image, or transferred, e.g., by etching, to a layer on a substrate that can represent the general shape of the structure while maintaining resolution, it is desirable for the structure (e.g., circuit features, alignment marks, or metrology target portions (e.g., grating features), etc.) generated or expected to be generated using a patterning process to have a mathematical representation. In the context of lithography or other patterning processes, the structure can be a device or a part thereof during manufacturing, and the image can be an SEM image of the structure. Optionally, the structure can be a feature of a semiconductor device, e.g., an integrated circuit. In this case, the structure may be referred to as a pattern or a desired pattern including a plurality of features of the semiconductor device. Optionally, the structure can be an alignment mark, or a part thereof (e.g., a grating of alignment marks), used in an alignment measurement process to determine the alignment between an object (e.g., a substrate) and another object (e.g., a patterning device), or a metrology target, or a part thereof (e.g., a grating of the metrology target), used to measure parameters of the patterning process (e.g., overlay, focus, dose amount, etc.). In one embodiment, the metrology target is a diffraction grating used, for example, to measure overlay.

[0083]

[0097] FIG. 10 schematically shows a further embodiment of an inspection apparatus. The system is used to inspect a sample 90 (such as a substrate) on a sample stage 88 and includes a charged particle beam generator 81, a condenser lens module 82, a probe forming objective lens module 83, a charged particle beam deflection module 84, a secondary charged particle detector module 85, and an image forming module 86.

[0084]

[0098] The charged particle beam generator 81 generates a primary charged particle beam 91. The condenser lens module 82 condenses the generated primary charged particle beam 91. The probe forming the objective lens module 83 focuses the condensed primary charged particle beam into a charged particle beam probe 92. The charged particle beam deflection module 84 scans the formed charged particle beam probe 92 across the surface of the area of interest on the sample 90 fixed on the sample stage 88. In one embodiment, the charged particle beam generator 81, the condenser lens module 82, and the probe forming objective lens module 83, or their equivalent designs, alternatives, or any combination thereof, together form a charged particle beam probe generator that generates a charged particle beam probe 92 for scanning.

[0085]

[0099] The secondary charged particle detector module 85 detects secondary charged particles 93 (which may also include other charged particles reflected or scattered from the sample surface) emitted from the sample surface and generates a secondary charged particle detection signal 94 at the time of collision by the charged particle beam probe 92. The image forming module 86 (e.g., a computer device) is coupled to the secondary charged particle detector module 85, receives the secondary charged particle detection signal 94 from the secondary charged particle detector module 85, and forms at least one scanned image in response thereto. In one embodiment, the secondary charged particle detector module 85 and the image forming module 86, or their equivalent designs, alternatives, or any combination thereof, together form an image forming device that forms a scanned image from the detected secondary charged particles emitted from the sample 90 collided by the charged particle beam probe 92.

[0086]

[0100] In one embodiment, the monitoring module 87 is coupled to the image forming module 86 of the image forming apparatus to monitor and control the patterning process and / or to derive parameters for the design, control, monitoring, etc. of the patterning process using the scanned image of the sample 90 received from the image forming module 86. Therefore, in one embodiment, the monitoring module 87 is configured or programmed to execute the methods described herein. In one embodiment, the monitoring module 87 includes a computer device. In one embodiment, the monitoring module 87 is a computer program for providing the functions herein, including a computer program encoded on a computer-readable medium forming or disposed within the monitoring module 87.

[0087]

[0101] In one embodiment, similar to the electron beam inspection tool of FIG. 9 that uses a probe to inspect a substrate, the electron current in the system of FIG. 10 is significantly larger than that of a CD-SEM as depicted, for example, in FIG. 9. As a result, the probe spot is large enough, and thus the inspection speed may be fast. However, the resolution may not be higher than that of a CD-SEM due to the large probe spot. In one embodiment, the inspection apparatus discussed above can be a single-beam or multi-beam apparatus without limiting the scope of the present disclosure.

[0088]

[0102] For example, the SEM images from the systems of FIGS. 9 and / or 10 can be processed to extract the contours depicting the edges of the objects representing the device structures within the image. These contours are then typically quantified using measurement criteria such as CD at user-defined cut lines. Therefore, typically, the images of the device structures are compared and quantified using measurement criteria such as the edge-to-edge distance (CD) measured on the extracted contours or the simple pixel difference between the images.

[0089]

[0103] In one embodiment, one or more steps of process 200 and / or 300 can be implemented as instructions (e.g., program code) in a processor of a computer system (e.g., process 104 of computer system 100). In one embodiment, to increase computational efficiency, the steps can be distributed among multiple processors (e.g., parallel computing). In one embodiment, a computer program product including a non-transitory computer-readable medium has instructions stored on the non-transitory computer-readable medium, and the instructions, when executed by a computer hardware system, implement the methods described herein.

[0090]

[0104] According to the present disclosure, combinations of the disclosed elements and sub-combinations constitute separate embodiments. For example, the first combination includes determining a noise removal model based on a simulated image related to a design pattern and a noise image. The sub-combination may include determining a noise removal image using the noise removal model. In another combination, the noise removal image can be used in an inspection process to determine OPC or SMO based on the dispersion data generated by the model. In another example, the combination includes determining process adjustments for a lithography process, a resist process, or an etching process based on inspection data based on the noise removal image to improve the yield of the patterning process.

[0091]

[0105] FIG. 11 is a block diagram showing a computer system 100 that can assist in implementing the methods, flows, or apparatuses disclosed herein. The computer system 100 includes a bus 102 or other communication mechanism for communicating information, and a processor 104 (or processors 104 and 105) coupled to the bus 102 for processing information. The computer system 100 also includes a main memory 106 coupled to the bus 102 for storing information and instructions to be executed by the processor 104, such as random access memory (RAM) or other dynamic storage devices. The main memory 106 may also be used to store temporary variables or other intermediate information during execution of instructions by the processor 104. The computer system 100 further includes a read only memory (ROM) 108 or other static storage device coupled to the bus 102 for storing static information and instructions for the processor 104. A storage device 110, such as a magnetic disk or optical disk, is provided and coupled to the bus 102 for storing information and instructions.

[0092]

[0106] The computer system 100 may be coupled via the bus 102 to a display 112, such as a cathode ray tube (CRT), flat panel, or touch panel display, for displaying information to a computer user. An input device 114, including alphanumeric and other keys, is coupled to the bus 102 for communicating information and command selections to the processor 104. Another type of user input device is a cursor control 116, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to the processor 104 and for controlling the movement of a cursor on the display 112. This input device generally has two degrees of freedom that enable the device to specify a position within a plane in two axes (a first axis (e.g., x) and a second axis (e.g., y)). A touch panel (screen) display may be used as the input device.

[0093]

[0107] According to one embodiment, portions of one or more of the methods herein may be performed by computer system 100 in response to one or more sequences of one or more instructions contained in main memory 106 and executed by processor 104. Such instructions may be read into main memory 106 from another computer-readable medium, such as storage device 110. Execution of the sequences of instructions contained in main memory 106 causes processor 104 to perform the process steps described herein. To execute the sequences of instructions contained in main memory 106, one or more processors in a multiprocessing configuration may be used. In an alternative embodiment, instead of software instructions, or in combination with software instructions, hardwired circuitry may be used. Accordingly, the description herein is not limited to a particular combination of hardware circuitry and software.

[0094]

[0108] As used herein, the term "computer-readable medium" refers to any medium that participates in providing instructions to processor 104 for execution. Such a medium may take many forms, including but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 110. Volatile media includes dynamic memory, such as main memory 106. Transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that make up bus 102. Transmission media can also take the form of acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, other magnetic media, CD-ROM, DVD, other optical media, punch cards, paper tape, other physical media with patterns of holes, RAM, PROM, and EPROM, FLASH-EPROM, other memory chips or cartridges, carrier waves as described hereinafter, or other media that can be read by a computer.

[0095]

[0109] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to processor 104 for execution. For example, the instructions may initially be on a magnetic disk of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 100 can receive the data on the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector coupled to bus 102 can receive the data carried in the infrared signal and place that data on bus 102. Bus 102 carries the data to main memory 106, from which processor 104 reads and executes the instructions. The instructions received by main memory 106 may optionally be stored in storage device 110 before or after execution by processor 104.

[0096]

[0110] Computer system 100 may also include a communication interface 118 coupled to bus 102. Communication interface 118 also provides bidirectional data communication coupling to network link 120 connected to local network 122. For example, communication interface 118 may be a digital integrated services network (ISDN) card or modem that provides a data communication connection to a corresponding type of telephone line. As another example, communication interface 118 may be a local area network (LAN) card that provides a data communication connection to a compatible LAN. A wireless link may be implemented. In such an implementation, communication interface 118 transmits and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.

[0097]

[0111] Network link 120 generally provides data communication to other data devices through one or more networks. For example, network link 120 can provide a connection to a data device operated by host computer 124 or Internet service provider (ISP) 126 through local network 122. ISP 126 then provides a data communication service via a worldwide packet data communication network (currently generally referred to as the "Internet" 128). Both local network 122 and Internet 128 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals passing through various networks to and from computer system 100, signals on network link 120, and signals through communication interface 118 are examples of carrier waves that carry information.

[0098]

[0112] Computer system 100 can send messages and receive data including program code through one or more networks, network link 120, and communication interface 118. In an Internet example, server 130 may send request code for an application program through Internet 128, ISP 126, local network 122, and communication interface 118. Some such downloaded applications can provide all or part of the methods described herein. The received code may be executed by processor 104 upon receipt and / or stored in storage device 110 or other non-volatile storage for later execution. In this way, computer system 100 may obtain application code in the form of a carrier wave.

[0099]

[0113] FIG. 12 schematically depicts an exemplary lithographic projection apparatus that can be used in combination with the technology described herein. The apparatus includes the following: - An illumination system IL for adjusting the radiation beam B. In this particular case, the illumination system also includes a radiation source SO; - A first object table (e.g., a patterning device table) MT equipped with a patterning device holder for holding a patterning device MA (e.g., a reticle), and connected to a first positioner for accurately positioning the patterning device with respect to the item PS; - A second object table (substrate table) WT equipped with a substrate holder for holding a substrate W (e.g., a resist-coated silicon wafer), and connected to a second positioner for accurately positioning the substrate with respect to the item PS; - A projection system (a "lens") PS (e.g., a refractive, reflective, or catadioptric optical system) that images the irradiated portion of the patterning device MA onto a target portion C (e.g., including one or more dies) of the substrate W.

[0100]

[0114] As depicted herein, the apparatus is transmissive (i.e., has a transmissive patterning device). However, generally, it may also be, for example, reflective (having a reflective patterning device). The apparatus may use different types of patterning devices than conventional masks, and examples include programmable mirror arrays or LCD matrices.

[0101]

[0115] The source SO (e.g., a mercury lamp, an excimer laser, an LPP (laser-produced plasma) EUV source) generates a radiation beam. This beam is supplied to the illumination system (illuminator) IL either directly or after passing through adjustment means such as a beam expander Ex, for example. The illuminator IL may include adjustment means AD for setting the outer and / or inner radius ranges of the beam intensity distribution (generally referred to as σ-outer and σ-inner, respectively). Furthermore, it generally includes various other components such as an integrator IN and a condenser CO. In this way, the beam B impinging on the patterning device MA has a desired uniformity and intensity distribution in cross-section.

[0102]

[0116] With respect to FIG. 12, the source SO may be located within the housing of the lithographic projection apparatus (in most cases, when the source SO is, for example, a mercury lamp), but is located at a position remote from the lithographic projection apparatus, and it should be noted that the radiation beam it generates may be introduced into the apparatus (for example, using appropriate guiding mirrors). This latter scenario is often the case when the source SO is an excimer laser (for example, based on KrF, ArF, or F2 lasing).

[0103]

[0117] Subsequently, the beam PB intersects the patterning device MA held on the patterning device table MT. After the beam B traverses the patterning device MA, it passes through a lens PL that focuses the beam B onto the target portion C of the substrate W. Using the second positioning means (and the interferometric means IF), the substrate table WT can be accurately moved, for example, to position different target portions C within the path of the beam PB. Similarly, for example, after a mechanical search for the patterning device MA from a patterning device library or during a scan, the patterning device MA can be accurately positioned with respect to the path of the beam B using the first positioning means. In general, the movement of the object tables MT, WT is realized using a long-stroke module (coarse positioning) and a short-stroke module (fine positioning) that are not explicitly depicted in FIG. 12. However, in the case of a stepper (in contrast to a step-and-scan tool), the patterning device table MT may be connected only to a short-stroke actuator or may be fixed.

[0104]

[0118] The tool depicted can be used in two different modes: - In step mode, the patterning device table MT remains essentially stationary, and the entire patterning device image is projected onto the target portion C in one go (i.e., a single “flash”). Then, the substrate table WT is shifted in the x and / or y directions so that different target portions C can be irradiated by the beam PB; - In scan mode, basically the same scenario applies, except that a given target portion C is not exposed in a single “flash”. Instead, the patterning device table MT is movable at a speed v in a given direction (the so-called “scan direction”, e.g., the y direction) such that the projection beam B is scanned over the patterning device image. In parallel, the substrate table WT is simultaneously moved in the same or opposite direction at a speed V = Mv (where M is the magnification of the lens PL (generally, M = 1 / 4 or 1 / 5)). In this way, a relatively large target portion C can be exposed without the need to compromise the resolution.

[0105]

[0119] FIG. 13 schematically shows another exemplary lithographic projection apparatus LA that can be used in conjunction with the technology described herein.

[0106]

[0120] The lithographic projection apparatus LA includes the following: - A source collector module SO - An illumination system (illuminator) IL configured to condition a radiation beam B (e.g., EUV radiation) - A support structure (e.g., a patterning device table) MT constructed to support a patterning device (e.g., a mask or reticle) MA and connected to a first positioner PM configured to accurately position the patterning device - A substrate table (e.g., a wafer table) WT constructed to hold a substrate (e.g., a resist-coated wafer) W and connected to a second positioner PW configured to accurately position the substrate, and - A projection system (e.g., a reflective projection system) PS configured to project a pattern imparted to a radiation beam B by a patterning device MA onto a target portion C (e.g., including one or more dies) of a substrate W.

[0107]

[0121] As depicted herein, the device LA is reflective (e.g., using a reflective patterning device). It should be noted that since most materials are absorptive within the EUV wavelength range, the patterning device can have, for example, a multilayer reflector including a multi-stack of molybdenum and silicon. In one example, the multi-stack reflector has 40 layer pairs of molybdenum and silicon, each layer having a thickness of a quarter wavelength. Even smaller wavelengths can be generated using X-ray lithography. Since most materials are absorptive at EUV and x-ray wavelengths, a thin patch of patterned absorptive material (e.g., TaN absorber on a multilayer reflector) on the patterning device topography defines where features are printed (positive resist) or not printed (negative resist).

[0108]

[0122] Referring to FIG. 13, illuminator IL receives an extreme ultraviolet (EUV) radiation beam from source collector module SO. The method of generating EUV radiation is not necessarily limited, but includes converting a material into a plasma state having at least one element (e.g., xenon, lithium, or tin) with one or more emission lines in the EUV range. In one such method, often referred to as laser-produced plasma (“LPP”), the plasma can be generated by irradiating a fuel such as a droplet, stream, or cluster of a material having a line-emitting element with a laser beam. Source collector module SO may be part of an EUV radiation system that includes a laser (not shown in FIG. 13) that provides a laser beam for exciting the fuel. The resulting plasma emits output radiation (e.g., EUV radiation), which is collected using a radiation collector disposed in the source collector module. The laser and the source collector module may be separate entities, for example, when a CO2 laser is used to provide a laser beam for fuel excitation.

[0109]

[0123] In such a case, the laser is not considered to form part of the lithographic apparatus, and the radiation beam is passed from the laser to the source collector module using a beam delivery system that includes, for example, suitable guiding mirrors and / or a beam expander. In other cases, for example, when the source is a discharge-produced plasma EUV generator, often referred to as a DPP source, the source may be an integral part of the source collector module.

[0110]

[0124] Illuminator IL may include an adjuster for adjusting the angular intensity distribution of the radiation beam. Generally, at least the outer and / or inner radius ranges of the intensity distribution of the pupil plane of the illuminator (commonly referred to as σ-outer and σ-inner, respectively) can be adjusted. Further, illuminator IL may include various other components such as a facet field and a pupil mirror device. The illuminator can be used to adjust the radiation beam to have a desired uniformity and intensity distribution in cross-section.

[0111]

[0125] The radiation beam B is incident on a patterning device (e.g., a mask) MA held on a support structure (e.g., a patterning device table) MT and is patterned by the patterning device. After being reflected from the patterning device (e.g., a mask) MA, the radiation beam B passes through a projection system PS that focuses the beam onto a target portion C of the substrate W. Using a second positioner PW and a position sensor PS2 (e.g., an interferometric device, a linear encoder, or a capacitance sensor), the substrate table WT can be accurately moved, for example, to position different target portions C within the path of the radiation beam B. Similarly, using a first positioner PM and another position sensor PS1, the patterning device (e.g., a mask) MA can be accurately positioned with respect to the path of the radiation beam B. The patterning device (e.g., a mask) MA and the substrate W may be aligned using patterning device alignment marks M1, M2 and substrate alignment marks P1, P2.

[0112]

[0126] The depicted apparatus LA can be used in at least one of the following modes: [[ID=X]] 1. In step mode, while the entire pattern imparted to the radiation beam is projected onto the target portion C in one go, the support structure (e.g., the patterning device table) MT and the substrate table WT remain essentially stationary (i.e., single static exposure). Then, the substrate table WT is shifted in the X and / or Y directions so that different target portions C can be exposed. 2. In scan mode, while the pattern imparted to the radiation beam is projected onto the target portion C, the support structure (e.g., the patterning device table) MT and the substrate table WT are scanned synchronously (i.e., single dynamic exposure). The speed and direction of the substrate table WT relative to the support structure (e.g., the patterning device table) MT can be determined by the reduction and image inversion characteristics of the projection system PS. 3. In another mode, while the pattern imparted to the radiation beam is projected onto the target portion C, the support structure (e.g., the patterning device table) MT holds the programmable patterning device and remains substantially stationary, and the substrate table WT is moved or scanned. In this mode, a pulsed radiation source is generally used, and the programmable patterning device is updated as necessary after each movement of the substrate table WT or between successive radiation pulses during scanning. This mode of operation can be readily applied to maskless lithography that utilizes a programmable patterning device such as a programmable mirror array of the type mentioned above.

[0113]

[0127] Figure 14 shows more details of an apparatus LA including a source collector module SO, an illumination system IL, and a projection system PS. The source collector module SO is constructed and arranged such that a vacuum environment can be maintained within the closed structure 220 of the source collector module SO. The EUV radiation emitting plasma 210 can be formed by a discharge generating plasma source. The EUV radiation can be generated by a gas or vapor (e.g., Xe gas, Li vapor, or Sn vapor that creates an ultra-high temperature plasma 210 to emit radiation within the EUV range of the electromagnetic spectrum). The ultra-high temperature plasma 210 is created, for example, by a discharge that produces at least a partially ionized plasma. A partial pressure of, for example, 10 Pa of Xe, Li, Sn vapor or any other suitable gas or vapor may be required for the efficient generation of the radiation. In one embodiment, a plasma of excited tin (Sn) is provided to generate EUV radiation.

[0114]

[0128] The radiation emitted by the high-temperature plasma 210 is passed from the source chamber 211, through an optional gas barrier or contaminant trap 230 (sometimes also called a contaminant barrier or foil trap) located within or behind the opening of the source chamber 211, into the collector chamber 212. The contaminant trap 230 can include a channel structure. The contaminant trap 230 can also include a gas barrier, or a combination of a gas barrier and a channel structure. The contaminant trap or contaminant barrier 230 further shown herein includes at least a channel structure, as is known in the art.

[0115]

[0129] The collector chamber 211 can include a radiation collector CO, which can be a so-called oblique incidence type collector. The radiation collector CO has an upstream radiation collector side 251 and a downstream radiation collector side 252. The radiation traversing the collector CO can be reflected by the grating spectral filter 240 and focused on a virtual light source point IF along the optical axis indicated by the dashed line “O”. The virtual light source point IF is generally called an intermediate focus, and the source collector module is arranged such that the intermediate focus IF is located at or near the opening 221 of the closed structure 220. The virtual light source point IF is an image of the radiation emitting plasma 210.

[0116]

[0130] Subsequently, the radiation traverses an illumination system IL that can include a facet field mirror device 22 and a facet pupil mirror device 24 arranged to provide a desired angular distribution of the radiation beam 21 and a desired uniformity of the radiation intensity in the patterning device MA. When the radiation beam 21 is reflected in the patterning device MA held by the support structure MT, a patterned beam 26 is formed, and the patterned beam 26 is imaged by the projection system PS, via the reflection elements 28, 30, onto a substrate W held by the substrate table WT.

[0117]

[0131] In general, more elements than shown may be present within the illumination optical system unit IL and the projection system PS. The grating spectral filter 240 may optionally be present, depending on the type of lithographic apparatus. Further, more mirrors than shown in the drawings may be present, for example, 1 to 6 additional reflective elements may be present in the projection system PS than shown in FIG. 14.

[0118]

[0132] A collector system CO such as shown in FIG. 14 is depicted as a nested collector comprising the obliquely incident reflectors 253, 254, and 255, as just one example of a collector (or collector mirror). The obliquely incident reflectors 253, 254, and 255 are arranged axially symmetrically with respect to the optical axis O, and this type of collector system CO can be used in combination with a discharge-generated plasma source, often referred to as a DPP source.

[0119]

[0133] Alternatively, the source collector module SO may be part of an LPP radiation system, as shown in FIG. 15. The laser LA is arranged to deposit laser energy onto a fuel such as xenon (Xe), tin (Sn), or lithium (Li) to generate a highly ionized plasma 210 with an electron temperature of several tens of eV. The energy radiation generated during the de-excitation and recombination of these ions is emitted from the plasma, collected by the near-normal incidence collector system CO, and focused onto the aperture 221 of the enclosure structure 220.

[0120]

[0134] The concepts disclosed herein can perform simulations or mathematical modeling of general imaging systems for imaging sub-wavelength features and can be useful, in particular, for new imaging techniques that can generate shorter wavelengths. New technologies already in use include EUV (extreme ultraviolet), DUV lithography, which can generate wavelengths of 193 nm using an ArF laser and even 157 nm using a fluorine laser. EUV lithography can also generate wavelengths in the range of 20 - 5 nm by using a synchrotron or by bombarding a material (solid or plasma) with high-energy electrons to generate photons within this range.

[0121]

[0135] Embodiments of the present disclosure can be further described by the following clauses. 1. One or more non-transitory computer-readable media that store a noise removal model and generate a noise removal image by noise removal model instructions that provide the noise removal model when executed by one or more processors, wherein the noise removal model converts a design pattern into a first set of simulated images, provides the first set of simulated images as input to a basic noise removal model to obtain a second set of simulated images, wherein the second set of simulated images is a noise removal image associated with the design pattern, and updates one or more configurations of the basic noise removal model using a reference noise removal image as feedback, wherein the one or more configurations are updated based on a comparison between the reference noise removal image and the second set of simulated images, and is generated thereby. 2. The medium according to clause 1, wherein each image of the first set of simulated images is a combination of a simulated SEM image and image noise. 3. Converting a design pattern includes using the design pattern as an input to execute a trained model to generate a simulated SEM image, the medium according to clause 2. 4. The trained model is trained based on a design pattern and a captured image of a patterned substrate, and each captured image is associated with the design pattern, the medium according to clause 3. 5. The captured image is an SEM image obtained via a scanning electron microscope (SEM), the medium according to clause 4. 6. The image noise is noise extracted from the captured image of the patterned substrate, the medium according to clause 5. 7. The trained model is a first machine learning model, the medium according to any one of clauses 3 to 6. 8. The trained model is a convolutional neural network or a deep convolutional neural network trained using an adversarial generation network, the medium according to clause 7. 9. The trained model is a generation model configured to generate a simulated SEM image for a given design pattern, the medium according to clause 8. 10. The image noise is Gaussian noise, white noise, or salt-and-pepper noise characterized by user-specified parameters, the medium according to any one of clauses 2 to 9. 11. The reference noise-removed image is a simulated SEM image associated with the design pattern, the medium according to any one of clauses 2 to 10. 12. The noise removal model is a second machine learning model, the medium according to any one of clauses 1 to 11. 13. The noise removal model is a convolutional neural network or a deep convolutional neural network, the medium according to any one of clauses 1 to 12. 14. The design pattern is in the graphic data signal (GDS) file format, the medium according to any one of clauses 1 to 13. 15. An electron beam optical system configured to capture an image of a patterned substrate, One or more processors configured to execute a noise removal model using a captured image as an input to generate a noise-removed image of a patterned substrate A system comprising 16. The system according to clause 15, wherein the noise removal model is a convolutional neural network 17. The system according to clause 15 or 16, wherein the one or more processors are further configured to execute a trained model using a design pattern provided in a Graphics Data Signal (GDS) file format to generate a simulated image 18. The system according to any one of clauses 15 to 17, wherein the one or more processors are further configured to update the noise removal model based on a captured image of a patterned substrate 19. The system according to any one of clauses 15 to 18, wherein the one or more processors are further configured to update one or more parameters of the noise removal model based on a comparison between the noise-removed image and a reference noise-removed image 20. The noise removal model is Converting a design pattern into a first set of simulated images Providing the first set of simulated images as an input to a basic noise removal model to obtain a second set of simulated images, wherein the second set of simulated images is a noise-removed image associated with the design pattern Using a reference noise-removed image as feedback to update one or more configurations of the basic noise removal model, wherein the one or more configurations are updated based on a comparison between the reference noise-removed image and the second set of simulated images Generated by, the system according to any one of clauses 15 to 19 21. The system according to clause 20, wherein each image of the first set of simulated images is a combination of a simulated SEM image and image noise 22. A method for training a noise removal model, comprising Converting a design pattern into a first set of simulated images Training a noise removal model based on a first set of simulated images and image noise, the noise removal model being operable to generate a noise-removed image of an input image, the training including the method. 23. Converting a design pattern into a first set of simulated images includes using the design pattern as an input to execute a trained model to generate a simulated image, the method according to clause 22. 24. The trained model is trained based on a design pattern and a captured image of a patterned substrate, each captured image being associated with the design pattern, the method according to clause 23. 25. The captured image is a SEM image obtained via a scanning electron microscope (SEM), the method according to clause 24. 26. Further including adding image noise to the first set of simulated images to generate a second set of simulated images, the image noise being extracted from a captured image of a patterned substrate, the method according to clause 25. 27. Training the noise removal model includes using the first set of simulated images, the image noise, and the captured image as training data, the method according to clause 26. 28. The trained model is a first machine learning model, the method according to any one of clauses 23 to 27. 29. The trained model is a convolutional neural network or a deep convolutional neural network trained using an adversarial generation network training method, the method according to clause 28. 30. The trained model is a generation model configured to generate a simulated SEM image for a given design pattern, the method according to clause 29. 31. The image noise is Gaussian noise, white noise, or salt-and-pepper noise characterized by user-specified parameters, the method according to any one of clauses 22 to 30. 32. The noise removal model is the method described in any one of clauses 23 to 31, which is a second machine learning model. 33. The noise removal model is the method described in any one of clauses 22 to 32, which is a convolutional neural network. 34. The design pattern is the method described in any one of clauses 22 to 33, which is a graphic data signal (GDS) file format. 35. Obtaining a SEM image of a patterned substrate via a measurement tool, using the SEM image as an input image to execute a trained noise removal model to generate a noise-removed SEM image, and further including the method described in any one of clauses 22 to 24. 36. One or more non-transitory computer-readable media that store a noise removal model and generate a noise-removed image by noise removal model instructions that provide the noise removal model when executed by one or more processors, where the noise removal model converts a design pattern into a first set of simulated images, and trains the noise removal model based on the first set of simulated images and image noise, where the noise removal model is operable to generate a noise-removed image of an input image, the training and is generated thereby. 37. The converting in clause 36 includes using a design pattern as an input to execute a trained model to generate a simulated image. 38. The trained model is trained based on a design pattern and a captured image of a patterned substrate, and each captured image is associated with the design pattern. 39. The captured image is a SEM image obtained via a scanning electron microscope (SEM). 40. Further comprising generating a second set of simulated images by adding image noise to the first set of simulated images, wherein the image noise is extracted from the captured image of the patterned substrate, the medium according to clause 39. 41. The trained model is a first machine learning model, the medium according to any one of clauses 37 to 40. 42. The trained model is a convolutional neural network or a deep convolutional neural network trained using an adversarial generation network training method, the medium according to clause 41. 43. The trained model is a generation model configured to generate a simulated SEM image for a given design pattern, the medium according to clause 42. 44. The image noise is Gaussian noise, white noise, or salt-and-pepper noise characterized by user-specified parameters, the medium according to any one of clauses 36 to 43. 45. The noise removal model is a second machine learning model, the medium according to any one of clauses 36 to 44. 46. Training the noise removal model includes using the first set of simulated images, the image noise, and the captured image as training data, the medium according to any one of clauses 38 to 45. 47. The noise removal model is a convolutional neural network, the medium according to any one of clauses 36 to 46. 48. The design pattern is in the graphic data signal (GDS) file format, the medium according to any one of clauses 36 to 47. 49. Obtaining an SEM image of the patterned substrate via a measurement tool, Using the SEM image as an input image to execute a trained noise removal model to generate a noise-removed SEM image, Further comprising, the medium according to any one of clauses 36 to 48.

[0122]

[0136] The concepts disclosed herein can be used for imaging on a substrate such as a silicon wafer, but it is understood that the disclosed concepts can be used in any type of lithographic imaging system (e.g., those used for imaging on substrates other than silicon wafers).

[0123]

[0137] The above description is intended to be illustrative rather than limiting. Thus, it will be apparent to those skilled in the art that modifications can be made as described without departing from the scope of the claims presented below.

Claims

1. A method for training a noise removal model, comprising: converting a design pattern, which is a design layout in data, into a first set of simulated SEM images; training the noise removal model based on the first set of simulated SEM images and image noise, wherein the noise removal model is operable to generate a noise-removed image of an input image; The method as described above.

2. The method according to claim 1, wherein converting the design pattern into the first set of simulated SEM images includes executing a trained model configured to generate the simulated SEM images using the design pattern as an input.

3. The method according to claim 2, wherein the trained model is trained based on the design pattern and a captured image of a patterned substrate, and each captured image is associated with a design pattern.

4. The method according to claim 3, wherein the captured image is an SEM image obtained via a scanning electron microscope (SEM).

5. The method according to claim 4, further comprising adding the image noise to the first set of simulated SEM images to generate a second set of simulated images, wherein the image noise is extracted from the captured images of the patterned substrate.

6. The method according to claim 5, wherein training the noise removal model includes using the first set of simulated SEM images, the image noise, and the captured images as training data.

7. The method according to claim 2, wherein the trained model includes a first machine learning model.

8. The method according to claim 7, wherein the trained model includes a convolutional neural network or a deep convolutional neural network trained using an adversarial generation network training method.

9. The method according to claim 8, wherein the trained model is a generation model configured to generate a simulated SEM image for a given design pattern.

10. The method according to claim 1, wherein the image noise is Gaussian noise, white noise, or salt-and-pepper noise characterized by user-specified parameters.

11. The method of claim 1 , wherein the denoising model comprises a second machine learning model.

12. The method of claim 1 , wherein the design pattern is in a graphics data signal (GDS) file format.

13. obtaining a captured SEM image of the patterned substrate; running a trained denoising model using the captured SEM image as an input image to generate a denoised SEM image; The method of claim 1 further comprising:

14. The method of claim 13 , further comprising updating the denoising model based on captured images of the patterned substrate.

15. One or more non-transitory computer-readable media storing instructions that, when executed by a processor, converting a design pattern, which is a design layout in the data, into a first set of simulated SEM images; One or more non-transitory computer-readable media that cause the processor to implement a method of training a denoising model based on the first set of simulated SEM images and image noise, the denoising model being operable to generate a denoised image of an input image.

Citation Information

Patent Citations

  • Noise removal device, and method and program therefor

    JP2012059118A

  • Image noise reduction method using forward propagation type neural network

    JP2019008599A

  • Image evaluation method and image evaluation device

    JP2019129169A

  • Generating simulated images from design information

    US20170148226A1