Defect detection system based on deep learning and computer implementation method
Through deep learning models, simulated design data images are generated from high-resolution images, and combined with high- and low-resolution imaging systems, the accuracy and efficiency of defect detection in semiconductor manufacturing are solved, and efficient defect recognition is achieved.
Patent Information
- Application Number
- CN202180043756.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-01-21
- Filing Date
- 2021-07-27
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2041-07-27
AI Technical Summary
In the prior art In the semiconductor manufacturing process, defect detection is difficult, especially due to the difficulty in identifying image noise and subtle defects, resulting in false alarms and missed detection, and the current method consumes additional scanning time.
Deep learning model is used to generate grayscale simulated design data images from high-resolution images, and detect defects on samples by simulating binary design data images, combining high-resolution and low-resolution imaging systems to improve detection accuracy.
Improve the accuracy and efficiency of defect detection, reduce false alarms and missed detection, and reduce the cost of additional scanning time.
Smart Images

Figure CN115769254B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to methods and systems for detecting defects on a sample. Background Art
[0002] The following description and examples are not admitted to be prior art by virtue of their inclusion in this section.
[0003] The fabrication of semiconductor devices, such as logic and memory devices, typically involves processing a substrate, such as a semiconductor wafer, using a number of semiconductor fabrication processes to form the various features and multiple levels of the semiconductor device. For example, photolithography is a semiconductor fabrication process that involves transferring a pattern from a mask to a resist disposed on a semiconductor wafer. Additional examples of semiconductor fabrication processes include, but are not limited to, chemical mechanical polishing (CMP), etching, deposition, and ion implantation. Multiple semiconductor devices can be fabricated in an arrangement on a single semiconductor wafer and then separated into individual semiconductor devices.
[0004] Inspection processes are used at various steps in the semiconductor manufacturing process to detect defects on the wafer, driving higher yields and, therefore, higher profits in the manufacturing process. Inspection has always been an important part of manufacturing semiconductor devices. However, as semiconductor device sizes decrease, inspection becomes even more critical to successfully manufacturing acceptable semiconductor devices, as even small defects can cause device failure.
[0005] Inspection results are typically reviewed using scanning electron microscope (SEM) images for defect classification. One of the key steps in this process is defect detection within the SEM image. Current methods for detecting defects in SEM inspection images use the SEM image itself and possibly a reference SEM image. When using machine learning algorithms, this can also include designing images to enhance detection and improve performance. Defect detection is often difficult due to factors such as image noise and the fineness of defect patterns relative to normal patterns. Consequently, one drawback of currently used methods includes issues caused by image and pattern noise, which can result in missed defects or false positives. Subtracting SEM reference images can assist, but at the cost of additional scanning time.
[0006] Accordingly, it would be advantageous to develop systems and methods for detecting defects on samples that do not suffer from one or more of the disadvantages described above. Summary of the Invention
[0007] The following description of various embodiments should not be construed in any way as limiting the subject matter of the appended claims.
[0008] One embodiment relates to a system configured to detect defects on a sample. The system includes one or more computer systems and one or more components executed by the one or more computer systems. The one or more components include a deep learning model configured to generate a grayscale analog design data image for a location on the sample from a high-resolution image generated at the location, wherein the high-resolution image is generated at the location by a high-resolution imaging system. The one or more computer systems are configured to generate an analog binary design data image for the location from the grayscale analog design data image. The one or more computer systems are further configured to detect defects at the location on the sample by subtracting the design data for the location from the analog binary design data image. The system may be further configured as described herein.
[0009] Another embodiment relates to a computer-implemented method for detecting defects on a sample. The method includes generating, for a location on the sample, a grayscale analog design data image from a high-resolution image generated at the location. The high-resolution image is generated at the location by a high-resolution imaging system. Generating the grayscale analog design data image is performed by a deep learning model included in one or more components executed by one or more computer systems. The method also includes generating a simulated binary design data image for the location from the grayscale analog design data image. Additionally, the method includes detecting the defect at the location on the sample by subtracting the design data for the location from the simulated binary design data image. Generating the simulated binary design data image and detecting the defect are performed by the one or more computer systems.
[0010] Each of the steps of the method described above may be further performed as described herein. In addition, embodiments of the method described above may include any other steps of any other method described herein. The method described above may be performed by any of the systems described herein.
[0011] Another embodiment relates to a non-transitory computer-readable medium storing program instructions, the program instructions being executable on one or more computer systems for performing a computer-implemented method for detecting defects on a sample. The computer-implemented method includes the steps of the method described above. The computer-readable medium may be further configured as described herein. The steps of the computer-implemented method may be executed as further described herein. Additionally, the computer-implemented method, on which the program instructions are executable, may include any other steps of any other method described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Further advantages of the present invention will become apparent to those skilled in the art with the benefit of the following detailed description of the preferred embodiments and upon reference to the accompanying drawings, in which:
[0013] Figure 1 and 1a is a schematic diagram illustrating a side view of an embodiment of a system configured as described herein;
[0014] Figure 2 and 3 is a block diagram illustrating an embodiment of a deep learning model that may be included in the systems described herein;
[0015] Figure 4 is a flowchart illustrating steps that may be performed by the embodiments described herein; and
[0016] Figure 5 is a block diagram illustrating one embodiment of a non-transitory computer-readable medium storing program instructions for causing a computer system to perform the computer-implemented methods described herein.
[0017] While the invention is susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and described in detail herein. The drawings may not be drawn to scale. However, it should be understood that the drawings and detailed description thereof are not intended to limit the invention to the particular forms disclosed, but on the contrary, the invention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the invention as defined by the appended claims. DETAILED DESCRIPTION
[0018] The terms "design," "design data," and "design information," as used interchangeably herein, generally refer to the physical design (layout) of an IC or other semiconductor device and data derived from the physical design through complex simulations or simple geometric shapes and Boolean operations. Additionally, an image of a reticle and / or its derivatives acquired by a reticle inspection system may be used as a "proxy" or "proxies" for the design. This reticle image or its derivatives may serve as a substitute for the design layout in any of the embodiments described herein using a design. The design may include any other design data or design data proxies described in commonly owned U.S. Patent No. 7,570,796, issued to Zafar et al. on August 4, 2009, and U.S. Patent No. 7,676,077, issued to Kulkarni et al. on March 9, 2010, the contents of which are incorporated by reference as if fully set forth herein. Additionally, the design data may be standard cell library data, integrated layout data, design data for one or more layers, derivatives of design data, and all or part of the chip design data.
[0019] In addition, the "design," "design data," and "design information" described herein refer to information and data generated by semiconductor device designers during the design process and thus can be printed in advance on any physical samples (such as masks and wafers) for good use in the embodiments described herein.
[0020] "Nuances" (which are sometimes used interchangeably with "nuisance defects"), as the term is used herein, are generally defined as defects that are of no concern to the user and / or events that are detected on a sample but are not actual defects on the sample. Nuisances that are not actually defects may be detected as events due to non-defect noise sources on the sample (e.g., particles in metal lines on the sample, signals from underlying layers or materials on the sample, line edge roughness (LER), relatively small critical dimension (CD) variations in patterned features, thickness variations, etc.) and / or due to marginalities of the imaging system itself or its configuration for imaging.
[0021] As used herein, the term "defect of interest (DOI)" can be defined as a defect detected on a sample that is, in fact, an actual defect on the sample. Therefore, DOIs are of interest to users because users are generally concerned about how many and what types of actual defects are on the inspected sample. In some contexts, the term "DOI" is used to refer to a subset of all actual defects on a sample that only includes the actual defects that the user is concerned about. For example, there may be multiple types of DOIs on any given sample, and the user may be more concerned about one or more of the DOIs than one or more other types. However, in the context of the embodiments described herein, the term "DOI" is used to refer to any and all actual defects on a sample.
[0022] Turning now to the drawings, it should be noted that the figures are not drawn to scale. In particular, the proportions of some elements of the figures are greatly exaggerated to emphasize the characteristics of the elements. It should also be noted that the figures are not drawn to the same scale. Elements shown in more than one figure that may be similarly configured have been indicated using the same reference numerals. Unless otherwise specified herein, any of the elements described and shown may comprise any suitable commercially available components.
[0023] One embodiment relates to a system configured to detect defects on a sample. In some embodiments, the sample is a wafer. The wafer can include any wafer known in the semiconductor art. Although some embodiments may be described herein with respect to one or more wafers, the embodiments are not limited to the samples to which they may be applied. For example, the embodiments described herein may be applied to samples such as reticles, flat panels, personal computer (PC) boards, and other semiconductor samples.
[0024] An embodiment of this system is shown in Figure 1The system includes one or more computer subsystems (e.g., computer subsystems 36 and 102) and one or more components 100 executed by the one or more computer subsystems. The one or more components include a deep learning (DL) model 104, which is configured as further described herein. Figure 1 In the embodiment shown in FIG, the system includes a tool 10, which may include a high-resolution imaging subsystem and / or a low-resolution imaging subsystem. In some embodiments, the tool is configured as an optical (light-based) inspection tool. However, the tool can be configured as another type of inspection or other imaging tool as further described herein.
[0025] As used herein, the term "low resolution" is generally defined as a resolution that is unable to resolve all patterned features on a sample. For example, if the size of some patterned features on a sample is large enough to make them resolvable, then they can be resolved at a "low" resolution. However, the term "low resolution" as used herein refers to a resolution that does not make all patterned features on the sample described herein resolvable. In this way, the term "low resolution" as used herein cannot be used to generate information about the patterned features on the sample that is sufficient for applications such as defect inspection (which may include defect classification and / or verification) and metrology. In addition, as those terms are used herein, a "low resolution" imaging system, subsystem, tool, etc. generally refers to an imaging system, subsystem, tool, etc. with a relatively low resolution (e.g., lower than a defect inspection and / or metrology system) in order to have a relatively fast throughput. In this way, a "low resolution image" may also be generally referred to as a high throughput or HT image. Different types of imaging systems can be configured for low resolution. For example, to produce images at a higher throughput, the e / p and frame rate of a scanning electron microscope (SEM) may be lower than its highest achievable resolution, thereby resulting in lower quality SEM images.
[0026] "Low resolution" can also be referred to as "low resolution" because it is lower than the "high resolution" described herein. As used herein, the term "high resolution" can generally be defined as a resolution that can resolve all patterned features of a sample with relatively high accuracy. In this way, all patterned features on a sample can be resolved at high resolution, regardless of their size. Thus, as used herein, "high resolution" can be used to generate information about the patterned features of a sample sufficient for applications such as defect inspection (which may include defect classification and / or verification) and metrology. Furthermore, as used herein, "high resolution" refers to a resolution that is not typically used by inspection systems during normal operation, as such inspection systems are configured to sacrifice resolution capability in order to increase throughput. "High-resolution images" may also be referred to in the art as "high-sensitivity images," which is another term for "high-quality images." Different types of imaging systems can be configured for high resolution. For example, to produce high-quality images, the e / p, frame rate, etc. of an SEM can be increased, which produces high-quality SEM images but significantly reduces throughput. These images are then considered "high-sensitivity" images because they can be used for high-sensitivity defect detection.
[0027] The high- and / or low-resolution imaging subsystem includes at least an energy source and a detector. The energy source is configured to generate energy directed toward a sample. The detector is configured to detect energy from the sample and generate an output (e.g., an image) in response to the detected energy. Various configurations of the high- and / or low-resolution imaging subsystem are further described herein.
[0028] In general, the high-resolution imaging subsystem and the low-resolution imaging subsystem may share some or none of the tool's image-forming elements. For example, the high-resolution imaging subsystem and the low-resolution imaging subsystem may share the same energy source and detector, and may vary one or more parameters of the energy source, detector, and / or other image-forming elements of the tool depending on whether the high-resolution imaging subsystem or the low-resolution imaging subsystem is generating an image of the sample. In another example, the high-resolution imaging subsystem and the low-resolution imaging subsystem may share some of the tool's image-forming elements (e.g., an energy source) and may have other non-shared image-forming elements (e.g., a separate detector). In a further example, the high-resolution imaging subsystem and the low-resolution imaging subsystem may share no common image-forming elements. In this example, the high-resolution imaging subsystem and the low-resolution imaging subsystem may each have their own energy source, detector, and any other image-forming elements not used or shared by the other imaging subsystem.
[0029] exist Figure 1 In the embodiment of the system shown in FIG, the high-resolution imaging subsystem includes an illumination subsystem configured to direct light to the sample 12. The illumination subsystem includes at least one light source. For example, Figure 1As shown in FIG, the illumination subsystem includes a light source 14. The illumination subsystem is configured to direct light to the sample at one or more incident angles, which may include one or more oblique angles and / or one or more normal angles. Figure 1 , light from light source 14 is directed through optical element 16 to beam splitter 18. Beam splitter 18 directs the light from optical element 16 to lens 20, which focuses the light at a normal angle of incidence to sample 12. The angle of incidence may include any suitable angle of incidence, which may vary depending on, for example, the characteristics of the sample.
[0030] The illumination subsystem may be configured to direct light to the sample at different angles of incidence at different times. For example, the tool may be configured to change one or more characteristics of one or more elements of the illumination subsystem so that the light may be directed differently to the sample. Figure 1 In this example, the tool may be configured to use one or more apertures (not shown) to control the angle at which light is directed from lens 20 to the sample.
[0031] In one embodiment, light source 14 may comprise a broadband light source, such as a broadband plasma (BBP) light source. In this manner, the light generated by the light source and directed toward the sample may comprise broadband light. However, the light source may comprise any other suitable light source, such as a laser, which may comprise any suitable laser known in the art and may be configured to generate light of any suitable wavelength known in the art. Additionally, the laser may be configured to generate monochromatic or nearly monochromatic light. In this manner, the laser may be a narrowband laser. The light source may also comprise a polychromatic light source, which generates light at multiple discrete wavelengths or bands.
[0032] The light from the beam splitter 18 can be focused onto the sample 12 by the lens 20. Although the lens 20 is Figure 1 2 is shown as a single refractive optical element, but it should be understood that, in practice, lens 20 may include several refractive and / or reflective optical elements that, in combination, focus light onto the sample. The illumination subsystem of the high-resolution imaging subsystem may include any other suitable optical elements (not shown). Examples of such optical elements include, but are not limited to, polarization components, spectral filters, spatial filters, reflective optical elements, apodizers, beam splitters, apertures, and the like, which may include any such suitable optical elements known in the art. Additionally, the tool may be configured to change one or more of the elements of the illumination subsystem based on the type of illumination used for imaging.
[0033] Although the high-resolution imaging subsystem is described above as including one light source and one illumination channel in its illumination subsystem, the illumination subsystem may include more than one illumination channel, one of which may include, for example, Figure 1, and the other of the illumination channels (not shown) may include similar elements, which may be configured differently or identically, or may include at least a light source and possibly one or more other components, such as those described further herein. If light from different illumination channels is directed to the sample simultaneously, one or more characteristics (e.g., wavelength, polarization, etc.) of the light directed to the sample by the different illumination channels may differ, such that light resulting from illumination of the sample by the different illumination channels can be distinguished from one another at a detector.
[0034] In another example, the lighting subsystem may include only one light source (e.g., Figure 1 ) and the light from the light source can be separated into different paths (e.g., based on wavelength, polarization, etc.) by one or more optical elements (not shown) of the illumination subsystem. The light in each of the different paths can then be directed to the sample. Multiple illumination channels can be configured to direct light to the sample simultaneously or at different times (e.g., when different illumination channels are used to sequentially illuminate the sample). In another example, the same illumination channel can be configured to direct light to samples with different characteristics at different times. For example, in some examples, the optical element 16 can be configured as a spectral filter and the properties of the spectral filter can be changed in a variety of different ways (e.g., by replacing the spectral filter) so that light of different wavelengths is not directed to the sample at the same time. The illumination subsystem can have any other suitable configuration known in the art for directing light with different or the same characteristics to the sample at different or the same angles of incidence, either sequentially or simultaneously.
[0035] The tool may also include a scanning subsystem configured to scan light over the sample. For example, the tool may include a stage 22 on which the sample 12 is positioned during imaging. The scanning subsystem may include any suitable mechanical and / or robotic assembly (including the stage 22) that may be configured to move the sample so that light can be scanned over the sample. Additionally or alternatively, the tool may be configured so that one or more optical elements of the high-resolution imaging subsystem perform some of the scanning of light over the sample. The light may be scanned over the sample in any suitable manner, such as in a serpentine path or in a spiral path.
[0036] The high-resolution imaging subsystem further includes one or more detection channels. At least one of the one or more detection channels includes a detector configured to detect light from the sample due to illumination of the sample by the illumination subsystem and to generate an output in response to the detected light. For example, Figure 1The high-resolution imaging subsystem shown in includes a detection channel formed by a lens 20, an element 26, and a detector 28. Although the high-resolution imaging subsystem is described herein as including a common lens for both illumination and collection / detection, the illumination subsystem and the detection channel may include separate lenses (not shown) for focusing in the case of illumination and collecting in the case of detection. The detection channel may be configured to collect and detect light at different collection angles. For example, one or more holes (not shown) positioned in the path of the light from the sample may be used to select and / or change the angle of light collected and detected by the detection channel. The light from the sample detected by the detection channel of the high-resolution imaging subsystem may include specularly reflected light and / or scattered light. In this way, Figure 1 The high-resolution imaging subsystem shown in can be configured for dark-field (DF) and / or bright-field (BF) imaging.
[0037] Element 26 may be a spectral filter, an aperture, or any other suitable element or combination of elements that can be used to control the light detected by detector 28. Detector 28 may include any suitable detector known in the art, such as a photomultiplier tube (PMT), a charge-coupled device (CCD), and a time-delay integration (TDI) camera. The detector may also include a non-imaging detector or an imaging detector. If the detector is a non-imaging detector, the detector may be configured to detect a specific characteristic of light (e.g., intensity), but may not be configured to detect such characteristic as a function of position within the imaging plane. Thus, the output generated by the detector may be a signal or data rather than an image signal or image data. A computer subsystem (e.g., computer subsystem 36) may be configured to generate an image of the sample from the non-imaging output of the detector. However, the detector may be configured as an imaging detector configured to generate an imaging signal or image data. Thus, the high-resolution imaging subsystem may be configured to generate the images described herein in a variety of ways.
[0038] The high-resolution imaging subsystem may also include another detection channel. For example, light from the sample collected by lens 20 may be directed through beam splitter 18 to beam splitter 24, which may transmit a portion of the light to optical element 26 and reflect another portion of the light to optical element 30. Optical element 30 may be a spectral filter, an aperture, or any other suitable element or combination of elements that can be used to control the light detected by detector 32. Detector 32 may include any of the detectors described above. Different detection channels of the high-resolution imaging subsystem may be configured to produce different images of the sample (e.g., images of the sample produced using light having different characteristics (e.g., polarization, wavelength, etc., or some combination thereof)).
[0039] In various embodiments, the detection channel formed by lens 20, optical element 30, and detector 32 may be part of a low-resolution imaging subsystem of the tool. In this case, the low-resolution imaging subsystem may include the same illumination subsystem as the high-resolution imaging subsystem described in detail above (e.g., an illumination subsystem including light source 14, optical element 16, and lens 20). Thus, the high-resolution imaging subsystem and the low-resolution imaging subsystem may share a common illumination subsystem. However, the high-resolution imaging subsystem and the low-resolution imaging subsystem may include different detection channels, each of which is configured to detect light from the sample due to illumination by the shared illumination subsystem. In this way, the high-resolution detection channel may include lens 20, optical element 26, and detector 28, and the low-resolution detection channel may include lens 20, optical element 30, and detector 32. In this way, the high-resolution detection channel and the low-resolution detection channel may share a common optical element (lens 20) but also have non-shared optical elements.
[0040] The detection channels of the high-resolution imaging subsystem and the low-resolution imaging subsystem can be configured to produce high-resolution and low-resolution sample images, respectively, even though they share an illumination subsystem. For example, optical elements 26 and 30 can be differently configured apertures and / or spectral filters that control the portion of light detected by detectors 28 and 32, respectively, thereby controlling the resolution of the images produced by detectors 28 and 32, respectively. In various examples, detector 28 of the high-resolution imaging subsystem can be selected to have a higher resolution than detector 32. The detection channels can be configured in any other suitable manner to have different resolution capabilities.
[0041] In another embodiment, the high-resolution imaging subsystem and the low-resolution imaging subsystem may share all of the same image-forming elements. For example, both the high-resolution imaging subsystem and the low-resolution imaging subsystem may share an illumination subsystem formed by light source 14, optical element 16, and lens 20. The high-resolution imaging subsystem and the low-resolution imaging subsystem may also share the same detection channel or channels (e.g., one formed by lens 20, optical element 26, and detector 28 and / or the other formed by lens 20, optical element 30, and detector 32). In this embodiment, one or more parameters or characteristics of any of these image-forming elements may be varied depending on whether a high-resolution image or a low-resolution image is to be generated for the sample. For example, the numerical aperture (NA) of lens 20 may be varied depending on whether a high-resolution image or a low-resolution image is to be generated for the sample.
[0042] In further embodiments, the high-resolution imaging subsystem and the low-resolution imaging subsystem may not share any image forming elements. For example, the high-resolution imaging subsystem may include the image forming elements described above, which may not be shared by the low-resolution imaging subsystem. Instead, the low-resolution imaging subsystem may include its own illumination and detection subsystem. In this example, Figure 1 , the low-resolution imaging subsystem may include an illumination subsystem including a light source 38, an optical element 40, and a lens 44. Light from the light source 38 passes through the optical element 40 and is reflected by the beam splitter 42 to the lens 44, which directs the light to the sample 12. Each of these image forming elements may be configured as described above. The illumination subsystem of the low-resolution imaging subsystem may be further configured as described herein. The sample 12 may be placed on the stage 22, which may be configured as described above to cause light to be scanned over the sample during imaging. In this way, even though the high-resolution imaging subsystem and the low-resolution imaging subsystem do not share any image forming elements, they may share other elements of the tool, such as the stage, scanning subsystem, power supply (not shown), housing (not shown), and the like.
[0043] The low-resolution imaging subsystem may also include a detection channel formed by a lens 44, an optical element 46, and a detector 48. Light from the sample due to illumination by the illumination subsystem may be collected by lens 44 and directed through a beam splitter 42, which transmits the light to the optical element 46. The light that passes through the optical element 46 is then detected by the detector 48. Each of these image forming elements may be further configured as described above. The detection channel and / or detection subsystem of the low-resolution imaging subsystem may be further configured as described herein.
[0044] It should be noted that this article provides Figure 1The present invention generally illustrates the configuration of a high-resolution imaging subsystem and a low-resolution imaging subsystem that may be included in a tool or that may generate images used by the systems or methods described herein. The configuration of the high-resolution imaging subsystem and the low-resolution imaging subsystem described herein may be varied to optimize the performance of the high-resolution imaging subsystem and the low-resolution imaging subsystem, as is typically done when designing commercial tools. In addition, the systems described herein may be implemented using existing systems (e.g., by adding the functionality described herein to an existing system), such as the Altair and 29xx / 39xx series tools available from KLA, Milpitas, Calif. For some such systems, the embodiments described herein may be provided as optional functionality of the system (e.g., in addition to other functionality of the system). Alternatively, the tools described herein may be designed "from the ground up" to provide a completely new inspection or other tool.
[0045] The system also includes one or more computer subsystems configured to acquire images of the sample generated by the high-resolution imaging subsystem and the low-resolution imaging subsystem. For example, the computer subsystem 36 can be coupled to the detector of the tool in any suitable manner (e.g., via one or more transmission media, which can include "wired" and / or "wireless" transmission media) so that the computer subsystem can receive the output or image generated by the detector for the sample. The computer subsystem 36 can be configured to use the output or image generated by the detector to perform a number of functions described further herein.
[0046] Figure 1 The computer subsystem shown in FIG. 1 (as well as other computer subsystems described herein) may also be referred to herein as a computer system. Each of the computer subsystems or systems described herein may take various forms, including personal computer systems, graphics computers, mainframe computer systems, workstations, network appliances, Internet appliances, or other devices. In general, the term "computer system" may be broadly defined to encompass any device having one or more processors that execute instructions from a storage medium. A computer subsystem or system may also include any suitable processor known in the art, such as a parallel processor. Additionally, a computer subsystem or system may include a computer platform with high-speed processing and software, as a stand-alone or networked tool.
[0047] If the system includes more than one computer subsystem, the different computer subsystems may be coupled to each other so that images, data, information, instructions, etc. may be transmitted between the computer subsystems. For example, computer subsystem 36 may be coupled to computer subsystem 102 by any suitable transmission medium, such as by Figure 1, the transmission medium may include any suitable wired and / or wireless transmission medium known in the art. Two or more such computer subsystems may also be effectively coupled by a shared computer-readable storage medium (not shown).
[0048] Although the high-resolution imaging subsystem and the low-resolution imaging subsystem are described above as optical or light-based imaging subsystems, the high-resolution imaging subsystem and the low-resolution imaging subsystem may also or instead include an electron beam imaging subsystem configured to produce an electron beam image of the sample. In this embodiment, the high-resolution imaging system is configured as an electron beam imaging system. The electron beam imaging system may be configured to direct electrons to the sample or scan electrons over the sample and detect electrons from the sample. Figure 1a In this embodiment shown in , the electron beam imaging system includes an electron column 122 coupled to a computer subsystem 124 .
[0049] Also like Figure 1a , the electron column includes an electron beam source 126 configured to generate electrons that are focused by one or more elements 130 to a sample 128. The electron beam source may include, for example, a cathode source or emitter tip, and the one or more elements 130 may include, for example, a gun lens, an anode, a beam limiting aperture, a gate valve, a beam current selecting aperture, an objective lens, and a scanning subsystem, all of which may include any such suitable elements known in the art.
[0050] Electrons returning from the sample (eg, secondary electrons) may be focused by one or more elements 132 onto a detector 134. One or more elements 132 may include, for example, a scanning subsystem, which may be the same scanning subsystem included in element 130.
[0051] The electron column may include any other suitable elements known in the art. Additionally, the electron column may be further configured as described in U.S. Patent No. 8,664,594 issued to Jiang et al. on April 4, 2014, U.S. Patent No. 8,692,204 issued to Kojima et al. on April 8, 2014, U.S. Patent No. 8,698,093 issued to Gubbens et al. on April 15, 2014, and U.S. Patent No. 8,716,662 issued to MacDonald et al. on May 6, 2014, which are incorporated by reference as if fully set forth herein.
[0052] Although the electron column Figure 1aThe electron beam imaging system is shown as being configured so that electrons are directed to the sample at an oblique angle of incidence and returned from the sample at another oblique angle, but it should be understood that the electron beam can be directed to and detected from the sample at any suitable angle. In addition, the electron beam imaging system can be configured to use multiple modes to generate images of the sample as further described herein (e.g., with different illumination angles, collection angles, etc.). The multiple modes of the electron beam imaging system can differ in any image generation parameter. Figure 1a The electron column shown in the figure can also be configured to be used as a high-resolution imaging subsystem and a low-resolution imaging subsystem in any suitable manner known in the art (for example, by changing one or more parameters or characteristics of one or more elements included in the electron column so that a high-resolution image or a low-resolution image can be produced for the sample).
[0053] As described above, the computer subsystem 124 can be coupled to a detector 134. The detector can detect electrons returning from the surface of the sample, thereby forming an electron beam image of the sample. The electron beam image can include any suitable electron beam image. The computer subsystem 124 can be configured to use the output generated by the detector 134 to perform one or more functions further described herein for the sample. Figure 1a The electron beam imaging system shown in FIG. 1 may be further configured as described herein.
[0054] It should be noted that this article provides Figure 1a To generally illustrate the configuration of an electron beam imaging system that may be included in the embodiments described herein. As with the optical imaging subsystem described above, the electron beam imaging system configuration described herein can be varied to optimize the performance of the imaging system, as is typically done when designing commercial imaging systems. In addition, the systems described herein can be implemented using existing systems (e.g., by adding the functionality described herein to an existing system) (e.g., tools available from KLA-Tencor). For some such systems, the embodiments described herein can be provided as optional functionality of the system (e.g., in addition to other functionality of the system). Alternatively, the systems described herein can be designed "from the ground up" to provide an entirely new system.
[0055] Although the imaging system is described above as a light or electron beam imaging system, the imaging system may be an ion beam imaging system. Figure 1a , except that the electron beam source may be replaced by any suitable ion beam source known in the art. Additionally, the imaging system may be any other suitable ion beam-based imaging system, such as those included in commercially available focused ion beam (FIB) systems, helium ion microscopy (HIM) systems, and secondary ion mass spectrometry (SIMS) systems.
[0056] Although the imaging system is described above as including a high-resolution imaging subsystem and a low-resolution imaging subsystem based on optics, electron beams, or charged particle beams, the high-resolution imaging system and the low-resolution imaging system do not necessarily need to use the same type of energy. For example, the high-resolution imaging system may be an electron beam imaging system, while the low-resolution imaging system may be an optical imaging system based on light. Imaging systems using different types of energy can be combined into a single tool in any suitable manner known in the art.
[0057] As mentioned above, the imaging system can be configured to direct energy (e.g., light, electrons) to and / or scan energy over a physical version of the sample, thereby producing an actual image of the physical version of the sample. In this way, the imaging system can be configured as a "real" imaging system rather than a "virtual" system. However, Figure 1 The storage medium (not shown) and computer subsystem 102 shown in the example can be configured as a "virtual" system. Systems and methods configured as a "virtual" inspection system are described in commonly assigned U.S. Patent No. 8,126,255, issued to Bhaskar et al. on February 28, 2012, and U.S. Patent No. 9,222,895, issued to Duffy et al. on December 29, 2015, both of which are incorporated by reference as if fully set forth herein. The embodiments described herein can be further configured as described in these patents.
[0058] As further mentioned above, the imaging system can be configured to generate images of a sample in a variety of modes. Generally speaking, a "mode" can be defined by the values of parameters of the imaging system used to generate the image of the sample, or the output used to generate the image of the sample. Thus, different modes can differ in the value of at least one of the imaging parameters of the imaging system. For example, in an optical imaging system, different modes can use different wavelengths of light for illumination. For different modes, the modes can differ in the illumination wavelength, as further described herein (e.g., by using different light sources, different spectral filters, etc.). Both high-resolution imaging systems and low-resolution imaging systems can be capable of generating outputs or images for samples having different modes.
[0059] Generally speaking, as further described herein, embodiments can be configured for defect detection using simulated design data images (e.g., database images) generated from high-resolution images (e.g., SEM inspection images) and defect information from inspection (e.g., from an optical inspector). A deep learning (DL) model is configured to generate a grayscale simulated design data image for a location on a sample from a high-resolution image generated at that location. The high-resolution image was generated at that location by a high-resolution imaging system. In this way, embodiments described herein are configured for image-to-image transformation to render a simulated design or database image directly from a defective high-resolution image (e.g., a defective SEM image).
[0060] In some embodiments, the locations are selected from locations on the sample where one or more defects were previously detected. For example, the embodiments described herein can be used for defect inspection, where defects previously detected on a sample can be actually re-detected (rather than simply detected) using high-resolution images generated by a defect inspection system. Specifically, the term "inspection" generally refers to a process in which an area on a sample is scanned to search for any defects that may be present but whose presence or absence was not interrogated prior to inspection. Thus, inspection is not performed based on any information regarding any defects previously detected on the layer of the sample being inspected. In contrast, "defect inspection," as a term commonly used in the art, refers to a process in which discrete locations on a sample where defects have been detected are revisited to gather more information (e.g., higher-resolution images) so that the presence of the defects can be verified and, if re-detected, so that additional information can be used to determine additional information about the defects, such as defect type, defect characteristics, and so on. Because the embodiments described herein can be performed for both defect detection (for locations where inspection was not previously performed) or re-inspection (for locations where defects were previously detected), inspection or re-inspection will be referred to simply as inspection.
[0061] In this manner, the steps described herein may be performed separately and independently for one or more locations on a sample where defects have been detected (e.g., by inspection, another defect review process, etc.). However, the steps described herein may also or alternatively be performed at locations where defects have not been previously detected and / or where the presence of defects is unknown despite inspection having been performed on the sample or because inspection has not yet been performed on the sample. In this manner, the embodiments described herein may be used in inspection-type applications where at least a portion of a sample (that has not been previously inspected) is inspected for defects and / or where defects are searched for in locations with otherwise unknown defects. In such embodiments, the high-resolution image used for the steps described herein may be a high-resolution image generated by a high-resolution inspection system (e.g., an electron beam system configured for high-resolution inspection of a sample or another system in a high-resolution mode for inspection). In any case, although some steps and / or embodiments are described herein with respect to locations, the steps described herein may be performed separately and independently for different locations on a sample.
[0062] The location or locations at which the steps described herein are performed can be selected from all defect locations in any suitable manner. For example, the steps described herein can be performed for all locations where defects are detected. However, due to the number of defects typically detected on a sample, this approach may be impractical. In this manner, the locations at which the steps described herein are performed can be selected to be fewer than (and typically substantially fewer than) all locations where defects are detected. The defect locations can be selected (i.e., sampled) in any suitable manner (e.g., randomly, based on one or more characteristics of the defect, based on the defect location, based on a predetermined distribution of locations as a function of the characteristics of the defect, the defect location, or any other information determined through inspection). The embodiments described herein can perform such defect location selection. However, the embodiments described herein can receive the selected defect locations from another method or system for selecting defect locations (e.g., an inspection tool not included in the system).
[0063] The location at which the steps described herein are performed may also be determined by the embodiments described herein and / or received from another system or method. For example, the location of a detected defect on a sample (and any information that may be generated for the defect and / or the location) may be obtained by the embodiments described herein from another method or system that determines the location (e.g., through inspection or by generating a sampling plan for the embodiments described herein) or a storage medium where the location has been stored by another method or system. However, the embodiments described herein may determine the location (e.g., if the embodiments described herein perform inspection of the sample to identify the location of the defect detected on the sample, or if the embodiments described herein generate a sampling plan for the location at which the steps described herein are performed on the sample).
[0064] In one such embodiment, one or more defects are detected by a light-based inspection system at a location on a sample. For example, the location at which the steps described herein are performed may be the location of a defect detected by light-based inspection of the sample. The embodiments described herein may or may not be configured to perform such inspection. For example, Figure 1 The light-based tools shown in and described above may be used in an inspection mode (e.g., one or more modes performed using a low-resolution imaging subsystem) in which the light-based tools scan light over a sample while detecting light from the sample. Figure 1 The computer system shown in [ 15 ] can use the output responsive to the detected light to detect defects on a sample. In one such example, the computer system can apply a defect detection method or algorithm to the output. In the simplest version of a suitable defect detection method or algorithm, the output can be compared to a defect detection threshold. Output above the threshold can be designated as a defect (or potential defect), and output below the threshold can be not designated as a defect. Of course, the computer system can use the output generated by any of the tools described herein to detect defects on a wafer in any other suitable manner.
[0065] In another embodiment, one or more defects are detected at locations on a sample by a high-resolution imaging system. For example, a high-resolution imaging system can be used in inspection-type applications where defects are detected in much smaller areas of the sample due to the much lower throughput of high-resolution imaging systems compared to systems typically used for inspection, and / or inspection is performed only at discrete areas on the sample. In this way, a high-resolution imaging system can be used to scan an area where the defect rate is unknown and determine whether a defect is located in that area. In this way, inspection of a sample can be performed using a high-resolution imaging system in a high-resolution imaging mode. However, as further described above, the imaging systems described herein can be configured for use in multiple modes. For example, both the light and electron beam tools described herein are capable of multi-mode imaging, one of which can be a low-resolution mode and the other a high-resolution mode. In this way, a high-resolution imaging system can perform inspection-type functions on a sample using a low-resolution mode and then use a high-resolution mode to acquire a high-resolution image of a selected location at which the steps described herein will be performed.
[0066] In some embodiments, the high-resolution imaging system is configured as an electron beam imaging system. For example, an electron beam imaging system (e.g. Figure 1a and described above) can produce high-resolution images at the locations on the sample where the steps described herein will be performed.
[0067] In some embodiments, the system includes a high-resolution imaging system. Figure 1 and 1aAs shown in , one or more computer systems of a system can be coupled to a high-resolution imaging system included in the system. In this manner, the steps described herein can be performed by a computer system on a tool (e.g., on an inspection tool that detects defects on a sample during an inspection process and / or on a defect inspection tool that re-detects defects on a sample during a defect review process). However, the one or more computer systems described herein need not be included in a system that includes an imaging system. For example, a computer system described herein can acquire high-resolution images of selected locations from a high-resolution imaging system that generates the images, or from a storage medium (such as described further herein), where the high-resolution imaging system stores the high-resolution images. In this manner, a computer system can be part of a system that does not have or need to have sample processing capabilities. Thus, the computer system can be configured to perform the steps described herein in a manner that is generally referred to as "off-tool" because the steps are not performed "on" the imaging tool.
[0068] A DL model can be a deep neural network with a set of weights determined based on the data used to train it. A neural network can generally be defined as a computational method based on a relatively large collection of neural units, loosely modeling the way biological brains use relatively large clusters of biological neurons connected by axons to solve problems. Each neural unit is connected to many other neural units, and links can enforce or inhibit their influence on the activation state of connected neural units. These systems are self-learning and trained, rather than explicitly programmed, and excel in areas where solutions or feature detection are difficult to express using traditional computer programs.
[0069] Neural networks typically consist of multiple layers, with signal paths traversing from front to back. These layers perform numerous algorithms or transformations. Generally speaking, the number of layers is unimportant and depends on the use case. For practical purposes, a suitable number of layers ranges from two to dozens. Modern neural network projects typically utilize thousands to millions of neural units and millions of connections. The goal of neural networks is to solve problems in the same way as the human brain, albeit at a much more abstract level.
[0070] The DL models described in this article belong to a class of computing commonly known as machine learning (ML). ML can generally be defined as a type of artificial intelligence (AI) that provides computers with the ability to learn without being explicitly programmed. ML focuses on the development of computer programs that can grow and change on their own when exposed to new data. In other words, ML can be defined as a subfield of computer science that "gives computers the ability to learn without being explicitly programmed." ML explores the study and construction of algorithms that can learn and predict from data - such algorithms overcome the problem of strictly adhering to static program instructions by building models from sample inputs and making data-driven predictions or decisions.
[0071] The DL models described herein may be further configured as described in "Introduction to Statistical Machine Learning," Sugiyama, Morgan Kaufmann, 2016, p. 534; "Discriminative, Generative, and Imitative Learning," Jebara, MIT Thesis, 2002, p. 212; and "Principles of Data Mining (Adaptive Computation and Machine Learning)," Hand et al., MIT Press, 2001, p. 578; the references are incorporated by reference as if fully set forth herein. The embodiments described herein may be further configured as described in these references.
[0072] Generally speaking, "DL" (also known as deep structured learning, hierarchical learning, or deep machine learning) is a branch of ML based on a set of algorithms that attempt to model high-level abstractions in data. In a simple case, there may be two groups of neurons: one that receives input signals and one that sends output signals. When the input layer receives input, it passes a modified version of the input to the next layer. In DL-based models, there are many layers between the input and output (and the layers are not composed of neurons, but it can help to think of it that way), allowing the algorithm to use multiple processing layers composed of multiple linear and nonlinear transformations.
[0073] In DL, observations (e.g., images) can be represented in many ways, such as as a vector of intensity values per pixel, or in a more abstract way as a set of edges, regions of a specific shape, etc. Some representations are better than others at simplifying learning tasks (e.g., face recognition or facial expression recognition). One of the promises of DL is to replace handcrafted features with efficient algorithms for unsupervised or semi-supervised feature learning and hierarchical feature extraction.
[0074] In one embodiment, the DL model is configured as a generative adversarial network. In this way, the image-to-image transformations described herein can be performed by a generative adversarial network (GAN). A "generative" model can generally be defined as a model that has a probabilistic nature. In other words, a "generative" model is not a model that performs forward simulation or rule-based methods, and thus, a physical model of the processes involved in generating the actual image or output (for which the simulated image is being generated) is not required. Instead, as further described herein, the generative model can be learned based on a suitable training data set (because its parameters can be learned). The generative model can be configured to have a DL architecture, which can include multiple layers that perform several algorithms or transformations.
[0075] GANs have recently achieved success in generating simulated realistic images. These networks have been generalized to allow pixel-to-pixel mapping, where an input image is transformed into an output image that is significantly different from the input. These networks have many applications, including scene rendering from primitives and shading. The GAN framework alternates between training two models. The generator is an encoder / decoder network that is trained to produce "fake" images that cannot be distinguished from "real" images by a discriminator network that is fed alternating fake and real images for classification. The discriminator network is in turn trained to correctly classify real and fake instances.
[0076] Two adversarial objectives are described in Figure 2 In particular, Figure 2 The training objectives of the GAN generator (G) and discriminator (D) networks are described using input (x) and true image (y). Figure 2 As shown in , an input image 200 may be input to G 202, which may be configured as further described herein. In one such example, G may include several convolutional and other layers having any suitable configuration and arrangement. The output of G 202 is a fake image 204, which is input to D 206 along with the input image 200. D 206 may also include several convolutional and other layers having any suitable configuration and arrangement. In a similar manner, a real image 210 may be input to D 206 along with the input image 200. Thus, different image tuples (real, input) and (fake, input) may be input to the discriminator at different times. D then learns to classify between real and fake images. For example, when a fake image and an input image are input to the discriminator, D learns to produce a Classification: Fake 208 result. Additionally, when a real image and an input image are input to the discriminator, D learns to produce a Classification: True 212 result. The objective function is expressed as follows, where the GAN generator (G) tries to minimize this function (maximum discriminator uncertainty) using fake images and the discriminator (D) tries to maximize the function (minimum uncertainty) using real and fake images.
[0077] G* =argmin G max D L GAN (G, D) + λL L l(G), where
[0078] L GAN (G, D)
[0079] =E x,y [logD(x,y)]+E x,z [(1-D(x,G(x,z)))]
[0080] Note that the input to the discriminator is conditioned on the input image x, and z is a random variable.Finally, the generator is also taught to minimize the L1 distance between the fake image and the ground truth (e.g., SEM) image.
[0081] The discriminator network architecture can be a standard convolutional neural network classifier. The generator network can be a U-Net architecture using skip connections that are intended to preserve details in the output image by bypassing the encoder / decoder bottleneck. An example of a U-Net architecture is shown in Figure 3 In. Figure 3 As shown in Figure 3 The encoder stack shown as a whole by layers 302 and 304 in FIG. Figure 3 The decoder stack is shown as a whole by layers 308 and 310 in FIG. Although the encoder and decoder stacks are Figure 3 , but the encoder and decoder stacks may include any suitable number of layers. Input image x 300 is input to layer 302, resulting in hidden representation 306 (also commonly referred to as the bottleneck layer) being produced by layer 304, or the last layer in the encoder stack. The hidden representation is input to layer 308, or the first layer of the decoder stack, resulting in layer 310, or the last layer of the decoder stack, producing output image y 312.
[0082] The skip connections in this configuration make the architecture a U-Net architecture. Specifically, the architecture includes skip connections between mirror layers in the encoder and decoder stacks. For example, in Figure 3 In the layer 302 and its mirror layer ( Figure 3 There is a jump connection between layer 304 and its mirror layer ( Figure 3 There are skip connections between layers (shown schematically as mirrored layers 316 in the decoder stack). The skip connections are configured to connect all channels at one encoder layer with all channels at the mirrored encoder layer in the decoder stack. In this way, low-level information shared between input and output can be directly shuttled across the network via the skip connections.
[0083] The DL models described herein may be further configured as Goodfellow et al., "Generative Adversarial Nets," arXiv:1406.2661 v1, June 10, 2014, pp. 1-9, and Isola et al., "Image-to-Image Translation with Conditional Adversarial Networks," arXiv:1611.07004 v2, November 22, 2017, p. 17, which are incorporated by reference as if fully set forth herein. The embodiments described herein may be further configured as described in these references. Although in some embodiments described herein, the DL model is described as a GAN, the DL model is not limited to a GAN and may be constructed using any suitable image-to-image transformation, such as a standalone autoencoder or a traditional model-based approach (e.g., solving a transformation function).
[0084] The GANs described herein perform an inverse function. Typically, a GAN presents a "fake" (i.e., simulated) high-resolution (e.g., SEM) image from a design input. Training is performed using design clips (binary or multi-grayscale of multiple design layers) as training data and high-resolution images corresponding to the design clips as labels (i.e., "ground truth") for the training data. The trained GAN then performs inference using the design clips as input to output a GAN "fake" high-resolution image. In contrast, in the embodiments described herein, the GAN may be a type of "inverse GAN" that presents a "fake" design image from a high-resolution image input. Training of the inverse GAN may be performed using reference high-resolution images (e.g., defect-free SEM images) as training data, with the design clips corresponding to the reference high-resolution images serving as labels for the training data. In this way, the inverse GAN may be trained using a set of defect-free high-resolution images and their corresponding design images as ground truth labels. Inverse GAN training may be performed in other ways, as further described herein or in any other manner known in the art. The trained inverse GAN can then use the high-resolution image (e.g., the defective high-resolution image) as input to perform inference to thereby output a GAN "fake" design cut (e.g., a GAN "fake" defective design cut). In this way, once the inverse GAN is trained for a particular process level, inference is performed by running the trained GAN on the high-resolution image generated at the location described herein. The output of the inverse GAN is a grayscale version of the rendered design image of the input high-resolution image. For example, Figure 4As shown in FIG, a high-resolution image 400 can be input to a reverse GAN (not shown in FIG). Figure 4 ) to thereby generate a grayscale simulation design data image 402.
[0085] In one embodiment, the DL model is configured to transfer artifacts of one or more defects in a high-resolution image to a grayscale analog design data image. For example, the DL model is not trained to make a high-resolution image from a sample look like a defect-free design data image. Instead, while the DL model transfers the high-resolution image from the sample to the grayscale analog design data image (e.g., by reducing the roundness of corners in the high-resolution image, by reducing relatively small line edge roughness, etc.), the grayscale analog design data image will deviate from the design data originally created by the designer due to any defects on the sample and appear in the high-resolution image. In this way, the simulated database image generated from a high-resolution sample image (e.g., a defective SEM image) will contain artifacts of the defects in the high-resolution image (typically missing or additional patterns in the design image). These artifacts can, in turn, be used to detect defects, as further described herein, rather than the original SEM or other high-resolution image. Therefore, one advantage of these embodiments is a significant reduction in image complexity (in the analog design data image compared to the high-resolution image), resulting in a much simpler defect detection algorithm (when performing defect detection using the analog design data image instead of the high-resolution image).
[0086] One or more computer systems are configured to generate an analog binary design data image of the location from a grayscale analog design data image. Figure 4 , the computer system may perform a binarization step 404 on a grayscale analog design data image 402 to thereby generate an analog binary design data image 406. Although one particularly suitable manner for generating the analog binary design data image is further described herein, the computer system may generate the analog binary design data image in any other suitable manner known in the art.
[0087] In one embodiment, generating a simulated binary design data image includes thresholding the grayscale simulated design data image to binarize the grayscale simulated design data image and match the nominal dimensions of pattern features in the grayscale simulated design data image with the nominal dimensions of pattern features in the design data for the sample. For example, the ground truth design image is typically binary or ternary (depending on whether the design image is for only one layer on the sample or two layers on the sample), and the rendered grayscale design image generated by the DL model may have a substantially large histogram peak corresponding to the ground truth image values. To perform inspection using the rendered design image generated by the DL model, the computer system may threshold the rendered design image to binary or ternary values corresponding to the ground truth values. In some such examples, the computer system may generate a histogram of grayscale values in the grayscale simulated design data image and then apply a threshold to the histogram to thereby binarize the image. The threshold is selected to match the pattern width (or other dimension) of the features in the design. In this way, the computer system can threshold the simulated database image to binarize and match the nominal real design pattern width (or other dimension) and then, as further described herein, subtract it from the real design pattern. In other words, in order to use the inverse GAN image for defect detection, the computer system can threshold the design clip generated by the inverse GAN to match the high-resolution image design rule (DR) width.
[0088] In another embodiment, the design data is for patterned features on only one level of the sample. When the design data is for only one level of the sample (e.g., a layer on a wafer or a layer on a reticle), the analog binary design data image may contain only binary values, one for the pattern on the layer (for both the intended pattern and any unintended patterns (i.e., defects or defective portions of the pattern)) and another for the background or unpatterned areas of the layer. For example, Figure 4 , in the simulated binary design data image 406, patterned features are shown with one binary value (as indicated by the white portion of the image) and unpatterned areas are shown with another binary value (as indicated by the dark gray portion of the image). Whether the patterned features or the background are indicated by one binary value or the other binary value is no different from the embodiments described herein.
[0089] In a further embodiment, the design data is for patterned features at different levels of the sample, and the simulated binary design data image includes different grayscales for the patterned features at different levels of the sample. For example, some high-resolution imaging tools can generate images responsive to patterned features at different levels of the sample, one of which is formed below another of the levels. In some instances, imaging patterned features at different levels of the sample can be advantageous. In other instances, it can be problematic (when patterned features at a level different from the level at which defects are detected affect the image in a way that can make defect detection more difficult). Regardless, the computer system can generate a simulated binary design data image such that the patterned features at different levels of the sample are indicated by different grayscales. In this manner, the simulated binary design data image described herein can be a three-level type image, where, in addition to the binary value, another grayscale is used to distinguish features formed at different levels of the sample.
[0090] In such examples, information about the design of the sample or the design data itself can be used in the binarization step to assign multiple grayscales to patterned features at different levels. For example, patterned features formed at different levels of the sample may have one or more different characteristics, such as different nominal (or as designed) sizes, different orientations, different shapes, different spatial relationships between patterned features at the same level or different levels, etc., which can be used to distinguish these patterned features. In this way, using the grayscale analog design data image and the design information or data, a computer system can identify patterned features at different levels of the sample and assign them different grayscales in the analog binary design data image. Once the patterned features at different levels have been identified, this information can be used to identify areas of interest (or areas of interest) at the levels of the sample for which defects are detected by the embodiments described herein.
[0091] The one or more computer systems are configured to detect defects at the locations on the sample by subtracting the design data at the locations from the simulated binary design data image. In this way, the binarized defective reverse GAN design image can be subtracted from the real design (or vice versa) for defect detection. For example, Figure 4 , the computer system may be configured to perform a subtraction step 412, in which design data 410 is subtracted from the simulated binary design data image 406. The result of the subtraction may include a difference image 414 that illustrates any differences between the design data and the simulated binary design data image.
[0092] Detecting defects may also include using the difference image to detect defects on the sample. For example, the computer system may apply a threshold to the difference image, and any pixels in the difference image with a value above the threshold may be identified as defects or defect candidates. The difference image may be used in any other manner to detect defects on the sample. For example, any defect detection method or algorithm that uses a difference image to detect defects on a sample may be applied to the difference image generated by the embodiments described herein.
[0093] Additionally or alternatively, one or more computer systems may be configured to detect defects on a sample in a simulated binary design data image using single image inspection (SID). In this manner, detecting a defect may or may not include subtracting the design data at the location from the simulated binary design data image. Specifically, in SID, defect detection does not require a reference image. Instead, only an image of the potential defect location (i.e., a "test" image) is input to the SID method or system. SID may be performed by embodiments described herein, such as those described in U.S. Patent Application Publication No. 2017 / 0140524 to Karsenti et al., published on May 18, 2017, which is incorporated by reference as if fully set forth herein. The embodiments described herein may be further configured as described in such publication.
[0094] The embodiments described herein perform defect detection in the analog design data (database) space, rather than using the high-resolution images acquired from the sample itself for defect detection. One advantage of this defect detection approach is that the complexity of the image space used for detection is reduced. As a result, simpler detection and nuisance reduction algorithms can be applied. The defect detection described herein can also be used as an alternative to existing defect detection methods or as an additional defect detection method to improve detectability and nuisance suppression. As a completely new defect detection method, the embodiments described herein can be used to extend the capabilities and performance of existing inspection platforms. In addition, by extending capabilities through algorithms (rather than hardware), the roadmap of existing inspection platforms can be expanded.
[0095] In some embodiments, the design data subtracted from the simulated binary design data image is not generated by simulation of the sample or a high-resolution imaging system. For example, because the sample information used for defect detection is a simulated image presented in the design data space, the information subtracted from the image is also an image in the design data space. In this way, unlike many inspection methods in which an image generated from the sample is used for inspection, thereby requiring a reference image in the sample space (e.g., an image from an adjacent area on the sample (e.g., a die, cell, field, etc.) or an image simulated from the design (e.g., based on the simulated characteristics of the sample and a high-resolution imaging system to present an image in the sample space), the embodiments described herein can use the design data for reference used in defect detection without having to obtain a reference image from another area on the sample or modify the design data to create a reference image.
[0096] In one embodiment, detecting defects includes limiting the region of interest in the subtracted result to the region of interest in the simulated binary design data image and applying the detection region threshold to the subtracted result only in the limited region of interest. For example, nuisance pixels in the difference image can be eliminated by limiting the region of interest to the region of interest and applying the region threshold within the region of interest. Defective pixel counts exceeding the region threshold are marked as defects. In this example, if Figure 4 As shown in , region 408 in simulated binary design data image 406 may correspond to a region of interest for the sample. The region of interest may be determined in any suitable manner and may have any suitable configuration. Placement or identification of the region of interest in the simulated binary design data image may be performed in any suitable manner using any of the images described herein (e.g., by aligning one or more of the images with the design data for the sample and using information about the location of the region of interest in the design data to identify the region of interest in one or more of the images aligned with the design data). A corresponding region 416 may be located in difference image 414 resulting from subtraction step 412, and the region of interest for defect detection may be limited to this corresponding region. In this manner, the rendered GAN, binarized design image may be subtracted from the ground-truth design image (or vice versa), and detection may be applied to the region of interest portion of the resulting difference image. Thus, using an inverse GAN for defect detection may include using a region-based threshold for verifying defect detection in the region of interest region.
[0097] In some embodiments, the DL model is configured to generate additional grayscale analog design data images for additional locations on the sample that are known to be defect-free from additional high-resolution images generated at the additional locations, and the one or more computer systems are configured to generate additional analog binary design data images for the additional locations from the additional grayscale analog design data images. The additional locations can be known to be defect-free in any suitable manner known in the art. For example, one or more locations where no defects were detected during an inspection process performed on the sample can be used as the one or more additional locations. In some examples, one or more additional locations can be confirmed as defect-free before being used as additional locations. This confirmation can be performed in any suitable manner, such as by acquiring high-resolution images at the potential additional locations and displaying the images to a user for review and acceptance or rejection as additional locations. Generating the additional grayscale analog design data images can be performed in other manners as further described herein. Alternatively, generating the additional analog binary design data images can be performed as described herein. If these steps are performed for more than one additional location, the steps can be performed separately and independently for each additional location. In some examples, if the steps are performed for more than one additional location on the sample that corresponds to the same location within the design locations, the additional simulated binary design data images may be combined in some manner (e.g., by averaging, median, average, etc.) to produce a combined additional simulated binary design data image, which may then be used in the additional manners described herein.
[0098] In one such embodiment, the design data subtracted from the simulated binary design data image used in the defect detection step includes additional simulated binary design data images. For example, the images generated by the GAN can have several applications to assist in defect review or inspection, including reference image generation for die-to-database type subtraction. In some currently used GAN applications, design clippings are used as input to the GAN, and the GAN generates simulated high-resolution images. However, in the embodiments described herein, the GAN is trained and used to perform the reverse function, i.e., the GAN generates simulated design data images from high-resolution sample image inputs. The additional simulated binary design data images can be used in the detection steps as further described herein.
[0099] In another such embodiment, detecting a defect includes aligning the design data with the simulated binary design data image generated for the location using an additional simulated binary design data image. For example, images generated by a GAN may have several applications to assist in defect inspection, including alignment image generation. In some currently used GAN applications, a design clip is used as input to the GAN, and the GAN generates a fake high-resolution image. However, in the embodiments described herein, the GAN is trained and used to perform the reverse function, i.e., the GAN generates a simulated design data image from a high-resolution sample image input. The additional simulated binary design data image may be aligned with the simulated binary design data image and / or the design data in any suitable manner known in the art, such as via pattern matching.
[0100] In additional embodiments of this invention, one or more computer systems are configured to train a defect classifier using additional simulated binary design data images. For example, images generated by a GAN can have several applications to assist in defect inspection, including training set augmentation for training defect classifiers based on high-resolution images, which is particularly important when generally few training samples are available. In some currently used GAN applications, design clips are used as input to the GAN, and the GAN generates fake high-resolution images. However, in the embodiments described herein, a GAN is trained and used to perform the reverse function: the GAN generates simulated design data images from high-resolution sample image inputs. In some instances, because the additional simulated binary design data images are for known defect-free locations on the sample, the computer system can be configured to augment the additional simulated binary design data images with artificial defects of known DOI and / or known nuisance to create defective additional simulated binary design data images. Modifying a non-defective image to include a defective image can be performed as described in U.S. Patent Application Publication No. 2019 / 0294923, published by Riley et al. on September 26, 2019, and U.S. Patent Application Publication No. 2019 / 0303717, published by Bhaskar et al. on October 3, 2019, which are incorporated by reference as if fully set forth herein. The embodiments described herein can be further configured as described in these publications. However, additional simulated binary design data images can also be used as non-defective images for defect classifier training without augmentation, so that the defect classifier can be trained to distinguish between images containing defects and non-defective images.
[0101] In some embodiments, one or more computer systems are configured to train a defect classifier using one or more of the grayscale analog design data images, the analog binary design data images, the design data, and the subtracted results. For example, it may be advantageous to train the defect classifier using an image stack of known defect locations, the image stack including any one or more of the images described herein for the defect locations. In some such examples, the images may be labeled with ground truth data (in this case, defect classifications) by a user or using another trained defect classifier, which may or may not be of the same type as the defect classifier being trained. For example, the defect classifier used to establish the defect classifications then used for training may be a non-ML-type defect classifier, while the defect classifier trained using this data may be an ML-type defect classifier. After the defect classifier is trained using any one or more of the images described herein, the same one or more types of images may be input to the trained classifier at runtime. For example, if training is performed using grayscale analog design data images and analog binary design data images, then when a defect classifier is used for classification, the images input to the defect classifier may be grayscale analog design data images and defect-detected analog binary design data images.
[0102] The defect classifier trained as described above may include any suitable defect classifier known in the art. The defect classifier may be configured and trained as described in U.S. Patent Application Publication No. 2019 / 0073566 to Brauer, published on March 7, 2019, and U.S. Patent Application Publication No. 2019 / 0073568 to He et al., published on March 7, 2019, which are incorporated by reference as if fully set forth herein. The embodiments described herein may be further configured as described in these publications.
[0103] In one embodiment, one or more computer systems are configured to use the results of the inspection to discover previously unknown defect types on a sample. For example, inspection results are often reviewed using SEM or other high-resolution images for defect classification. This is particularly important during defect discovery due to unknown defects on a sample. Defect discovery can include classifying defects that a trained classifier cannot classify (meaning the trained classifier cannot classify the defect). This classification can be performed using any one or more of the images described herein to locate the unclassifiable defect and in any suitable manner known in the art.
[0104] The embodiments described herein may be further configured as described in the following patent application publications: commonly owned U.S. Patent Application Publication No. 2017 / 0140524 to Karlsandy et al., published on May 18, 2017; U.S. Patent Application Publication No. 2017 / 0148226 to Zhang et al., published on May 25, 2017; U.S. Patent Application Publication No. 2017 / 0193400 to Bhaskar et al., published on July 6, 2017; U.S. Patent Application Publication No. 2017 / 0193680 to Zhang et al., published on July 6, 2017; U.S. Patent Application Publication No. 2017 / 0194126 to Bhaskar et al., published on July 6, 2017; No. 2017 / 0200260, published on July 13, 2017, by Park et al., U.S. Patent Application Publication No. 2017 / 0200264, published on July 13, 2017, by Bhaskar et al., U.S. Patent Application Publication No. 2017 / 0200265, published on July 13, 2017, by Zhang et al., U.S. Patent Application Publication No. 2017 / 0345140, published on November 30, 2017, by Broe et al., U.S. Patent Application Publication No. 2019 / 0073566, published on March 7, 2019, and U.S. Patent Application Publication No. 2019 / 0073568, published on March 7, 2019, by He et al., are incorporated by reference as if fully set forth herein. In addition, the embodiments described herein can be configured to perform any of the steps described in these publications.
[0105] All embodiments described herein may be configured to store the results of one or more steps of the embodiments in a computer-readable storage medium. The results may include any of the results described herein and may be stored in any manner known in the art. The storage medium may include any storage medium described herein or any other suitable storage medium known in the art. After the results have been stored, the results may be accessed in the storage medium and used by any of the method or system embodiments described herein, formatted for display to a user, used by another software module, method or system, etc. to perform one or more functions for a sample or another sample.
[0106] Such functions include, but are not limited to, modifying a process, such as a manufacturing process or step, that has been or will be performed on a sample, in a feedback or feedforward manner, or the like. For example, the computer system may be configured to determine one or more changes to a process performed on a sample and / or to be performed on the sample based on the detected defect. The changes to the process may include any suitable changes to one or more parameters of the process. The computer system preferably determines those changes so that defects can be reduced or prevented on other samples on which the modified process is performed, defects on the sample can be corrected or eliminated in another process performed on the sample, defects can be compensated for in another process performed on the sample, and so on. The computer system may determine such changes in any suitable manner known in the art.
[0107] Those changes can then be sent to Figure 1 A storage medium (not shown) accessible to the semiconductor manufacturing system 108 shown in FIG. 1 or both the computer system and the semiconductor manufacturing system. Figure 1 ). A semiconductor manufacturing system may or may not be part of the system embodiments described herein. For example, the high-resolution imaging system and computer system described herein can be coupled to a semiconductor manufacturing system, for example, via one or more common elements, such as a housing, a power supply, a sample handling device or mechanism, and the like. The semiconductor manufacturing system can include any semiconductor manufacturing system known in the art, such as a lithography tool, an etching tool, a chemical mechanical polishing (CMP) tool, a deposition tool, and the like.
[0108] Each of the embodiments of each of the systems described above may be combined together into a single embodiment.
[0109] Another embodiment relates to a computer-implemented method for detecting defects on a sample. The method includes: for a location on the sample, generating a grayscale analog design data image from a high-resolution image generated at the location. The high-resolution image is generated at the location by a high-resolution imaging system. Generating the grayscale analog design data image is performed by a DL model included in one or more components executed by one or more computer systems, all of which can be configured as further described herein. The method also includes: generating a simulated binary design data image of the location from the grayscale analog design data image. In addition, the method includes: detecting the defect at the location on the sample by subtracting the design data for the location from the simulated binary design data image. Generating the simulated binary design data image and detecting the defect are performed by the one or more computer systems.
[0110] Each of the steps of the method may be performed as further described herein. The method may also include any other steps that can be performed by the systems, computer systems, components and / or DL models described herein. In addition, the method described above may be performed by any of the system embodiments described herein.
[0111] Additional embodiments relate to a non-transitory computer-readable medium storing program instructions executable on one or more computer systems to perform a computer-implemented method for detecting defects on a sample. One such embodiment is shown in Figure 5 In particular, Figure 5 , non-transitory computer-readable medium 500 includes program instructions 502 executable on a computer system 504. The computer-implemented method may include any steps of any method described herein.
[0112] Program instructions 502 implementing methods such as those described herein may be stored on a computer-readable medium 500. The computer-readable medium may be a storage medium such as a magnetic or optical disk, tape, or any other suitable non-transitory computer-readable medium known in the art.
[0113] Program instructions may be implemented in any of a variety of ways, including process-based techniques, component-based techniques, and / or object-oriented techniques, etc. For example, program instructions may be implemented using ActiveX controls, C++ objects, JavaBeans, Microsoft Foundation Classes ("MFC"), SSE (Streaming SIMD Extensions), or other techniques or methods as needed.
[0114] Computer system 504 may be configured according to any of the embodiments described herein.
[0115] In view of this description, further modifications and alternative embodiments of various aspects of the present invention will be apparent to those skilled in the art. For example, methods and systems for detecting defects on samples are provided. Therefore, this description is to be interpreted as illustrative only and is intended to teach those skilled in the art the general manner of implementing the present invention. It should be understood that the forms of the invention shown and described herein are to be considered as currently preferred embodiments. Elements and materials may be substituted for those illustrated and described herein, parts and processes may be reversed, and specific features of the present invention may be used independently, all of which will be apparent to those skilled in the art having the benefit of the description of the present invention. Changes may be made to the elements described herein without departing from the spirit and scope of the present invention as described in the appended claims.
Claims
1. A system configured to detect defects on a sample, comprising: one or more computer systems; and one or more components executed by the one or more computer systems, wherein the one or more components include a deep learning model configured to, for a location on a sample, generate a grayscale analog design data image for the location on the sample from a high-resolution image generated at the location, and wherein the high-resolution image is generated at the location by a high-resolution imaging system; wherein the one or more computer systems are configured to generate an analog binary design data image of the location using the grayscale analog design data image; and The one or more computer systems are further configured to detect defects at the locations on the sample by subtracting the design data at the locations from the simulated binary design data image to obtain a subtracted result.
2. The system of claim 1, wherein the deep learning model is further configured to generate an adversarial network.
3. The system of claim 1 , wherein the deep learning model is further configured to transfer artifacts of one or more defects in the high-resolution image to the grayscale analog design data image.
4. The system of claim 1 , wherein generating the analog binary design data image comprises thresholding the grayscale analog design data image to binarize the grayscale analog design data image and match nominal sizes of pattern features in the grayscale analog design data image with nominal sizes of pattern features in the design data for the sample.
5. The system of claim 1 , wherein the detecting comprises limiting an area of interest in the subtracted result to an area of interest in the simulated binary design data image, the detecting further comprising applying a detection area threshold only to the subtracted result in the limited area of interest.
6. The system of claim 1 , wherein the deep learning model is further configured to generate additional grayscale analog design data images using additional high resolution images generated at additional locations, the additional grayscale analog design data images being for the additional locations on the sample that are known to be defect-free, wherein the one or more computer systems are further configured to generate additional analog binary design data images for the additional locations using the additional grayscale analog design data images, and wherein the design data for the inspection that is subtracted from the analog binary design data images includes the additional analog binary design data images.
7. The system of claim 1 , wherein the deep learning model is further configured to generate additional grayscale analog design data images using additional high resolution images generated at additional locations, the additional grayscale analog design data images being for the additional locations on the sample that are known to be defect-free, wherein the one or more computer systems are further configured to generate additional analog binary design data images for the additional locations using the additional grayscale analog design data images, and wherein the inspecting comprises aligning the design data with the analog binary design data images generated at the locations using the additional analog binary design data images.
8. The system of claim 1 , wherein the deep learning model is further configured to generate additional grayscale analog design data images using additional high resolution images generated at additional locations, the additional grayscale analog design data images being for additional locations on the sample that are known to be defect-free, wherein the one or more computer systems are further configured to generate additional analog binary design data images for the additional locations using the additional grayscale analog design data images, and wherein the one or more computer systems are further configured to train a defect classifier using the additional analog binary design data images.
9. The system of claim 1, wherein the one or more computer systems are further configured to use the subtracted results to train a defect classifier.
10. The system of claim 1, wherein the one or more computer systems are further configured to train a defect classifier using one or more of the grayscale analog design data image, the analog binary design data image, the design data, and the subtracted result.
11. The system of claim 1, wherein the one or more computer systems are further configured to use the results of the inspection to discover previously unknown defect types on the sample.
12. The system of claim 1, wherein the design data subtracted from the simulated binary design data image is not generated by simulation of the sample or the high resolution imaging system.
13. The system of claim 1, wherein the design data and the simulated binary design data image illustrate optical proximity correction features in the design data, the design data being subtracted from the simulated binary design data image.
14. The system of claim 1, wherein the design data and the simulated binary design data image do not account for optical proximity correction features in the design data, the design data being subtracted from the simulated binary design data image.
15. The system of claim 1, wherein the location is selected from a location on the sample where one or more defects were previously detected, and wherein the one or more defects are detected by a light-based inspection system at the location on the sample.
16. The system of claim 1, wherein the location is selected from a location on the sample where one or more defects were previously detected, and wherein the one or more defects are detected by the high-resolution imaging system at the location on the sample.
17. The system of claim 1, wherein the high-resolution imaging system is configured as an electron beam imaging system.
18. The system of claim 1, wherein the system comprises the high-resolution imaging system.
19. The system of claim 1, wherein the design data is used to indicate patterned features at only one level of the sample.
20. The system of claim 1, wherein the design data is used to indicate pattern features at different levels of the sample, and wherein the simulated binary design data image comprises different grayscales of the pattern features at the different levels of the sample.
21. A non-transitory computer-readable medium storing program instructions executable on one or more computer systems to perform a computer-implemented method for detecting defects on a sample, wherein the computer-implemented method comprises: generating, for a location on a sample, a grayscale analog design data image from a high-resolution image generated at the location, wherein the high-resolution image is generated at the location by a high-resolution imaging system, and wherein generating the grayscale analog design data image is performed by a deep learning model included in one or more components executed by the one or more computer systems; generating an analog binary design data image of the location from the grayscale analog design data image; and Defects at the locations on the sample are detected by subtracting the design data for the locations from the simulated binary design data image, wherein generating the simulated binary design data image and detecting the defects are performed by the one or more computer systems.
22. A computer-implemented method for detecting defects on a sample, comprising: generating, for a location on a sample, a grayscale analog design data image from a high-resolution image generated at the location, wherein the high-resolution image is generated at the location by a high-resolution imaging system, and wherein generating the grayscale analog design data image is performed by a deep learning model included in one or more components executed by the one or more computer systems; generating an analog binary design data image of the location from the grayscale analog design data image; and Defects at the locations on the sample are detected by subtracting the design data for the locations from the simulated binary design data image, wherein generating the simulated binary design data image and detecting the defects are performed by the one or more computer systems.
Citation Information
Patent Citations
Single image detection
US20170140524A1
Generating simulated images from design information
US20170148226A1
Accelerated training of a machine learning based model for semiconductor applications
US20170193400A1
Generating high resolution images from low resolution images for semiconductor applications
US20170193680A1
Hybrid inspectors
US20170194126A1