Learnable Defect Detection for Semiconductor Applications
Through the deep metric learning model and the learning low-rank reference image generator, the shortcomings of defect detection methods in the existing semiconductor manufacturing process are solved, and efficient and sensitive detection of defects on samples is achieved.
Patent Information
- Application Number
- CN202080027481.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-02
- Filing Date
- 2020-04-07
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-04-07
AI Technical Summary
There are shortcomings in the existing defect detection methods in semiconductor manufacturing, including the inability to effectively handle process differences between dies, the constructed reference images are invalid for sub-regions and may destroy defect signals, unsupervised detection relies on reference image quality, lacks selective sensitivity enhancement for specific defect types, and supervised detection requires a large amount of labeled data, which is time-consuming.
Using a deep metric learning (DML) defect detection model and a learningable low-rank reference image generator, the distance between different parts is calculated to detect defects on the sample by projecting the test image and the reference image into the latent space.
It realizes efficient detection of defects on samples, can handle process differences between dies, improves detection sensitivity for small defects, reduces dependence on labeled data, and can learn to generate appropriate reference images to support defect detection.
Smart Images

Figure CN113678236B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to learnable defect detection methods and systems for semiconductor applications. Certain embodiments relate to systems and methods for detecting defects on a sample using a deep metric learning defect detection model and / or a learnable low-rank reference image generator. Background Art
[0002] The following description and examples are not admitted to be prior art by virtue of their inclusion in this section.
[0003] Fabricating semiconductor devices such as logic and memory devices generally involves processing a substrate such as a semiconductor wafer using numerous semiconductor fabrication processes to form various features and multiple levels of the semiconductor device. For example, lithography is a semiconductor fabrication process that involves transferring a pattern from a photomask to a resist disposed on a semiconductor wafer. Additional examples of semiconductor fabrication processes include, but are not limited to, chemical mechanical polishing (CMP), etching, deposition, and ion implantation. Multiple semiconductor devices can be fabricated in an arrangement on a single semiconductor wafer and then separated into individual semiconductor devices.
[0004] Inspection processes are used at various steps during semiconductor manufacturing to detect defects on the wafer to facilitate a higher yield rate and thus higher profit in the manufacturing process. Inspection has always been an important part of fabricating semiconductor devices. However, as the size of semiconductor devices decreases, inspection becomes even more important for the successful fabrication of acceptable semiconductor devices because smaller defects can cause the device to fail.
[0005] Most inspection methods include two main steps: generating a reference image, followed by performing defect detection. There are many different ways to perform each step. For example, a reference image can be generated by determining the median or average of multiple images. In another example, the reference image can be a constructed reference and substitute based on low-rank approximation. Defect detection can also be performed in several different ways. For example, defect detection can be unsupervised, using subtraction-based detection algorithms (e.g., MDAT, LCAT, etc.). Alternatively, defect detection can be supervised, using pixel-level detection algorithms (e.g., single image detection performed using a deep learning (DL) model and an electron beam image).
[0006] However, the various defect detection methods currently in use have several drawbacks. For example, generating a reference image by calculating the median or average is often insufficient to handle inter-die process variations, although the constructed reference and surrogate based on low-rank approximation partially address this issue. However, the constructed reference and surrogate may be ineffective for sub-regions on the wafer and may corrupt the defect signal. The currently used unsupervised defect detection methods are disadvantageous because the detection depends on the quality of the reference image and the test image. Such unsupervised defect detection methods also do not provide selective sensitivity enhancement for the target defect type or relatively small defects. Supervised defect detection methods require a large number of labeled defect candidates for training, which can be time-consuming for recipe setup in practice.
[0007] Accordingly, it would be advantageous to develop learnable defect detection systems and methods for semiconductor applications that do not have one or more of the drawbacks described above. SUMMARY
[0008] The following description of various embodiments should not be construed in any way as limiting the subject matter of the appended claims.
[0009] One embodiment relates to a system configured to detect defects on a sample. The system includes one or more computer systems and one or more components executed by the one or more computer systems. The one or more components include a deep metric learning (DML) defect detection model configured to project a test image and a corresponding reference image generated for the sample into a latent space. For one or more different portions of the test image, the DML defect detection model is further configured to determine distances in the latent space between the one or more different portions and corresponding one or more portions of the corresponding reference image. Additionally, the DML defect detection model is configured to detect defects in the one or more different portions of the test image based on the distances determined for the one or more different portions of the test image. The system can be further configured as described herein.
[0010] Another embodiment relates to a computer-implemented method for detecting defects on a sample. The method includes projecting a test image and a corresponding reference image generated for the sample into a latent space. The method further includes, for one or more different portions of the test image, determining distances between the one or more different portions and corresponding one or more portions of the corresponding reference image in the latent space. Additionally, the method includes detecting defects in the one or more different portions of the test image based on the distances determined for the one or more different portions of the test image, respectively. The projecting, determining, and detecting steps are performed by a DML defect detection model included in one or more components executed by one or more computer systems.
[0011] Each of the steps of the method described above may be further performed as described herein. Additionally, embodiments of the method described above may include any other steps of any other method described herein. Furthermore, the method described above may be performed by any one of the systems described herein.
[0012] Additional embodiments relate to a non-transitory computer-readable medium storing program instructions executable on one or more computer systems to perform a computer-implemented method for detecting defects on a sample. The computer-implemented method includes the steps of the method described above. The computer-readable medium may be further configured as described herein. The steps of the computer-implemented method may be performed as further described herein. Additionally, the computer-implemented method for which the program instructions are executed may include any other steps of any other method described herein.
[0013] Yet another embodiment relates to a system configured to generate a reference image of a sample. The system includes one or more computer systems and one or more components executed by the one or more computer systems. The one or more components include a low-rank reference image generator. The one or more computer systems are configured to input one or more test images of the sample into the learnable low-rank reference image generator. The one or more test images are generated for different locations on the sample corresponding to the same location in the design of the sample. The learnable low-rank reference image generator is configured to remove noise from the one or more test images, thereby generating one or more reference images corresponding to the one or more test images. A defect detection component detects defects on the sample based on the one or more test images and their corresponding one or more reference images. The system may be further configured as described herein.
[0014] Another embodiment relates to a computer-implemented method for generating a reference image of a sample. The method includes inputting one or more test images of the sample into a learnable low-rank reference image generator. The learnable low-rank reference image generator is included in one or more components executed by one or more computer systems. The one or more test images are generated for different locations on the sample corresponding to the same location in the design of the sample. The learnable low-rank reference image generator is configured to remove noise from the one or more test images, thereby generating one or more reference images corresponding to the one or more test images. A defect detection component detects defects on the sample based on the one or more test images and their corresponding one or more reference images.
[0015] Each of the steps of the method described above can be further performed as described herein. Additionally, embodiments of the method described above may include any other steps of any other method described herein. Furthermore, the method described above can be performed by any of the systems described herein.
[0016] Additional embodiments relate to a non-transitory computer-readable medium storing program instructions executable on one or more computer systems to perform a computer-implemented method for generating a reference image of a sample. The computer-implemented method includes the steps of the method described above. The computer-readable medium can be further configured as described herein. The steps of the computer-implemented method can be performed as further described herein. Additionally, the computer-implemented method for which the program instructions are executed can include any other steps of any other method described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Those skilled in the art will appreciate additional advantages of the present invention in view of the following detailed description of the preferred embodiments and with reference to the accompanying drawings, in which:
[0018] Figure 1 and 1a is a schematic side view illustrating an embodiment of a system configured as described herein;
[0019] Figure 2 and 3 is a schematic diagram illustrating an embodiment of a network architecture that can be used for one or more components described herein; and
[0020] Figure 4 is a block diagram illustrating an embodiment of a non-transitory computer-readable medium storing program instructions for causing a computer system to perform the computer-implemented method described herein.
[0021] While it is easy to make various modifications and alternative forms to the present invention, specific embodiments thereof are shown by way of example in the drawings and will be described in detail herein. The drawings may not be to scale. However, it should be understood that the drawings and the detailed description thereof are not intended to limit the present invention to the particular forms disclosed, but on the contrary, the present invention will cover all modifications, equivalent forms and alternative forms falling within the spirit and scope of the present invention as defined by the appended claims. Detailed Description
[0022] Turning now to the drawings, it should be noted that the figures are not drawn to scale. Specifically, the proportions of some of the elements of the figures are greatly enlarged to emphasize the characteristics of the elements. It should also be noted that the figures are not drawn to the same scale. The same reference numerals have been used to indicate elements shown in more than one figure that may be similarly configured. Unless otherwise mentioned herein, any of the elements described and shown may include any suitable commercially available elements.
[0023] One embodiment relates to a system configured to detect defects on a sample. The embodiments described herein generally relate to deep learning (DL) and machine learning (ML) based defect detection methods for tools such as optical inspection tools. Some embodiments are generally configured for learnable low-rank defect detection for semiconductor inspection and metrology applications.
[0024] The system described herein may include an imaging system that includes at least an energy source and a detector. The energy source is configured to generate energy that is directed to the sample. The detector is configured to detect energy from the sample and generate an image in response to the detected energy. In one embodiment, the imaging system is configured as an optical imaging system. Figure 1 An embodiment of this imaging system is shown in.
[0025] In one embodiment, the sample is a wafer. The wafer may include any wafer known in the semiconductor art. Although some embodiments may be described herein with respect to one wafer or several wafers, the embodiments are not limited to the samples on which they can be used. For example, the embodiments described herein can be used for samples such as photomasks, flat panels, personal computer (PC) boards, and other semiconductor samples. In another embodiment, the sample is a photomask. The photomask may include any photomask known in the art.
[0026] Figure 1 The imaging system shown in generates an optical image by directing light to the sample or scanning light over the sample and detecting the light from the sample. In the embodiment of the system shown in Figure 1 In the embodiment of the system shown in, the imaging system 10 includes an illumination subsystem configured to direct light to the sample 14. The illumination subsystem includes at least one light source. For example, as inFigure 1 As shown, the illumination subsystem includes a light source 16. The illumination subsystem can be configured to direct light to the sample at one or more incident angles, which can include one or more oblique angles and / or one or more normal angles. For example, as Figure 1 shown, light from the light source 16 is directed through the optical element 18 at an oblique incident angle and then through the lens 20 to the sample 14. The oblique incident angle can include any suitable oblique incident angle, which can vary depending on, for example, the characteristics of the sample.
[0027] The imaging system can be configured to direct light to the sample at different incident angles at different times. For example, the imaging system can be configured to change one or more characteristics of one or more elements of the illumination subsystem such that light can be directed to the sample at an incident angle different from the Figure 1 incident angle shown. In one such example, the imaging system can be configured to move the light source 16, the optical element 18, and the lens 20 such that light is directed to the sample at a different oblique incident angle or a normal (or near-normal) incident angle.
[0028] In some examples, the imaging system can be configured to direct light to the sample at more than one incident angle simultaneously. For example, the illumination subsystem can include more than one illumination channel, one of the illumination channels can include the light source 16, the optical element 18, and the lens 20 as Figure 1 shown, and another (not shown) of the illumination channels can include similar elements that can be configured differently or the same, or can include at least the light source and possibly one or more other components (such as those further described herein). If this light is directed to the sample simultaneously with other light, then one or more characteristics (e.g., wavelength, polarization, etc.) of the light directed to the sample at different incident angles can be different such that the light caused by the illumination of the sample at different incident angles can be distinguished from each other at the detector.
[0029] In another example, the illumination subsystem can include only one light source (e.g., Figure 1), and the light from the light source can be separated into different light paths (e.g., based on wavelength, polarization, etc.) by one or more optical elements (not shown) of the illumination subsystem. The light in each of the different light paths can then be directed to the sample. Multiple illumination channels can be configured to direct light to the sample simultaneously or at different times (e.g., when different illumination channels are used to sequentially illuminate the sample). In another example, the same illumination channel can be configured to direct light with different characteristics to the sample at different times. For example, the optical element 18 can be configured as a spectral filter, and the properties of the spectral filter can be changed in a variety of different ways (e.g., by swapping out the spectral filter) so that light of different wavelengths can be directed to the sample at different times. The illumination subsystem can have any other suitable configuration known in the art for directing light with different or the same characteristics to the sample sequentially or simultaneously at different or the same angles of incidence.
[0030] In one embodiment, the light source 16 is a broadband plasma (BBP) light source. In this way, the light generated by the light source and directed to the sample may include broadband light. However, the light source may include any other suitable light source known in the art that is configured to generate light at any suitable wavelength, such as any suitable laser. In addition, the laser may be configured to generate monochromatic or nearly monochromatic light. In this way, the laser may be a narrow-band laser. The light source may also include a polychromatic light source that generates light at multiple discrete wavelengths or bands.
[0031] Light from optical element 18 may be focused onto sample 14 by lens 20. Although lens 20 is Figure 1 20 is shown as a single refractive optical element, but it should be understood that in practice, lens 20 may include several refractive and / or reflective optical elements that, in combination, focus light from the optical element to the sample. The illumination subsystem may include any other suitable optical elements (not shown). Examples of such optical elements include, but are not limited to, polarizing components, spectral filters, spatial filters, reflective optical elements, apodizers, beam splitters, apertures, etc., which may include any such suitable optical elements known in the art. In addition, the imaging system can be configured to change one or more of the elements of the illumination subsystem based on the type of illumination to be used for imaging.
[0032] The imaging system may further include a scanning subsystem configured to cause light to scan over the sample. For example, the imaging system may include a stage 22 on which the sample 14 is disposed during imaging. The scanning subsystem may include any suitable mechanical and / or robotic assembly (which includes the stage 22) configured to move the sample such that light can scan over the sample. Additionally or alternatively, the imaging system may be configured such that one or more optical elements of the imaging system perform a certain scan of light over the sample. The light may be scanned over the sample in any suitable manner, such as in a serpentine path or in a helical path.
[0033] The imaging system further includes one or more detection channels. At least one of the one or more detection channels includes a detector configured to detect light from the sample attributable to illumination of the sample and to generate an output in response to the detected light. For example, Figure 1 the imaging system shown in includes two detection channels, one of which is formed by a condenser 24, an element 26, and a detector 28 and the other of which is formed by a condenser 30, an element 32, and a detector 34. As Figure 1 shown, the two detection channels are configured to collect and detect light at different collection angles. In some examples, both detection channels are configured to detect scattered light, and the detection channels are configured to detect light scattered from the sample at different angles. However, one or more of the detection channels may be configured to detect another type of light from the sample (e.g., reflected light).
[0034] As Figure 1 further shown, both detection channels are shown to be located in the plane of the paper, and the illumination subsystem is also shown to be located in the plane of the paper. Thus, in this embodiment, both detection channels are located in the plane of incidence (e.g., centered therein). However, one or more of the detection channels may be located out of the plane of incidence. For example, the detection channel formed by the condenser 30, the element 32, and the detector 34 may be configured to collect and detect light scattered out of the plane of incidence. Thus, this detection channel may generally be referred to as a "side" channel, and this side channel may be centered in a plane substantially perpendicular to the plane of incidence.
[0035] Although Figure 1An embodiment of an imaging system is shown that includes two detection channels, although the imaging system may include a different number of detection channels (e.g., only one detection channel or two or more than two detection channels). In one such example, the detection channel formed by the condenser 30, element 32, and detector 34 may form one side channel as described above, and the imaging system may include additional detection channels (not shown) as the other side channel located on the opposite side of the incident plane. Thus, the imaging system may include a detection channel that includes the condenser 24, element 26, and detector 28 and is centered in the incident plane and configured to collect and detect light at a scattering angle normal or near normal to the sample surface. This detection channel may thus generally be referred to as the "top" channel, and the imaging system may also include two or more side channels configured as described above. As such, the imaging system may include at least three channels (i.e., one top channel and two side channels), and each of the at least three channels has its own condenser, each of which is configured to collect light at a different scattering angle compared to each of the other condensers.
[0036] As further described above, each of the detection channels included in the imaging system may be configured to detect scattered light. Thus, Figure 1 the imaging system shown in may be configured for dark field (DF) imaging of a sample. However, the imaging system may also or alternatively include a detection channel configured for bright field (BF) imaging of the sample. In other words, the imaging system may include at least one detection channel configured to detect light specularly reflected from the sample. Thus, the imaging systems described herein may be configured for only DF imaging, only BF imaging, or both DF imaging and BF imaging. Although each of the condensers is shown in Figure 1 as a single refractive optical element, it should be understood that each of the condensers may include one or more refractive optical elements and / or one or more reflective optical elements.
[0037] One or more detectors may include a photomultiplier tube (PMT), a charge-coupled device (CCD), a time delay integration (TDI) camera, and any other suitable detector known in the art. The detectors may also include non-imaging detectors or imaging detectors. If the detectors are non-imaging detectors, then each of the detectors may be configured to detect certain characteristics (such as intensity) of the scattered light but not configured to detect such characteristics based on the position within the imaging plane. Thus, the output generated by each of the detectors may be a signal or data, but not an image signal or image data. In such examples, a computer subsystem (such as computer subsystem 36) may be configured to generate an image of the sample based on the non-imaging output of the detector. However, in other examples, the detectors may be configured as imaging detectors, which are configured to generate imaging signals or image data. Thus, the imaging subsystem may be configured to generate the images described herein in several ways.
[0038] Note that the Figure 1 optical imaging system configurations provided herein are generally diagrammatic to illustrate what may be included in the system embodiments described herein or what may generate the images used by the system embodiments described herein. Obviously, the optical imaging system configurations described herein may be modified to optimize the performance of the system, as is typically done when designing a commercial imaging system. Additionally, the systems described herein may be implemented using existing systems such as the 29xx / 39xx series of tools commercially available from KLA (Milpitas, California) (e.g., by adding the functionality described herein to an existing system). For some such systems, the embodiments described herein may be provided as optional functionality of the system (e.g., as a supplement to other functionality of the system). Alternatively, the optical imaging systems described herein may be designed "from scratch" to provide a brand new optical imaging system.
[0039] The computer subsystem 36 may be coupled to the detector of the imaging system in any suitable manner (e.g., via one or more transmission media, which may include "wired" and / or "wireless" transmission media) such that the computer subsystem may receive the output generated by the detector for the sample. The computer subsystem 36 may be configured to perform several functions further described herein using the output of the detector.
[0040] The system may also include more than one computer subsystem (e.g., Figure 1 the computer subsystem 36 and the computer subsystem 102 shown in Figure 1The computer subsystems (and other computer subsystems described herein) shown in [Figure X] may also be referred to as computer systems. Each of the computer subsystems or systems may take various forms, including personal computer systems, image computers, mainframe computer systems, workstations, network appliances, Internet appliances, or other devices. Generally speaking, the term "computer system" may be broadly defined to cover any device having one or more processors that execute instructions from a memory medium. The computer subsystem or system may also include any suitable processor known in the art, such as a parallel processor. Additionally, the computer subsystem or system may include a computer platform with high-speed processing and software as a stand-alone tool or a networked tool.
[0041] If the system includes more than one computer subsystem, the different computer subsystems may be coupled to each other such that images, data, information, instructions, etc. can be sent between the computer subsystems, as further described herein. For example, computer subsystem 36 may be coupled to computer subsystem 102 (as shown by the dashed line in [Figure X]) through any suitable transmission medium that may include any suitable wired and / or wireless transmission media known in the art. Figure 1 Two or more of such computer subsystems may also be effectively coupled through a shared computer-readable storage medium (not shown).
[0042] Although the imaging system is described above as an optical or light-based imaging system, in another embodiment, the imaging system is configured as an electron beam imaging system. For example, the system may also or alternatively include an electron beam imaging system configured to produce an electron beam image of a sample. The electron beam imaging system may be configured to direct electrons to the sample or scan electrons above the sample and detect electrons from the sample. In Figure 1a one such embodiment shown in [Figure X], the electron beam imaging system includes an electron column 122 coupled to computer subsystem 124.
[0043] Also as Figure 1a shown in [Figure X], the electron column includes an electron beam source 126 configured to produce electrons that are focused onto the sample 128 through one or more elements 130. The electron beam source may include, for example, a cathode source or an emitter tip, and the one or more elements 130 may include, for example, a gun lens, an anode, a beam-limiting aperture, a gate valve, a beam current selection aperture, an objective lens, and a scanning subsystem, all of which may include any such suitable elements known in the art.
[0044] Electrons returning from the sample (e.g., secondary electrons) may be focused onto the detector 134 through one or more elements 132. The one or more elements 132 may include, for example, a scanning subsystem, which may be the same scanning subsystem included in elements 130.
[0045] The electron column may include any other suitable elements known in the art. Additionally, the electron column may be further configured as described in the following U.S. patents: U.S. Patent No. 8,664,594 issued to Jiang et al. on April 4, 2014, U.S. Patent No. 8,692,204 issued to Kojima et al. on April 8, 2014, U.S. Patent No. 8,698,093 issued to Gubbens et al. on April 15, 2014, and U.S. Patent No. 8,716,662 issued to MacDonald et al. on May 6, 2014, which patents are incorporated herein by reference in their entirety as if fully set forth.
[0046] Although the electron column is shown in Figure 1a configured such that electrons are directed to the sample at an oblique incident angle and scattered from the sample at another oblique angle, it should be understood that the electron beam can be directed to the sample and scattered from the sample at any suitable angle. Additionally, the electron beam imaging system may be configured to generate an image of the sample using multiple modes (e.g., at different illumination angles, collection angles, etc.) as further described herein. The multiple modes of the electron beam imaging system may differ in any image generation parameter of the electron beam imaging system.
[0047] The computer subsystem 124 may be coupled to the detector 134 as described above. The detector may detect electrons returning from the surface of the sample, thereby forming an electron beam image of the sample. The electron beam image may include any suitable electron beam image. The computer subsystem 124 may be configured to perform one or more functions further described herein for the sample using the output generated by the detector 134. The computer subsystem 124 may be configured to perform any additional steps described herein. The system including the Figure 1a electron beam imaging system shown in
[0048] It should be noted that the Figure 1a provided herein generally illustrates the configuration of an electron beam imaging system that may be included in the embodiments described herein. Similar to the optical imaging system described above, the electron beam imaging system described herein may be modified to optimize the performance of the imaging system, as is typically done when designing a commercial imaging system. Additionally, the systems described herein may be implemented using existing systems such as tools commercially available from KLA (e.g., by adding the functionality described herein to an existing system). For some such systems, the embodiments described herein may be provided as optional functionality of the system (e.g., as a supplement to other functionality of the system). Alternatively, the systems described herein may be designed "from scratch" to provide a brand new system.
[0049] Although the imaging system was described above as a light beam or electron beam imaging system, the imaging system can be an ion beam imaging system. This imaging system can be configured as shown in Figure 1a , except that the electron beam source can be replaced with any suitable ion beam source known in the art. Additionally, the imaging system can be any other suitable ion beam imaging system, such as those included in commercially available focused ion beam (FIB) systems, helium ion microscope (HIM) systems, and secondary ion mass spectrometry (SIMS) systems.
[0050] As mentioned above, the imaging system can be configured to direct energy (e.g., light, electrons) to a physical version of the sample and / or scan the energy over the physical version of the sample, thereby generating an actual image of the physical version of the sample. In this way, the imaging system can be configured as an "actual" imaging system, rather than a "virtual" system. However, the storage medium (not shown) and Figure 1 the computer subsystem 102 shown in can be configured as a "virtual" system. Specifically, the storage medium and the computer subsystem are not part of the imaging system 10 and do not have any ability to handle the physical version of the sample, but can be configured to perform functions similar to inspection using the stored detector output as a virtual inspector. Systems and methods configured as "virtual" inspection systems are described in the following co-owned U.S. patents: U.S. Patent No. 8,126,255 issued to Bhaskar et al. on February 28, 2012, U.S. Patent No. 9,222,895 issued to Duffy et al. on December 29, 2015, and U.S. Patent No. 9,816,939 issued to Duffy et al. on November 14, 2017, which are incorporated herein by reference in their entirety as if fully set forth. The embodiments described herein can be further configured as described in these patents. For example, one or more of the computer subsystems described herein can be further configured as described in these patents.
[0051] As further mentioned above, the imaging system can be configured to generate images of the sample in multiple modes. Generally speaking, a "mode" can be defined by the parameter values of the imaging system used to generate the image of the sample or the output used to generate the image of the sample. Therefore, different modes can be different in at least one of the imaging parameters of the imaging system. For example, in an optical imaging system, different modes can use light of different wavelengths for illumination. The modes can be different in terms of the illumination wavelength, as further described herein for different modes (e.g., by using different light sources, different spectral filters, etc.). In another embodiment, different modes use different illumination channels of the imaging system. For example, as mentioned above, the imaging system can include more than one illumination channel. Thus, different illumination channels can be used for different modes.
[0052] The imaging system described herein can be configured as an inspection subsystem. If so, the computer subsystem can be configured to receive an output from the inspection subsystem as described above (e.g., from the detector of the imaging system), and can be configured to detect defects on a sample based on the output, as further described herein.
[0053] The imaging system described herein can be configured as another type of semiconductor-related process / quality control system, such as a defect re-inspection system and a metrology system. For example, the embodiments of the imaging system described and Figure 1 and 1a shown herein can be modified in one or more parameters to provide different imaging capabilities depending on the application for which it will be used. In one embodiment, the imaging system is configured as an electron beam defect re-inspection system. For example, if Figure 1a the imaging system shown herein will be used for defect re-inspection or metrology rather than inspection, then it can be configured to have a higher resolution. In other words, Figure 1 and 1a the embodiments of the imaging system shown herein describe some general and various configurations of the imaging system, which can be modified in several ways obvious to those skilled in the art to produce imaging systems with different imaging capabilities more or less suitable for different applications.
[0054] A system configured to detect defects on a sample includes one or more computer systems and one or more components executed by the one or more computer systems. The one or more computer systems can be configured as described above. The one or more components include a deep metric learning (DML) defect detection model. The DML defect detection model can have several different architectures further described herein. The one or more components can be executed by the computer system in any suitable manner known in the art.
[0055] The DML defect detection model is configured to project a test image and a corresponding reference image generated for a sample into a latent space. In this way, when provided with two images (test and reference), the DML defect detection model will project the images into the latent space. For example, as described below, the DML defect detection model may include different CNNs for different images input into the model. Each of the different CNNs may project its input image into the latent space. As used herein, the term "latent space" refers to a hidden layer in the DML defect detection model that contains a hidden representation of the input. Additional explanation of the term latent space (as it is commonly used in this technical field) can be found in "Latent Variable Modeling for Generative Concept Representations and Deep Generative Models" (18 pages), published by Chang on arXiv in December 2018, which is incorporated herein by reference as if fully set forth. Regarding Figure 2 block A shown in
[0056] is further described such CNNs. Although the projection step (and other steps) may be described with respect to a "test image", these steps may be performed for more than one test image generated for a sample. For example, for each test image generated for a sample (or at least one or more test images), the projection step may be performed independently and separately. Test images may be generated for any suitable test (e.g., inspection, re-inspection, metrology) area on the sample. The test images may have different sizes (e.g., tile images, die images, job frames, etc.) and may be measured in pixels (e.g., 32 pixels × 32 pixels for a tile image) or in any other suitable manner. The other steps described herein may also be performed independently and separately for different test images.
[0057] In one embodiment, the test image and the corresponding reference image are for corresponding locations in different dies on a sample. In another embodiment, the test image and the corresponding reference image are for corresponding locations in different cells on a sample. For example, the DML described herein can be used for inter-die type inspection or inter-cell type inspection. Traditional inter-die inspection requires reference images that are collected from a tool and may already be processed. The embodiments described herein can be performed using die images for the test image and the corresponding reference image or using cell images for the test image and the corresponding reference image. Corresponding images from different dies and different cells can be generated and / or acquired as further described herein. The die images and cell images can be further configured as described herein. In this way, both the test image and the reference image can be generated using the sample.
[0058] In some embodiments, the test image and the corresponding reference image are generated for a sample without using the design data of the sample. For example, the defect detection described herein can be performed without design data, and the images used for defect detection can all be generated by imaging the sample. In such examples, the defect detection can also be performed independently of the design data of the sample. In other words, any steps performed on the images for defect detection purposes can be performed without using or not based on the design data of the sample. For example, the defect detection described herein can be performed in the same manner (using the same parameters) in all test images generated for the sample (and for all pixels in any test image generated for the sample), regardless of the design data of the sample. In this way, the defect detection can be performed in the same manner, regardless of the design data corresponding to the test image (or the pixels in the test image).
[0059] In additional embodiments, a test image is generated for a sample by an imaging system that directs energy to the sample and detects energy from the sample, and a corresponding reference image is generated without using the sample. In one such embodiment, the corresponding reference image is obtained from a database containing design data of the sample. For example, the DML described herein can be used for die-to-database defect detection. In this way, the embodiments described herein can be performed using the die image of the test image and the database image of the corresponding reference image rather than the reference image from the sample. The test image can be generated for the sample as further described herein and can be further configured as described herein. Generating the reference image without using the sample can be performed by simulating the reference image based on the design or design information of the sample or in any other suitable manner known in the art. The reference image preferably simulates what the test image of the sample would be like if the portion of the sample for which the test image is generated is defect-free. The design, design data, or design information can include any suitable design, design data, or design information known in the art, and these terms are used interchangeably herein. The reference image can be obtained from the database in any suitable manner, and the database can have any suitable configuration.
[0060] In another embodiment, one or more computer systems are configured to input design data of a sample into a DML defect detection model, and the DML defect detection model is configured to perform detection using the design data. In this way, the presented design can be part of the input to the DML defect detection model. Additionally, the design data can be regarded by the DML defect detection model as an additional image channel. Then, the DML defect detection model can perform defect detection using the presented design. The DML defect detection model can use the presented design in several different ways. For example, the DML defect detection model can be configured to use the presented design to align different images with each other (e.g., by aligning multiple images with the design, thereby aligning the images with a common reference). The presented design can also be used as a reference image as further described herein or for generating the reference images described herein. The DML defect detection model can also use the presented design to set or adjust defect detection parameters and classify defects in the defect detection step. The DML defect detection model can also be configured to use the design data in any other manner known in the field of defect detection.
[0061] In additional embodiments, the detection is performed using one or more parameters determined from a region of interest of a sample. In some such embodiments, one or more computer systems are configured to input information about the region of interest into a DML defect detection model. In this way, the DML defect detection model can treat the region of interest as an additional image channel. The region of interest can be determined from the design data of the sample in any suitable manner. Additionally, the region of interest can be generated by a commercially available system (such as those from KLA) configured with the ability to align test images and / or reference images with the design data with substantially high accuracy and precision. The region of interest can define the area on the sample that will be inspected, thereby inherently also defining the area on the sample that will not be inspected. The region of interest can also be designed or configured to indicate which areas on the sample will be inspected using different parameters (e.g., different detection sensitivities). The region of interest of the sample can be determined by one or more computer systems from the design in any suitable manner and then input by the one or more computer systems into the DML defect detection model such that the DML defect detection model can identify the region of interest in the test image and the corresponding reference image. The one or more parameters for performing the detection can also be determined by one or more computer systems and input into the DML defect detection model. Thus, when the region of interest in the test image is identified by the DML defect detection model using the information about the region of interest input by the computer system, the DML defect detection model can determine the one or more parameters for defect detection in the region of interest based on the input from the computer system and perform the defect detection accordingly.
[0062] In another embodiment, the detection is performed without information about the region of interest of the sample. For example, although many inspection processes described herein for samples use a region of interest, the defect detection described herein need not use any region of interest to be performed. In such cases, the defect detection described herein can be performed for all test images generated for the sample. Additionally, the defect detection described herein can be performed using the same parameters for all pixels in all test images generated for the sample. For example, the defect detection sensitivity for all pixels in all test images can be the same.
[0063] In some embodiments, the test images are generated in the logic region of the sample. In another embodiment, the test images are generated in the array region of the sample. For example, the embodiments described herein can be used and configured for defect detection in both the logic region and the array region of the sample. The logic region and the array region of the sample can include any such regions known in the art.
[0064] For one or more different parts of a test image, the DML defect detection model is configured to determine the distance in the latent space between the one or more different parts and the corresponding one or more parts of a corresponding reference image. For example, both the test features and the reference features are measured for the distance between them in the latent space based on the output from block B, which is further described herein. In other words, the DML defect detection model determines the degree of similarity between the test image and the reference image based on the degree of difference between the features of the test image and the reference image (measured by the distance in the latent space), and the degree of similarity can then be used, as further described herein, to determine which parts contain defects (or defect candidates) or are defective (or potentially defective).
[0065] The distance in the latent space determined by the DML defect detection model can be various different distances. For example, the distance metrics that can be determined and used by the DML defect detection model include Euclidean distance, L1 distance, L_infinity distance, Person’s distance (i.e., cross-correlation), Manhattan distance, generalized Lp-norm, cosine distance, etc. Such distances can be determined in any suitable manner known in the art.
[0066] In one embodiment, the different parts of the test image include different pixels in the test image. For example, the different parts of the test image can include any suitable parts, such as individual pixels or relatively small pixel arrays (e.g., 9×9 pixel neighborhoods). The different parts of the test image can also be measured in terms of pixels or any other suitable measure. In this way, the DML defect detection model determines the distance in the latent space part by part of the test image. In other words, the distance in the latent space can be determined separately and independently for different parts of the test image, such that those distances can be used to detect defects in each of the different parts of the test image.
[0067] The DML defect detection model is further configured to detect defects in one or more different portions of a test image based on distances determined for one or more different portions of the test image, respectively. For example, distances in a latent space are used to determine whether each portion (e.g., each pixel) in the test image is defective relative to a reference image. In this way, the DML defect detection model can make a binary decision as to whether each pixel is a defect. However, the DML defect detection model can also determine the type of defect for each defective pixel. Thus, the DML defect detection model can perform both detection (yes or no) and classification (which type of defect). In the art, this application is referred to as defect detection, and the method is referred to as classification. Defect detection based on distances in a latent space can be further performed as described herein.
[0068] In one embodiment, the DML defect detection model has a siamese network architecture. In another embodiment, the DML defect detection model has a triplet network architecture. In additional embodiments, the DML defect detection model has a quadruplet network architecture. For example, DML can be constructed from a siamese network, a triplet network, a quadruplet network, etc.
[0069] The classical siamese classification model can be extended to work with die-to-die detection use cases and other use cases described herein. A siamese network is generally defined in the art as a neural network containing two identical sub-network components. The inventors utilize this concept for the embodiments described herein because the siamese model is useful for tasks involving finding the similarity or relationship between two comparable things, which naturally applies to the die-to-die comparison scheme used in some inspection systems commercially available from KLA and used in the multi-die auto-threshold (MDAT) defect detection algorithm. By using a non-linear model (e.g., a convolutional neural network (CNN)), both the test image and the reference (patch) image can be transformed into a latent space, and as further described herein, the distance therebetween (i.e., a measure of their similarity) can be constructed as an indicator of defectiveness.
[0070] By Figure 2Illustrated is a construction of a twin detection model that can be used in the embodiments described herein. A test image can include N BBP images 202 and 204 from N adjacent dies and a design image 200. These images are optionally selected from the same die coordinates with the same field of view (FOV) (i.e., the same die coordinates centered on the same die within the die positions of multiple dies). Blocks A, B, and C are three different depth CNNs. The two networks shown in block B have the same architecture configuration. In addition to having the same architecture, the weights of the two networks in block B must be shared by the networks to make the networks have a twin architecture. First, the test image passes through block A to calculate a reference feature, which is the average of N outputs from block A. Second, the test feature and the reference feature pass through block B to measure the distance therebetween based on the output from block B. Third, block C is applied to generate a final label (defective versus non-defective) for each image pixel position.
[0071] For example, as Figure 2 shown, the design image 200 is input into the first CNN 206 in block A 208, the optical image 202 is input into the CNN 210 in block A, and the optical image 204 is input into the CNN 212 in block A. The CNNs 206, 210, and 212 respectively generate reference features for each of the input images. These CNNs can have any suitable configuration known in the art.
[0072] The outputs of CNNs 206 and 210 can be input into the slicing layer 214 and the outputs of CNNs 210 and 212 can be input into the slicing layer 216. The slicing layer can have any suitable configuration known in the art. The outputs of the slicing layers 214 and 216 can be respectively input into the batch normalization (BN) layers 218 and 222, and the output of CNN 210 can be input into the BN layer 220. Batch normalization can be performed as described, for example, in “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift” by Ioffe et al. (arXiv:1502.03167, March 2, 2015, 11 pages), which is incorporated herein by reference as if fully set forth. The embodiments described herein can be further configured as described in such reference. The outputs of the BN layers 218 and 220 can be input into the concatenation layer 224, and the outputs of the BN layers 220 and 222 can be input into the concatenation layer 226. The concatenation layer can have any suitable configuration known in the art.
[0073] The outputs of cascading layers 224 and 226 can be respectively input into networks 228 and 230 in block B 232. Although networks 228 and 230 are schematically shown in Figure 2 as a residual neural network (ResNet) of a network that is typically defined in the art as having shortcuts between some of its layers where layers are skipped in between, the networks can have any suitable configuration known in the art, including ordinary networks (networks without shortcuts) and various types of networks with shortcuts (e.g., highway networks and DenseNet). Examples of some such suitable network configurations are described in Huang et al.'s "Densely Connected Convolutional Networks" (arXiv:1608.06993, January 28, 2018, 9 pages) and He et al.'s "Deep Residual Learning for Image Recognition" (arXiv:1512.03385, December 10, 2015, 12 pages), which are incorporated herein by reference as if fully set forth. As described above, in order for networks 228 and 230 to have a twin configuration, the networks will have the same configuration and share the same weights.
[0074] As also mentioned above, the networks in block B can have a triplet or quadruplet network configuration. Examples of suitable triplet network architectures that can be used for the networks in block B can be found in Hoffer et al.'s "Deep Metric Learning Using Triplet Network" (arXiv:1412.6622, December 4, 2018, 8 pages), which is incorporated herein by reference as if fully set forth. Examples of suitable quadruplet network architectures that can be used for the networks in block B can be found in Dong et al.'s "Quadruplet Network with One-Shot Learning for Fast Visual Object Tracking" (arXiv:1705.07222, 12 pages), which is incorporated herein by reference as if fully set forth. The embodiments described herein can be further configured as described in these references.
[0075] The outputs of networks 228 and 230 are input into the fully connected layer and SoftMax layer combination 234 in block C 236, and the fully connected layer and SoftMax layer combination 234 produces an output 238 that includes the final label (i.e., defective or non-defective) described above. The fully connected layer and the SoftMax layer can have any suitable configuration known in the art. For example, the fully connected layer and the SoftMax layer can be configured as described in U.S. Patent Application Publication No. 2019 / 0073568, published on March 7, 2019, by He et al., which is incorporated herein by reference in its entirety as if fully set forth.
[0076] In another embodiment, the DML defect detection model includes one or more DL convolutional filters, and one or more computer systems are configured to determine the configuration of the one or more DL convolutional filters based on the physical phenomena involved in generating the test image. The DL convolutional filters (not shown) can be located at the beginning of the neural network or at any other layer of the neural network. For example, the DL convolutional filters can be located Figure 2 before block A shown in. The physical phenomena involved in generating the image can include any known, simulated, measured, estimated, calculated, etc. parameters involved in the imaging process for generating the test image, possibly in combination with any known or simulated characteristics of the sample (e.g., material, dimensions, etc.). For example, the physical phenomena involved in generating the test image can include the hardware configuration of the imaging tool used to generate the test image and any hardware settings of the components of the imaging tool used to generate the test image. The known or simulated characteristics of the sample can be obtained, generated, determined, or simulated in any suitable manner known in the art.
[0077] In yet another embodiment, the DML defect detection model includes one or more DL convolutional filters, and one or more computer systems are configured to determine the configuration of the one or more DL convolutional filters based on the imaging hardware used to generate the test image. For example, the DL convolutional filters can be designed based on optical (or other, e.g., electron beam, ion beam) hardware (including parameters of the hardware, such as aperture, wavelength, numerical aperture (NA), etc.). Such DL convolutional filters can be configured in other ways as described herein.
[0078] In one such embodiment, determining the configuration of the DL convolutional filter comprises determining one or more parameters of one or more DL convolutional filters based on the point spread function (PSF) of the imaging hardware. For example, the optical hardware information can be used to determine one or more parameters of the DL convolutional filter based on the measured or simulated PSF of the tool. The PSF of the tool can be measured or simulated in any suitable manner known in the art. In this way, the embodiments described herein can be configured for PSF-based defect detection. Different from detection algorithms that depend on pixel-level information (e.g., MDAT), the embodiments described herein can rely on PSF-level signatures to perform defect detection. The basic assumption is that the defect signal (i.e., the information content) is confined within the local neighborhood context dominated by the optical interaction approximately defined by the PSF. By using the filter size in the CNN to control the FOV, the inventors studied the detection accuracy relative to the cut-off radius on the PSF. The inventors found that as the FOV is enlarged until it reaches an area consistent with the calculated PSF area, the detection accuracy increases. The area can be determined in units of pixels or any other suitable measure.
[0079] In one such embodiment, one or more parameters of the DL convolutional filter comprise one or more of filter size, filter symmetry, and filter depth. For example, the imaging hardware information can be used to determine the filter size, filter symmetry, filter depth, etc. (based on the measured or simulated PSF of the tool). In one such instance, the filter size can be set to be equal to or roughly equal to the PSF of the imaging tool.
[0080] In another such embodiment, determining one or more parameters of one or more DL convolutional filters comprises learning the one or more parameters by optimizing a loss function. For example, one or more parameters such as those described above can be determined by learning the parameters (by optimizing a loss function) based on information such as the information described above. Optimizing the loss function can be performed in any suitable manner known in the art, and the loss function can be any suitable loss function known in the art.
[0081] In some such embodiments, determining the configuration includes selecting one or more DL convolutional filters from a predetermined set of DL convolutional filters based on the PSF of the imaging hardware. For example, determining the configuration may include determining the filters themselves (e.g., based on the measured or simulated PSF of the tool). The predetermined set of DL convolutional filters may include any or all suitable known DL convolutional filters. In one such embodiment, one or more parameters of one or more of the DL convolutional filters in the predetermined set are fixed. For example, the filter parameters may be fixed. In other words, the DL convolutional filters can be used without any adjustment to the predetermined parameters. In another such embodiment, determining the configuration includes fine-tuning one or more initial parameters of one or more DL filters by optimizing a loss function. For example, the filter parameters may be fine-tuned by optimizing a loss function. Optimizing the loss function can be performed in any suitable manner, and the loss function may include any suitable loss function known in the art. The filter parameters being optimized may include any suitable filter parameters, including those described above.
[0082] In yet another embodiment, one or more components include a learnable low-rank reference image generator configured to generate a corresponding reference image, one or more computer systems are configured to input one or more test images generated for a sample into the learnable low-rank reference image generator, the one or more test images being generated for different positions on the sample, the different positions corresponding to the same position in the design for the sample, and the learnable low-rank reference image generator is configured to remove noise from the one or more test images, thereby generating a corresponding reference image. The learnable low-rank reference image generator may be further configured as described herein.
[0083] In another embodiment, the DML defect detection model is configured to project additional corresponding reference images into the latent space and determine the average of the corresponding reference images and the additional corresponding reference images with respect to a reference region in the latent space, and one or more portions of the corresponding reference images for determining the distance include the reference region. For example, if a test image and multiple reference images (i.e., 1+N images) are provided, then the DML defect detection model will project all the images into the latent space and use the N reference points in the latent space to estimate the "averaged" reference point and the possible reference region, which are used to determine whether each portion in the test image is a defect with respect to the reference region. The multiple reference images may include any combination of the reference images described herein (e.g., images from multiple dies or cells adjacent to the test die or cell) and / or any other suitable reference images known in the art.
[0084] In yet another embodiment, the corresponding reference image includes a defect-free test image of the sample, the projected corresponding reference image includes a reference region in the learned latent space, and one or more portions of the corresponding reference image used to determine the distance include the reference region. For example, if only test images are provided, the DML defect detection model can project both the "defective" test images and the "defect-free" test images into the latent space, learn the reference region in the latent space, and use it to determine whether one or more portions of the test image are defects relative to the reference region. The "defect-free" test image can be generated using a sample, for example, by imaging a known defect-free portion of the sample. The "defect-free" test image can also or alternatively be generated from a test image of unknown defectiveness (using a network or model configured to generate a reference image from a test image of unknown defectiveness). An example of such a network or model is described in U.S. Patent No. 10,360,477, issued July 23, 2019, to Bhaskar et al., which is incorporated herein by reference in its entirety as if fully set forth. The embodiments described herein can be further configured as described in the reference.
[0085] In some embodiments, different modes of the imaging system are utilized to generate a test image and an additional test image for the sample, respectively; the DML defect detection model is configured to project the test image and the corresponding reference image into a first latent space, project the additional test image and the additional corresponding reference image into a second latent space, and combine the first latent space and the second latent space into a joint latent space; and the latent space used to determine the distance is the joint latent space. For example, the embodiments described herein can be configured for multimodal DML. Specifically, multiple mode images can be input into the DML defect detection model as additional channels. Additionally, independent DML can be applied to each mode image to construct multiple independent latent spaces and combine them into a joint latent space for distance calculation. The different multiple modes of the imaging system can include any of the modes described herein (e.g., BF and DF, different detection channels, such as a top channel and one or two side channels, etc.).
[0086] In another embodiment, one or more computer systems are configured to input design data of a sample into a DML defect detection model; generate a test image and additional test images for the sample using different modes of an imaging system, respectively; the DML defect detection model is further configured to project the test image and the corresponding reference image into a first latent space, project the additional test image and the additional corresponding reference image into a second latent space, project the design data into a third latent space, and combine the first latent space, the second latent space, and the third latent space into a joint latent space; and the latent space for determining distances is the joint latent space. For example, a design or region of interest can be treated as an additional image "mode" by configuring the DML defect detection model to project the design data, region of interest, etc. into its own latent space, and then combining the latent space with the latent spaces of the test images (and their corresponding reference images) from two or more modes to generate a joint latent space for defect detection. Specifically, independent DMLs can be applied individually to each mode image and each other input (design data, region of interest, etc.) to construct multiple independent latent spaces, and then the independent latent spaces are combined into a joint latent space for distance calculation. The different modes of the imaging system used in this embodiment can include any of the modes described herein. The steps of this embodiment can be performed in other ways as further described herein.
[0087] In an additional embodiment, one or more computer systems are configured to input design data of a sample into a DML defect detection model; the DML defect detection model is further configured to project the test image and the corresponding reference image into a first latent space, project the design data into a second latent space, and combine the first latent space and the second latent space into a joint latent space; and the latent space for determining distances is the joint latent space. For example, even if only one mode is used for defect detection, the DML defect detection model can project the design data. In this way, the DML defect detection model can be configured for single-mode defect detection in which the design data is projected into a separate latent space, which is combined with the latent space into which the test images are projected to generate a joint latent space that is then used for defect detection. The steps of this embodiment can be performed in other ways as described herein.
[0088] In yet another embodiment, one or more computer systems are configured to input design data of a sample into a DML defect detection model; generate a test image and additional test images for the sample using different modes of an imaging system, respectively; the DML defect detection model is further configured to project a first set including one or more of the test image and corresponding reference image, additional test images and additional corresponding reference images, and design data into a first latent space, project a second set including one or more of the test image and corresponding reference image, additional test images and additional corresponding reference images, and design data into a second latent space, and combine the first latent space and the second latent space into a joint latent space; and the latent space for determining distances is the joint latent space. For example, the DML defect detection model may project images and design data of a first mode into a first latent space, project images and design data of a second mode into a second latent space, combine the two latent spaces into a joint latent space, and use the joint latent space for defect detection. In this way, the first and second sets for the projection step may be different combinations of inputs to the DML defect detection model, and some of the different combinations may include one or more of the same inputs. Alternatively, the first and second sets for the projection step may be mutually exclusive because none of the different combinations include the same input. For example, the first set may include images of two or more modes, and the second set may include only design data. In other aspects, the steps of this embodiment may be performed as further described herein.
[0089] Thus, in general, in the embodiments described herein, there are several different ways in which a DML defect detection model may be configured to use various possible inputs. For example, the DML defect detection model may treat design data, regions of interest, etc. and other non-sample images as additional image channels. If design data, regions of interest, etc. are treated as additional image channels for single-mode inspection, then the DML defect detection model may be configured to combine single-channel images with design, regions of interest, etc. and then project them into a latent space. If design data, regions of interest, etc. are treated as additional image channels for multi-mode inspection, then the DML defect detection model may be configured to combine multiple channels of images with design data, regions of interest, etc. into a multi-channel tensor and then project it into a latent space. In this way, multiple channels of data may be combined before projecting the multiple channels of data into a latent space.
[0090] In an alternative embodiment, different channels of data can be individually projected into separate latent spaces, whereby different channels are treated as different modalities. For example, in the absence of a design, region of interest, etc., the DML defect detection model can project images of different modalities into their respective latent spaces, combine the latent spaces, and compute distances. When a design, region of interest, etc. is available, the DML defect detection model can project each imaging modality image into its respective latent space, project the design into another latent space, and then combine all the imaging latent spaces and the design latent space for distance calculation.
[0091] In the above embodiments, there is only one latent space into which all channels are projected, or the number of latent spaces into which different channels are projected individually and independently is equal to the number of input channels (e.g., the number of latent spaces combined into a joint latent space = the number of modalities for inspection + the number of design-related inputs). However, as described above, in some embodiments, the DML defect detection model can be configured such that groups (images and / or designs, regions of interest, etc.) are defined as multi-channel inputs, and each group can be projected into one latent space. In this way, the number of latent spaces combined to form a joint latent space can be greater than one and different from the total number of inputs.
[0092] There can be two phases in using the DML defect detection model: setup and runtime. At setup, defect candidates can be provided in a pixel-labeled manner. For a given training image, a defect is predetermined for each pixel. These pixel-level training data are used to train the DML model at the pixel level. At runtime, the DML defect detection model is used to determine "whether each pixel is a defect" and possibly the "defect type" for each job frame of the region to be inspected. The steps described herein can also be performed for pixels, job frames, or any other test image part described herein.
[0093] The embodiments described herein can be configured to train the DML defect detection model. However, another system or method can alternatively be configured to train the DML defect detection model. In this way, the embodiments described herein may or may not perform the training of the DML defect detection model.
[0094] In one embodiment, one or more computer systems are configured to train a DML defect detection model using one or more training images and pixel-level ground truth information of the one or more training images. In one such embodiment, the one or more training images and the pixel-level ground truth information are generated from a Process Window Qualification (PWQ) wafer. For example, a PWQ wafer is a wafer on which different device regions are formed using one or more different parameters (such as exposure and dose of a lithography process). The one or more different parameters used to form the different device regions may be selected such that defect detection-related data generated using this wafer simulates process variations and drifts. Then, the collected data from the PWQ wafer can be used as training data to train the DML defect detection model to achieve model stability for process variations. In other words, by training the DML defect detection model using a training data set that captures possible process variations and drifts, the DML defect detection model will be more stable for those process variations compared to a situation where it is not trained using PWQ-type data. The training using this data can be performed in other aspects as described herein.
[0095] The PWQ method can be performed as described in the following U.S. patents: U.S. Pat. No. 6,902,855, issued Jun. 7, 2005 to Peterson et al.; U.S. Pat. No. 7,418,124, issued Aug. 26, 2008 to Peterson et al.; U.S. Pat. No. 7,729,529, issued Jun. 1, 2010 to Wu et al.; U.S. Pat. No. 7,769,225, issued Aug. 3, 2010 to Kekare et al.; U.S. Pat. No. 8,041,106, issued Oct. 18, 2011 to Pak et al.; U.S. Pat. No. 8,111,900, issued Feb. 7, 2012 to Wu et al.; and U.S. Pat. No. 8,213,704, issued Jul. 3, 2012 to Peterson et al., which patents are incorporated herein by reference in their entirety as if fully set forth. The embodiments described herein may include any steps of any methods described in these patents and may be further configured as described in these patents. The PWQ wafer can be printed as described in these patents.
[0096] In another embodiment, one or more computer systems are configured to perform active learning to train a DML defect detection model. For example, the training process of the DML defect detection model can be combined with active learning for low defect candidate cases or in-line learning cases within a semiconductor factory. Active learning can be performed as described in U.S. Patent Application Publication No. 2019 / 0370955, published Dec. 5, 2019 by Zhang et al., which patent application publication is incorporated herein by reference in its entirety as if fully set forth. The embodiments described herein can be further configured as described in this publication.
[0097] Training data for training a DML defect detection model may also include any other in-situ data, such as information generated by an electron beam system, a light-based system, user, and physical simulations, including such data further described herein. The training data may also include any combination of such data.
[0098] The input to the trained DML model may vary as described herein. For example, the input image may include 1 test frame image and 1 reference frame image per mode, whether there is only one mode or multiple modes. In another example, the input image may include 1 test frame image and N reference frame images per mode, whether there is only one mode or multiple modes. In an additional example, the input image may include 1 test frame per mode (and no reference images), whether there is only one mode or multiple modes. The input to the DML defect detection model may also optionally include a region of interest (i.e., the region where an inspection or another test function will be performed). Another optional input to the DML defect detection model includes the design information of the sample. If the DML defect detection model is used in conjunction with learnable principal component analysis (LPCA) reference image generation (further described herein), then the test and reference images are used to form the LPCA reference image, which is then fed into the DML defect detection model. The output of the DML defect detection model may include a decision for each pixel in the test frame (or another suitable test image portion) as to whether the pixel is a defect and what type of defect the pixel is (in the case of a decision that the pixel is a defect). The output may otherwise have any suitable form or format known in the art.
[0099] Another embodiment relates to a system configured to generate a reference image of a sample. The system includes one or more computer systems and one or more components executed by the one or more computer systems. The one or more computer systems may be configured as further described herein. The one or more components may be executed by the one or more computer systems in any suitable manner. In these systems, the one or more components include a learnable low-rank reference image (LLRI) generator.
[0100] One or more computer systems are configured to input one or more test images of the sample into the LLRI generator. The one or more test images are generated for different locations on the sample corresponding to the same location in the design of the sample. For example, one or more test images may be generated at corresponding locations in different die on the sample, different cells on the sample, etc. Thus, in addition to the same FOV (region on the sample or in the image), the corresponding locations may also have the same die coordinates, the same cell coordinates, etc. The computer system may input the test images into the LLRI generator in any suitable manner.
[0101] The LLRI generator is configured to remove noise from one or more test images, thereby generating one or more reference images corresponding to the one or more test images. For example, portions of a test image corresponding to defects may appear to have more noise than other portions of the test image (e.g., it may have a signal that is an outlier relative to other portions of the image, and whether a signal is an outlier can be defined in several different ways known in the art). Thus, by identifying and removing noise from one or more test images, the resulting images may be suitable for use as references for defect detection (and other) purposes. The one or more reference images may then be used for defect detection as further described herein. The LLRI generator may have one of the various configurations described herein and may perform noise removal as further described herein.
[0102] Techniques referred to in the art as “low-rank constraints” may be used in outlier detection algorithms such as Computational Reference (CR) and Tensor Decomposition (TD). The embodiments described herein may extend low-rank constraint techniques to spatial context and utilize this concept in DL classification.
[0103] Principal Component Analysis (PCA) is a tool developed by the inventors for low-rank constraints in CR. When applying PCA to a multi-image reconstruction problem, the problem statement can be summarized as follows. Given that X is a 3D tensor of dimensions (w*h,c) (usually with the DC component removed) where w, h, and c are the image width, height, and number respectively, PCA attempts to find the principal component vector w of dimensions (c,1) that maximizes the variance estimate,
[0104]
[0105] where T is the matrix transpose operator. The principal component is defined as:
[0106] p.c. = Xw (Equation 2)
[0107] Multiple principal component vectors can be calculated via iterative PCA, matrix orthonormalization, Singular Value Decomposition (SVD), etc. Low-rank approximation is typically applied by filtering the principal components based on their eigenvalues. Larger eigenvalues correspond to directions of larger variance in the image data. According to the low-rank approximation, the reconstruction of X can be achieved by the following equation:
[0108] X′ = (XW)W T (Equation 3)
[0109] where W is a 2D matrix composed of selected principal component vectors in column format (i.e., {w1,w2,…}).
[0110] The above problem statement is very classical and known in the art. One drawback of this problem statement is that it treats the pixels in an image as a group of independent values. Therefore, if the pixels in the same image are rearranged, PCA will produce exactly the same principal components and reconstructions, which is generally considered insufficient to describe the spatial information in a 2D image.
[0111] In another embodiment, the LLRI generator includes a spatially low-rank neural network model. For example, LPCA can be Spatially Neural PCA. To extend PCA to capture spatial correlation, PCA with spatial context or SpatialPCA is introduced below. First, X is redefined as a 3D tensor with dimensions of (w, h, c), and the principal component vector w is defined as a set of spatial kernels of X with dimensions of (w’, h’, c, 1). Similar to PCA, the principal components can be calculated via a 2D convolutional layer as
[0112]
[0113] The goal of SpatialPCA is
[0114]
[0115] Equivalently, by using the Auto-Correlation AC(·),
[0116]
[0117] The auto-correlation function AC(·) calculates the covariance matrix of the input X and its shifted version of itself. The principal component vector can be solved similarly by iterating PCA, orthogonalization, and SVD.
[0118] For learnable PCA, SpatialPCA is extended to incorporate a supervised classifier to adapt the low-rank constraint in DL classification.
[0119] Before demonstrating the combined solution, the SpatialPCA implementation needs to be adapted using 2D convolutional operations. There are three computations that need to be mapped to conv2d:
[0120] · Calculate the auto-correlation of X.
[0121] · Calculate the principal components given X and w.
[0122] · Calculate the reconstruction of X given the truncated p.c. and w.
[0123] Given an input image of X with dimensions (n, w, h, c) and a learnable principal component vector w with dimensions (w’, h’, c, o), where w and h (or w’ and h’) are the width and height of the input (or filter), n is the size of the mini-batch, c is the number of channels for the input (the third dimension of w equals c), and o is the output dimension of conv2D, which in untruncated PCA satisfies o = w’ * h’ * c.
[0124] The principal components and the reconstruction of X can be computed via TensorFlow or an alternative DL framework.
[0125] Given the filter w truncated to i, X can be reconstructed as follows.
[0126]
[0127] Therefore, SpatialPCA can be solved by the following optimization
[0128]
[0129] s.t. w T · w = I (Equation 8b)
[0130] Several observations made by the inventors in numerical experiments include:
[0131] · The orthogonality constraint is strong to keep the model closer to PCA.
[0132] · The L1 loss of reconstruction is better than L2.
[0133] · The model can be used to enhance the difference signal depending on the target.
[0134] In one embodiment, the LLRI generator includes a learnable principal component analysis (PCA) model. Learnable PCA (LPCA) is introduced in the embodiments described herein to enhance weak signals in the presence of color differences. Traditionally, PCA is a method of selectively constructing a low-frequency reference by removing higher-order principal components. The focus of LPCA is slightly broader than PCA; in addition to removing color differences, LPCA is also expected to enhance significant signals simultaneously.
[0135]
[0136] LPCA is derived from the original PCA reconstruction formula by extending it to spatial 2D PCA, as demonstrated in Equation 9. (Note that the * operator in Equation 9 is the convolution operator.) T is the matrix transpose operator. Thus, PCA reconstruction can be expressed as a shallow CNN with two convolutional layers.
[0137]
[0138] such that ω T ·ω = 1 (Equation 10)
[0139] The LPCA solves for the low - rank filter via optimization (see Equation 10) rather than diagonalization. This method provides the freedom to directly link the LPCA network and subsequently link any detection or classification network.
[0140] In another embodiment, the LLRI generator includes a learnable independent component analysis (ICA) model or a learnable canonical correlation analysis (CCA) model. In this way, the LLRI generator can be a model of types such as PCA, ICA, CCA, etc. Additionally, the LLRI generator can be a tensor decomposition model. An example of a suitable ICA model configuration for use in the embodiments described herein can be found in Hyvärinen et al.'s “Independent Component Analysis: Algorithms and Applications” (Neural Networks, 13(4 - 5):411 - 430, 2000), which is incorporated herein by reference as if fully set forth. A description of an example of a CCA model suitable for use in the embodiments described herein is described in Uurtio et al.'s “A Tutorial on Canonical Correlation Methods” (arXiv:1711.02391, November 7, 2017, 33 pages), which is incorporated herein by reference as if fully set forth. An example of a suitable tensor decomposition model for use in the embodiments described herein is described in Rabanser et al.'s “Introduction to Tensor Decompositions and their Applications in Machine Learning” (arXiv:1711.10781, November 29, 2017, 13 pages), which is incorporated herein by reference as if fully set forth. The embodiments described herein can be further configured as described in these references.
[0141] In some embodiments, the LLRI generator includes a linear or non - linear regression model. Any suitable linear or non - linear regression model known in the art can be adapted for use as the LLRI generator described herein.
[0142] In yet another embodiment, the LLRI generator includes a spatially low-rank probability model. For example, the LLRI generator can be a Bayesian CNN. An example of a Bayesian CNN can be found in "A Comprehensive guide to Bayesian Convolutional Neural Network with Variational Inference" by Shridhar et al. (arXiv:1901.02731, January 8, 2019, 38 pages), which is incorporated herein by reference as if fully set forth. The spatially low-rank probability model can also be a probabilistic PCA model. Examples of probabilistic PCA models suitable for use in the embodiments described herein can be found in "Probabilistic Principal Component Analysis" by Tipping and Bishop (Journal of the Royal Statistical Society, Series B, 61, Part 3, pp. 611-622, September 27, 1999) and "Probabilistic Principal Component Analysis for 2D data" by Zhao et al. (Institute of Statistics: 58th World Statistics Congress, August 2011, Dublin, pp. 4416-4421), which are incorporated herein by reference as if fully set forth. The embodiments described herein can be further configured as described in these references.
[0143] In one embodiment, the different locations include locations in different dies on the sample. In another embodiment, the different locations include multiple locations in only one die on the sample. For example, one or more test images can be generated for corresponding locations in different dies on the sample, corresponding locations in different cells on the sample, etc. Multiple cells can be located in each die on the sample. Thus, test images can be generated at one or more locations of one or more cells in only one die on the sample.
[0144] In yet another embodiment, one or more test images correspond to job frames generated by an imaging system for a sample, and one or more computer systems are configured to repeat the input for one or more other test images corresponding to different job frames generated by the imaging system for the sample, such that the LLRI generator generates one or more reference images separately for the job frames and the different job frames. For example, there may be two phases using the LLRI generator: setup and runtime. At setup, defect candidates may be used to learn LLRI generation for each die. If no defect candidates are available, LPCA is reduced to normal PCA. At runtime, LLRI is generated for each job frame of the inspection area. In other words, different reference images may be generated separately for different job frames. The job frames for which the reference images are generated may include all job frames in which testing (inspection, defect re-detection, metrology, etc.) is performed. In this way, the input to the trained LLRI generator may include N frame images at the same relative die location for different dice in the case of die-to-die type inspection or N cell images from different cells for cell-to-cell type inspection, and the trained LLRI generator may output N reference images.
[0145] In some embodiments, the imaging system generates one or more test images for the sample using only a single mode of the imaging system. For example, the test images for which the reference images are generated may include only test images generated using the same mode. The modes of the imaging system may include any of the modes described herein.
[0146] In another embodiment, the imaging system generates one or more test images for the sample using only a single mode of the imaging system, and one or more computer systems are configured to repeat the input for one or more other test images generated by the imaging system for the sample using different modes of the imaging system, such that the LLRI generator generates one or more reference images for the one or more other test images. In this way, different reference images may be generated for test images generated in different modes. Generating reference images separately and independently for test images generated in different modes will be important for multi-mode inspection and other tests because the test images and the noise in the test images can vary significantly from one mode to another. Thus, the generated reference images suitable for use with one mode may not be equally suitable for use with a different mode. The different modes for which the reference images are generated may include any of the plurality of modes described herein.
[0147] The defect detection component detects defects on a sample based on one or more test images and their corresponding one or more reference images. Thus, the reference images generated by the LLRI generator can be used to detect defects on the sample. The defect detection performed using the generated reference images can include the defect detection described herein or any other suitable defect detection known in the art. In other words, the reference images generated by the LLRI generator can be used in any defect detection method in the same manner as any other reference image.
[0148] In one embodiment, the defect detection component is included in one or more components executed by one or more computer systems. In this way, the embodiment can include a combination of the LLRI generator and the defect detection component, and the defect detection component can be one of the supervised or unsupervised detectors / classifiers described herein. The defect detection component can thus be included in components executed by one or more computer systems included in the system. In other words, the systems described herein can perform defect detection using the generated reference images. Alternatively, the defect detection component can be included in another system that performs defect detection. For example, the reference images generated as described herein can be stored in a computer-readable medium accessible by another system or otherwise transmitted to or used by another system, such that the other system can perform defect detection using the generated reference images.
[0149] In one such embodiment, one or more computer systems are configured to jointly train the LLRI generator and the defect detection component using one or more training images and pixel-level ground truth information of the one or more training images. The pixel-level ground truth information can include labels generated by manual classification, an electron beam detection model, or a hybrid inspector. The hybrid inspection can be performed as described in U.S. Patent Application Publication No. 2017 / 0194126, published on July 6, 2017, by Bhaskar et al., which is incorporated herein by reference in its entirety. Then, the combined LLRI generator and defect detection component can be trained using the images (e.g., optical images) and the pixel-level ground truth information. The training of the combined LLRI generator and / or defect detection component can be performed in any suitable manner known in the art (e.g., modifying one or more parameters of the generator and / or defect detection component until the detection and / or classification results generated by the generator and / or defect detection component match the input ground truth information).
[0150] In some such embodiments, one or more training images include images of one or more defect categories selected by a user. For example, a user may determine the critical defect types in a given sample layer; otherwise, by default, all defect types are considered equally important. Candidate defect samples of the selected critical defect types may be obtained, for example, from user-acquired selected critical defect types through BBP defect discovery or electron beam inspection defect discovery. In this way, selected (not all) or all defect types may be assigned to the model for learning to achieve a target sensitivity enhancement. In a similar manner, selected (or not all) or all defect types may be assigned to the model for learning to achieve a target disruption reduction.
[0151] In another such embodiment, one or more training images include images of one or more hotspots on a sample selected by a user. A "hotspot" is generally defined in the art as a location in the design of a sample that is known (or suspected) to be more prone to defects. The hotspots may be selected by the user in any suitable manner. Hotspots may also be defined and discovered as described in the following U.S. patents: U.S. Patent No. 7,570,796 issued to Zafar et al. on August 4, 2009, and U.S. Patent No. 7,676,077 issued to Kulkarni et al. on March 9, 2010, which are hereby incorporated by reference in their entirety. The embodiments described herein may be further configured as described in these patents.
[0152] In additional such embodiments, one or more training images include images of one or more weak patterns in the design of a sample selected by a user. A "weak pattern" is generally defined in the art as a pattern in the design of a sample that is known (or suspected) to be more prone to defects than other patterns in the design. The weak patterns may be selected by the user in any suitable manner. Weak patterns may also be defined and discovered as described in the patents of Zafar and Kulkarni referenced above. In some examples, weak patterns may also be designated as hotspots (and vice versa), although this is not always true (i.e., a weak pattern may be identified as a hotspot in the design, but a hotspot does not necessarily have to be defined at a weak pattern, and vice versa).
[0153] In one embodiment, pixel-level ground truth information is generated by an electron beam imaging system. For example, the training ground truth data may come from an electron beam system, such as an electron beam inspection system, an electron beam defect re-inspection system, a SEM, a transmission electron microscope (TEM), etc. The electron beam imaging system may be further configured as described herein and may or may not be part of the system. For example, the systems described herein may be configured to use electron beam imaging to generate pixel-level ground truth information, and the computer systems described herein may generate ground truth information of electron beam images. Alternatively, another system or method may generate electron beam ground truth information, and this information may be obtained through the embodiments described herein.
[0154] In another embodiment, pixel-level ground truth information is generated by an optical-based system. For example, training ground truth data can be from an optical-based system, such as an optical inspection system (possibly configured for substantially high resolution or for use in high-resolution mode), an optical-based defect re-detection system, etc. The optical-based system can be further configured as described herein and may or may not be part of the system. For example, the systems described herein can be configured to use optical-based imaging to generate pixel-level ground truth information, and the computer systems described herein can generate ground truth information for optical-based images. Alternatively, another system or method can generate optical-based ground truth information, and this information can be obtained by the embodiments described herein.
[0155] In some embodiments, pixel-level ground truth information includes information received from a user. For example, the system can receive from a user ground truth information for one or more training images generated for a training sample. In one such instance, an electron beam image of the training sample can be displayed to the user, and the user can input information about the electron beam image, such as whether the image contains a defect and possibly what type of defect is shown in the image. This information can be obtained by the systems described herein by displaying the image to the user and providing the user with the ability to input information. This information can also or alternatively be obtained from another method or system described herein that obtains information from the user.
[0156] In yet another embodiment, pixel-level ground truth information includes information for one or more training images generated from the results of a physical simulation performed using one or more training images. The physical simulation can include any simulation known in the art. For example, for a defect shown in a training image, the physical simulation can include simulating how the defect will affect the physical phenomena of a device formed using the sample on which the defect is located. Such simulations can be performed in any suitable manner known in the art. Then, the results of the physical simulation can be used to generate additional information about the defect, and this additional information is used as pixel-level ground truth information. For example, the results of the physical simulation can show that the defect will cause a type of problem (e.g., short circuit, disconnection, etc.) in the device, and then a classification indicating that type of problem can be assigned to the defect. Then, these classifications can be used as the pixel-level ground truth information for the defect. Any other information that can be generated from this physical simulation can also or alternatively be used as pixel-level ground truth information. This pixel-level ground truth information can be performed by the embodiments described herein. Alternatively, pixel-level ground truth information generated using a physical simulation can be obtained from another method or system that generates this pixel-level ground truth information.
[0157] In yet another embodiment, the pixel-level ground truth information includes information converted from known defect locations in a second format different from the first format to the first format. For example, the known defect locations can be converted to pixel-level ground truth data. Additionally, defect information in one format (e.g., a KLARF file (which is a proprietary file format used by tools commercially available from KLA), a result file generated by Klarity which is a tool commercially available from KLA, batch results, etc.) can be converted to pixel-level ground truth data. The format to which the defect information is converted (i.e., the pixel-level ground truth information) can be the format of the images that will be input into the DML defect detection model during use and the output that will be generated by the DML defect detection model from the images (i.e., the input images and the labels that will be generated by the DML defect detection model for the input images).
[0158] The known defect locations can be "known" in several different ways, via inspection and re-inspection, via programmed or synthetic defects, via simulation, etc. The known defect locations can be converted to pixel-level ground truth data in any suitable manner. The known defect locations can be converted to pixel-level ground truth data by one or more of the computer systems of the embodiments described herein. Alternatively, another system or method can convert the known defect information to pixel-level ground truth data, and the embodiments described herein can obtain the pixel-level ground truth data from other methods or systems in any suitable manner. Some examples of systems and methods for obtaining known defect location data that can be converted to pixel-level ground truth data for use in the embodiments described herein are described in U.S. Patent Application Publication No. 2019 / 0303717, published on October 3, 2019, by Bhaskar et al., which patent application publication is incorporated herein by reference in its entirety as if fully set forth. The embodiments described herein can be further configured as described in such publication.
[0159] In one embodiment, the defect detection component includes a DL defect detection component. In this embodiment, the DL detection component may or may not be the DML defect detection model further described herein. In this way, in one embodiment of the LLRI-based detection, the LLRI generator can be combined with a DL CNN to form an end-to-end learning system to simultaneously learn the "optimal" approximate low-rank transformation and the pixel-level detector / classifier for the defect classes selected by the user. This embodiment thus combines the stability of low rank with the capabilities of the DL detector / classifier. The DL defect detection component can also be a detection model based on machine learning (ML) features, such as decision trees, random forests, support vector machines (SVMs), etc. The DL defect detection component can also be a DL-based detection model, such as CNN, Bayesian CNN, metric CNN, memory CNN, etc. Examples of memory CNNs are described in Braman et al., "Disease Detection in Weakly Annotated Volumetric Medical Images using a Convolutional LSTM Network" (arXiv:1812.01087, December 3, 2018, 4 pages) and Luo et al., "LSTM Pose Machines" (arXiv:1712.06316, March 9, 2018, 9 pages), which are incorporated herein by reference as if fully set forth. A metric CNN is a type of CNN that uses a similarity metric (metric) to determine whether two things (e.g., images) match. Examples of metric CNNs are described in Bell et al., "Learning visual similarity for product design with convolutional neural networks" (ACM Transactions on Graphics (TOG), Vol. 34, No. 4, July 2015, Article No.: 98, pp. 1-10) and Wang et al., "DARI: Distance Metric and Representation Integration for Person Verification" (Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, pp. 3611-3617, February 2016, published by AAAI Press, Palo Alto, California), which are incorporated herein by reference as if fully set forth.The embodiments described herein may be further configured as described in any or all of the above references.
[0160] In some embodiments, the defect detection component is not configured as a defect classifier. In other words, the defect detection component may detect an event on a sample but not identify the event as a defect of any type. This defect detection component may also or may not perform disruptive filtering of the detected event. However, the defect detection component may also perform defect classification and disruptive filtering as further described herein. Alternatively, a defect classification component (which may or may not be included in a component executed by one or more computer systems included in the system) may perform classification of the defects detected by the defect detection component. Some examples of defect classification and ML-based defect detection components are described in U.S. Patent Application Publication No. 2019 / 0073568, published Mar. 7, 2019, by He et al., which is incorporated herein by reference in its entirety as if fully set forth. Some examples of ML-based defect detectors are described in U.S. Patent No. 10,186,026, issued Jan. 22, 2019, to Karsenti et al., which is incorporated herein by reference in its entirety as if fully set forth. The embodiments described herein may be further configured as described in these references.
[0161] In yet another embodiment, the defect detection component includes a non-DL defect detection component. For example, the defect detection component may perform classical defect detection, such as subtracting the LLRI from a corresponding test image to produce a difference image and then using the difference image to detect defects on the sample. In one such example, the defect detection component may be a threshold-based defect detection component, where a reference image produced as described herein is subtracted from a corresponding test image, and the resulting difference image is compared to a threshold. In the simplest version of the threshold algorithm, any signal or output having a value above the threshold may be identified as a potential defect or defect candidate, and any signal or output not having a value above the threshold may not be identified as a potential defect or defect candidate. However, the threshold algorithm may be relatively complex compared to that described above in the MDAT algorithm and / or LCAT algorithm that may be available on some systems commercially available from KLA.
[0162] In another embodiment, the defect detection component includes a DML defect detection model. In this way, the LLRI generator may be combined with DML detection. In one such embodiment, the LLRI generator is shown as Figure 3 an LPCA block 300 in Figure 3As shown, the input to the LPCA block 300 can include N die images 302, each of the N die images 302 being acquired at the same relative die position but at different dies. The output of the LPCA block is N reference images after removing both low-frequency and high-frequency noise. For example, as Figure 3 shown, the output of the LPCA block is N reference images 304, one reference image 304 for each of the input die images.
[0163] In this configuration (other configurations are possible), both the original die images and the LPCA'd reference images are input to the DML detection model. For example, as Figure 3 shown, a pair of a test and the LPCA-generated reference image (e.g., a pair of test images 306 and the LPCA-generated reference image 308) can be input to the DL feature finder 310, which can output the features 312 of the test image and the features 314 of the reference image. A distance in the latent space is determined to decide whether each pixel in the test image is a defect. For example, as Figure 3 shown, the features 312 and 314 can be input to the DL latent projection 316 that can project the features into the latent space 318. Then, defect detection can be performed based on the distance between the features in the latent space, as further described herein. The layers in the detection block are different from the layers in the LPCA block. The loss function is a combination of the LPCA loss and the DML loss (e.g., siamese loss). Although using the LPCA-generated reference images with DML detection can provide relatively high sensitivity compared to other detection methods and systems, the LPCA-generated reference images can be used with any other defect detection algorithms (e.g., those further described herein).
[0164] In one embodiment, the LLRI generator and the defect detection component are configured for in-line defect detection. For example, the reference images can be generated while the sample is being scanned by the imaging system (i.e., in real-time when generating the test images) and the defect detection can be performed using the generated reference images. In this way, the reference images are not generated before the sample is scanned.
[0165] In another embodiment, the LLRI generator and the defect detection component are configured for off-line defect detection. For example, the reference images can be generated after the imaging system has scanned the sample (i.e., after the test images have been generated) and the defect detection can be performed using the generated reference images. In one such instance, the test images can be generated and stored by the system including the imaging system that generated the test images in a computer-readable storage medium. The embodiments described herein can then access the test images in the storage medium and use the test images as described herein.
[0166] In some embodiments, one or more components include a defect classification component configured to classify detected defects into two or more types, and the defect classification component is a DL defect classification component. The DL defect classification component may be configured as described in U.S. Patent No. 10,043,261, issued Aug. 7, 2018 to Bhaskar et al., U.S. Patent No. 10,360,477, issued Jul. 23, 2019 to Bhaskar et al., and U.S. Patent Application Publication No. 2019 / 0073568, published Mar. 7, 2019 to He et al., which patents are incorporated herein by reference in their entirety. The embodiments described herein may be further configured as described in these publications.
[0167] In one embodiment, the sample is a wafer on which a layer of patterned features has been formed using multiple lithographic exposure steps. For example, the sample may be a double-patterned wafer or other patterned wafer where a first set of patterned features is formed on one layer of the wafer in one lithographic exposure step and a second set of patterned features is formed on the same layer of the wafer in another lithographic exposure step. The multiple lithographic exposure steps may be performed in any suitable manner known in the art. In another embodiment, the sample is a wafer on which a layer of patterned features has been formed using extreme ultraviolet (EUV) lithography. The EUV lithography may be performed in any suitable manner known in the art. For example, the embodiments described herein provide an approximate low-rank supervised / semi-supervised defect detection algorithm with enhanced sensitivity for target defects, particularly for double / quadruple patterning and smaller defects in EUV lithography on optical and other inspection tools.
[0168] Implementations of the architectures of the embodiments described herein may be implemented entirely on GPU acceleration. For example, the LPCA block may be implemented as two convolutional layers that may be configured as relatively small neural networks. Both training (setup) and inference may run directly on the GPU. The architecture may also run on a CPU or other accelerator chip.
[0169] Each of the embodiments of each of the systems described above may be combined together into a single embodiment. The embodiments described herein may also be further configured as described in the following U.S. Patent Application Publications: U.S. Patent Application Publication No. 2017 / 0194126, published Jul. 6, 2017 to Bhaskar et al., and U.S. Patent Application Publication No. 2018 / 0342051, published Nov. 19, 2018 to Sezginer et al., which patent application publications are incorporated herein by reference in their entirety.
[0170] The embodiments described herein have several advantages over other methods and systems for defect detection. For example, the embodiments described herein can handle intra-wafer and (possibly) inter-wafer process variations (e.g., because the reference images generated as described herein will be substantially unaffected by intra-wafer and inter-wafer process variations). Another advantage is that the reference images generated by the embodiments described herein can be learned to achieve optimal defect detection for user-specified defects, hotspots, or weak patterns. This method will limit the degradation of unpredictable sensitivity. Additionally, the embodiments described herein can enhance the defect sensitivity for target defect types, thereby extending the BBP defect sensitivity limit. Further, the embodiments described herein can greatly enhance the BBP sensitivity and reduce the acquisition cost of BBP tools by providing better availability of the target sensitivity in both research and development and high-volume manufacturing use cases. The embodiments described herein can also advantageously combine defect classifier learning with reference generation, which permits selective sensitivity enhancement. Additionally, when applying an approximate low-rank constraint, the embodiments described herein have a much lower requirement for labeled defect candidates. An additional advantage of the embodiments described herein is that they can be used for a variety of different inspection types. For example, as further described herein, the system can be configured for die-to-die inspection, cell-to-cell inspection, and standard reference inspection, each of which can be performed in only a single optical mode or multiple optical modes.
[0171] Another embodiment relates to a computer-implemented method for detecting defects on a sample. The method includes projecting a test image and a corresponding reference image generated for the sample into a latent space. The method also includes determining, for one or more different portions of the test image, the distance in the latent space between the one or more different portions and the corresponding one or more portions of the corresponding reference image. Additionally, the method includes detecting defects in the one or more different portions of the test image based, respectively, on the distances determined for the one or more different portions of the test image. The projecting, determining, and detecting steps are performed by a DML defect detection model included in one or more components executed by one or more computer systems.
[0172] Each of the steps of the method can be performed as further described herein. The method can also include any other steps described herein. The computer system can be configured according to any of the embodiments described herein, e.g., computer subsystem 102. Additionally, the method described above can be performed by any of the system embodiments described herein.
[0173] Another embodiment relates to a method for generating a reference image of a sample. The method includes removing noise from one or more test images of the sample by inputting the one or more test images into an LLRI generator, thereby generating one or more reference images corresponding to the one or more test images. The one or more test images are generated for different positions on the sample corresponding to the same position in the design of the sample. The LLRI generator is included in one or more components executed by one or more computer systems. The method may further include detecting a defect on the sample based on the one or more test images and their corresponding one or more reference images. The detection may be performed by a defect detection component that may or may not be included in the one or more components executed by the one or more computer systems.
[0174] Each of the steps of this method may be performed as further described herein. This method may further include any other steps described herein. These computer systems may be configured according to any of the embodiments described herein, such as computer subsystem 102. Additionally, the methods described above may be performed by any of the system embodiments described herein.
[0175] An additional embodiment relates to a non-transitory computer-readable medium that stores program instructions that may be executed on one or more computer systems to perform a computer-implemented method for detecting a defect on a sample and / or generating a reference image of the sample. Figure 4 One such embodiment is shown in Figure 4 As shown in
[0176] The program instructions 402 for implementing a method such as those described herein may be stored on the computer-readable medium 400. The computer-readable medium may be a storage medium such as a magnetic disk or optical disk, a magnetic tape, or any other suitable non-transitory computer-readable medium known in the art.
[0177] The program instructions may be implemented in any of a variety of ways including techniques based on program steps, component-based techniques, and / or object-oriented techniques, as well as other techniques. For example, ActiveX controls, C++ objects, JavaBeans, Microsoft Foundation Classes (“MFC”), SSE (Streaming SIMD Extensions), or other techniques or methods may be used to implement the program instructions as needed.
[0178] The computer system 404 may be configured according to any of the embodiments described herein.
[0179] In view of this description, those skilled in the art will appreciate additional modifications and alternative embodiments of various aspects of the present invention. For example, a learnable defect detection method and system for semiconductor applications are provided. Accordingly, this description should be regarded as merely illustrative and is for the purpose of teaching those skilled in the art the general manner of implementing the present invention. It should be understood that the forms of the present invention shown and described herein should be regarded as presently preferred embodiments. As will be fully appreciated by those skilled in the art after benefiting from this description of the present invention, the elements and materials may be substituted for those illustrated and described herein, the components and processes may be reversed, and certain features of the present invention may be utilized independently. Changes may be made to the elements described herein without departing from the spirit and scope of the present invention as described in the appended claims.
Claims
1. A system configured to detect defects on a sample, comprising: One or more computer systems; and One or more components, which are executed by the one or more computer systems, wherein the one or more components include a deep metric learning defect detection model, and the deep metric learning defect detection model is configured to: Project a test image and a corresponding reference image generated for a sample into a latent space; For one or more different portions of the test image, determine distances between the one or more different portions and corresponding one or more portions of the corresponding reference image in the latent space; and Detect defects in the one or more different portions of the test image respectively based on the distances determined for the one or more different portions of the test image.
2. The system according to claim 1, wherein the test image and the corresponding reference image are for corresponding positions in different dies on the sample.
3. The system according to claim 1, wherein the test image and the corresponding reference image are for corresponding positions in different cells on the sample.
4. The system according to claim 1, wherein the test image and the corresponding reference image are generated for the sample without using the design data of the sample.
5. The system according to claim 1, wherein the test image is generated for the sample by an imaging system that directs energy to the sample and detects energy from the sample, and wherein the corresponding reference image is generated without using the sample.
6. The system according to claim 5, wherein the corresponding reference image is obtained from a database containing the design data of the sample.
7. The system according to claim 1, wherein the one or more computer systems are configured to input the design data of the sample into the deep metric learning defect detection model, and wherein the deep metric learning defect detection model is further configured to perform the detection using the design data.
8. The system according to claim 1, wherein the detection is performed using one or more parameters determined from the region of interest of the sample.
9. The system according to claim 8, wherein the one or more computer systems are configured to input the information of the region of interest into the deep metric learning defect detection model.
10. The system according to claim 1, wherein the detection is performed without information of the region of interest of the sample.
11. The system according to claim 1, wherein the test image is generated in the logic region of the sample.
12. The system according to claim 1, wherein the test image is generated in the array region of the sample.
13. The system according to claim 1, wherein the different parts of the test image include different pixels in the test image.
14. The system according to claim 1, wherein the deep metric learning defect detection model is further configured to project additional corresponding reference images into the latent space and determine an average value of the corresponding reference image and the additional corresponding reference images with a reference region in the latent space, and wherein the one or more parts of the corresponding reference image for determining the distance include the reference region.
15. The system according to claim 1, wherein the corresponding reference image includes a defect-free test image of the sample, wherein projecting the corresponding reference image includes learning the reference region in the latent space, and wherein the one or more parts of the corresponding reference image for determining the distance include the reference region.
16. The system according to claim 1, wherein the deep metric learning defect detection model has a siamese network architecture.
17. The system according to claim 1, wherein the deep metric learning defect detection model has a triplet network architecture.
18. The system according to claim 1, wherein the deep metric learning defect detection model has a quadruplet network architecture.
19. The system according to claim 1, wherein the deep metric learning defect detection model includes one or more deep learning convolutional filters, and wherein the one or more computer systems are configured to determine a configuration of the one or more deep learning convolutional filters based on a physical phenomenon involved in generating the test image.
20. The system according to claim 1, wherein the deep metric learning defect detection model includes one or more deep learning convolutional filters, and wherein the one or more computer systems are configured to determine a configuration of the one or more deep learning convolutional filters based on imaging hardware used to generate the test image.
21. The system according to claim 20, wherein determining the configuration includes determining one or more parameters of the one or more deep learning convolutional filters based on a point spread function of the imaging hardware.
22. The system according to claim 21, wherein the one or more parameters of the one or more deep learning convolutional filters include one or more of a filter size, a filter symmetry, and a filter depth.
23. The system according to claim 21, wherein determining the one or more parameters of the one or more deep learning convolutional filters includes learning the one or more parameters by optimizing a loss function.
24. The system according to claim 20, wherein determining the configuration includes selecting the one or more deep learning convolutional filters from a predetermined set of deep learning convolutional filters based on a point spread function of the imaging hardware.
25. The system according to claim 24, wherein one or more parameters of the one or more deep learning convolutional filters in the predetermined set are fixed.
26. The system according to claim 24, wherein determining the configuration further includes fine-tuning one or more initial parameters of the one or more deep learning convolutional filters by optimizing a loss function.
27. The system according to claim 1, wherein the one or more components further include a learnable low-rank reference image generator configured to generate the corresponding reference image, wherein the one or more computer systems are configured to input one or more test images generated for the sample into the learnable low-rank reference image generator, wherein the one or more test images are generated for different positions on the sample, the different positions corresponding to the same position in the design of the sample, and wherein the learnable low-rank reference image generator is further configured to remove noise from the one or more test images, thereby generating the corresponding reference image.
28. The system according to claim 1, wherein the test image and an additional test image are generated for the sample using different modes of the imaging system respectively; wherein the deep metric learning defect detection model is further configured to project the test image and the corresponding reference image into a first latent space, project the additional test image and an additional corresponding reference image into a second latent space, and combine the first latent space and the second latent space into a joint latent space; and wherein the latent space for determining the distance is the joint latent space.
29. The system according to claim 1, wherein the one or more computer systems are configured to input design data of the sample into the deep metric learning defect detection model; wherein the test image and an additional test image are generated for the sample using different modes of the imaging system respectively; wherein the deep metric learning defect detection model is further configured to project the test image and the corresponding reference image into a first latent space, project the additional test image and an additional corresponding reference image into a second latent space, project the design data into a third latent space, and combine the first latent space, the second latent space and the third latent space into a joint latent space; and wherein the latent space for determining the distance is the joint latent space.
30. The system according to claim 1, wherein the one or more computer systems are configured to input design data of the sample into the deep metric learning defect detection model; wherein the deep metric learning defect detection model is further configured to project the test image and the corresponding reference image into a first latent space, project the design data into a second latent space, and combine the first latent space and the second latent space into a joint latent space; and wherein the latent space for determining the distance is the joint latent space.
31. The system according to claim 1, wherein the one or more computer systems are configured to input design data of the sample into the deep metric learning defect detection model; wherein the test image and an additional test image are generated for the sample using different modes of an imaging system; wherein the deep metric learning defect detection model is further configured to project a first set including one or more of the test image and the corresponding reference image, the additional test image and the additional corresponding reference image, and the design data into a first latent space, project a second set including one or more of the test image and the corresponding reference image, the additional test image and the additional corresponding reference image, and the design data into a second latent space, and combine the first latent space and the second latent space into a joint latent space; and wherein the latent space for determining the distance is the joint latent space.
32. The system according to claim 1, wherein the one or more computer systems are configured to train the deep metric learning defect detection model using one or more training images and pixel-level ground truth information of the one or more training images.
33. The system according to claim 32, wherein the one or more training images and the pixel-level ground truth information are generated from a process window qualification wafer.
34. The system according to claim 1, wherein the one or more computer systems are configured to perform active learning to train the deep metric learning defect detection model.
35. The system according to claim 1, wherein the sample is a wafer.
36. The system according to claim 1, wherein the sample is a reticle.
37. A system configured to generate a reference image of a sample, comprising: One or more computer systems; and One or more components, which are executed by the one or more computer systems, wherein the one or more components include a learnable low-rank reference image generator, and the one or more computer systems are configured to input one or more test images of a sample into the learnable low-rank reference image generator, wherein the one or more test images are generated for different positions on the sample, the different positions corresponding to the same position in the design of the sample, and wherein the learnable low-rank reference image generator is configured to remove noise from the one or more test images, thereby generating one or more reference images corresponding to the one or more test images; and wherein a defect detection component detects defects on the sample based on the one or more test images and their corresponding one or more reference images.
38. The system according to claim 37, wherein the defect detection component includes a deep learning defect detection component.
39. The system according to claim 37, wherein the defect detection component includes a deep metric learning defect detection model.
40. The system according to claim 37, wherein the defect detection component includes a non-deep learning defect detection component.
41. The system according to claim 37, wherein the sample is a wafer, and a layer of patterned features has been formed on the wafer using a plurality of lithographic exposure steps.
42. The system according to claim 37, wherein the sample is a wafer, and a layer of patterned features has been formed on the wafer using extreme ultraviolet lithography.
43. The system according to claim 37, wherein the learnable low-rank reference image generator includes a learnable principal component analysis model.
44. The system according to claim 37, wherein the learnable low-rank reference image generator includes a learnable independent component analysis model or a learnable canonical correlation analysis model.
45. The system according to claim 37, wherein the learnable low-rank reference image generator includes a linear or non-linear regression model.
46. The system according to claim 37, wherein the learnable low-rank reference image generator includes a spatial low-rank neural network model.
47. The system according to claim 37, wherein the learnable low-rank reference image generator includes a spatial low-rank probability model.
48. The system according to claim 37, wherein the defect detection component is included in the one or more components executed by the one or more computer systems.
49. The system according to claim 48, wherein the one or more computer systems are further configured to jointly train the learnable low-rank reference image generator and the defect detection component using one or more training images and pixel-level ground truth information of the one or more training images.
50. The system according to claim 49, wherein the one or more training images include images of one or more defect categories selected by a user.
51. The system according to claim 49, wherein the one or more training images include images of one or more hotspots on the sample selected by a user.
52. The system according to claim 49, wherein the one or more training images include images of one or more weak patterns in the design of the sample selected by a user.
53. The system according to claim 49, wherein the pixel-level live information is generated by an electron beam imaging system.
54. The system according to claim 49, wherein the pixel-level live information is generated by an optical-based system.
55. The system according to claim 49, wherein the pixel-level live information includes information received from a user.
56. The system according to claim 49, wherein the pixel-level live information includes information of the one or more training images generated based on results of physical simulations performed using the one or more training images.
57. The system according to claim 49, wherein the pixel-level live information includes information that is converted from known defect positions in a second format to a first format, the second format being different from the first format.
58. The system according to claim 37, wherein the learnable low-rank reference image generator and the defect detection component are further configured for in-line defect detection.
59. The system according to claim 37, wherein the learnable low-rank reference image generator and the defect detection component are further configured for out-of-line defect detection.
60. The system according to claim 37, wherein the one or more components further include a defect classification component configured to classify the detected defects into two or more types, and wherein the defect classification component is a deep learning defect classification component.
61. The system according to claim 37, wherein the different positions include positions in different dies on the sample.
62. The system according to claim 37, wherein the different positions include multiple positions in only one die on the sample.
63. The system according to claim 37, wherein the one or more test images correspond to operation frames generated by an imaging system for the sample, and wherein the one or more computer systems are further configured to repeat the input for one or more other test images corresponding to different operation frames generated by the imaging system for the sample, such that the learnable low-rank reference image generator generates the one or more reference images separately for the operation frames and the different operation frames.
64. The system according to claim 37, wherein the one or more test images are generated by the imaging system for the sample using only a single mode of the imaging system.
65. The system according to claim 37, wherein the one or more test images are generated for the sample by the imaging system using only a single mode of the imaging system, and wherein the one or more computer systems are further configured to repeat the input for one or more other test images generated for the sample by the imaging system using different modes of the imaging system, such that the learnable low-rank reference image generator generates the one or more reference images for the one or more other test images.
Citation Information
Patent Citations
Generating simulated output for a specimen
US10043261B2
Single image detection
US10186026B2
Accelerating semiconductor-related computations using learning based models
US10360477B2
Hybrid inspectors
US20170194126A1
Unified neural network for defect detection and classification
US20190073568A1