Direct edge placement error measurement using a high voltage scanning charged particle microscope and machine learning

By integrating a high voltage scanning charged particle microscope (HV-SCPM) with machine learning techniques to correct and enhance lower layer images, the method addresses the inaccuracies in current EPE measurement techniques, achieving improved precision in edge placement error measurements for semiconductor manufacturing.

WO2025103678A1PCT designated stage expired Publication Date: 2025-05-22ASML NETHERLANDS BV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/078864
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-14
Filing Date
2024-10-14
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Current methods for measuring edge placement error (EPE) in semiconductor wafer manufacturing are either indirect and prone to errors due to combining different metrology results, or they result in blurry images of lower layers using high voltage scanning charged particle microscopes (HV-SCPMs), leading to inaccurate critical dimension (CD) and placement error (PE) measurements.

Method used

The use of a high voltage scanning charged particle microscope (HV-SCPM) in conjunction with machine learning to obtain accurate edge placement error measurements. This involves obtaining images of the lower and upper layers, correcting distortions, generating a mask, and training a machine learning model to enhance the sharpness of the lower layer image, thereby improving the accuracy of EPE measurements.

Benefits of technology

This approach allows for more accurate edge placement error measurements by enhancing the sharpness of lower layer images, thereby improving the precision of critical dimension and placement error measurements, which is crucial for ensuring the fidelity of semiconductor patterns and predicting yield.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024078864_22052025_PF_FP_ABST
    Figure EP2024078864_22052025_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus for obtaining an edge placement error measurement includes a memory storing a set of instructions and at least one processor configured to execute the set of instructions to cause the apparatus to perform: obtain an image of a lower layer of a sample before a upper layer is deposited during a manufacturing process; obtain a combined image of the upper layer and the lower layer of the sample after the lower layer image is obtained; correct distortions in the lower layer image; generate a mask based on the combined image; apply the mask to the lower layer image; and train a machine learning model using the masked lower layer image and the combined image, wherein the trained machine learning model is configured to determine the edge placement error measurement for the combined image.
Need to check novelty before this filing date? Find Prior Art

Description

DIRECT EDGE PLACEMENT ERROR MEASUREMENT USING A HIGH VOLTAGE SCANNING CHARGED PARTICLE MICROSCOPE AND MACHINE LEARNINGCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority of EP application 23209568.7 which was filed on 14 November 2023 and which is incorporated herein in its entirety by reference.TECHNICAL FIELD

[0002] The embodiments provided herein relate to measuring edge placement error during semiconductor wafer manufacturing, and more particularly to using a high voltage scanning charged particle microscope and machine learning to measure edge placement error.BACKGROUND

[0003] Edge placement error (EPE) is an important measurement for wafer process control and monitoring. EPE is the difference between an intended feature of a semiconductor layout and the actual printed feature of the semiconductor layout. The fidelity of the final pattern is described in terms of EPE, which is defined as the relative displacement of the edges of a feature from their intended target position on the wafer. The EPE measures multi-layer geometry information and consists of multiple layers of information to provide an early indication of yield prediction, and helps to shorten the process control feedback loop from weeks to days. There are two commonly used ways to measure multi-layer EPE. One way is to measure each individual layer EPE with a scanning charged-particle microscope (SCPM) and construct the multi-layer EPE by combining overlay information (e.g., using information from an optical wafter metrology system for measuring overlay). A second way is to directly measure multi-layer EPE by see-through SCPM (e.g., high voltage (HV) SCPM).SUMMARY

[0004] Some embodiments provide an apparatus for performing operations for obtaining an edge placement error measurement of a sample. The apparatus can include a memory storing a set of instructions and at least one processor configured to execute the set of instructions to cause the apparatus to perform: obtaining an image of a lower layer of a sample before a upper layer is deposited during a manufacturing process; obtaining a combined image of the upper layer and the lower layer of the sample after the lower layer image is obtained; correcting distortions in the lower layer image; generating a mask based on the combined image; applying the mask to the lower layer image; and training a machine learning model using the masked lower layer image and the combined image, wherein the trained machine learning model is configured to determine the edge placement error measurement for the combined image.

[0005] Other advantages of the embodiments of the present disclosure will become apparent from the following description taken in conjunction with the accompanying drawings wherein are set forth, by way of illustration and example, certain embodiments of the present invention.BRIEF DESCRIPTION OF FIGURES

[0006] The above and other aspects of the present disclosure will become more apparent from the description of exemplary embodiments, taken in conjunction with the accompanying drawings.

[0007] Figs. 1A and IB are schematic diagrams illustrating ways to measure EPE, consistent with some embodiments of the present disclosure.

[0008] Fig. 2 is a schematic diagram illustrating an example charged-particle beam inspection (CPBI) system, consistent with some embodiments of the present disclosure.

[0009] Fig. 3 is a schematic diagram illustrating an example charged-particle beam tool, consistent with some embodiments of the present disclosure that may be a part of the example charged-particle beam inspection system of Fig. 2.

[0010] Fig. 4 is a schematic diagram illustrating an example neural network, consistent with some embodiments of the present disclosure.

[0011] Fig. 5 is a diagram of an example workflow for training a machine learning model, consistent with embodiments of the present disclosure.

[0012] Fig. 6 is a diagram of an example method for using a HV-SCPM and machine learning to construct EPE, consistent with embodiments of the present disclosure.

[0013] Fig. 7 is a flowchart of an example method for EPE measurement, consistent with embodiments of the present disclosure.DETAILED DESCRIPTION

[0014] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which the same numbers in different drawings represent the same or similar elements unless otherwise represented. The implementations set forth in the following description of exemplary embodiments do not represent all implementations. Instead, they are merely examples of apparatuses and methods consistent with aspects related to the disclosed embodiments as recited in the appended claims. For example, although some embodiments are described in the context of utilizing electron beams, the disclosure is not so limited. Other types of charged-particle beams (e.g., including protons, ions, muons, or any other particle carrying electric charges) may be similarly applied. Furthermore, other imaging systems may be used, such as optical imaging, photon detection, x-ray detection, ion detection, etc.

[0015] Relative dimensions of components in drawings may be exaggerated for clarity. Within the following description of drawings, the same or like reference numbers refer to the same or likecomponents or entities, and only the differences with respect to the individual embodiments are described. As used herein, unless specifically stated otherwise, the term “or” encompasses all possible combinations, except where infeasible. For example, if it is stated that a component may include A or B, then, unless specifically stated otherwise or infeasible, the component may include A, or B, or A and B. As a second example, if it is stated that a component may include A, B, or C, then, unless specifically stated otherwise or infeasible, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0016] Electronic devices are constructed of circuits formed on a piece of semiconductor material called a substrate. The semiconductor material may include, for example, silicon, gallium arsenide, indium phosphide, or silicon germanium, or the like. Many circuits may be formed together on the same piece of silicon and are called integrated circuits or ICs. The size of these circuits has decreased dramatically so that many more of them can be fit on the substrate. For example, an IC chip in a smartphone can be as small as a thumbnail and yet may include over 2 billion transistors, the size of each transistor being less than l / 1000th the size of a human hair.

[0017] Making these ICs with extremely small structures or components is a complex, time-consuming, and expensive process, often involving hundreds of individual steps. Errors in even one step have the potential to result in defects in the finished IC, rendering it useless. Thus, one goal of the manufacturing process is to avoid such defects to maximize the number of functional ICs made in the process; that is, to improve the overall yield of the process.

[0018] One component of improving yield is monitoring the chip-making process to ensure that it is producing a sufficient number of functional integrated circuits. One way to monitor the process is to inspect the chip circuit structures at various stages of their formation. Inspection can be carried out using a scanning charged-particle microscope (SCPM). For example, an SCPM may be a scanning electron microscope (SEM). An SCPM can be used to image these extremely small structures, in effect, taking a “picture” of the structures of the wafer. The image can be used to determine if the structure was formed properly in the proper location. If the structure is defective, then the process can be adjusted, so the defect is less likely to recur.

[0019] The working principle of an SCPM (e.g., an SEM) is similar to a camera. A camera takes a picture by receiving and recording intensity of light reflected or emitted from people or objects. An SCPM takes a “picture” by receiving and recording energies or quantities of charged particles (e.g., electrons) reflected or emitted from the structures of the wafer. Typically, the structures are made on a substrate (e.g., a silicon substrate) that is placed on a platform, referred to as a stage, for imaging. Before taking such a “picture,” a charged-particle beam may be projected onto the structures, and when the charged particles are reflected or emitted (“exiting”) from the structures (e.g., from the wafer surface, from the structures underneath the wafer surface, or both), a detector of the SCPM may receive and record the energies or quantities of those charged particles to generate an inspection image. To take such a “picture,” the charged-particle beam may scan through the wafer (e.g., in a line-by-line or zig-zag manner), and the detector may receive exiting charged particles coming from a region under charged particle -beam projection (referred to as a “beam spot”). The detector may receive and record exiting charged particles from each beam spot one at a time and join the information recorded for all the beam spots to generate the inspection image. Some SCPMs use a single charged-particle beam (referred to as a “single-beam SCPM,” such as a single-beam SEM) to take a single “picture” to generate the inspection image, while some SCPMs use multiple charged-particle beams (referred to as a “multi-beam SCPM,” such as a multi-beam SEM) to take multiple “sub-pictures” of the wafer in parallel and stitch them together to generate the inspection image. By using multiple charged-particle beams, the SCPM may provide more charged-particle beams onto the structures for obtaining these multiple “sub-pictures,” resulting in more charged particles exiting from the structures. Accordingly, the detector may receive more exiting charged particles simultaneously and generate inspection images of the structures of the wafer with higher efficiency and faster speed.

[0020] As the physical sizes of IC components continue to shrink, accuracy and yield in defect detection become more important. Metrology tools can be used to determine whether the ICs are correctly manufactured by identifying a number of defects on each wafer, including at different levels of detail, such as a pattern level, an image (field of view) level, a die level, a care area level, or a wafer level.

[0021] There are two commonly used ways to measure multi-layer EPE, as shown in Figs. 1A and IB. A first way to measure multi-layer EPE is indirect measurement, as shown in Fig. 1A, by combining the SCPM results, overlay measurement results, and statistics to construct the EPE. The multi-layer EPE accuracy is impacted by combining different metrology errors and matching between different types of tools. Overlay information is added for the constructed EPE, because pattern level overlay information is missing, which may limit the EPE accuracy at the pattern level. The indirect measurement lacks 3D information and takes extra metrology measurements to provide feedback for the process control loop compared with the direct measurements described in connection with Fig. IB.

[0022] For example, in the first way, the lower layer of the semiconductor is manufactured, and a first SCPM image is taken of the lower layer. The upper layer of the semiconductor is manufactured on top of the lower layer, and a second SCPM image is taken. The two SCPM images are aligned and the EPE is calculated based on the differences between the two images and overlay information. Overlay may be impacted by local pattern density variations between the upper layer and the lower layer, providing additional error in calculating the EPE.

[0023] A second way to measure multi-layer EPE, as shown in Fig. IB, directly sees two layers at one time by one tool, providing a one-time measurement (i.e., “looking” through an upper layer to capture an image of the upper layer and a lower layer). However, the contours of the lower layer are blurred or may not be visible in this image (e.g., because the upper layer may be physically on top of the edges of the lower layer), leading to inaccurate EPE measurements.

[0024] For example, in the second way, a HV-SCPM image is taken after the lower layer and the upper layer have been manufactured. A clear, sharp image is made of the upper layer because it is a direct measurement of the upper layer. For the lower layer, the quality of the image depends on the power of the e-beam used and how much penetration of the upper layer is achieved. The energy level of the e- beam used depends on the layer thickness (e.g., a higher power e-beam is more likely to achieve better penetration of the upper layer). However, using a high energy e-beam results in e-beam scattering, such as shown by the circle-like object in Fig. IB. The e-beam scattering leads to blurry images of the lower layer (e.g., the bottom portion of Fig. IB). Because of the blurry images, the critical dimension (CD) measurement and placement error (PE) measurement of the lower layer may be inaccurate. For example, the CD measurement may indicate whether the feature was manufactured at the correct size (e.g., not too large) and the PE measurement may indicate whether the feature was manufactured at the intended location (e.g., whether there is a difference between the intended location of the feature and the actual location of the manufactured feature). The CD measurement and the PE measurement may be used as inputs to determine the EPE measurement. A drawback to using the HV-SCPM is that the charging to get to the high voltage influences the blurring of the image by creating noise (e.g., the signal to noise ratio is lower with more charging). Also, the charge dissipates over time and can have different effects on the image depending on an amount of time that passes between charging the HV-SCPM and when the image is taken.

[0025] Current ways of obtaining an EPE measurement during semiconductor manufacturing are described in connection with Figs. 1A and IB. For example, by taking separate SCPM images of the upper layer and the lower layer or by taking a HV-SCPM image of the combined upper layer and lower layer.

[0026] Embodiments of the present disclosure can provide a way to measure EPE using a HV-SCPM image and machine learning, by using machine learning to “correct” the lower layer of the HV-SCPM image of the combined upper layer and lower layer. According to some embodiments of the present disclosure, a scanning charged-particle microscope (SCPM) image (such as an SEM image) of a lower layer of a sample is obtained before an upper layer is deposited during a manufacturing process. A high voltage SCPM (HV-SCPM) image of the upper layer and the lower layer of the sample is obtained after the SCPM image is obtained. Distortions in the SCPM image of the lower layer are corrected. A mask is generated based on the HV-SCPM image and the mask is applied to the SCPM image to obtain a clear, sharp image of the lower layer. A machine learning model is trained using the masked SCPM image and the HV-SCPM image to improve the sharpness of the lower layer in the HV-SCPM image. The trained machine learning model is applied to the HV-SCPM image to determine the edge placement error measurement.

[0027] Fig. 2 illustrates an exemplary charged-particle beam inspection (CPBI) system 200 consistent with some embodiments of the present disclosure. CPBI system 200 may be used for imaging. For example, CPBI system 200 may use an electron beam for imaging. As shown in Fig. 2, CPBI system200 includes a main chamber 201, a load / lock chamber 202, a beam tool 204, and an equipment front end module (EFEM) 206. Beam tool 204 is located within main chamber 201. EFEM 206 includes a first loading port 206a and a second loading port 206b. EFEM 206 may include additional loading port(s). First loading port 206a and second loading port 206b receive wafer front opening unified pods (FOUPs) that contain wafers (e.g., semiconductor wafers or wafers made of other material(s)) or samples to be inspected (the terms “wafers” and “samples” may be used interchangeably). A “lot” is a plurality of wafers that may be loaded for processing as a batch.

[0028] One or more robotic arms (not shown) in EFEM 206 may transport the wafers to load / lock chamber 202. Load / lock chamber 202 is connected to a load / lock vacuum pump system (not shown) which removes gas molecules in load / lock chamber 202 to reach a first pressure below the atmospheric pressure. After reaching the first pressure, one or more robotic arms (not shown) may transport the wafer from load / lock chamber 202 to main chamber 201. Main chamber 201 is connected to a main chamber vacuum pump system (not shown) which removes gas molecules in main chamber 201 to reach a second pressure below the first pressure. After reaching the second pressure, the wafer is subject to inspection by beam tool 204. Beam tool 204 may be a single-beam system or a multi-beam system.

[0029] A controller 209 is electronically connected to beam tool 204. Controller 209 may be a computer that may execute various controls of CPBI system 200. While controller 209 is shown in Fig. 2 as being outside of the structure that includes main chamber 201, load / lock chamber 202, and EFEM 206, it is appreciated that controller 209 may be a part of the structure.

[0030] In some embodiments, controller 209 may include one or more processors (not shown). A processor may be a generic or specific electronic device capable of manipulating or processing information. For example, the processor may include any combination of any number of a central processing unit (or “CPU”), a graphics processing unit (or “GPU”), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property (IP) core, a Programmable Logic Array (PLA), a Programmable Array Logic (PAL), a Generic Array Logic (GAL), a Complex Programmable Logic Device (CPLD), a Field- Programmable Gate Array (FPGA), a System On Chip (SoC), an Application-Specific Integrated Circuit (ASIC), and any type circuit capable of data processing. The processor may also be a virtual processor that includes one or more processors distributed across multiple machines or devices coupled via a network.

[0031] In some embodiments, controller 209 may further include one or more memories (not shown). A memory may be a generic or specific electronic device capable of storing codes and data accessible by the processor (e.g., via a bus). For example, the memory may include any combination of any number of a random-access memory (RAM), a read-only memory (ROM), an optical disc, a magnetic disk, a hard drive, a solid-state drive, a flash drive, a security digital (SD) card, a memory stick, a compact flash (CF) card, or any type of storage device. The codes may include an operating system (OS) and one or more application programs (or “apps”) for specific tasks. The memory may also be a virtualmemory that includes one or more memories distributed across multiple machines or devices coupled via a network.

[0032] Fig. 3 illustrates an example imaging system 300 consistent with some embodiments of the present disclosure. Beam tool 204 of Fig. 3 may be configured for use in CPBI system 200. Beam tool 204 may be a single beam apparatus or a multi-beam apparatus. As shown in Fig. 3, beam tool 204 includes a motorized sample stage 301, and a wafer holder 302 supported by motorized sample stage 301 to hold a wafer 303 to be inspected. Beam tool 204 further includes an objective lens assembly 304, a charged-particle detector 306 (which includes charged-particle sensor surfaces 306a and 306b), an objective aperture 308, a condenser lens 310, a beam limit aperture 312, a gun aperture 314, an anode 316, and a cathode 318. Objective lens assembly 304, in some embodiments, may include a modified swing objective retarding immersion lens (SORIL), which includes a pole piece 304a, a control electrode 304b, a deflector 304c, and an exciting coil 304d. Beam tool 204 may additionally include an Energy Dispersive X-ray Spectrometer (EDS) detector (not shown) to characterize the materials on wafer 303.

[0033] A primary charged-particle beam 320 (or simply “primary beam 320”), such as an electon beam, is emitted from cathode 318 by applying an acceleration voltage between anode 316 and cathode 318. Primary beam 320 passes through gun aperture 314 and beam limit aperture 312, both of which may determine the size of charged-particle beam entering condenser lens 310, which resides below beam limit aperture 312. Condenser lens 310 focuses primary beam 320 before the beam enters objective aperture 308 to set the size of the charged-particle beam before entering objective lens assembly 304. Deflector 304c deflects primary beam 320 to facilitate beam scanning on the wafer. For example, in a scanning process, deflector 304c may be controlled to deflect primary beam 320 sequentially onto different locations of top surface of wafer 303 at different time points, to provide data for image reconstruction for different parts of wafer 303. Moreover, deflector 304c may also be controlled to deflect primary beam 320 onto different sides of wafer 303 at a particular location, at different time points, to provide data for stereo image reconstruction of the wafer structure at that location. Further, in some embodiments, anode 316 and cathode 318 may generate multiple primary beams 320, and beam tool 204 may include a plurality of deflectors 304c to project the multiple primary beams 320 to different parts / sides of the wafer at the same time, to provide data for image reconstruction for different parts of wafer 303.

[0034] Exciting coil 304d and pole piece 304a generate a magnetic field that begins at one end of pole piece 304a and terminates at the other end of pole piece 304a. A part of wafer 303 being scanned by primary beam 320 may be immersed in the magnetic field and may be electrically charged, which, in turn, creates an electric field. The electric field reduces the energy of impinging primary beam 320 near the surface of wafer 303 before it collides with wafer 303. Control electrode 304b, being electrically isolated from pole piece 304a, controls an electric field on wafer 303 to prevent micro-arching of wafer 303 and to ensure proper beam focus.

[0035] A secondary charged-particle beam 322 (or “secondary beam 322”), such as secondary electron beams, may be emitted from the part of wafer 303 upon receiving primary beam 320. Secondary beam 322 may form a beam spot on sensor surfaces 306a and 306b of charged-particle detector 306. Charged- particle detector 306 may generate a signal (e.g., a voltage, a current, or the like) that represents an intensity of the beam spot and provide the signal to an image processing system 350. The intensity of secondary beam 322, and the resultant beam spot, may vary according to the external or internal structure of wafer 303. Moreover, as discussed above, primary beam 320 may be projected onto different locations of the top surface of the wafer or different sides of the wafer at a particular location, to generate secondary beams 322 (and the resultant beam spot) of different intensities. Therefore, by mapping the intensities of the beam spots with the locations of wafer 303, the processing system may reconstruct an image that reflects the internal or surface structures of wafer 303.

[0036] Imaging system 300 may be used for inspecting a wafer 303 on motorized sample stage 301 and includes beam tool 204, as discussed above. Imaging system 300 may also include an image processing system 350 that includes an image acquirer 360, storage 370, and controller 209. Image acquirer 360 may include one or more processors. For example, image acquirer 360 may include a computer, server, mainframe host, terminals, personal computer, any kind of mobile computing devices, and the like, or a combination thereof. Image acquirer 360 may connect with a detector 306 of beam tool 204 through a medium such as an electrical conductor, optical fiber cable, portable storage media, IR, Bluetooth, internet, wireless network, wireless radio, or a combination thereof. Image acquirer 360 may receive a signal from detector 306 and may construct an image. Image acquirer 360 may thus acquire images of wafer 303. Image acquirer 360 may also perform various post-processing functions, such as generating contours, superimposing indicators on an acquired image, and the like. Image acquirer 360 may perform adjustments of brightness and contrast, or the like of acquired images. Storage 370 may be a storage medium such as a hard disk, cloud storage, random access memory (RAM), other types of computer readable memory, and the like. Storage 370 may be coupled with image acquirer 360 and may be used for saving scanned raw image data as original images, post-processed images, or other images assisting of the processing. Image acquirer 360 and storage 370 may be connected to controller 209. In some embodiments, image acquirer 360, storage 370, and controller 209 may be integrated together as one control unit.

[0037] In some embodiments, image acquirer 360 may acquire one or more images of a sample based on an imaging signal received from detector 306. An imaging signal may correspond to a scanning operation for conducting charged particle imaging. An acquired image may be a single image including a plurality of imaging areas. The single image may be stored in storage 370. The single image may be an original image that may be divided into a plurality of regions. Each of the regions may include one imaging area containing a feature of wafer 303.

[0038] Consistent with some embodiments of this disclosure, a computer-implemented method of training a machine learning model for defect detection may include obtaining training data that includesan inspection image of a fabricated integrated circuit (IC) and design layout data of the IC. The obtaining operation, as used herein, may refer to accepting, taking in, admitting, gaining, acquiring, retrieving, receiving, reading, accessing, collecting, or any operation for inputting data. An inspection image, as used herein, may refer to an image generated as a result of an inspection process performed by a charged-particle inspection apparatus (e.g., system 200 of Fig. 2 or system 300 of Fig. 3). For example, an inspection image may be an SCPM image generated by image processing system 350 in Fig. 3. A fabricated IC in this disclosure may refer to an IC manufactured on a sample (e.g., a wafer) in a semiconductor manufacturing process (e.g., a photolithography process). For example, the fabricated IC may be manufactured in a die of the sample. Design layout data of an IC, as used herein, may refer to data representing a designed layout of the IC. In some embodiments, the design layout data may include a design layout file in a GDS format (e.g., a GDS layout file). The design layout file may be visualized (also referred to as “rendered”) to be a 2D image (referred to as a “rendered image” herein) that presents the layout of the IC. The rendered image may include various geometric features (e.g., vertices, edges, corners, polygons, holes, bridges, vias, or the like) of the IC.

[0039] In some embodiments, the design layout data of the IC may include an image (e.g., the rendered image) rendered based on GDS clip data of the IC. GDS clip data of an IC, as used herein, may refer to design layout data of the IC that is to be fabricated in a die, which is of the GDS format. In some embodiments, the design layout data of the IC may include only a design layout file (e.g., the GDS clip data) of the IC. In some embodiments, the design layout data of the IC may include only the rendered image of the IC. In some embodiments, the design layout data of the IC may include only a golden image of the IC. In some embodiments, the design layout data may include any combination of the design layout file, the golden image, and the rendered image of the IC.

[0040] Consistent with some embodiments of this disclosure, the computer-implemented method of obtaining an edge placement error measurement may also include training a machine learning model using the obtained training data. In some embodiments, the machine learning model may be trained by a computer hardware system. In some embodiments, as described elsewhere in this disclosure, the training data may include previously obtained images of the lower layer or may include simulated images of the lower layer.

[0041] In some embodiments, machine learning may be employed in the generation of inspection images, reference images or other images associated with apparatus 200 or apparatus 300. For example, in some embodiments, a machine learning system may be operated in association with, e.g., controller 209, image processing system 350, image acquirer 360, or storage 370 of FIGs. 2-3. In some embodiments, machine learning may be employed in the edge placement error measurement method, e.g., method 700 of FIG. 7. In some embodiments, a machine learning system may include a discriminative model. In some embodiments, a machine learning system may include a generative model. For example, learning can feature two types of mechanisms: discriminative learning that may be used to create classification and detection algorithms, and generative learning that may be used toactually create models that, in the extreme, can render images. For example, as described further below, a generative model may be configured for generating an image from a design clip that resembles a corresponding location on a wafer in an SCPM image. This may be performed by 1) training the generative model with design clips and the associated actual SCPM images from those locations on the wafer; and 2) using the model in inference mode to feed the model design clips in locations for which simulated SCPM images are desired. Such simulated images can be used as reference images in, e.g., die-to-database inspection.

[0042] If the model(s) include one or more discriminative models, the discriminative model(s) may have any suitable architecture and / or configuration known in the art. Discriminative models, also called conditional models, are a class of models used in machine learning for modeling the dependence of an unobserved variable “y” on an observed variable “x.” Within a probabilistic framework, this may be done by modeling a conditional probability distribution P(ylx), which can be used for predicting y based on x. Discriminative models, as opposed to generative models, may not allow one to generate samples from the joint distribution of x and y. However, for tasks such as classification and regression that do not require the joint distribution, discriminative models may yield superior performance. On the other hand, generative models are typically more flexible than discriminative models in expressing dependencies in complex learning tasks. In addition, most discriminative models are inherently supervised and cannot easily be extended to unsupervised learning. Application specific details ultimately dictate the suitability of selecting a discriminative versus generative model.

[0043] A generative model can be generally defined as a model that is probabilistic in nature. In other words, a “generative” model is not one that performs forward simulation or rule-based approaches and, as such, it may not be necessary to model the physics of the processes involved in generating an actual image or output (for which a simulated image or output is being generated). Instead, the generative model can be learned (in that its parameters can be learned) based on a suitable training set of data. Such generative models may have a number of advantages for the embodiments described herein. In addition, the generative model may be configured to have a deep learning architecture in that the generative model may include multiple layers, which may perform a number of algorithms or transformations. The number of layers included in the generative model may depend on the particular use case. For practical purposes, a suitable range of layers is from two layers to a few tens of layers.

[0044] Deep learning is a type of machine learning. Machine learning can be generally defined as a type of artificial intelligence (Al) that provides computers with the ability to learn without being explicitly programmed. Machine learning focuses on the development of computer programs that can teach themselves to grow and change when exposed to new data. Machine learning explores the study and construction of algorithms that can learn from and make predictions on data — such algorithms overcome following strictly static program instructions by making data driven predictions or decisions, through building a model from sample inputs.

[0045] The machine learning described herein may be further performed as described in “Introduction to Statistical Machine Learning,” by Sugiyama, Morgan Kaufmann, 2016, 534 pages; “Discriminative, Generative, and Imitative Learning,” by Jebara, MIT Thesis, 2002, 212 pages; and “Principles of Data Mining (Adaptive Computation and Machine Learning)” by Hand et al., MIT Press, 2001, 578 pages; which are incorporated by reference as if fully set forth herein. The embodiments described herein may be further configured as described in these references.

[0046] In some embodiments, a machine learning system may comprise a neural network. For example, a model may be a deep neural network with a set of weights that model the world according to the data that it has been fed to train it. Neural networks can be generally defined as a computational approach which is based on a relatively large collection of neural units loosely modeling the way a biological brain solves problems with relatively large clusters of biological neurons connected by axons. Each neural unit is connected with many others, and links can be enforcing or inhibitory in their effect on the activation state of connected neural units. These systems are self-learning and trained rather than explicitly programmed and excel in areas where the solution or feature detection is difficult to express in a traditional computer program.

[0047] Neural networks typically consist of multiple layers, and the signal path traverses from front to back. The goal of the neural network is to solve problems in the same way that the human brain would, although several neural networks are much more abstract. Modern neural network projects typically work with a few thousand to a few million neural units and millions of connections. The neural network may have any suitable architecture and / or configuration known in the art.

[0048] In a further embodiment, a model may comprise a convolutional and deconvolution neural network. For example, the embodiments described herein can take advantage of learning concepts such as a convolution and deconvolution neural network to solve the normally intractable representation conversion problem (e.g., rendering). The model may have any convolution and deconvolution neural network configuration or architecture known in the art.

[0049] A neural network, as used herein, may refer to a computing model for analyzing underlying relationships in a set of input data by way of mimicking human brains. Similar to a biological neural network, the neural network may include a set of connected units or nodes (referred to as “neurons”), structured as different layers, where each connection (also referred to as an “edge”) may obtain and send a signal between neurons of neighboring layers in a way similar to a synapse in a biological brain. The signal may be any type of data (e.g., a real number). Each neuron may obtain one or more signals as an input and output another signal by applying a non-linear function to the inputted signals. Neurons and edges may typically be weighted by corresponding weights to represent the knowledge the neural network has acquired. During a training process (similar to a learning process of a biological brain), the weights may be adjusted (e.g., by increasing or decreasing their values) to change the strengths of the signals between the neurons to improve the performance accuracy of the neural network. Neurons may apply a thresholding function (referred to as an “activation function”) to its output values of the non-linear function such that a signal is outputted only when an aggregated value (e.g., a weighted sum) of the output values of the non-linear function exceeds a threshold determined by the thresholding function. Different layers of neurons may transform their input signals in different manners (e.g., by applying different non-linear functions or activation functions). The output of the last layer (referred to as an “output layer”) may output the analysis result of the neural network, such as, for example, a categorization of the set of input data (e.g., as in image recognition cases), a numerical result, or any type of output data for obtaining an analytical result from the input data.

[0050] Training of the neural network, as used herein, may refer to a process of improving the accuracy of the output of the neural network. Typically, the training may be categorized into three types: supervised training, unsupervised training, and reinforcement training. In the supervised training, a set of target output data (also referred to as “labels” or “ground truth”) may be generated based on a set of input data using a method other than the neural network. The neural network may then be fed with the set of input data to generate a set of output data that is typically different from the target output data. Based on the difference between the output data and the target output data, the weights of the neural network may be adjusted in accordance with a rule. If such adjustments are successful, the neural network may generate another set of output data more similar to the target output data in a next iteration using the same input data. If such adjustments are not successful, the weights of the neural network may be adjusted again. After a sufficient number of iterations, the training process may be terminated in accordance with one or more predetermined criteria (e.g., the difference between the final output data and the target output data is below a predetermined threshold, or the number of iterations reaches a predetermined threshold). The trained neural network may be applied to analyze other input data.

[0051] In the unsupervised training, the neural network is trained without any external gauge (e.g., labels) to identify patterns in the input data rather than generating labels for them. Typically, the neural network may analyze shared attributes (e.g., similarities and differences) and relationships among the elements of the input data in accordance with one or more predetermined rules or algorithms (e.g., principal component analysis, clustering, anomaly detection, or latent variable identification). The trained neural network may extrapolate the identified relationships to other input data.

[0052] In the reinforcement learning, the neural network is trained without any external gauge (e.g., labels) in a trial-and-error manner to maximize benefits in decision making. The input data sets of the neural network may be different in the reinforcement training. For example, a reward value or a penalty value may be determined for the output of the neural network in accordance with one or more rules during training, and the weights of the neural network may be adjusted to maximize the reward values (or to minimize the penalty values). The trained neural network may apply its learned decision-making knowledge to other input data.

[0053] During the training of a neural network, a loss function (or referred to as a “cost function”) may be used to evaluate the output data. The loss function, as used herein, may map output data of a machine learning model (e.g., the neural network) onto a real number (referred to as a “loss” or a “cost”) thatintuitively represents a loss or an error (e.g., representing a difference between the output data and target output data) associated with the output data. The training of the neural network may seek to maximize or minimize the loss function (e.g., by pushing the loss towards a local maximum or a local minimum in a loss curve). For example, one or more parameters of the neural network may be adjusted or updated purporting to maximize or minimize the loss function. After adjusting or updating the one or more parameters, the neural network may obtain new input data in a next iteration of its training. When the loss function is maximized or minimized, the training of the neural network may be terminated.

[0054] By way of example, Fig. 4 is a schematic diagram illustrating an example neural network 400, consistent with some embodiments of the present disclosure. As depicted in Fig. 4, neural network 400 may include an input layer 420 that receives inputs, including input 410-1, . . ., input 410-m (m being an integer). For example, an input of neural network 400 may include any structure or unstructured data (e.g., an image). In some embodiments, neural network 400 may obtain a plurality of inputs simultaneously. For example, in Fig. 4, neural network 400 may obtain m inputs simultaneously. In some embodiments, input layer 420 may obtain m inputs in succession such that input layer 420 receives input 410-1 in a first cycle (e.g., in a first inference) and pushes data from input 410-1 to a hidden layer (e.g., hidden layer 430-1), then receives a second input in a second cycle (e.g., in a second inference) and pushes data from input the second input to the hidden layer, and so on. Input layer 420 may obtain any number of inputs in the simultaneous manner, the successive manner, or any manner of grouping the inputs.

[0055] Input layer 420 may include one or more nodes, including node 420-1, node 420-2, . . ., node 420-a (a being an integer). A node (also referred to as a “machine perceptron” or a “neuron”) may model the functioning of a biological neuron. Each node may apply an activation function to received inputs (e.g., one or more of input 410-1, . . ., input 410-m). An activation function may include a Heaviside step function, a Gaussian function, a multiquadratic function, an inverse multiquadratic function, a sigmoidal function, a rectified linear unit (ReLU) function (e.g., a ReLU6 function or a Leaky ReLU function), a hyperbolic tangent (“tanh”) function, or any non-linear function. The output of the activation function may be weighted by a weight associated with the node. A weight may include a positive value between 0 and 1 , or any numerical value that may scale outputs of some nodes in a layer more or less than outputs of other nodes in the same layer.

[0056] As further depicted in Fig. 4, neural network 400 includes multiple hidden layers, including hidden layer 430-1, . . ., hidden layer 430-n (n being an integer). When neural network 400 includes more than one hidden layer, it may be referred to as a “deep neural network” (DNN). Each hidden layer may include one or more nodes. For example, in Fig. 4, hidden layer 430-1 includes node 430-1-1, node 430-1-2, node 430-1-3, . . ., node 430- 1-b (b being an integer), and hidden layer 430-n includes node 430-n-l, node 430-n-2, node 430-n-3, . . ., node 430-n-c (c being an integer). Similar to nodes of input layer 420, nodes of the hidden layers may apply the same or different activation functions to outputsfrom connected nodes of a previous layer, and weight the outputs from the activation functions by weights associated with the nodes.

[0057] As further depicted in Fig. 4, neural network 400 may include an output layer 440 that finalizes outputs, including output 450-1, output 450-2, . . ., output 450-d (d being an integer). Output layer 440 may include one or more nodes, including node 440-1, node 440-2, . . ., node 440-d. Similar to nodes of input layer 420 and of the hidden layers, nodes of output layer 440 may apply activation functions to outputs from connected nodes of a previous layer and weight the outputs from the activation functions by weights associated with the nodes.

[0058] Although nodes of each hidden layer of neural network 400 are depicted in Fig. 4 to be connected to each node of its previous layer and next layer (referred to as “fully connected”), the layers of neural network 400 may use any connection scheme. For example, one or more layers (e.g., input layer 420, hidden layer 430-1, . . ., hidden layer 430-n, or output layer 440) of neural network 400 may be connected using a convolutional scheme, a sparsely connected scheme, or any connection scheme that uses fewer connections between one layer and a previous layer than the fully connected scheme as depicted in Fig. 4.

[0059] Moreover, although the inputs and outputs of the layers of neural network 400 are depicted as propagating in a forward direction (e.g., being fed from input layer 420 to output layer 440, referred to as a “feedforward network”) in Fig. 4, neural network 400 may additionally or alternatively use backpropagation (e.g., feeding data from output layer 440 towards input layer 420) for other purposes. For example, the backpropagation may be implemented by using long short-term memory nodes (LSTM). Accordingly, although neural network 400 is depicted similar to a convolutional neural network (CNN), neural network 400 may include a recurrent neural network (RNN) or any other neural network.

[0060] Some embodiments of the present disclosure use HV-SCPM and machine learning (ML) to accurately construct the EPE. The HV-SCPM produces an image with sharp upper layer information but blurred lower layer information. For example, the lower layer may only be partially visible. An ML algorithm may be trained to improve the sharpness of the lower layer image. In this way, more accurate CD measurements and center of gravity (CoG; or alignment of the lower layer image to the upper layer image) may be retrieved to construct a more accurate EPE.

[0061] Clear and sharp images from the lower layer are needed for the ML training. Such images can come from simulation and from SCPM measurements taken when the lower layer is manufactured (i.e., lower layer measurements taken before the upper layer is manufactured). 3D profiling information may also be added during training of the ML model to provide enhanced 3D EPE results. As the blurring effect of the lower layer image varies with material, thickness, pattern density, etc., it is possible that a new layer may need new training for the ML model.

[0062] Fig. 5 is a diagram of an example workflow 500 for training the machine learning model, consistent with embodiments of the present disclosure. In some embodiments, the workflow 500 may be performed by image processing system 350 of Fig. 3.

[0063] A HV-SCPM image of both the lower layer and the upper layer, Un, is acquired (step 502). As used herein, the subscript “1” refers to a lower layer and the subscript “2” refers to an upper layer. The HV-SCPM image is provided to a neural network 504 to produce an enhanced HV-SCPM image of the lower layer and the upper layer, u™h506. In some embodiments, the neural network 504 is trained with previously collected images of the lower layer and may also be trained with simulated images of the lower layer. It is noted that while some embodiments described herein relate to two layers (e.g., a lower layer and an upper layer), the concepts described herein may be similarly applied to samples having more than two layers.

[0064] A SCPM image of the lower layer, u> is acquired (step 510). The SCPM image of the lower layer may be acquired after the lower layer is manufactured and before the upper layer is manufactured. Distortions in the lower layer image are corrected using a mathematical distortion field 512 (step 514) to produce a corrected SCPM image of the lower layer, uorr516. The mathematical distortion field 512 is an image used to correct distortions and is not another SCPM image. A mask of the upper layer is extracted from Un and is applied to uorr516 to produce masked SCPM lower layer, u™ask518. In some embodiments, the u™ask518 may be based on a wafer design layout in Graphic Database System (GDS) format, Graphic Database System II (GDS II) format including a graphical representation of the features on the wafer surface, or an Open Artwork System Interchange Standard (OASIS) format. For example, the u™ask518 represents the desired output of the semiconductor manufacturing process for the lower layer. As shown in Fig. 5, in u™ask518, the circle is the drawing data (e.g., the GDS data) and what is shown inside the circle is the actual data obtained from the corrected SCPM lower layer,uc°rr51 insome embodiments, uorr516 may be aligned with the lower layer of the combined image to compensate for tool alignment issues.

[0065] The machine learning model is trained using the pairsh506 and u^ask518 to minimize the differences between the pairs of images, between what is expected to be in the image (e.g., based on the GDS data) and the actual image (step 520). It is noted that steps 502-506 may be performed before steps 510-518, after steps 510-518, or simultaneously with steps 510-518.

[0066] Fig. 6 is a diagram of an example method 600 for using HV-SCPM and machine learning to construct EPE, consistent with embodiments of the present disclosure. In general, the method 600 uses the trained ML model (as trained as described in connection with Fig. 5) to correct the blurred image of the lower layer.

[0067] An HV-SCPM image of the combined lower layer and upper layer is obtained (step 602). The HV-SCPM image 602 is provided to a trained machine learning model (trained as described in connection with Fig. 5) along with an optional information file 606. For example, the information file606 may include a wafer design layout in Graphic Database System (GDS) format, Graphic Database System II (GDS II) format including a graphical representation of the features on the wafer surface, or an Open Artwork System Interchange Standard (OASIS) format. The wafer design layout may be based on a pattern layout for constructing the wafer. The wafer design layout may correspond to one or more photolithography masks or reticles used to transfer features from the photolithography masks or reticles to wafer 303, for example. GDS information file or OASIS information file may include feature information stored in a binary file format representing planar geometric shapes, text, and other information related to wafer design layout. OASIS format may help reduce data volume, resulting in a more efficient data transfer process. A large amount of GDS or OASIS format images may have been collected and may make up a large dataset of comparison features. In some embodiments, the information file 606 may include the same or similar information as used in connection with creating the mask u™ask518 as described in connection with Fig. 5.

[0068] An SCPM image of each layer is taken along with an overlay measurement (a combined lower layer and upper layer image). The trained ML model 604 is used to link a sharp lower layer image (taken via SCPM) with the blurry combined image (taken via HV-SCPM). Distortion correction is performed on the lower layer image, wherein the distortion has a low placement error for each pattern. Any suitable distortion correction algorithm may be applied, including algorithms that may be physicsbased or mathematical-based. Examples of physics-based correction algorithms may include a charging model or a hardware compensation model. Examples of mathematical-based correction algorithms may include 2-par, 6-par, 20-par modeling, and variants of 20-par modeling. Image processing is performed on the HV-SCPM image (i.e., the combined lower layer and upper layer image) to remove the upper layer , to just keep the blurred patterns from the lower layer. The result of the image processing is a clear, sharp image of the lower layer. The trained ML model 604 is applied to improve the sharpness of the lower layer (image 608). By using the sharp image 608 of the lower layer, it is possible to obtain an accurate EPE measurement including CD and placement error (PE), and then obtain the EPE from the more accurate CD and PE measurements (step 610).

[0069] Fig. 7 is a flowchart of an example method 700 for EPE measurement, consistent with embodiments of the present disclosure. In some embodiments, the method 700 may be performed by image processing system 250 of Fig. 2.

[0070] At step 702, an image of a lower layer of a sample is obtained. A sample is a portion of a semiconductor (e.g., wafer 303) to be inspected. The image of the lower layer may be obtained after the lower layer is manufactured and before the upper layer is manufactured. In some embodiments, the image of the lower layer may be an SCPM image of the lower layer (e.g., the image u> of Fig. 5).

[0071] At step 704, a combined image of an upper layer and the lower layer of the sample is obtained. In some embodiments, the combined image may be an HV-SCPM image of the upper layer and the lower layer (e.g., the image Un of Fig. 5).

[0072] At step 706, distortions in the lower layer image of the sample are corrected to produce a corrected image of the lower layer (e.g., image uorrof Fig. 5). Any suitable distortion correction algorithm may be applied to correct distortions in the lower layer image. In some embodiments, the distortions may be corrected in a manner described in connection with step 514 of Fig. 5. Example distortion correction algorithms may include physics-based correction algorithms, such as a charging model or a hardware compensation model, or mathematical-based correction algorithms, such as 2-par, 6-par, 20-par modeling, and variants of 20-par modeling.

[0073] At step 708, a mask based on the combined image is generated and is used to remove the upper layer from the combined image to retain only the blurred patterns from the lower layer.

[0074] At step 710, the mask is applied to the lower layer image to obtain a clear, sharp image of the lower layer (e.g., image u™askof Fig. 5).

[0075] At step 712, a ML model is trained using the lower layer image and the combined image, to minimize the differences between the pairs of images, between what is expected to be in the image (e.g., based on GDS data) and the actual image. In some embodiments, step 712 may be performed in a similar manner as step 520.

[0076] At step 714, the trained ML model is applied to the combined image to improve the sharpness of the lower layer (e.g., image 608 of Fig. 6). By using the sharper image of the lower layer, an accurate CD measurement of the lower layer can be obtained. At step 716, the EPE for the sample can be constructed from the accurate CD measurement of the lower layer. In some embodiments, step 716 may be performed in a similar manner as step 610.

[0077] A non-transitory computer readable medium may be provided that stores instructions for a processor of a controller (e.g., controller 209 of FIG. 2) to carry out, among other things, image inspection, image acquisition, stage positioning, beam focusing, electric field adjustment, beam bending, condenser lens adjusting, activating charged particle source, beam deflecting, and operations 500 and 600 and method 700. Common forms of non-transitory media include, for example, a floppy disk, a flexible disk, hard disk, solid state drive, magnetic tape, or any other magnetic data storage medium, a Compact Disc Read Only Memory (CD-ROM), any other optical data storage medium, any physical medium with patterns of holes, a Random Access Memory (RAM), a Programmable Read Only Memory (PROM), and Erasable Programmable Read Only Memory (EPROM), a FLASH-EPROM or any other flash memory, Non-Volatile Random Access Memory (NVRAM), a cache, a register, any other memory chip or cartridge, and networked versions of the same.

[0078] The embodiments may further be described using the following clauses:LA non-transitory computer readable medium that stores a set of instructions that is executable by at least one processor of a computing device to cause the computing device to perform operations for obtaining an edge placement error measurement, the operations comprising: obtaining an image of a lower layer of a sample before an upper layer is deposited during a manufacturing process;obtaining a combined image of the upper layer and the lower layer of the sample after the lower layer image is obtained; correcting distortions in the lower layer image; generating a mask based on the combined image; applying the mask to the lower layer image; and training a machine learning model using the masked lower layer image and the combined image, wherein the trained machine learning model is configured to determine the edge placement error measurement for the combined image.2. The non-transitory computer readable medium of clause 1, wherein: the lower layer image is a scanning charged particle microscope (SCPM) image; and the combined image is a high voltage SCPM (HV-SCPM) image of the upper layer and the lower layer.3. The non-transitory computer readable medium of clauses 1 or 2, wherein correcting distortions includes applying a distortion correction algorithm to the lower layer image, the distortion correction algorithm including any one of a charging model, a hardware compensation model, 2-par, 6-par, 20-par modeling, or variants of 20-par modeling.4. The non-transitory computer readable medium of any one of clauses 1-3, wherein generating the mask includes using design layout information to generate the mask.5. The non-transitory computer readable medium of any one of clauses 1-4, wherein generating the mask includes removing the upper layer of the combined image.6. The non-transitory computer readable medium of any one of clauses 1-5, wherein training the machine learning model includes training the machine learning model to minimize a difference between the masked lower layer image and the combined image.7. The non-transitory computer readable medium of any one of clauses 1-6, wherein the trained machine learning model is configured to apply the combined image to improve sharpness of the lower layer.8. The non-transitory computer readable medium of any one of clauses 1-7, wherein the trained machine learning model is configured to obtain a critical dimension measurement of the sharpened lower layer.9. The non-transitory computer readable medium of any one of clauses 1-8, wherein the trained machine learning model is configured to determine the edge placement error measurement based on the critical dimension measurement.10. The non-transitory computer readable medium of any one of clauses 1-9, wherein the trained machine learning model is further configured to determine overlay.11. The non-transitory computer readable medium of any one of clauses 1-10, wherein: the machine learning model is trained by a computer hardware system; and the trained machine learning model is configured to determine the edge placement error measurement by a computer hardware system processing the trained machine learning model.12. An apparatus for obtaining an edge placement error measurement, comprising:a memory storing a set of instructions; and at least one processor configured to execute the set of instructions to cause the apparatus to perform operations comprising: obtain an image of a lower layer of a sample before an upper layer is deposited during a manufacturing process; obtain a combined image of the upper layer and the lower layer of the sample after the lower layer image is obtained; correct distortions in the lower layer image; generate a mask based on the combined image; apply the mask to the lower layer image; and train a machine learning model using the masked lower layer image and the combined image, wherein the trained machine learning model is configured to determine the edge placement error measurement for the combined image.13. The apparatus of clause 12, wherein: the lower layer image is a scanning charged particle microscope (SCPM) image; and the combined image is a high voltage SCPM (HV-SCPM) image of the upper layer and the lower layer.14. The apparatus of clauses 12 or 13, wherein the at least one processor is further configured to correct distortions by applying a distortion correction algorithm to the lower layer image, the distortion correction algorithm including any one of a charging model, a hardware compensation model, 2-par, 6- par, 20-par modeling, or variants of 20-par modeling.15. The apparatus of any one of clauses 12-14, wherein the at least one processor is further configured to generate the mask by using design layout information to generate the mask.16. The apparatus of any one of clauses 12-15, wherein the at least one processor is further configured to generate the mask by removing the upper layer of the combined image.17. The apparatus of any one of clauses 12-16, wherein the at least one processor is further configured to train the machine learning model by minimizing a difference between the masked lower layer image and the combined image.18. The apparatus of any one of clauses 12-17, wherein the trained machine learning model is configured to apply the combined image to improve sharpness of the lower layer.19. The apparatus of any one of clauses 12-18, wherein the trained machine learning model is configured to obtain a critical dimension measurement of the sharpened lower layer.20. The apparatus of any one of clauses 12-19, wherein the trained machine learning model is configured to determine the edge placement error measurement based on the critical dimension measurement.21. The apparatus of any one of clauses 12-20, wherein the trained machine learning model is further configured to determine overlay.22. The apparatus of any one of clauses 12-21, wherein: the machine learning model is trained by a computer hardware system; and the trained machine learning model is configured to determine the edge placement error measurement by a computer hardware system processing the trained machine learning model.23. A method for obtaining an edge placement error measurement, comprising: obtaining an image of a lower layer of a sample before an upper layer is deposited during a manufacturing process; obtaining a combined image of the upper layer and the lower layer of the sample after the lower layer image is obtained; correcting distortions in the lower layer image; generating a mask based on the combined image; applying the mask to the lower layer image; and training a machine learning model using the masked lower layer image and the combined image, wherein the trained machine learning model is configured to determine the edge placement error measurement for the combined image.24. The method of clause 23, wherein: the lower layer image is a scanning charged particle microscope (SCPM) image; and the combined image is a high voltage SCPM (HV-SCPM) image of the upper layer and the lower layer.25. The method of clauses 23 or 24, wherein correcting distortions includes applying a distortion correction algorithm to the lower layer image, the distortion correction algorithm including any one of a charging model, a hardware compensation model, 2-par, 6-par, 20-par modeling, or variants of 20- par modeling.26. The method of any one of clauses 23-25, wherein generating the mask includes using design layout information to generate the mask.27. The method of any one of clauses 23-26, wherein generating the mask includes removing the upper layer of the combined image.28. The method of any one of clauses 23-27, wherein training the machine learning model includes training the machine learning model to minimize a difference between the masked lower layer image and the combined image.29. The method of any one of clauses 23-28, wherein the trained machine learning model is configured to apply the combined image to improve sharpness of the lower layer.30. The method of any one of clauses 23-29, wherein the trained machine learning model is configured to obtain a critical dimension measurement of the sharpened lower layer.31. The method of any one of clauses 23-30, wherein the trained machine learning model is configured to determine the edge placement error measurement based on the critical dimension measurement.32. The method of any one of clauses 23-31, wherein the trained machine learning model is further configured to determine overlay.33. The method of any one of clauses 23-32, wherein: the machine learning model is trained by a computer hardware system; and the trained machine learning model is configured to determine the edge placement error measurement by a computer hardware system processing the trained machine learning model.

[0079] Block diagrams in the figures may illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer hardware or software products according to various exemplary embodiments of the present disclosure. In some embodiments, a non-transitory computer-readable medium is provided and can include instructions to perform the functions described in connection with any one or more of Figs. 5-7. In this regard, each block in a schematic diagram may represent certain arithmetical or logical operation processing that may be implemented using hardware such as an electronic circuit. Blocks may also represent a module, segment, or portion of code that comprises one or more executable instructions for implementing the specified logical functions. It should be understood that in some alternative implementations, functions indicated in a block may occur out of the order noted in the figures. For example, two blocks shown in succession may be executed or implemented substantially concurrently, or two blocks may sometimes be executed in reverse order, depending upon the functionality involved. Some blocks may also be omitted. It should also be understood that each block of the block diagrams, and combination of the blocks, may be implemented by special purpose hardware-based systems that perform the specified functions or acts, or by combinations of special purpose hardware and computer instructions.

[0080] It will be appreciated that the embodiments of the present disclosure are not limited to the exact construction that has been described above and illustrated in the accompanying drawings, and that various modifications and changes may be made without departing from the scope thereof. The present disclosure has been described in connection with various embodiments, and other embodiments will be apparent to those skilled in the art from consideration of the specification and practice of the technology disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the invention being indicated by the following claims.

Claims

CLAIMS1. A non-transitory computer readable medium that stores a set of instructions that is executable by at least one processor of a computing device to cause the computing device to perform operations for obtaining an edge placement error measurement, the operations comprising: obtaining an image of a lower layer of a sample before an upper layer is deposited during a manufacturing process; obtaining a combined image of the upper layer and the lower layer of the sample after the lower layer image is obtained; correcting distortions in the lower layer image; generating a mask based on the combined image; applying the mask to the lower layer image; and training a machine learning model using the masked lower layer image and the combined image, wherein the trained machine learning model is configured to determine the edge placement error measurement for the combined image.

2. The non-transitory computer readable medium of claim 1, wherein: the lower layer image is a scanning charged particle microscope (SCPM) image; and the combined image is a high voltage SCPM (HV-SCPM) image of the upper layer and the lower layer.

3. The non-transitory computer readable medium of claim 1, wherein correcting distortions includes applying a distortion correction algorithm to the lower layer image, the distortion correction algorithm including any one of a charging model, a hardware compensation model, 2-par, 6-par, 20-par modeling, or variants of 20-par modeling.

4. The non-transitory computer readable medium of claim 1 , wherein generating the mask includes using design layout information to generate the mask.

5. The non-transitory computer readable medium of claim 1, wherein generating the mask includes removing the upper layer of the combined image.

6. The non-transitory computer readable medium of claim 1 , wherein training the machine learning model includes training the machine learning model to minimize a difference between the masked lower layer image and the combined image.

7. The non-transitory computer readable medium of claim 1, wherein the trained machine learning model is configured to apply the combined image to improve sharpness of the lower layer.

8. The non-transitory computer readable medium of claim 1, wherein the trained machine learning model is configured to obtain a critical dimension measurement of the sharpened lower layer.

9. The non-transitory computer readable medium of claim 1, wherein the trained machine learning model is configured to determine the edge placement error measurement based on the critical dimension measurement.

10. The non-transitory computer readable medium of claim 1, wherein the trained machine learning model is further configured to determine overlay.

11. The non-transitory computer readable medium of claim 1, wherein: the machine learning model is trained by a computer hardware system; and the trained machine learning model is configured to determine the edge placement error measurement by a computer hardware system processing the trained machine learning model.

12. An apparatus for obtaining an edge placement error measurement, comprising: a memory storing a set of instructions; and at least one processor configured to execute the set of instructions to cause the apparatus to perform operations comprising: obtain an image of a lower layer of a sample before an upper layer is deposited during a manufacturing process; obtain a combined image of the upper layer and the lower layer of the sample after the lower layer image is obtained; correct distortions in the lower layer image; generate a mask based on the combined image; apply the mask to the lower layer image; and train a machine learning model using the masked lower layer image and the combined image, wherein the trained machine learning model is configured to determine the edge placement error measurement for the combined image.

13. The apparatus of claim 12, wherein: the lower layer image is a scanning charged particle microscope (SCPM) image; andthe combined image is a high voltage SCPM (HV-SCPM) image of the upper layer and the lower layer.

14. A method for obtaining an edge placement error measurement, comprising: obtaining an image of a lower layer of a sample before an upper layer is deposited during a manufacturing process; obtaining a combined image of the upper layer and the lower layer of the sample after the lower layer image is obtained; correcting distortions in the lower layer image; generating a mask based on the combined image; applying the mask to the lower layer image; and training a machine learning model using the masked lower layer image and the combined image, wherein the trained machine learning model is configured to determine the edge placement error measurement for the combined image.

15. The method of claim 14, wherein: the lower layer image is a scanning charged particle microscope (SCPM) image; and the combined image is a high voltage SCPM (HV-SCPM) image of the upper layer and the lower layer.

Citation Information

Patent Citations

  • Multilayer optical proximity correction (OPC) model for OPC correction

    US20210072635A1

  • Apparatus and methods to generate deblurring model and deblur image

    WO2022078740A1

  • Systems, methods, and software for multilayer metrology

    WO2023213534A1