Training a model to generate predictive data

A generator model predicts post-etching sample states from pre-etching data, addressing throughput and damage issues in semiconductor inspection by correlating data distributions, enabling efficient and accurate defect detection.

JP2025527982APending Publication Date: 2025-08-26ASML NETHERLANDS BV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024569067
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-19
Filing Date
2023-07-13
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing sample evaluation methods in semiconductor manufacturing suffer from low throughput and sample damage, particularly during pattern defect inspection before and after etching processes.

Method used

A generator model is trained to process measurement data from a sample before etching, generating predicted data that simulates the sample's post-etching state, using a classifier to determine data distribution correlations and reduce the need for direct post-etching scans.

Benefits of technology

This approach enhances throughput by minimizing sample damage and improving accuracy in defect detection by predicting post-etching patterns without requiring additional scanning, thus optimizing lithographic processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025527982000001_ABST
    Figure 2025527982000001_ABST
Patent Text Reader

Abstract

1. A method for training a generator model, comprising: using a generator model to generate predicted data based on first measured data, the first measured data and the predicted data usable to form an image of the sample; pairing a subset of the first measured data with a subset of predicted data, the subset corresponding to a location within the image of the sample formable from the first measured data and the predicted data; using a classifier to determine a likelihood that the predicted data is from the same data distribution as second measured data measured from the sample after an etch process; and training the generator model based on correlations between the paired subsets of data, the correlations of pairs corresponding to the same location compared to correlations of pairs corresponding to different locations, the correlations being between the paired subsets of data, and the likelihood determined by the classifier.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to European Application No. 22186636.1 filed on July 25, 2022, and European Application No. 22196424.0 filed on September 19, 2022, both of which are incorporated by reference in their entirety into this specification. [Background technology]

[0002]

[0002] The embodiments disclosed herein relate to generating predicted data for a sample after an etching process based on measurement data of the sample after an exposure process but before the etching process, for example, to enable optimization of sample parameters before etching.

[0003]

[0003] When manufacturing semiconductor integrated circuit (IC) chips, unwanted pattern defects inevitably occur on substrates (i.e., wafers) or masks during the fabrication process, for example as a result of optical effects and by-particles, thereby reducing yield. Therefore, monitoring the degree of unwanted pattern defects is an important process in the manufacture of IC chips. More generally, inspection and / or measurement of the surface of a substrate or other object / material is an important process during and / or after its manufacture.

[0004]

[0004] Pattern inspection apparatuses using charged particle beams have been used to inspect objects (sometimes called samples), for example, to detect pattern defects. These apparatuses typically use electron microscopy techniques such as scanning electron microscopes (SEMs). In an SEM, a primary electron beam of relatively high-energy electrons is targeted with a final deceleration step to land on the sample with a relatively low landing energy. The electron beam is focused as a probing spot on the sample. Interactions between material structures at the probing spot and landing electrons from the electron beam cause signal electrons, such as secondary electrons, backscattered electrons, or Auger electrons, to be emitted from the surface. The signal electrons can be emitted from the material structures of the sample. By scanning the primary electron beam as a probing spot across the sample surface, signal electrons can be emitted across the surface of the sample. By collecting these signal electrons from the sample surface, the pattern inspection apparatus can obtain an image representative of the material structure characteristics of the surface of the sample.

[0005]

[0005] It may be desirable to scan a sample at different stages of its processing. For example, the sample may be scanned after a lithographic exposure process has been performed on it, but before a subsequent etching process. This allows a so-called after-development inspection (ADI) image (often known as a post-lithography inspection image) of the sample to be measured. The sample may then be scanned after the etching process. This allows a so-called after-etch inspection (AEI) image of the sample to be formed. Data on these images can be used, for example, to optimize sample parameters before etching and / or to improve subsequent lithographic processing of the substrate. The process of scanning the sample may damage the sample (e.g., by damaging the resist material of the sample). The scanning process may also reduce the throughput of the sample evaluation method.

[0006]

[0006] There is a general need to increase the throughput of sample evaluation methods and / or reduce the damage caused by sample evaluation methods and / or increase the accuracy. Summary of the Invention

[0007]

[0007] It is an object of the present disclosure to provide embodiments that support increased throughput in sample evaluation methods and / or reduced damage caused by sample evaluation methods.

[0008]

[0008] According to one aspect of the present invention, there is provided a method for training a generator model that processes first measurement data measured from a sample before an etching process to generate predicted data that predicts the sample after the etching process, the method including: using the generator model to generate predicted data based on the first measurement data, wherein the first measurement data and the predicted data can be used to form an image of the sample; pairing a subset of the first measurement data with a subset of predicted data, wherein the subset corresponds to a location in the image of the sample that can be formed from the measurement data and the predicted data; using a classifier to determine a likelihood that the predicted data is from the same data distribution as second measurement data measured from the sample at a different location after the etching process; and training the generator model based on a correlation between the paired subsets of data, the correlation of pairs corresponding to the same location when compared to the correlation of pairs corresponding to different locations, and the likelihood determined by the classifier.

[0009]

[0009] According to one aspect of the present invention, there is provided a generator model training apparatus for training a generator model that processes first measurement data measured from a sample before an etching process to generate predicted data that predicts the sample after the etching process, the apparatus comprising a processor configured to: use the generator model to generate predicted data based on the first measurement data, wherein the first measurement data and the predicted data can be used to form an image of the sample; pair a subset of the first measurement data with a subset of predicted data, wherein the subset corresponds to a position in the image of the sample that can be formed from the first measurement data and the predicted data; determine using a classifier a likelihood that the predicted data is from the same data distribution as second measurement data measured from the sample at a different position after the etching process; and train the generator model based on a correlation between the paired subsets of data, the correlation being the correlation of pairs corresponding to the same position when compared to the correlation of pairs corresponding to different positions, and the likelihood determined by the classifier.

[0010]

[0010] According to one aspect of the present invention, there is provided a computer-readable medium storing instructions configured to control a processor to train a generator model that processes first measurement data measured from a sample before an etching process to generate predicted data that predicts the sample after the etching process, the computer-readable medium storing instructions configured to control a processor to: use the generator model to generate predicted data based on the first measurement data, wherein the first measurement data and the predicted data can be used to form an image of the sample; pairing a subset of the first measurement data with a subset of predicted data, wherein the subset corresponds to a position in the image of the sample that can be formed from the measurement data and the predicted data; determining, using a classifier, a likelihood that the predicted data is from the same data distribution as second measurement data measured from the sample at a different position after the etching process; and training the generator model based on a correlation between the paired subsets of data, the correlation of pairs corresponding to the same position when compared to the correlation of pairs corresponding to different positions, and the likelihood determined by the classifier.

[0011]

[0011] According to one aspect of the present invention, there is provided a method for training a generator model that processes paired measurement data measured after the etching process from a sample that was previously measured before the etching process to generate hypothesis data that simulates what the sample would be like after the etching process if the sample had not been previously measured before the etching process, the method including: using the generator model to generate hypothesis data based on the paired measurement data, wherein the paired measurement data and the hypothesis data can be used to form an image of the sample; using a classifier to determine the likelihood that the hypothesis data is from the same data distribution as actual measurement data measured after the etching process from a sample that was not previously measured before the etching process; and training the generator model based on a function indicating the level of correlation between the paired measurement data and the hypothesis data and the likelihood determined by the classifier.

[0012]

[0012] According to one aspect of the present invention, there is provided a generator model training apparatus for training a generator model that processes paired measurement data measured after the etching process from a sample that was previously measured before the etching process to generate hypothesis data that simulates what the sample would be like after the etching process if the sample had not been previously measured before the etching process, the apparatus comprising a processor configured to: use the generator model to generate hypothesis data based on the paired measurement data, wherein the paired measurement data and the hypothesis data can be used to form an image of the sample; use a classifier to determine the likelihood that the hypothesis data is from the same data distribution as actual measurement data measured after the etching process from a sample that was not previously measured before the etching process; and train the generator model based on a function indicating the level of correlation between the paired measurement data and the hypothesis data and the likelihood determined by the classifier.

[0013]

[0013] According to one aspect of the present invention, there is provided a computer-readable medium storing instructions configured to control a processor to train a generator model that processes paired measurement data measured after the etching process from a sample that was previously measured before the etching process to generate hypothesis data that simulates what the sample would be like after the etching process if the sample had not been previously measured before the etching process, the computer-readable medium storing instructions configured to control a processor to: use the generator model to generate hypothesis data based on the paired measurement data, wherein the paired measurement data and the hypothesis data can be used to form an image of the sample; use an identifier to determine the likelihood that the hypothesis data is from the same data distribution as actual measurement data measured after the etching process from a sample that was not previously measured before the etching process; and train the generator model based on a function indicating the level of correlation between the paired measurement data and the hypothesis data and the likelihood determined by the identifier. [Brief explanation of the drawings]

[0014]

[0014] The above and other aspects of the present disclosure will become more apparent from the description of exemplary embodiments taken in conjunction with the accompanying drawings. [Figure 1]

[0015] FIG. 1 is a schematic diagram illustrating an exemplary charged particle beam inspection system. [Figure 2]

[0016] 2 is a schematic diagram illustrating an exemplary multi-beam charged particle characterization apparatus that is part of the exemplary charged particle beam inspection system of FIG. 1. [Figure 3]

[0017] 1 is a schematic diagram of an exemplary single beam electron optical column. [Figure 4]

[0018] 1 shows images of the sample before and after the etching process taken at different locations on the sample. [Figure 5]

[0019] 1 shows a patch of an image of a sample before the etching process paired with a patch of a predicted image of the same sample after the etching process.

[0015]

[0020] These schematic diagrams show the components described below, however, the components shown in the figures are not to scale. DETAILED DESCRIPTION OF THE INVENTION

[0016]

[0021] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following description of exemplary embodiments do not represent all implementations consistent with the present invention. Instead, the implementations are merely examples of apparatus and methods consistent with aspects related to the present invention as set forth in the appended claims.

[0017]

[0022] Increased computing power in electronic devices, which reduces the physical size of devices, can be achieved by significantly increasing the packing density of circuit components such as transistors, capacitors, and diodes on IC chips. This has been made possible by improvements in resolution, which allow for the creation of ever-smaller structures. For example, an IC chip in a smartphone the size of a thumbnail and available before 2019 can contain over 2 billion transistors, each less than 1 / 1000 the size of a human hair. It is not surprising, then, that semiconductor IC manufacturing is a complex and time-consuming process involving hundreds of individual steps. An error in even a single step can dramatically affect the functionality of the final product. Just one "killer defect" can cause device failure. The goal of a manufacturing process is to improve the overall yield of the process. For example, to achieve a 75% yield for a 50-step process (where a step can refer to the number of layers formed on a wafer), each individual step must have a yield greater than 99.4%. If each individual step has a 95% yield, the overall process yield is as low as 7%.

[0018]

[0023] In IC chip manufacturing facilities, while high process yields are desirable, maintaining high substrate (i.e., wafer) throughput, defined as the number of substrates processed per hour, is also essential. High process yields and high substrate throughput can be affected by the presence of defects. This is especially true when operator intervention is required to investigate the defects. Therefore, high-throughput detection and identification of microscale and nanoscale defects by inspection devices (such as scanning electron microscopes ("SEMs")) is essential to maintaining high yields and low costs.

[0019]

[0024] An SEM includes a scanning device and a detector system. The scanning device includes an illumination system, which includes an electron source for generating primary electrons, and a projection system for scanning a sample, such as a substrate, with one or more focused beams of primary electrons. At least the illumination system or illumination system and the projection system or projection system may be collectively referred to as an electron optical system or apparatus. The primary electrons interact with the sample and generate secondary electrons. The detector system captures the secondary electrons from the sample as it is scanned, allowing the SEM to create an image of the scanned area of ​​the sample. For high-throughput inspection, some inspection systems use multiple focused beams of primary electrons (i.e., multibeams). The component beams of a multibeam may be called subbeams or beamlets. A multibeam can simultaneously scan different portions of the sample. Therefore, a multibeam inspection system can inspect a sample at much faster speeds than a single-beam inspection system.

[0020]

[0025] Known multi-beam inspection system implementations are described below.

[0021]

[0026] Although the description and drawings are directed to electron-optical systems, it will be understood that the embodiments are not used to limit the present disclosure to particular charged particles, and thus throughout this document, references to electrons can be considered more generally as references to charged particles that are not necessarily electrons.

[0022]

[0027] Reference is now made to Figure 1, which is a schematic diagram illustrating an exemplary charged particle beam inspection system 100, which may also be referred to as a charged particle beam evaluation system or simply an evaluation system. The charged particle beam inspection system 100 of Figure 1 includes a main chamber 10, a load lock chamber 20, an electron beam system 40, a front-end equipment module (EFEM) 30, and a controller 50. The electron beam system 40 is positioned within the main chamber 10.

[0023]

[0028] The EFEM 30 includes a first loading port 30a and a second loading port 30b. The EFEM 30 may include additional loading port(s). The first loading port 30a and the second loading port 30b can receive, for example, substrate front-opening integrated pods (FOUPs) containing substrates (e.g., semiconductor substrates or substrates made of other materials) or samples to be inspected (hereinafter, substrates, wafers, and samples are collectively referred to as "samples"). One or more robotic arms (not shown) within the EFEM 30 transport the samples to the load lock chamber 20.

[0024]

[0029] The load lock chamber 20 is used to remove gas from around the sample. This creates a vacuum, a local gas pressure lower than the pressure of the surrounding environment. The load lock chamber 20 can be connected to a load lock vacuum pumping system (not shown), which removes gas particles from the load lock chamber 20. Operation of the load lock vacuum pumping system allows the load lock chamber to reach a first pressure below atmospheric pressure. After the first pressure is reached, one or more robotic arms (not shown) transport the sample from the load lock chamber 20 to the main chamber 10. The main chamber 10 is connected to the main chamber vacuum pumping system (not shown). The main chamber vacuum pumping system removes gas particles from the main chamber 10 so that the pressure around the sample reaches a second pressure below the first pressure. After the second pressure is reached, the sample is transported to the electron beam system, where it can be inspected. The electron beam system 40 may include a multi-beam electron optical device.

[0025]

[0030] The controller 50 is electronically connected to the electron beam system 40. The controller 50 may be a processor (such as a computer) configured to control the charged particle beam inspection apparatus 100. The controller 50 may also include processing circuitry configured to perform various signal and image processing functions. While FIG. 1 illustrates the controller 50 as being external to the structure including the main chamber 10, the load lock chamber 20, and the EFEM 30, it is understood that the controller 50 may be part of the structure. The controller 50 may be located within one of the component elements of the charged particle beam inspection apparatus, or the controller 50 may be distributed among at least two of the component elements. While this disclosure provides an example in which the main chamber 10 houses an electron beam system, it should be noted that aspects of this disclosure are not limited to chambers housing electron beam systems in their broadest sense. Rather, it is understood that the principles described above may also be applied to other devices and other arrangements of apparatus operating under a second pressure.

[0026]

[0031] Reference is now made to FIG. 2, which is a schematic diagram illustrating an exemplary electron beam system 40 including a multi-beam electron optical system 41 that is part of the exemplary charged particle beam inspection system 100 of FIG. 1. The electron beam system 40 includes an electron source 201 and a projection device 230. The electron beam system 40 further includes a motorized stage 209 and a sample holder 207. The electron source 201 and the projection device 230 together may be referred to as the electron optical system 41 or an electron optical column. The sample holder 207 is supported by the motorized stage 209 to hold a sample 208 (e.g., a substrate or a mask) for inspection. The multi-beam electron optical system 41 further includes a detector 240 (e.g., an electron detection device).

[0027]

[0032] The electron source 201 may include a cathode (not shown) and an extractor or anode (not shown). In operation, the electron source 201 is configured to emit electrons from the cathode as primary electrons. The primary electrons are extracted or accelerated by the extractor and / or anode to form a primary electron beam 202.

[0028]

[0033] The projection device 230 is configured to convert the primary electron beam 202 into multiple sub-beams 211, 212, 213 and direct each sub-beam onto the sample 208. Although three sub-beams are shown for simplicity, there may be tens, hundreds, thousands, tens of thousands, or hundreds of thousands of sub-beams. The sub-beams may be referred to as beamlets.

[0029]

[0034] 1, such as the electron source 201, the detector 240, the projection system 230, and the motorized stage 209. The controller 50 may perform various image and signal processing functions. The controller 50 may also generate various control signals to govern the operation of the charged particle beam inspection system, including the charged particle multi-beam system.

[0030]

[0035] The projection device 230 can be configured to focus the sub-beams 211, 212, and 213 onto the sample 208 for inspection, forming three probe spots 221, 222, and 223 on the surface of the sample 208. The projection device 230 can be configured to deflect the primary sub-beams 211, 212, and 213 to scan the probe spots 221, 222, and 223 over respective scan areas of a section of the surface of the sample 208. In response to the incidence of the primary sub-beams 211, 212, and 213 on the probe spots 221, 222, and 223 on the sample 208, electrons, including secondary electrons and backscattered electrons, which may also be referred to as signal particles, are generated from the sample 208. Secondary electrons typically have electron energies of 50 eV or less. While actual secondary electrons may have energies less than 5 eV, energies less than 50 eV are typically considered secondary electrons. Backscattered electrons typically have electron energies between 0 eV and the landing energy of the primary sub-beams 211, 212 and 213. Detected electrons with energies less than 50 eV are usually treated as secondary electrons, so a proportion of the actual backscattered electrons are counted as secondary electrons.

[0031]

[0036] The detector 240 is configured to detect signal particles, such as secondary electrons and / or backscattered electrons, and generate corresponding signals that are sent to a signal processing system 280, for example, to construct an image of a corresponding scanned area of ​​the sample 208. The detector 240 may be integrated into the projection device 230.

[0032]

[0037] The signal processing system 280 may include circuitry (not shown) configured to process signals from the detector 240 to form an image. The signal processing system 280 may otherwise be referred to as an image processing system. The signal processing system may be incorporated into a component of the electron beam system 40, such as the detector 240 (as shown in FIG. 2). However, the signal processing system 280 may be incorporated into any number of components of the inspection apparatus 100 or the electron beam system 40, such as part of the projection apparatus 230 or the controller 50. The signal processing system 280 may include an image acquirer (not shown) and a storage device (not shown). For example, the signal processing system may include a processor, a computer, a server, a mainframe host, a terminal, a personal computer, any type of mobile computing device, etc., or a combination thereof. The image acquirer may include at least a portion of the processing functionality of the controller. As such, the image acquirer may include at least one or more processors. The image acquirer may be communicatively coupled to the detector 240 via a signal communication mechanism, such as electrical conductors, fiber optic cables, portable storage media, IR, Bluetooth, the Internet, a wireless network, a wireless radio, or a combination thereof, among others. The image acquirer may receive signals from the detector 240, process the data contained in the signals, and construct an image therefrom. Thus, the image acquirer may acquire an image of the sample 208. The image acquirer may also perform various post-processing functions, such as contour generation and overlaying indicators on the acquired image. The image acquirer may be configured to adjust the brightness and contrast of the acquired image, etc. The storage may be a storage medium, such as a hard disk, a flash drive, cloud storage, random access memory (RAM), or other type of computer-readable memory. The storage may be coupled to the image acquirer and may be used to save the raw scanned image data as the original image and to save post-processed images.

[0033]

[0038] The signal processing system 280 may include measurement circuitry (e.g., an analog-to-digital converter) to obtain a distribution of detected secondary electrons. The electron distribution data collected within the detection time window, in combination with the corresponding scan path data of each of the primary sub-beams 211, 212, and 213 incident on the sample surface, can be used to reconstruct an image of the sample structure under inspection. The reconstructed image can be used to reveal various features of the internal or external structure of the sample 208. The reconstructed image can thereby be used to reveal defects that may be present in the sample. The above functions of the signal processing system 280 may be performed in the controller 50 or may be shared between the signal processing system 280 and the controller 50 as appropriate.

[0034]

[0039] The controller 50 can control the motorized stage 209 to move the sample 208 during inspection of the sample 208. The controller 50 can enable the motorized stage 209 to move the sample 208 in a certain direction, e.g., at a constant speed, preferably continuously, at least during inspection of the sample, which can be referred to as a type of scanning. The controller 50 can control the movement of the motorized stage 209 such that the motorized stage 209 varies the speed of movement of the sample 208 depending on various parameters. For example, the controller 50 can control the stage speed (including its direction) depending on the characteristics of the inspection step of the scanning process and / or the scan of the scanning process, at least as far as the combined step and scan strategy of the stage is concerned, as disclosed, for example, in European Application No. A21171877.0, filed May 3, 2021, which is incorporated herein by reference.

[0035]

[0040] Known multi-beam systems, such as the electron beam system 40 and charged particle beam inspection apparatus 100 described above, are disclosed in U.S. Patent Application Publication No. 2020118784, U.S. Patent Application Publication No. 20200203116, U.S. Patent Application Publication No. 2019 / 0259570, and U.S. Patent Application Publication No. 2019 / 0259564, which are incorporated herein by reference.

[0036]

[0041] The electron beam system 40 may include a projection assembly for illuminating the sample 208 and thereby regulating charge buildup on the sample.

[0037]

[0042] FIG. 3 is a schematic diagram of an exemplary single-beam electron beam system 41''' according to one embodiment. As shown in FIG. 3, in one embodiment, the electron beam system includes a sample holder 207 supported by a motorized stage 209 to hold a sample 208 to be inspected. The electron beam system includes an electron source 201. The electron beam system further includes a gun aperture 122, a beam-limiting aperture 125, a condenser lens 126, a column aperture 135, an objective lens assembly 132, and an electron detector 144. The objective lens assembly 132 in some embodiments may be a modified swinging objective deceleration immersion lens (SORIL) including a pole piece 132a, a control electrode 132b, a deflector 132c, and an excitation coil 132d. The control electrode 132b has an aperture formed therein for passage of the electron beam. The control electrode 132b forms a facing surface 72, which will be described in detail below.

[0038]

[0043] In the imaging process, the electron beam from the electron source 201 may pass through the gun aperture 122, the beam-limiting aperture 125, the condenser lens 126, be focused by a modified SORIL lens into a probe spot, and then impinge on the surface of the sample 208. The probe spot may be scanned across the surface of the sample 208 by the deflector 132c or other deflectors in the SORIL lens. Secondary electrons emerging from the sample surface may be collected by the electron detector 144 to form an image of the area of ​​interest on the sample 208.

[0039]

[0044] The collection and illumination optics of the electron optical system 41 may include or be supplemented by electromagnetic quadrupole electron lenses. For example, as shown in FIG. 3 , the electron optical system 41 may include a first quadrupole lens 148 and a second quadrupole lens 158. In one embodiment, the quadrupole lenses are used to control the electron beam. For example, the first quadrupole lens 148 may be controlled to adjust the beam current, and the second quadrupole lens 158 may be controlled to adjust the beam spot size and beam shape.

[0040]

[0045] As mentioned in the introduction to this description, sample evaluation methods can be used to evaluate the degree of undesired visible defects in a sample. Such methods can include scanning the sample (or at least a portion of the sample) at one or more stages of the process of forming a pattern. When manufacturing IC chips, the fabrication process can include a lithographic exposure process. The lithographic exposure process can include irradiating the sample (i.e., substrate) with radiation. For example, a resist (e.g., photoresist) can be irradiated. The fabrication process can include an etching process. The etching process can include etching irradiated or non-irradiated portions of the resist.

[0041]

[0046] The sample may be scanned after the exposure process and before the etching process (so-called after-development inspection (ADI)). This scanning may generate information about the extent of unwanted visible defects in the sample. This information may be used to form an image of the sample. In some cases, it may not be necessary to generate an image. For example, the information (from which an image may be formed) may be used in subsequent processing steps without actually generating an image. A data set generated by scanning the sample after the exposure process and before the etching process may be referred to as an ADI image. Additionally or alternatively, the sample may be scanned after the etching process (and therefore also after the exposure process) (so-called after-etch inspection (AEI)). This may generate data about the extent of unwanted visible defects in the sample. This data may be referred to as an AEI image. In some cases, it may not be necessary to generate a visual image. Instead, the data that may be used to form an image may be used in subsequent processing steps without actually generating an image.

[0042]

[0047] A method for training a generator model is disclosed. The generator model is configured to process first measurement data 60 to generate predicted data 81. The first measurement data 60 is measured from a sample 208 before an etching process. The etching process is a process of etching the sample 208 (e.g., etching a resist layer on the sample 208). In one embodiment, the first measurement data includes data that can be used to form an ADI image. The predicted data 81 predicts the sample 208 after the etching process. The predicted data 81 can be data that can be used to form an AEI image.

[0043]

[0048] In one embodiment, the method is for training a deep learning model (i.e., a generator model) capable of mapping between SEM images of ADI and AEI. In one embodiment, the method includes measuring first measurement data 60 from the sample 208. For example, the first measurement data 60 can be measured from the sample 208 by scanning the sample 208 using the charged particle beam inspection system 100 (e.g., SEM). Alternatively, the first measurement data 60 can be already measured from the sample 208. The method can use first measurement data 60 that has already been measured from the sample 208.

[0044]

[0049] Figure 4 illustrates a data set used in a method according to an embodiment of the present invention. The data set shown on the left-hand side of Figure 4 corresponds to first measurement data 60. As shown in Figure 4, in one embodiment, first measurement data 60 includes one or more ADI images 61-63. In the example shown in Figure 4, ADI images 61-63 show a contact hole 65 in sample 208. As shown in Figure 4, contact hole 65 may have a generally round shape when viewed from above. Sample 208 may have additional or alternative features. For example, sample 208 may have a linear feature.

[0045]

[0050] In one embodiment, the method includes using the generator model to generate predicted data 81 based on the first measured data 60. The first measured data 60 and the predicted data 81 can be used to form an image of the sample 208.

[0046]

[0051] 5 is a diagram illustrating an ADI image 61 of the first measured data 60 and a predicted AEI image that can be formed from the predicted data 81. Referring to FIG. 5, in one embodiment, the method includes pairing subsets 67-69 of the first measured data 60 with subset 87 of the predicted data 81. As shown in FIG. 5, subsets 67-69, 87 correspond to locations within an image of sample 208 that can be formed from the first measured data 60 and the predicted data 81.

[0047]

[0052] Three possible pairs can be made from the subsets 67-69, 87 of data shown in FIG. 5 . A first pair can be formed between subset 67 and subset 87. These two subsets 67, 87 correspond to the same location of the sample 208, as shown by a comparison between the ADI image 61 shown in FIG. 5 and the AEI image formed from the predicted data 81. In contrast, subsets 68, 69 correspond to a different location than subset 87. A second possible pair can be formed using subset 68 and subset 87. A third possible pair can be formed between subset 69 and subset 87. Each pair consists of a subset of the first measured data 60 and a subset of the predicted data 81. Some pairs, such as the first pair of subsets 67, 87, correspond to the same location of the sample 208. Other pairs, such as the second and third pairs mentioned above, correspond to different locations of the sample 208.

[0048]

[0053] In one embodiment, the method includes training a generator model based on the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations. The correlation is the correlation between two subsets of data in a pair. For example, in one embodiment, the method includes determining the correlation between a subset 67 from the first measured data 60 and a subset 87 from the predicted data 81. The correlation may be determined by calculating the cross-entropy of the paired subsets of data. In one embodiment, the method includes determining the correlation of other paired subsets of data corresponding to different locations (e.g., a second pair of subsets 68, 87 and a third pair of subsets 69, 87). Determining the correlation may include determining the cross-entropy of each pair.

[0049]

[0054] In one embodiment, the generator model is trained to increase the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations. The generator model is trained to generally improve the similarity between the ADI image 61 from the first measured data 60 and the predicted AEI image formed from the predicted data 81. In practice, the sample 208 changes during the etching process. Therefore, some differences are expected to exist between the ADI image 61 and the predicted AEI image formed from the predicted data 81. However, the measured ADI image 61 and the predicted AEI image are expected to generally show the same physical structure, e.g., the same features (e.g., contact holes or linear features) of the sample 208.

[0050]

[0055] As shown in FIG. 5, contact holes 65 visible in the ADI image 61 may also be visible in the predicted AEI image as post-etch contact holes 75. Feature parameters and dimensions may differ between the two images. For example, the critical dimensions (CDs) may differ. However, the general shape of the physical structure may remain the same. By training a generator model based on the correlation of pairs corresponding to the same locations compared to the correlation of pairs corresponding to different locations, the generator model may be expected to improve.

[0051]

[0056] By generating predicted data 81, it may be unnecessary to scan the sample 208 to form AEI images at locations corresponding to the measured ADI images 61-63. An embodiment of the present invention is expected to reduce damage to the sample 208. It is not necessary to measure ADI and AEI images at the same locations on the sample 208. This is desirable because, if not, inspecting the sample 208 may damage the sample. For example, ADI may have a damaging effect on the resist, which may affect what is measured after the etching process. An embodiment of the present invention is expected to improve the accuracy of forming patterns on the sample 208.

[0052]

[0057] The subsets of data 67-69, 87 may be referred to as patches. In one embodiment, the method includes comparing patches from the ADI image 61 and the generated AEI image. In one embodiment, the method includes setting the ADI patch that corresponds to the AEI patch as a positive example. For example, the first pair 67, 87 of the subset is considered a positive example because the ADI patch 67 corresponds to the AEI patch 87 in that the ADI patch 67 and the AEI patch 87 correspond to the same location on the sample. In one embodiment, all other pairs of patches are set as negative examples. For example, the second pair 68, 87 of the subset and the third pair 69, 87 of the subset may be set as negative examples. In one embodiment, the method includes reducing or minimizing the cross-entropy of the positive examples.

[0053]

[0058] In one embodiment, the method includes using a classifier to determine the likelihood that the predicted data 81 is from the same data distribution as the second measured data 70. The second measured data 70 is measured from the sample 208 after the etching process. The second measured data 70 is shown on the right-hand side of FIG. 4. The second measured data 70 can be used to form AEI images 71-73. The AEI images 71-73 show post-etch views of the contact holes 75. In one embodiment, the method includes measuring the second measured data 70 from the sample 208. For example, the second measured data 70 can be measured by scanning the sample 208 after the etching process using a charged particle inspection system 100 (e.g., a SEM). Alternatively, the second measured data 70 may have been measured in advance.

[0054]

[0059] In one embodiment, the present invention uses a generative adversarial network (GAN). The classifier is configured to evaluate the likelihood that a predicted AEI image formed from the prediction data 81 matches the real AEI images 71-73 formed from the second measurement data 70. In one embodiment, the method includes training a generator model based on the likelihood determined by the classifier. For example, the generator model can be trained to increase the likelihood determined by the classifier. By taking into account the decisions made by the classifier, the generator model can be expected to improve.

[0055]

[0060] In one embodiment, the output of the generator model is constrained to generate images (or a set of data that can be used to form images) that are from the same distribution as the AEI images 71-73 in the training set. The training set may include the ADI images 61-63 of the first measured data 60 and the AEI images 71-73 of the second measured data 70. In one embodiment, a classifier network is used to take the predicted data 81 as input and output a value (e.g., a value between 0 and 1). In one embodiment, the higher the value output by the classifier, the more the image formed from the predicted data 81 appears to be from the desired distribution (i.e., the AEI images 71-73 of the second measured data 70).

[0056]

[0061] In one embodiment, the generator model is trained based on the computed patch contrast loss and the classifier. It is expected that one embodiment of the present invention will enable the generator model to be valid over a wide range of possible mappings between ADI images and AEI images. For example, the mapping between the first measurement data 60 and the second measurement data may be irreversible. This means that there is no one-to-one correspondence between, for example, the ADI image 61 from the first measurement data 60 and a corresponding image of the same location on the sample 208 after the etching process. As an example, a differently sized contact hole on the sample 208 before the etching process may subsequently have the same dimensions after the etching process. As a result, it is not possible to map back from the AEI image to determine the dimensions of the contact hole before the etching process. This is an example of an irreversible mapping. It is expected that one embodiment of the present invention will enable the generation of accurate prediction data without relying on the ADI / AEI mapping being invertible.

[0057]

[0062] In one embodiment, the method includes calculating one or more parameter values ​​from the predicted data 81 and from the second measured data 70 for one or more parameters of the features of the sample 208. For example, in one embodiment, the parameters may include one or more of CD, local critical dimension uniformity (LCDU), local edge placement error (LEPE), line edge roughness, and line width roughness. LCDU relates to the uniformity of CD values ​​of features such as contact holes or linear features. CD values ​​can be calculated locally, and their standard deviation is calculated to determine LCDU. LEPE relates to the placement of the edges of the features. LEPE can be a combination of CD and overlay. Line edge roughness relates to the uniformity of the position of the edges of lines in linear features. Line edge roughness can be a measure of how straight the edges of a linear feature are. Line width roughness relates to the uniformity of the width of a linear feature along its length. One or more key performance indicators such as these parameters can be taken into account when training the generator model.

[0058]

[0063] In one embodiment, the method includes comparing one or more parameter values ​​calculated from the predicted data 81 with one or more parameter values ​​calculated from the second measured data 70. In one embodiment, the determination by the classifier relies on a comparison of the one or more parameter values. The classifier may receive the parameter values ​​and perform parameter distribution matching. For example, the classifier may take into account the likelihood that the CD calculated from the predicted data 81 is from a distribution of CD values ​​calculated from the second measured data 70. One embodiment of the present invention is expected to improve a generator model applied to images of a sample 208 during a lithography process.

[0059]

[0064] In one embodiment, the smaller the difference between one or more parameter values ​​calculated from the predicted data 81 and one or more parameter values ​​calculated from the second measured data 70, the greater the likelihood determined by the classifier. Such a small difference may indicate a good match between the predicted data 81 and the second measured data 70. A larger difference may indicate a less realistic match between the predicted data 81 and the second measured data 70.

[0060]

[0065] An embodiment of the present invention is expected to reduce the requirement for ADI / AEI mapping to be reversible. In practice, ADI / AEI mapping is unlikely to be perfectly reversible. An embodiment of the present invention is expected to enable more accurate edge bias prediction across a wider application space. An embodiment of the present invention is expected to improve lithography processes.

[0061]

[0066] In one embodiment, the generator model is trained by minimizing the loss L:

number

[0062]

[0067] The overall loss function L is made up of two contributing losses L GAN and L patch It is the sum of L GAN L relates to how well the predicted AEI image fits the same data distribution as the measured AEI image. GAN refers to the loss function corresponding to the decision made by the classifier. patch L relates to the similarity between the measured ADI image and the predicted AEI image. patchis the loss function for the patch contrast loss. α is a controllable parameter that controls the degree to which the generator model is trained based on increasing the similarity between the ADI and AEI images, and the degree to which the generator model is trained based on improving the similarity between the predicted and measured AEI images. A larger value of α means that the generator model is trained to a greater extent based on increasing the similarity between the measured ADI and predicted AEI images.

[0063]

[0068] The loss function is a function of G, D, X, Y, and H. G refers to the generator model. D refers to the discriminator model. X refers to the first measured data 60. Y refers to the second measured data 70. H refers to a network that extracts features from the first measured data 60 to compress the first measured data 60 into relevant features of the samples 208.

[0064]

[0069] The contribution loss function can be mathematically defined as follows:

number

number

[0065]

[0070] E refers to the expected value over the data distribution. The z component relates to a condensed representation of the patch (i.e., a subset of data from the first measured data 60 and predicted data 81).

[0066]

[0071] In one embodiment, the generator model includes an encoder and a decoder. As indicated above, in one embodiment, the method includes determining the cross entropy of extracted features of the paired subsets of data to determine the correlation between the paired subsets of data. In one embodiment, the method includes encoding the paired subsets of data with an encoder of the generator model so that the features can be extracted.

[0067]

[0072] As indicated above, in one embodiment, the patch contrast loss is calculated by summing over multiple patches. In one embodiment, the patch contrast loss is determined by summing over multiple layers of the sample 208. The sample 208 may include multiple layers. Each layer may include a set of features.

[0068]

[0073] In one embodiment, a method is provided for processing first measurement data 60 measured from a sample 208 before an etching process to generate predicted data 81 that predicts the sample 208 after the etching process. In one embodiment, the method includes using a generator model to generate the predicted data 81 based on the first measurement data 60. In one embodiment, the generator model has been trained according to the method described above.

[0069]

[0074] In one embodiment, the method includes applying a deep learning model (i.e., a generator model) that can perform a mapping between SEM images of the ADI and the AEI. In one embodiment, an SEM image of a given ADI is converted to a corresponding SEM image of a predicted AEI. Alternatively, images may not need to be generated. In one embodiment, the method includes converting given ADI data to corresponding predicted AEI data.

[0070]

[0075] As explained above, it is possible to generate predicted data without using measure-etch-measure (MEM) data, which is data measured when both ADI and AEI are performed at the relevant locations. In an alternative embodiment, a method is provided for applying corrections to the MEM data using unpaired image mapping.

[0071]

[0076] As explained above, patch contrast loss and classifiers can be used. In one embodiment, these techniques, or alternatively, cycleGAN techniques, can be used in combination with MEMs data.

[0072]

[0077] In one embodiment, there is a method for training a generator model that processes paired measurement data measured after the etch process from a sample that was previously measured before the etch process to generate hypothetical data that simulates what the sample would have been like after the etch process if the sample had not been previously measured before the etch process.

[0073]

[0078] In one embodiment, the method includes generating paired measurement data. In one embodiment, the method includes inspecting a portion of the sample after a development process and before an etching process (e.g., generating an ADI image). The inspection may be performed with an electron beam. In one embodiment, the charged particle beam inspection system 100 is used to inspect the sample with one or more charged particle beams (e.g., an electron beam). The method further includes inspecting the same portion of the sample after the etching process (e.g., generating an AEI image). The ADI image and the AEI image may correspond to the same location on the sample. The data measured after the etching process is paired measurement data.

[0074]

[0079] During the process of inspecting the sample before the etching process, the sample may be damaged. In particular, the electron beam(s) used to inspect the sample may damage the resist. The damage affects the paired measurement data recorded after the etching process. Therefore, the paired measurement data differs from the data that would have been measured after etching if the sample had not been inspected before the etching process. Of course, there may be some stochastic variation in the measurement data. Such stochastic variation can be accounted for by repeating measurements and averaging. However, sample damage caused by inspection performed after development and before etching results in a difference that remains even when repeat measurements are averaged. Due to resist damage caused by ADI, the AEI pattern is affected and not completely realistic.

[0075]

[0080] In an alternative embodiment, the paired measurement data is pre-generated, and the method can be performed using such pre-provided paired measurement data, and thus the step of generating the paired measurement data can be omitted.

[0076]

[0081] In one embodiment, the method includes training a network (such as a cycleGAN, or a network using a patch contrast loss and a classifier) ​​to learn how to map AEI data from the MEM data (i.e., paired measurement data inspected after development) to AEI data that was not damaged during ADI (e.g., hypothetical data that simulates the sample after the etching process if the sample had not been pre-measured before the etching process).

[0077]

[0082] Alternatively, such a network may be pre-trained. A pre-trained network can be provided and used to perform the mapping. The training step can be omitted.

[0078]

[0083] In one embodiment, the method includes applying the trained model to new MEMs data. The output (i.e., the hypothetical data) can be used in related studies. For example, the output can be used to monitor whether there is variation (e.g., drift) in the etch process over time. The output can be used to monitor one or more other aspects (e.g., defects) of the etch process.

[0079]

[0084] The present invention can be embodied as a training scheme for using MEM experimental data while reducing / avoiding the effects of ADI loss on AEI data. One embodiment of the present invention is expected to generate hypothetical data that is closer to reality (i.e., closer to patterns that were not subjected to ADI) than the actual patterns examined during the MEM experiment.

[0080]

[0085] In one embodiment, the method includes using a generator model to generate hypothetical data based on the paired measurement data. The paired measurement data may be used to form an image of the sample (i.e., an AEI image). The hypothetical data may be used to form an image of the sample (i.e., a hypothetical AEI image). It is not necessary to actually generate an image. In one embodiment, the data may remain in a non-image form. Alternatively, an image may be generated and displayed.

[0081]

[0086] In one embodiment, the method includes using a classifier to determine the likelihood that the hypothetical data is from the same data distribution as actual measured data measured after the etching process from samples that were not previously measured before the etching process.

[0082]

[0087] In one embodiment, the method includes generating real measurement data. The real measurement data may be used to form an image of the sample, i.e., an AEI image (or multiple images of one or more locations on the sample). Alternatively, the real measurement data may be pre-generated. The step of generating the real measurement data may be omitted.

[0083]

[0088] In one embodiment, the actual measurement data corresponds to one or more locations on the sample that are physically similar (e.g., have similar features) to locations corresponding to the paired measurement data. For example, if the paired measurement data can be used to form an image of a contact hole, the actual measurement data desirably corresponds to locations having a similarly positioned contact hole.

[0084]

[0089] In one embodiment, the actual measurement data corresponds to one or more locations of the sample that are located near the locations corresponding to the paired measurement data. For example, preferably, the actual measurement data corresponds to locations adjacent to the locations corresponding to the paired measurement data. This can help explain systematic variations (which may be called fingerprints) in data measured at different locations of the sample. For example, there may be systematic effects related to the distance of the locations from the center of the sample. In one embodiment, the actual measurement data corresponds to one or more locations that are the same distance from the center of the sample as the paired measurement data.

[0085]

[0090] In one embodiment, the actual measurement data corresponds to one or more locations on the sample having a surrounding pattern density similar to the locations corresponding to the paired measurement data. For example, the locations of the paired measurement data may be surrounded by contact holes in a regular hexagonal pattern. Desirably, the actual measurement data similarly corresponds to one or more locations on the sample surrounded by regularly spaced contact holes.

[0086]

[0091] In one embodiment, the method includes training a generator model based on the likelihood determined by the classifier. In one embodiment, the generator model is trained to increase the likelihood determined by the classifier. This helps increase the proximity between the generated hypothetical data and the data that would have been measured if the sample location had not been subjected to ADI.

[0087]

[0092] As mentioned above, in one embodiment, a contrastive loss may be used. In one embodiment, the method includes training a generator model based on a function indicative of a level of correlation between paired measurement data and hypothetical data. In one embodiment, the method includes pairing a subset of paired measurement data with a subset of hypothetical data, the subsets corresponding to locations in an image of a sample formable from the paired measurement data and the hypothetical data. The function (indicative of a level of correlation between the paired measurement data and the hypothetical data) is the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations, and the correlation is the correlation between the paired subsets of data.

[0088]

[0093] In one embodiment, the generator model is trained to increase the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations. In one embodiment, the method includes determining the cross-entropy of extracted features of the paired subsets of data to determine the correlation between the paired subsets of data. In one embodiment, the method includes encoding the paired subsets of data with an encoder of the generator model so that features can be extracted.

[0089]

[0094] Alternatively, as mentioned above, in one embodiment, cycleGAN can be used. In one embodiment, the method includes using a posterior generator model to generate simulated paired data based on hypothetical data, where the simulated paired data can be used to form an image of the sample. The function (indicative of the level of correlation between the paired measured data and the hypothetical data) is the similarity between the paired measured data and the simulated paired data. In one embodiment, the generator model and the posterior generator model are trained to increase the similarity between the paired measured data and the simulated paired data.

[0090]

[0095] In one embodiment, there is provided a processor apparatus comprising a processor configured to perform the methods described above. For example, the processor may be configured to perform the method of training a generator model. Additionally or alternatively, the processor may be configured to perform the method of applying a generator model.

[0091]

[0096] In some embodiments, an evaluation method is provided that includes generating a sample map according to any of the methods described above. The evaluation method includes inspecting the sample 208 using the generated sample map to locate one or more features of interest. The evaluation method includes assessing the degree to which the one or more features of interest contain defects. This may be accomplished by comparing the image of the feature of interest to reference images elsewhere on the sample, on other samples, or in a database.

[0092]

[0097] In some embodiments, inspection is performed by directing one or more beams of charged particles (e.g., electrons) onto the sample 208 and detecting one or more charged particles (e.g., electrons) emitted from the sample 208. The inspection can use any of the electron optical arrangements described above with reference to Figures 1-5, or any other suitable arrangement for inspecting a sample using a beam of charged particles. The inspection can also be performed using other techniques, such as optical techniques based on electromagnetic radiation.

[0093]

[0098] References to upper and lower, up and down, above and below should be understood as referring to directions parallel to the (typically, but not always perpendicular) up-beam and down-beam directions of the electron beam or multi-beam impinging on the sample 208. As such, references to up-beam and down-beam are intended to refer to directions relative to the beam path, independent of any prevailing gravitational field.

[0094]

[0099] The embodiments described herein may take the form of a series of aperture arrays or electron-optical elements arranged in an array along a beam path or multi-beam path. Such electron-optical elements may be electrostatic. In one embodiment, for example, all electron-optical elements from the beam-limiting aperture array to the final electron-optical element in the sub-beam path before the sample may be electrostatic and / or in the form of aperture arrays or plate arrays. In some arrangements, one or more of the electron-optical elements are fabricated as microelectromechanical systems (MEMS) (i.e., by using MEMS fabrication techniques). The electron-optical elements may include magnetic and electrostatic elements. For example, a compound array lens may feature a macro-magnetic lens arranged along the multi-beam path, containing the multi-beam path with upper and lower pole plates within the magnetic lens. Within the pole plates may be an array of apertures for the multi-beam beam paths. Electrodes may be present above, below, or between the pole plates to control and optimize the electromagnetic field of the compound lens array.

[0095]

[0100] An evaluation tool or evaluation system according to the present disclosure may include a device that performs a qualitative evaluation (e.g., pass / fail) of a sample, a device that performs a quantitative measurement (e.g., size of features) of a sample, or a device that generates an image of a map of a sample. Examples of evaluation tools or systems are inspection tools (e.g., to identify defects), review tools (e.g., to classify defects), and metrology tools, or tools that can perform any combination of evaluation functions associated with an inspection tool, review tool, or metrology tool (e.g., metro inspection tool).

[0096]

[0101] Reference to a component or system of components or elements that is controllable to manipulate a charged particle beam in a particular manner includes configuring a controller or control system or control unit to control the component to manipulate the charged particle beam in the manner described, and optionally using other controllers or devices (e.g., voltage sources) to control the component to manipulate the charged particle beam in this manner. For example, a voltage source may be electrically connected to one or more components, such as electrodes of control lens array 250 and objective lens array 241, to apply a potential to the component under the control of a controller or control system or control unit. A drivable component, such as a stage, may be actuated using one or more controllers, control systems, or control units to control the actuation of the component, and thus controllable to move relative to another component, such as the beam path.

[0097]

[0102] The functions provided by a controller or control system or control unit may be computer-implemented. Any suitable combination of elements may be used to provide the required functionality, including, for example, a CPU, RAM, SSD, motherboard, network connection, firmware, software, and / or other elements known in the art that enable the required computing operations to be performed. The required computing operations may be defined by one or more computer programs. The one or more computer programs may be provided in the form of a medium, optionally a non-transitory medium, that stores computer-readable instructions. When the computer-readable instructions are read by a computer, the computer performs the required method steps. The computer may consist of a self-contained unit or a distributed computing system with multiple different computers connected to each other via a network.

[0098]

[0103] The terms "sub-beam" and "beamlet" are used interchangeably herein and are both understood to encompass any radiation beam derived from a parent radiation beam by splitting or separating the parent radiation beam. The term "manipulator" is used to encompass any element that affects the path of a sub-beam or beamlet, such as a lens or deflector. References to elements that are aligned along a beam path or sub-beam path are understood to mean that the respective element is positioned along the beam path or sub-beam path. References to optics are understood to mean electron optics.

[0099]

[0104] The methods of the present invention can be performed by a computer system including one or more computers. A computer used to implement the present invention can include one or more processors, including a general-purpose CPU, a graphics processing unit (GPU), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other special-purpose processor. As noted above, in some cases, a particular type of processor can offer advantages in terms of reduced cost and / or increased processing speed, and the methods of the present invention can be adapted for use with a particular processor type. Some steps of the methods of the present invention involve parallel processing, which is amenable to implementation on a processor capable of parallel processing (e.g., a GPU).

[0100]

[0105] As used herein, the term "image" is intended to refer to any array of values, where each value relates to a sample at a location, and the arrangement of values ​​in the array corresponds to the spatial arrangement of the sampled locations. An image may include a single layer or multiple layers. In the case of a multiple layer image, each layer, which may also be called a channel, represents a different sample at several locations. The term "pixel" is intended to refer to a single value in the array, or, in the case of a multiple layer image, to a group of values ​​corresponding to a single location.

[0101]

[0106] Embodiments of the present disclosure are set forth in the following numbered clauses: 1. A method of training a generator model that processes first measurement data measured from a sample before an etching process to generate prediction data that predicts the sample after the etching process, comprising: generating predicted data based on the first measured data using the generator model, the first measured data and the predicted data being usable to form an image of the sample; pairing a first subset of measured data with a subset of predicted data, the subset corresponding to a location within an image of the sample formable from the measured data and the predicted data; using the classifier to determine the likelihood that the predicted data is from the same data distribution as second measured data measured from the sample at a different location after the etching process; The generator model is correlations between pairs of data corresponding to the same location compared to correlations between pairs of data corresponding to different locations; and the likelihood determined by the classifier, and training based on the 2. The method of clause 1, wherein the generator model is trained to increase the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations. 3. The method of clause 1 or 2, wherein the generator model is trained to increase the likelihood determined by the classifier. 4. Calculating one or more parameter values ​​for one or more parameters of the feature of the sample from the predicted data and from the second measured data; comparing one or more parameter values ​​calculated from the predicted data with one or more parameter values ​​calculated from the second measured data; 4. The method of any one of clauses 1 to 3, wherein the determination by the classifier relies on a comparison of one or more parameter values. 5. The method of clause 4, wherein the parameters include one or more of critical dimension, local critical dimension uniformity, local edge placement error, line edge roughness, and line width roughness. 6. The method according to clause 4 or 5, wherein the smaller the difference between one or more parameter values ​​calculated from the prediction data and one or more parameter values ​​calculated from the second measurement data, the greater the likelihood determined by the classifier. 7. The method of any one of clauses 1 to 6, wherein the second measurement data corresponds to a different position than the first measurement data. 8. The method of any one of clauses 1 to 7, wherein the mapping between the first measurement data and the second measurement data is irreversible. 9. The method of any one of clauses 1 to 8, wherein the generator model includes an encoder and a decoder. 10. The method of clause 9, comprising determining the cross entropy of extracted features of the paired subsets of data to determine the correlation between the paired subsets of data. 11. The method of clause 10, comprising encoding the paired subsets of data with an encoder so that features can be extracted. 12. The method of any one of clauses 1 to 11, wherein the first measurement data is measured from the sample before the etching process and after the lithography exposure process. 13. A method of processing first measurement data measured from a sample before an etching process to generate prediction data predicting the sample after an etching process, comprising: 13. A method comprising: generating predicted data based on first measurement data using a generator model, the generator model having been trained by a method according to any one of clauses 1 to 12. 14. A processing device comprising a processor configured to perform the method of any one of clauses 1 to 13. 15. A computer program comprising instructions configured to control a processor to perform a method according to any one of clauses 1 to 14. 16. A generator model training apparatus for training a generator model that processes first measurement data measured from a sample before an etching process to generate prediction data that predicts the sample after the etching process, the apparatus comprising: a processor; generating predicted data based on the first measured data using the generator model, the first measured data and the predicted data being usable to form an image of the sample; pairing a subset of the first measured data with a subset of predicted data, the subset corresponding to a location within an image of the sample formable from the first measured data and the predicted data; using the classifier to determine the likelihood that the predicted data is from the same data distribution as second measured data measured from the sample at a different location after the etching process; The generator model is correlations between pairs of data corresponding to the same location compared to correlations between pairs of data corresponding to different locations; and the likelihood determined by the classifier, and training the generator model based on the 17. The generator model training device of clause 16, wherein the processor is configured to train the generator model to increase the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations. 18. A generator model training device according to clause 16 or 17, wherein the processor is configured to train the generator model to increase the likelihood determined by the classifier. 19. The processor calculating one or more parameter values ​​for one or more parameters of the feature of the sample from the predicted data and from the second measured data; comparing one or more parameter values ​​calculated from the predicted data with one or more parameter values ​​calculated from the second measured data; 19. The generator model training apparatus of any one of clauses 16 to 18, wherein the determination by the classifier relies on a comparison of one or more parameter values. 20. The generator model training apparatus of clause 19, wherein the parameters include one or more of critical dimension, local critical dimension uniformity, local edge placement error, line edge roughness, and line width roughness. 21. A generator model training device as described in clause 19 or 20, wherein the smaller the difference between one or more parameter values ​​calculated from the prediction data and one or more parameter values ​​calculated from the second measurement data, the greater the likelihood determined by the classifier. 22. A generator model training apparatus according to any one of clauses 16 to 21, wherein the second measurement data corresponds to a different position than the first measurement data. 23. A generator model training device according to any one of clauses 16 to 22, wherein the mapping between the first measurement data and the second measurement data is irreversible. 24. A generator model training device according to any one of clauses 16 to 23, wherein the generator model includes an encoder and a decoder. 25. The generator model training apparatus of clause 24, wherein the processor is configured to determine cross entropy of extracted features of the paired subsets of data to determine correlation between the paired subsets of data. 26. The generator model training apparatus of clause 25, wherein the processor is configured to encode the paired subsets of data with an encoder so that features can be extracted. 27. A generator model training apparatus according to any one of clauses 16 to 26, wherein the first measurement data is measured from the sample before an etching process and after a lithography exposure process. 28. A prediction data generating device for processing first measurement data measured from a sample before an etching process to generate prediction data for predicting the sample after the etching process, the device comprising: a processor; 13. A predictive data generating apparatus, wherein the processor is configured to generate predictive data based on first measurement data using a generator model, the generator model having been trained by a method according to any one of clauses 1 to 12. 29. A computer-readable medium storing instructions configured to control a processor to train a generator model to process first measurement data measured from a sample before an etching process to generate prediction data that predicts the sample after the etching process, the computer-readable medium comprising: generating predicted data based on the first measured data using the generator model, the first measured data and the predicted data being usable to form an image of the sample; pairing a first subset of measured data with a subset of predicted data, the subset corresponding to a location within an image of the sample formable from the measured data and the predicted data; using the classifier to determine the likelihood that the predicted data is from the same data distribution as second measured data measured from the sample at a different location after the etching process; The generator model is correlations between pairs of data corresponding to the same location compared to correlations between pairs of data corresponding to different locations; and the likelihood determined by the classifier, and a computer-readable medium storing instructions configured to control a processor to: 30. The computer-readable medium of clause 29, storing instructions configured to control a processor to train a generator model to increase the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations. 31. A computer-readable medium according to clause 29 or 30, storing instructions configured to control a processor to train a generator model to increase the likelihood determined by the classifier. 32. Calculating one or more parameter values ​​for one or more parameters of the feature of the sample from the predicted data and from the second measured data; comparing one or more parameter values ​​calculated from the predicted data with one or more parameter values ​​calculated from the second measured data; 32. The computer-readable medium of any one of clauses 29-31, wherein the determination by the classifier relies on a comparison of one or more parameter values. 33. The computer-readable medium of clause 32, wherein the parameters include one or more of critical dimension, local critical dimension uniformity, local edge placement error, line edge roughness, and line width roughness. 34. A computer-readable medium according to clause 32 or 33, wherein the smaller the difference between one or more parameter values ​​calculated from the predicted data and one or more parameter values ​​calculated from the second measured data, the greater the likelihood determined by the classifier. 35. The computer-readable medium of any one of clauses 29 to 34, wherein the second measurement data corresponds to a different position than the first measurement data. 36. The computer-readable medium of any one of clauses 29 to 35, wherein the mapping between the first measurement data and the second measurement data is irreversible. 37. The computer-readable medium of any one of clauses 29 to 36, wherein the generator model includes an encoder and a decoder. 38. The computer-readable medium of clause 37, storing instructions configured to control a processor to determine a cross-entropy of extracted features of paired subsets of data to determine correlation between the paired subsets of data. 39. The computer-readable medium of clause 38 storing instructions configured to control a processor to encode with an encoder the paired subsets of data such that features can be extracted. 40. The computer-readable medium of any one of clauses 29-39, wherein the first measurement data is measured from the sample before an etching process and after a lithography exposure process. 41. A method of training a generator model that processes paired measurement data measured after an etching process from a sample that was previously measured before the etching process to generate hypothetical data that simulates what the sample would be like after the etching process if the sample had not been previously measured before the etching process, comprising: generating hypothesis data based on the paired measurement data using the generator model, wherein the paired measurement data and the hypothesis data can be used to form an image of the sample; using a classifier to determine the likelihood that the hypothetical data is from the same data distribution as actual measured data measured after the etching process from a sample that was not previously measured before the etching process; The generator model is a function indicating the level of correlation between the paired measurement data and the hypothesized data; the likelihood determined by the classifier, and training based on the 42. Pairing a subset of the paired measurement data with a subset of the hypothesis data, the subset corresponding to a location in the image of the sample formable from the paired measurement data and the hypothesis data; 42. The method of claim 41, wherein the function is the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations, the correlation being between paired subsets of data. 43. The method of clause 42, wherein the generator model is trained to increase the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations. 44. The method of clause 42 or 43, comprising determining the cross entropy of extracted features of paired subsets of data to determine the correlation between the paired subsets of data. 45. The method of clause 44, comprising encoding the paired subsets of data with an encoder of a generator model so that features can be extracted. 46. ​​Using a backward generator model to generate simulated paired data based on hypothetical data, the simulated paired data being usable to form an image of the sample; 42. The method of claim 41, wherein the function is a similarity between the paired measurement data and the simulated paired data. 47. The method of clause 46, wherein the generator model and the backward generator model are trained to increase the similarity between the paired measurement data and the simulated paired data. 48. The method of any one of clauses 41 to 47, wherein the generator model is trained to increase the likelihood determined by the classifier. 49. The method of any one of clauses 41 to 48, wherein the paired measurement data is measured after the etching process from a sample that has been previously measured before the etching process and after the lithography exposure process. 50. A method of processing paired measurement data measured after an etching process from a sample pre-measured before the etching process to generate hypothetical data simulating what the sample would be like after the etching process if the sample had not been pre-measured before the etching process, comprising: 50. A method comprising generating hypothesis data based on paired measurement data using a generator model, the generator model having been trained by a method according to any one of clauses 41 to 49. 51. A processing device comprising a processor configured to perform the method according to any one of clauses 41 to 50. 52. A computer program comprising instructions configured to control a processor to perform a method according to any one of clauses 41 to 51. 53. A generator model training apparatus for training a generator model that processes paired measurement data measured after an etching process from a sample that has been previously measured before the etching process to generate hypothetical data simulating the sample after the etching process if the sample had not been previously measured before the etching process, the apparatus comprising: a processor; generating hypothesis data based on the paired measurement data using the generator model, wherein the paired measurement data and the hypothesis data can be used to form an image of the sample; using a classifier to determine the likelihood that the hypothetical data is from the same data distribution as actual measured data measured after the etching process from a sample that was not previously measured before the etching process; The generator model is a function indicating the level of correlation between the paired measurement data and the hypothesized data; the likelihood determined by the classifier, and training the generator model based on the 54. The processor is configured to pair a subset of the paired measurement data with a subset of the hypothesis data, the subset corresponding to a location in the image of the sample formable from the paired measurement data and the hypothesis data; 54. The generator model training apparatus of clause 53, wherein the function is the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations, and the correlation is the correlation between paired subsets of the data. 55. The generator model training device of clause 54, wherein the processor is configured to train the generator model to increase the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations. 56. A generator model training apparatus as described in clause 54 or 55, wherein the processor is configured to determine the cross entropy of extracted features of the paired subsets of data to determine the correlation between the paired subsets of data. 57. A generator model training apparatus as described in clause 56, wherein the processor is configured to encode the paired subsets of data with an encoder of the generator model so that features can be extracted. 58. The processor is configured to generate simulated paired data based on hypothetical data using a posterior generator model, the simulated paired data being usable to form an image of the sample; 54. The generator model training apparatus of clause 53, wherein the function is a similarity between the paired measurement data and the simulated paired data. 59. A generator model training device as described in clause 58, wherein the generator model and the backward generator model are trained to increase the similarity between the paired measurement data and the simulated paired data. 60. A generator model training device as described in any one of clauses 53 to 59, wherein the processor is configured to train the generator model to increase the likelihood determined by the classifier. 61. A generator model training apparatus as described in any one of clauses 53 to 60, wherein the paired measurement data is measured after the etching process from a sample that has been previously measured before the etching process and after the lithography exposure process. 62. A hypothetical data generating device for processing paired measurement data measured after an etching process from a sample previously measured before the etching process to generate hypothetical data simulating the sample after the etching process if the sample had not been previously measured before the etching process, the device comprising: a processor; 50. A hypothesis data generation device configured to generate hypothesis data based on paired measurement data using a generator model, the generator model having been trained by a method according to any one of clauses 41 to 49. 63. A computer-readable medium comprising instructions configured to control a processor to train a generator model that processes paired measurement data measured after an etching process from a sample that was previously measured before the etching process to generate hypothetical data that simulates what the sample would be like after the etching process if the sample had not been previously measured before the etching process, the computer-readable medium comprising: generating hypothesis data based on the paired measurement data using the generator model, wherein the paired measurement data and the hypothesis data can be used to form an image of the sample; using a classifier to determine the likelihood that the hypothetical data is from the same data distribution as actual measured data measured after the etching process from a sample that was not previously measured before the etching process; The generator model is a function indicating the level of correlation between the paired measurement data and the hypothesized data; the likelihood determined by the classifier, and a computer-readable medium storing instructions configured to control a processor to: 64. Store instructions configured to control a processor to: pair a subset of paired measurement data with a subset of hypothesis data, the subset corresponding to a location within an image of a sample formable from the paired measurement data and the hypothesis data; 64. The computer-readable medium of clause 63, wherein the function is the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations, and the correlation is the correlation between paired subsets of the data. 65. The computer-readable medium of clause 64, storing instructions configured to control a processor to train a generator model to increase the correlation of pairs corresponding to the same location as compared to the correlation of pairs corresponding to different locations. 66. The computer-readable medium of clause 64 or 65 storing instructions configured to control a processor to determine the cross-entropy of extracted features of paired subsets of data to determine the correlation between the paired subsets of data. 67. The computer-readable medium of clause 66, storing instructions configured to control a processor to encode the paired subsets of data with an encoder of a generator model so that features can be extracted. 68. Store instructions configured to control a processor to: generate simulated paired data based on hypothetical data using a backward generator model, the simulated paired data usable to form an image of the sample; 64. The computer-readable medium of clause 63, wherein the function is a similarity between the paired measurement data and the simulated paired data. 69. The computer-readable medium of clause 68, wherein the generator model and backward generator model are trained to increase the similarity between the paired measurement data and the simulated paired data. 70. A computer-readable medium according to any one of clauses 63 to 69, storing instructions configured to control a processor to train a generator model to increase the likelihood determined by the classifier. 71. The computer-readable medium of any one of clauses 63 to 70, wherein the paired measurement data is measured after the etching process from a sample that was previously measured before the etching process and after the lithography exposure process.

[0102]

[0107] The computer used to implement the present invention may be physical or virtual. The computer used to implement the present invention may be a server, client, or workstation. Multiple computers used to implement the present invention may be distributed and interconnected via a local area network (LAN) or a wide area network (WAN). The results of the methods of the present invention may be displayed to a user or stored in any suitable storage medium. The present invention may be embodied in a non-transitory computer-readable storage medium that stores instructions for performing the methods of the present invention. The present invention may be embodied in a computer system that includes one or more processors and memory or storage that stores instructions for performing the methods of the present invention.

[0103]

[0108] While the invention has been described in connection with various embodiments, other embodiments of the invention will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the invention being indicated by the following claims.

Claims

1. 1. A method of training a generator model that processes first measurement data measured from a sample before an etching process to generate prediction data that predicts said sample after an etching process, comprising: using the generator model to generate the predicted data based on the first measured data, the first measured data and the predicted data being usable to form an image of the sample; and pairing the first subset of measurement data with a subset of prediction data, the subset corresponding to a location within the image of the sample formable from the measurement data and the prediction data; using a classifier to determine the likelihood that the predicted data is from the same data distribution as second measured data measured from the sample at a different location after the etching process; The generator model is correlations between pairs corresponding to the same location compared to correlations between pairs corresponding to different locations, the correlations being between the paired subsets of data; and the likelihood determined by the classifier; and and training based on the

2. The method of claim 1 , wherein the generator model is trained to increase the correlation of pairs corresponding to the same location compared to the correlation of pairs corresponding to different locations.

3. The method of claim 1 , wherein the generator model is trained to increase the likelihood determined by the classifier.

4. calculating one or more parameter values ​​for one or more parameters of features of the sample from the predicted data and from the second measured data; comparing the one or more parameter values ​​calculated from the predicted data with the one or more parameter values ​​calculated from the second measured data; The method of claim 1 , wherein the determination by the classifier depends on the comparison of the one or more parameter values.

5. The method of claim 4 , wherein the parameters include one or more of a critical dimension, a local critical dimension uniformity, a local edge placement error, a line edge roughness, and a line width roughness.

6. 5. The method of claim 4, wherein the smaller the difference between the one or more parameter values ​​calculated from the predicted data and the one or more parameter values ​​calculated from the second measured data, the greater the likelihood determined by the classifier.

7. The method of claim 1 , wherein the second measurement data corresponds to a different location than the first measurement data.

8. The method of claim 1 , wherein the mapping between the first measurement data and the second measurement data is irreversible.

9. The method of claim 1 , wherein the generator model includes an encoder and a decoder.

10. 10. The method of claim 9, wherein determining the correlation between the paired subsets of data comprises determining a cross-entropy of extracted features of the paired subsets of data.

11. The method of claim 10 , comprising encoding the paired subsets of data with the encoder so that the features can be extracted.

12. The method of claim 1 , wherein the first measurement data is measured from the sample before an etching process and after a lithography exposure process.

13. 1. A method of processing first measurement data measured from a sample before an etching process to generate prediction data predicting said sample after an etching process, comprising:

13. A method comprising: generating the predicted data based on the first measurement data using a generator model, the generator model having been trained by a method according to any one of claims 1 to 12.

14. A processing device comprising a processor configured to perform the method of any one of claims 1 to 13.

15. A computer program comprising instructions adapted to control a processor to perform a method according to any one of claims 1 to 14.