Non-destructive estimation of structural properties of sample via x-ray modeling based on ground truth measurements
Through X-ray measurement and ground truth data analysis, combined with electron beam and X-ray detector, the destructive problem of three-dimensional sample analysis in the prior art is solved, and non-destructive three-dimensional characterization and structural parameter estimation are realized, which is suitable for the semiconductor industry.
Patent Information
- Application Number
- CN202510029424.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-16
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is mostly destructive when analyzing samples of three-dimensional internal structures, making it difficult to achieve non-destructive three-dimensional detection and characterization in large-scale manufacturing.
Using an analysis method based on X-ray measurement and ground truth data, electron beams are projected to the sample through an electron beam source, combined with an X-ray detector and processing circuit system, key features are extracted and sample structural parameters are extrapolated, and non-destructive three-dimensional characterization is performed using loss function and k-NN regression algorithm.
Non-destructive three-dimensional detection and characterization of samples is realized, and the structural parameters of samples can be accurately estimated, and it is suitable for patterned wafers and semiconductor devices in the semiconductor industry.
Smart Images

Figure CN120334272A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to non-destructive three-dimensional probing and characterization of samples based on X-ray measurements. Background Art
[0002] “Three-dimensional” structures are increasingly being used in the semiconductor industry, especially for manufacturing logic and memory components. Thus, as part of quality control, “three-dimensional” data of the structures within a sample must typically be obtained. Currently, most techniques for profiling samples that include three-dimensional internal structures are destructive and may involve extracting thin slices or shaving off sections from the sample, and subsequently examining the thin slices or sections using, for example, transmission electron microscopy (TEM). The challenge remains to develop non-destructive techniques for profiling samples in combination with three-dimensional internal structures, which allow high-volume manufacturing (HVM). Summary of the Invention
[0003] In some embodiments according to the present disclosure, aspects of the present disclosure relate to non-destructive three-dimensional probing and characterization of samples based on X-ray measurements and analysis based on ground truth (GT) data associated with a GT (i.e., actual) sample. More specifically but not exclusively, in some embodiments according to the present disclosure, aspects of the present disclosure relate to non-destructive three-dimensional probing and characterization of samples based on the measurement of characteristic X-rays and modeling extrapolated from GT data.
[0004] In some embodiments, the present invention provides a system for non-destructive characterization of a sample, the system comprising:
[0005] an electron beam (e-beam) source configured to project an e-beam onto the sample being tested at one or more e-beam landing energies;
[0006] an X-ray detector configured to sense X-rays emitted from the sample being tested; and
[0007] a processing circuitry configured to:
[0008] receive X-ray measurement data related to one or more e-beam landing energies from the X-ray detector;
[0009] extract a vector of values of key features that specify the X-ray measurement data from the X-ray measurement data
[0010] and
[0011] based on and a set of vectors estimate values of one or more structural parameters of the sample being tested
[0012] The vector set includes vectors of key features corresponding to ground truth (GT) reference samples, and for each 1 ≤ n ≤ N, specifies values of one or more actually measured structural parameters of the n-th sample, and is obtained by actually measuring the emission of X-rays from the n-th reference sample The emission of X-rays from the n-th reference sample is caused by bombarding the n-th reference sample with an e-beam at each of one or more landing energies.
[0013] In some embodiments of the system, minimize a loss function that is a function of at least and a vector valued function where the vector valued function is extrapolated from In some embodiments, the processing circuitry is configured to estimate by calculating the minimum distance between and In some embodiments, the processing circuitry is configured to estimate by calculating the distance between and
[0014] In some embodiments of the system, the GT samples include samples having the same or a similar intended design as the samples being tested. In some embodiments, the GT samples include specially prepared samples that exhibit selected variations relative to the intended design. In some embodiments, the samples are selected to include expected variations of one or more structural parameters. In some embodiments, one or more structural parameters include any geometric parameters and / or compositional parameters that characterize the samples being tested, and modifications to the any geometric parameters and / or compositional parameters affect at least some of the values of the key features.
[0015] In some embodiments of the system, one or more structural parameters include one or more of the following: the total concentration of at least one material included in the samples being tested; and optionally, when the samples being tested include structures embedded in or on the samples being tested, the width of the embedded structures.
[0016] In some embodiments of the system, the sample under test includes multiple layers. In some embodiments, one or more structural parameters include one or more of the following: (i) the respective at least one thickness of at least one of these layers; (ii) the respective combined thickness of at least two or more of these layers; (iii) the respective at least one mass density of at least one of these layers; and (vi) the respective at least one relative concentration of at least one material in one or more of these layers.
[0017] In some embodiments of the system, one or more e-beam landing energies cause the emission of X-rays regarding one or more characteristic X-ray lines, and the one or more characteristic X-ray lines are related to one or more target substances included in the sample under test. In some embodiments, the X-ray detector is configured to separately sense at least one measured spectrum of the separately emitted X-rays in at least one photon energy range, and the at least one photon energy range includes at least one of the characteristic X-ray lines. In some embodiments, the X-ray measurement data includes the measured spectrum.
[0018] In some embodiments of the system, one or more e-beam landing energies cause the emission of X-rays from at least two of the multiple layers. In some embodiments, one or more e-beam landing energies cause the emission of X-rays from each of the multiple layers.
[0019] In some embodiments of the system, the key feature is the intensity of the characteristic X-ray line and / or the intensity of the background radiation, including the intensity of the characteristic X-ray line and / or the intensity of the background radiation, or is a function of the intensity of the characteristic X-ray line and / or the intensity of the background radiation.
[0020] In some embodiments of the system, the key feature is or includes the intensity of the characteristic X-ray line, and each characteristic X-ray line is normalized by the mean of the background intensity regarding the corresponding characteristic X-ray line. In some embodiments, the key feature is extracted from the difference spectrum obtained for the sample under test by subtracting the control spectrum of the control sample from the measured spectrum of the sample under test.
[0021] In some embodiments of the system, where the double vertical bars indicate the vector norm.
[0022] In some embodiments of the system, where specifies the nominal value of one or more structural parameters, specifies the deviation from the nominal value, is the vector of the values of the key features corresponding to and A is a matrix.
[0023] In some embodiments of the system, matrix A is equal to where the double vertical bars indicate the matrix norm, and for each 1 ≤ n ≤ N,
[0024] In some embodiments of the system, obtain and matrix A as the solution of, where the double vertical bars indicate the matrix norm, and for each 1 ≤ n ≤ N,
[0025] In some embodiments of the system, matrix A is equal to where the double vertical bars indicate the vector norm; for each 1 ≤ n ≤ N, and between at least two data points, α i has different values.
[0026] In some embodiments of the system, obtain and matrix A as the solution of, where the double vertical bars indicate the vector norm; for each 1 ≤ n ≤ N, and between at least two data points, α i has different values.
[0027] In some embodiments of the system, the processing circuitry is configured to apply a k-nearest neighbor (k-NN) regression algorithm relative to as part of estimating to in order to determine the k vectors that are closest to of the k vectors.
[0028] In some embodiments of the system, is the mean or median (optionally weighted) of corresponding to the k closest vectors. In some embodiments, is the weighted mean or median of corresponding to the k closest vectors.
[0029] In some embodiments of the system, the processing circuitry is further configured to obtain by subjecting to a (k = N)-NN classifier relative to a set of vectors of key features where N′ > N, and includes and obtain the additional N′ - N vectors by actually measuring additional GT samples.
[0030] In some embodiments of the system, to derive the intensity of characteristic X-ray spectral lines, the processing circuitry is configured to fit a free curve to each interval of the measured spectrum, each interval being approximately centered on a corresponding characteristic X-ray spectral line and consisting of the vicinity of the characteristic X-ray spectral line, thereby obtaining a corresponding optimized curve.
[0031] In some embodiments of the system, where the free curve is a sum of functions, the sum of functions includes a convex function and a second function, and the second function is a polynomial.
[0032] In some embodiments of the system, the processing circuitry is configured to fit a convex function to the peak of the characteristic X-ray spectral line of the corresponding measured spectrum and fit the second function as part of fitting the free curve, so as to account for the background intensity component of the corresponding measured spectrum.
[0033] In some embodiments of the system, the processing circuitry is further configured to extrapolate from extrapolate
[0034] In some embodiments of the system, the X-ray detector is an energy-dispersive X-ray spectrometer or a wavelength-dispersive X-ray spectrometer.
[0035] In some embodiments of the system, optionally, in one of the manufacturing stages of the patterned wafer, the test sample is a patterned wafer or a part of the patterned wafer.
[0036] In some embodiments, the present invention provides a method for non-destructive characterization of a sample, the method comprising:
[0037] a measurement operation, the measurement operation including sub-operations for each of one or more landing energies:
[0038] projecting an e-beam onto the test sample; and
[0039] obtaining measurement data by measuring the intensity of X-rays emitted from the test sample due to the penetration of the e-beam through the test sample; and
[0040] a measurement data analysis operation, the measurement data analysis operation including sub-operations:
[0041] extracting a vector of values of specified key features from the measurement data and
[0042] estimating the value of one or more structural parameters of the test sample based on and the vector set
[0043] The vector set includes vectors of key features corresponding to ground truth (GT) samples, and for each 1 ≤ n ≤ N, specifies the values of one or more actually measured structural parameters of the n-th sample and is obtained by measuring the emission of X-rays from the n-th sample The emission of the X-rays from the n-th sample is generated by impinging the n-th sample with an e-beam at each of one or more landing energies.
[0044] In some embodiments of the method, a loss function is minimized, the loss function being at least of a key feature and a vector quantization function The vector quantization function is extrapolated from In some embodiments, the method further includes estimating by calculating the minimum distance between and In some embodiments, the method further includes estimating by calculating the distance between and
[0045] In some embodiments of the method, the GT samples include samples having the same or similar intended design as the samples being tested. In some embodiments, the GT samples include specially prepared samples that exhibit selected variations relative to the intended design. In some embodiments, the samples are selected to include expected variations of one or more structural parameters. In some embodiments, the one or more structural parameters include any geometric parameters and / or compositional parameters that characterize the samples being tested, and modification of the any geometric parameters and / or compositional parameters affects at least some of the values of the key features.
[0046] In some embodiments of the method, the one or more structural parameters include one or more of the following: the total concentration of at least one material included in the samples being tested; and optionally, when the samples being tested include structures embedded in or on the samples being tested, the width of the embedded structures.
[0047] In some embodiments of the method, the test sample includes a plurality of layers. In some embodiments, one or more structural parameters include one or more of the following: (i) the respective at least one thickness of at least one of these layers; (ii) the respective combined thickness of at least two or more of these layers; (iii) the respective at least one mass density of at least one of these layers; and (vi) the respective at least one relative concentration of at least one material in one or more of these layers.
[0048] In some embodiments of the method, one or more e-beam landing energies are selected to cause the emission of X-rays regarding one or more characteristic X-ray spectral lines, where the one or more characteristic X-ray spectral lines are related to one or more target substances included in the test sample. In some embodiments, in each implementation of the sub-operations of projecting the e-beam and obtaining measurement data, what is obtained is at least one measured spectrum of the X-rays respectively emitted in at least one photon energy range, where the at least one photon energy range includes at least one of the characteristic X-ray spectral lines. In some embodiments, the X-ray measurement data includes the measured spectrum.
[0049] In some embodiments of the method, one or more e-beam landing energies are selected to cause the emission of X-rays from at least two of the plurality of layers. In some embodiments, one or more e-beam landing energies cause the emission of X-rays from each of the plurality of layers.
[0050] In some embodiments of the method, a key feature is the intensity of the characteristic X-ray spectral line and / or the intensity of the background radiation, includes the intensity of the characteristic X-ray spectral line and / or the intensity of the background radiation, or is a function of the intensity of the characteristic X-ray spectral line and / or the intensity of the background radiation.
[0051] In some embodiments of the method, a key feature is or includes the intensity of the characteristic X-ray spectral line, where each characteristic X-ray spectral line is normalized by the mean of the background intensity regarding the corresponding characteristic X-ray spectral line. In some embodiments, the key feature is extracted from the difference spectrum obtained from the test sample by subtracting the control spectrum of the control sample from the measured spectrum of the test sample.
[0052] In some embodiments of the method, where the double vertical bars indicate the vector norm.
[0053] In some embodiments of the method, where specifies the nominal value of one or more structural parameters, specifies the deviation from the nominal value, is the vector of the values of the key features corresponding to and A is a matrix.
[0054] In some embodiments of the method, matrix A equals where the double vertical bars indicate the matrix norm, and for each 1 ≤ n ≤ N,
[0055] In some embodiments of the method, obtain and matrix A as the solution of, where the double vertical bars indicate the matrix norm, and for each 1 ≤ n ≤ N,
[0056] In some embodiments of the method, matrix A equals where the double vertical bars indicate the vector norm; for each 1 ≤ n ≤ N, and between at least two data points, α i has different values.
[0057] In some embodiments of the method, obtain and matrix A as the solution of, where the double vertical bars indicate the vector norm; for each 1 ≤ n ≤ N, and between at least two data points, α i has different values.
[0058] In some embodiments of the method, as part of estimating the method further includes applying a k-nearest neighbor (k-NN) regression algorithm to with respect to in order to determine the k vectors closest to of the closest.
[0059] In some embodiments of the method, is the mean or median (optionally weighted) of the corresponding to the k closest vectors. In some embodiments, is the weighted mean or median of the corresponding to the k closest vectors.
[0060] In some embodiments of the method, the method further includes obtaining by subjecting to a (k = N)-NN classifier with respect to a set of vectors of key features where N′ > N, and includes and obtaining an additional N′ - N vectors by actually measuring additional GT samples.
[0061] In some embodiments of the method, to derive the intensity of characteristic X-ray spectral lines, the method fits a free curve to each interval of the measured spectrum, each interval being approximately centered on a corresponding characteristic X-ray spectral line and consisting of a vicinity region of the characteristic X-ray spectral line, thereby obtaining a corresponding optimized curve.
[0062] In some embodiments of the method, where the free curve is a sum of functions, the sum of functions includes a convex function and a second function, and the second function is a polynomial.
[0063] In some embodiments of the method, the method further includes, as part of fitting the free curve, fitting a convex function to the peak of the characteristic X-ray spectral line of the corresponding measured spectrum and fitting the second function to account for the background intensity component of the corresponding measured spectrum.
[0064] In some embodiments of the method, the method further includes extrapolating from extrapolate
[0065] In some embodiments of the method, optionally, in one of the manufacturing stages of the patterned wafer, the test sample is the patterned wafer or a part of the patterned wafer.
[0066] The processes and displays presented herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct a more specialized apparatus to perform the desired method(s). The following description shows various structures desired by the various systems. Additionally, embodiments of the present disclosure are not described with reference to any particular programming language. It should be understood that various programming languages may be used to implement the teachings of the present disclosure as described herein.
[0067] Aspects of the present disclosure may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The disclosed embodiments may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including memory storage devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] This disclosure describes some embodiments with reference to the accompanying drawings. The specification, together with the drawings, enables those skilled in the art to understand how to practice some embodiments. The drawings are for illustrative purposes and do not attempt to show the structural details of the embodiments in more detail than is necessary for a basic understanding of the disclosure. For clarity, some of the objects shown in the figures are not drawn to scale. In addition, two different objects in the same figure may be drawn to different scales. In particular, the scale of some objects may be greatly magnified compared to other objects in the same figure.
[0069] In the drawings:
[0070] Figure 1 A flowchart of a method for non-destructive three-dimensional detection and characterization of a test sample based on X-ray measurements of the test sample and analysis based on GT data according to some embodiments is presented;
[0071] Figures 2A to 2D Schematically depicts a test sample depth-probed as part of the characterization of a sample according to a method according to some embodiments Figure 1 ;
[0072] Figure 3 A flowchart of the measurement data analysis operation of some specific embodiments of the method according to Figure 1 is presented;
[0073] Figure 4A A flowchart of the measurement operation of some embodiments of the method according to Figure 1 is presented, showing the X-ray emission spectrum of a sample obtained by implementing the measurement operation;
[0074] Figure 4B A flowchart of the measurement data analysis operation of specific embodiments of the method according to Figure 1 is presented, showing an optimized curve fitted to the Figure 4A X-ray emission spectrum;
[0075] Figure 4C An optimized curve superimposed on the Figure 4A X-ray emission spectrum of Figure 4B is presented;
[0076] Figure 4D A fitted Gaussian included in the optimized curve according to some embodiments included in Figure 4B is presented;
[0077] Figure 4E A fitted polynomial included in the optimized curve according to some embodiments included in Figure 4B is presented, with the fitted polynomial accounting for bremsstrahlung;
[0078] Figure 5Schematically depicts a system for non-destructive three-dimensional detection and characterization of a test sample based on X-ray measurements of the test sample and analysis based on GT data. Detailed Description
[0079] The principles, uses, and implementations taught herein can be better understood with reference to the accompanying description and drawings. After carefully reading the description and drawings herein, those skilled in the art will be able to implement the teachings herein without undue effort or experimentation. In the drawings, like reference numerals always denote like parts.
[0080] This application relates to methods and systems for non-destructive three-dimensional detection and characterization of a test sample (e.g., a semiconductor sample). In principle, the method of the present invention is based on analyzing the X-ray emission profile measured for the test sample (with unknown structural parameters) and a reference data set including ground truth (GT) emission profiles of GT samples (with actually measured, i.e., non-simulated, structural parameters) in order to estimate the values of the structural parameters of the test sample.
[0081] More specifically, according to some embodiments, an e-beam is projected onto the test sample to be depth profiled at each of a plurality of (e-beam) landing energies. Each e-beam penetrates the sample and excites the emission of characteristic X-rays from the sample (and accompanying bremsstrahlung, i.e., background radiation). The greater the e-beam landing energy, the greater the depth to which the e-beam penetrates the sample.
[0082] The spectrum of the emitted X-rays depends on the internal geometry of the sample as well as the material composition of the sample, in particular the distribution of each material (i.e., substance) that makes up the sample. As the e-beam passes through the sample, the e-beam "probes" different regions that the e-beam passes through. The contribution of each penetrated region to the spectrum of the emitted X-rays depends not only on the concentration of each material included in the penetrated region, but also on the energy of the e-beam entering the region, which in turn decreases with depth.
[0083] Since the spectrum of the emitted X-rays and the key features of the spectrum generally depend on the structural parameters of the test sample, the unknown structural parameters of the test sample can be estimated by processing the values of the key features of the spectrum of the emitted X-rays and the values of the key features of the spectrum of the emitted X-rays obtained from reference samples with physically measured structural parameters in a similar experiment.
[0084] As used herein, the term "key feature" relates to selected features of the X-ray measurement, such as the intensity of characteristic X-ray spectral lines, and the selected features are characteristic of or related to the values of the structural parameters of the corresponding sample.
[0085] Accordingly, the present application teaches how to characterize the internal structural parameters of a test sample without destroying the sample based on the values of key features extracted from the measured spectrum of the test sample. The measured spectrum can be analyzed using GT measurements based on reference samples and an optimization tool for the measurement settings used to obtain the measured spectrum.
[0086] Method
[0087] In accordance with one aspect of some embodiments, there is provided a computerized method for non-destructively, three-dimensionally characterizing a test sample, the method comprising: estimating values of structural parameters of the test sample by measuring characteristic X-rays returned from the test sample; and processing the measurement data, as explained in more detail below. Figure 1 A flowchart of an exemplary method 100 in accordance with some embodiments is presented, the method comprising:
[0088] - A measurement operation 110, the measurement operation comprising, for each of one or more e-beam landing energies (i.e., the landing energy of the e-beam):
[0089] ■ A sub-operation 110a, in which the e-beam is projected onto a sample (also referred to as the "test sample").
[0090] ■ A sub-operation 110b, in which measurement data is obtained by measuring the intensity of X-rays emitted from the test sample due to the penetration of the e-beam.
[0091] - A data analysis operation 120, the data analysis operation comprising:
[0092] ■ A sub-operation 120a, in which values of key features of the test sample specified by a vector are extracted from the measurement data.
[0093] ■ A sub-operation 120b, in which values of one or more structural parameters of the test sample are estimated based on and a set of reference vectors The set of reference vectors includes vectors of key features obtained by implementing the measurement operation 110 with respect to a GT reference sample. In some embodiments, the set of reference vectors
[0094] consists of vectors of key features obtained with respect to a GT reference sample. In some embodiments, the set of reference vectors does not include vectors of key features obtained with respect to a reference sample of a computer simulation.
[0095] Figure 5 The method of the present invention can be implemented by the system described in the description) or a system similar to the system.
[0096] The terms "bremsstrahlung" and "background radiation" are used interchangeably herein.
[0097] According to some embodiments, optionally, in one of the manufacturing stages of the patterned wafer, the test sample is selected from the patterned wafer, a part of the patterned wafer, a semiconductor device embedded in the patterned wafer (such as a gate stack), or on the patterned wafer. According to some embodiments, the test sample is or includes a structure that includes one or more semiconductor materials. According to some embodiments, the structure can be configured as part of the manufacturing process of a semiconductor device and / or a component of a semiconductor device. According to some embodiments, the structure can be an auxiliary structure that is configured as part of the manufacturing process of a semiconductor device and / or a component of a semiconductor device. According to some embodiments, optionally, in one of the manufacturing stages of the test sample, the test sample can be or include one or more logic components (such as fin field-effect transistors (FinFETs), gate-all-around (GAA) FETs), memory components (such as dynamic RAM and / or vertical NAND (V-NAND)).
[0098] According to some embodiments, the test sample includes a single layer. According to some embodiments, the test sample includes two or more layers.
[0099] In some embodiments, both the test sample and the reference sample have the same or similar structure. For example, the test sample and the reference sample may both include the same number of layers or be the same type of structure, such as a patterned wafer or a semiconductor device. In some embodiments, both the test sample and the reference sample are made of the same material. "Same or similar" means that the test sample and the reference sample are defined by the same type of structure (as defined above, such as a patterned wafer, a part of the patterned wafer, a semiconductor device embedded in the patterned wafer (such as a gate stack), or a semiconductor device on the patterned wafer, etc.), but may have different measured values (values of structural parameters), such as height, length, thickness of the layer, concentration of the material in each layer, etc.
[0100] As used herein, the term "structural parameter" relates to the physical properties of the sample. According to some embodiments, one or more structural parameters may include one or more of the following: the total concentration of at least one material included in the test sample; and optionally, when the test sample includes a structure embedded in or on the test sample, the width of the embedded structure (such as the width of a gate, fin, or depletion layer).
[0101] More generally, as described in more detail above and below, one or more structural parameters may include any geometric and / or compositional parameters of the sample being tested, modifications to which measurably affect at least some of the components of
[0102] such that values of one or more structural parameters characterizing the sample being tested are allowed to be estimated. Additionally or alternatively, according to some embodiments in which the sample being tested includes one or more layers, one or more structural parameters may include one or more of the following: (i) at least one thickness of at least one of these layers; (ii) a combined thickness of at least two or more of these layers; (iii) at least one mass density of at least one of these layers; and (iv) at least one relative concentration of at least one material in one or more of these layers. As a non-limiting example, with respect to item (iv), one or more structural parameters may include the relative concentration of a first material in a subset of these layers, such as adjacent layers or layers of a first type (which may not be adjacent). The relative concentration of the first material in the subset may correspond to the total number of particles of the first material included in the subset divided by the total number of particles (of all materials, including the first material) included in the subset.
[0103] According to some such embodiments, one or more structural parameters include structural parameters of at least one of two or more layers.
[0104] The first step of the method (illustrated in measurement operation 110 of Figure 1 involves a measurement that includes measuring the intensity of X-rays emitted from the sample being tested due to penetration by an e-beam.
[0105] As used herein, the term "e-beam" represents "electron beam". The term "characteristic X-ray regime" refers to the energy range of photons (i.e., the energy range or equivalent frequency range) within an X-ray spectrum that includes characteristic X-ray spectral lines.
[0106] Parameters characterizing the e-beam, particularly the e-beam landing energy, are selected such that in sub-operation 110a, characteristic X-rays are emitted by particles (specifically, particles of at least one target material) in a detection region centered at a corresponding depth that depends on the e-beam landing energy. More precisely, each detection region may correspond to a respective volume of the sample being tested, where electrons from the respective e-beam may cause ejection of electrons in the inner shells of atoms (in the detection region), leaving each of these atoms with an inner shell vacancy. The inner shell vacancy may be filled by relaxation of outer shell electrons to the inner shell. The relaxation may be accompanied by emission of photons (the energy of the photons being equal to the energy lost by the electrons in transitioning from the outer shell to the inner shell).
[0107] According to some embodiments, the number of e-beam landing energies and the minimum and maximum e-beam landing energies can be selected to ensure probing of the sample under test over a depth range. According to some such embodiments, the number of e-beam landing energies and the minimum and maximum e-beam landing energies can be selected to ensure probing of the sample under test along the depth dimension of the sample.
[0108] According to some embodiments, in each implementation of sub-operation 110b, at least one measured spectrum is obtained. Each of the measured spectra can correspond to a photon energy range or a corresponding photon energy range, the photon energy ranges respectively including at least one characteristic X-ray spectral line related to at least one target substance. The measurement data includes the measured spectra (e.g., consists of the measured spectra).
[0109] As used herein, the term "target substance" is used to refer to a substance (i.e., material) that is included in the sample under test and the spectrum of the substance, at least about one characteristic X-ray spectral line of the substance, is measured as part of measurement operation 110 (i.e., when method 100 is applied to the sample under test).
[0110] According to some embodiments, (a vector of values of key features of the sample under test) is the measured spectrum of characteristic X-ray spectral lines, includes the measured spectrum of characteristic X-ray spectral lines, or is obtained or extracted from the measured spectrum of characteristic X-ray spectral lines (i.e., is dependent on the measured spectrum of characteristic X-ray spectral lines). More specifically, according to some embodiments, each component of can be derived based on an extracted set of parameters characterizing the shape of the spectral peak with respect to the corresponding characteristic X-ray spectral line. According to some embodiments, the key features are the intensity of the characteristic X-ray spectral lines and / or the intensity of the background radiation, include the intensity of the characteristic X-ray spectral lines and / or the intensity of the background radiation, or are functions of the intensity of the characteristic X-ray spectral lines and / or the intensity of the background radiation. According to some embodiments, the key features include the so-called "energy signal" (e.g., consists of the so-called "energy signal"). According to some embodiments, each component of the energy signal can correspond to the absolute, normalized, or relative intensity of the corresponding characteristic X-ray spectral line. Each possibility corresponds to a separate embodiment. According to some embodiments, each component of the energy signal can correspond to the intensity of the corresponding characteristic X-ray spectral line normalized by the mean background intensity with respect to the characteristic X-ray spectral line. The various ways by which the energy signal can be derived therefrom are described below in Figures 4A to 4E the description of
[0111] According to some embodiments, and as extended below, for example in Figures 4A to 4EIn the description, for each of at least one target substance and for each e-beam landing energy, in sub-operation 110b, an X-ray emission spectrum is measured in a photon energy range that includes the characteristic X-ray spectral lines of the target substance. According to some embodiments, sub-operation 110b can be implemented using an energy-dispersive X-ray (EDX) spectrometer and / or a wavelength-dispersive X-ray (WDX) spectrometer. According to some embodiments, and as elaborated in detail below in the description of sub-operation 120a, in order to derive A corresponding curve is fitted to each of the X-ray emission spectra.
[0112] According to some embodiments, the photon energy range of the measured X-ray emission spectrum can be narrow in the sense of being limited to a region adjacent to (e.g., about three times, about five times, or even about ten times the width) or in the immediate vicinity of the characteristic X-ray spectral lines of the target substance. Related non-limiting examples include embodiments in which a WDX spectrometer is used to obtain the measured spectrum. Alternatively, according to some embodiments, and as described in more detail below, an X-ray detector and an optical filter can be employed to measure the intensity of the emitted X-rays at or near the characteristic X-ray spectral lines of the target substance.
[0113] For ease of description, reference is also made to Figures 2A to 2D . Figures 2A to 2D Schematically depicts an implementation of the operation according to some embodiments of operation 110 of method 100. Figure 2A A cross-sectional view of a sample 20 probed by an e-beam according to measurement operation 110 is shown. To make the description more specific, it is assumed that the sample 20 includes a plurality of transverse (i.e., horizontal) layers 22, where at least some of the layers 22 are different from each other in terms of material composition (e.g., the concentration of one or more of the target substances). According to some embodiments, at least some of the layers 22 can be different from each other in thickness.
[0114] As a non-limiting example, in Figures 2A to 2D , the sample 20 is shown as including three layers arranged one above the other: a first layer 22′, a second layer 22″, and a third layer 22″′. The first layer 22′ is disposed above the second layer 22″. The second layer 22″ is sandwiched between the first layer 22′ and the third layer 22″′. The top surface of the first layer 22′ constitutes the outer surface 24 of the sample 20. Also shown is an e-beam source 202 and the resulting e-beam 205 so as to impinge (e.g., normally impinge) on the outer surface 24. The e-beam source 202 can be configured to project the e-beam (one at a time) at each of a plurality of e-beam landing energies, thereby implementing sub-operation 110a.
[0115] The greater the landing energy of the e-beam 205, the greater the depth to which electrons from the e-beam 205 will penetrate the sample 20 (on average). Additionally, the greater the landing energy of the e-beam 205, the larger the volume within the sample where electrons from the e-beam 205 interact with the material in the sample 20, thereby causing the emission of characteristic X-rays. This is illustrated in Figure 2A by way of three detection regions 26: The first detected region 26a corresponds to a volume in which, due to the e-beam penetrating the sample 20 at a first e-beam landing energy E1, substantially all (e.g., at least 80%, at least 90%, or at least 95%) of the characteristic X-ray (i.e., electromagnetic X-ray radiation) emission interactions will occur. The second detected region 26b corresponds to a volume in which, due to the e-beam penetrating the sample 20 at a second e-beam landing energy E2, substantially all of the characteristic X-ray emission interactions will occur. The third detection region 26c corresponds to a volume in which, due to the e-beam penetrating the sample 20 at a third e-beam landing energy E3, substantially all of the characteristic X-ray emission interactions will occur. The first detected region 26a is centered at a first point P1 at a depth u1, the second detected region 26b is centered at a second point P2 at a depth u2, and the third detection region 26c is centered at a third point P3 at a depth u3. E1 < E2 < E3. Accordingly, u1 < u2 < u3. According to some embodiments, as Figure 2A shown, the size of the third detected region 26c is greater than the size of the second detected region 26b, and the size of the second detected region is greater than the size of the first detected region 26a.
[0116] Figure 2B illustrates a first e-beam 205a generated by an e-beam source 202 and having a first e-beam landing energy E1 incident on the sample 20. The first detected region 26a (where substantially all of the characteristic X-ray emission interactions caused by the first e-beam 205a occur) is also labeled. X-rays can be emitted in all directions, as illustrated by X-rays 215a. X-rays 215a' represent the X-rays (from X-rays 215a) that reach the X-ray detector 204 (such as Figure 5 the X-ray detector).
[0117] Figure 2C illustrates a second e-beam 205b generated by an e-beam source 202 and having a second e-beam landing energy E2 incident on the sample 20. The second detected region 26b (where substantially all of the characteristic X-ray emission interactions caused by the second e-beam 205b occur) is also labeled. X-rays can be emitted in all directions, as shown by X-rays 215b. X-rays 215b' represent the X-rays (from X-rays 215b) that reach the X-ray detector 204.
[0118] Figure 2D Shows a third electron beam 205c generated by an e-beam source 202 and having a third e-beam landing energy E3 incident on a sample 20. Also marked is a third detected region 26c (where substantially all characteristic X-ray emission interactions caused by the third e-beam 205c occur). X-rays can be emitted in all directions, as shown by X-rays 215c. X-rays 215c' represent the X-rays (from the emitted X-rays 215c) that reach the X-ray detector 204.
[0119] Although in Figures 2B to 2D the layers 22 are depicted as being different from each other in their respective refractive indices (as evidenced by the refraction of X-rays as they transition from one layer to another), it should be understood that method 100 is equally applicable without such a difference.
[0120] For each of the e-beam landing energies (e.g., e-beam landing energies E1, E2, and E3), corresponding measurement data of the emitted X-rays can be obtained by the X-ray detector 204, thus enabling sub-operation 110b. In particular, for each of the e-beam landing energies, the corresponding measurement data can include a corresponding X-ray emission spectrum in a photon energy range (the X-ray detector 204 is configured to measure the corresponding X-ray emission spectrum), the photon energy range including at least one characteristic X-ray spectral line related to a corresponding target substance of at least one target substance. The intensity of the characteristic X-ray spectral line corresponding to the target substance indicates the average concentration (i.e., particle density) of the target substance in the corresponding detected region. More specifically, each substance is characterized by a unique set of characteristic X-ray spectral lines (i.e., spectral lines within the characteristic X-ray boundaries) corresponding to the energy differences between the orbits of the elements that make up the substance. The greater the concentration of the substance, the greater the measured intensity of each characteristic X-ray spectral line related to the concentration of the substance.
[0121] In the next step, (illustrated in sub-operation 120b of Figure 1 ), based on the values of the above key features and the values of the key features derived from the reference data (having corresponding structural feature values ), as explained in detail below, the value of the structural parameter of the test sample is estimated The reference data includes the emission profile of the measured reference sample (with the actually measured structural parameters). The reference data is obtained by actual measurement (as specified in operation 110).
[0122] Generally, the process included in sub-operation 120b can be performed by applying an estimator to the values of the key features of the test sample (extracted in sub-operation 120a) As used herein, the term "estimator" refers to a means for based on and to obtain values of the structural parameters of a sample under test. The algorithm may include minimizing a loss function of key features and a vector-valued function where the vector-valued function is extrapolated from According to some embodiments, the loss function may include a term indicating the distance between and According to some embodiments, the algorithm may include a k-NN algorithm. According to some embodiments, the algorithm may include minimizing the loss function after applying the k-NN algorithm.
[0123] The term "extrapolate" is used herein in an extended sense and generally refers to deriving a continuous function from a plurality of data points.
[0124] In some embodiments, a vector of values of key features including reference data, and optionally an estimator or extrapolation, is obtained before measuring the sample under test ("offline"). However, it is contemplated that the vector, estimator, and / or extrapolation may be obtained after measuring the sample under test.
[0125] The term "ground truth (GT)" or "GT" when used herein with respect to a sample refers to an actual or "true" sample (as opposed to a simulated sample). When using these terms with respect to data, measurements, vectors, etc., they refer to data, measurements, vectors, etc. derived from a GT sample (i.e., an actual sample) by actual (i.e., non-simulated) measurements. In the present application, the terms "GT" and "reference" may be used interchangeably with respect to a sample or data retrieved from a sample.
[0126] The reference data may include values of structural parameters obtained by actual (i.e., non-simulated) measurements, possibly (but not necessarily) involving the destruction of the corresponding sample. Non-limiting examples of such physical measurements include time-of-flight secondary ion mass spectrometry (ToF-SIMS) and transmission electron microscopy energy-dispersive X-ray (TEM-EDX).
[0127] In addition, the reference data may include measurements, where the measurements refer to actual X-ray measurements, such as measurements of X-ray emissions of a reference sample or background radiation detected during an actual experiment.
[0128] In some embodiments, the reference sample has values of structural parameters that are close to nominal values or a gold standard. For example, it may have values of structural parameters similar or identical to those of the intended design. In some embodiments, the reference sample includes a sample having the same or similar intended design as the sample under test.
[0129] In some embodiments, the reference sample includes a specially prepared sample that exhibits selected variations relative to the intended design.
[0130] In some embodiments, characteristic X-rays for each reference sample are obtained by projecting an e-beam onto the reference sample and obtaining measurement data by measuring the intensity of X-rays emitted from the reference sample due to the penetration of the e-beam, as described herein for the sample under test.
[0131] In some embodiments, under conditions that are the same or similar to those used for the sample under test, such as by using the same or similar equipment in the same or similar settings, characteristic X-rays for each reference sample are measured.
[0132] Thus, in some embodiments, the method further includes a measurement operation for each of the reference samples, the measurement operation including performing for each of the e-beam landing energies: a sub-operation in which the e-beam is projected onto the reference sample; and a sub-operation in which measurement data is obtained by measuring the intensity of X-rays emitted from the reference sample due to the penetration of the e-beam. In some embodiments, the measurement operation for the reference sample is performed before operation 110. In some embodiments, the measurement operation for the reference sample is performed after operation 110.
[0133] In some embodiments, the measured X-ray intensity of the reference sample includes a background radiation intensity similar to that measured for the sample under test.
[0134] In some embodiments, the e-beam landing energy used to project onto the reference sample includes the e-beam landing energy used to project onto the sample under test. In some embodiments, the e-beam landing energy used to project onto the reference sample is the same as the e-beam landing energy used to project onto the sample under test.
[0135] In some embodiments, the method of the present invention further includes obtaining structural parameters for the reference sample. As described above, the structural parameters can be obtained by various destructive or non-destructive methods (including, by way of non-limiting example, ToF-SIMS and TEM-EDX).
[0136] It should be noted that when performed by a method that may be destructive to the reference sample (i.e., cause any change to the sample), the structural parameters must be obtained after the measurement operation.
[0137] In some embodiments, as described herein for a test sample, values of key features are extracted from measurement data for a reference sample. In some embodiments, the key features are the same as the key features for the test sample. In some embodiments, the values of the key features are extracted from the measurement data for the reference sample by the same method as the method used to extract the values of the key features from the measurement data for the test sample.
[0138] Thus, in some embodiments, the method further includes a data analysis operation for each of the reference samples, wherein values of key features specified by a vector are extracted from the reference measurement data. In some embodiments, the reference data analysis operation is performed before operation 120a. In some embodiments, the reference data analysis operation is performed after operation 120a.
[0139] According to some embodiments, the reference samples may be selected to reflect the expected variation (e.g., due to manufacturing defects) in the values of one or more structural parameters among samples (having the same intended design). In some embodiments, the values of the structural parameters are used to "sample" a selected hypervolume centered in a K-dimensional vector space defined by the one or more structural parameters. p In the vector space. The nominal values of one or more structural parameters are specified, and K p is the number of structural parameters. In some embodiments, K p is about 2 to 6 or about 3 to 5 structural parameters. The size and boundaries of the selected hypervolume may be chosen to include the expected variation in the one or more structural parameters. According to some embodiments, the number N2 of reference samples is greater than 2K p (or at least strictly greater than K p ). In some embodiments, the number N2 of reference samples is about 3K p to 7K p or about 5K p to 7K p . As a non-limiting example, according to some embodiments, may be selected such that, with respect to the deviation from the nominal value , includes each vector in a set of 2K p vectors. For each 1 ≤ k ≤ K , p , is approximately equal to σ k is the expected standard deviation of the k-th structural parameter, and is a vector pointing to the K p -dimensional vector space defined by the K pThe unit vector of the k-th axis line of the vector space.
[0140] In some embodiments, all of those associated with the reference sample are within the selected hypervolume. In some embodiments, at least one of those associated with the reference sample is within the selected hypervolume. In some embodiments, at least one of those in the reference sample is not within the selected hypervolume.
[0141] Note that in embodiments where one or more of the structural parameters consist of a single structural parameter, and each of which is a one-dimensional vector (i.e., a scalar). Non-limiting examples include embodiments where only the thickness of a single layer or the total mass density of a single target substance is determined.
[0142] Note that any particular implementation of a model or algorithm, such as those used in the present invention, can be applied to the complete vector set or to a partial vector set
[0143] According to some embodiments, and as extended in the following description in Figure 3 in sub-operation 120b, the minimum distance between and the vector-valued function extrapolated from can be calculated, thereby obtaining Each component of the vector-valued function quantifies the dependence of the corresponding (extrapolated) key feature on one or more structural parameters (values). One or more structural parameters are parameterized, where is a vector of free parameters (i.e., variables corresponding to the structural parameters respectively). As described below, according to some embodiments, the minimum distance can be obtained by minimizing , where each component of varies within the corresponding continuous value range.
[0144] According to some embodiments, and as described in detail in the following description in Figure 3 , minimize the loss function of the key feature and the vector-valued function where the vector-valued function is extrapolated from According to some embodiments, in addition to depending on and In addition to the first term for both, the loss function may further include (one or more) regularization terms, e.g., in order to stabilize the solution or as (one or more) constraints reflecting some prior knowledge about one or more structural parameters. According to some embodiments, the first term corresponds to and the mathematical distance between. According to some embodiments, standard local and / or global optimization tools may be used to minimize the loss function and thereby estimate Alternatively, according to some embodiments where the optimization problem defined by the minimization on the loss function allows a known analytical solution, may be directly computed from the analytical solution (the function defining the analytical solution)
[0145] According to some embodiments, and as further extended below, is estimated by computing the distance from to
[0146] Alternatively, according to some embodiments, may be obtained by applying a k-nearest neighbor (k-NN) regression algorithm (k < N) with respect to towards According to some embodiments, includes (e.g., a vector specifying the nominal values of one or more structural parameters). Note that the k-NN regression algorithm may be weighted or unweighted. That is, may be regarded as equal to the mean or weighted mean of the corresponding to the k closest .
[0147] In some embodiments, when weighting the k-NN algorithm, the weights are determined through multiple iterations of the k-NN algorithm, each time excluding a different reference sample from the calculation and using the excluded reference sample as the tested sample, attempting to predict the value of the structural parameter of the excluded GT sample. The weights selected for the closest prediction are used for the estimator.
[0148] In some embodiments, the k-NN algorithm is applied to the full reference dataset. In some embodiments, the k-NN algorithm is applied to a partial reference dataset.
[0149] Alternatively, may be regarded as equal to the median of the k closest corresponding to the closest . Generally, that is, when one or more structural parameters include two or more structural parameters, the term "median" should be understood to refer to a multivariate extension of the median (a one-dimensional concept), such as the marginal median or the geometric median. More generally, can be substantially closest to k corresponding to any function. In some embodiments, the median is a weighted median, assigning different weights to different data points.
[0150] Still according to some other embodiments, is estimated as the output of a neural network. The neural network is configured to receive a vector of key features (i.e., ) as input and output The neural network is trained using a training set including N pairs of vectors such that for each 1 ≤ n ≤ N, is used as the (training) input, and is used as the corresponding (training) output. In some embodiments, the neural network algorithm is configured to process different data points differently, for example, by using non-uniform weights in a loss function.
[0151] Figure 3 Presents a flowchart of a measurement data analysis operation 300, which corresponds to a specific embodiment of the measurement data analysis operation 120 of method 100. The measurement data analysis operation 300 includes:
[0152] - Sub-operation 310, where key features specified by a vector are extracted from the measurement data.
[0153] - Sub-operation 320, where is obtained as the (numerical or analytical) solution that minimizes a loss function that depends at least on and . (That is, minimizes the loss function.) is a vector-valued function of the key features obtained by extrapolating from .
[0154] Sub-operations 310 and 320 respectively correspond to specific embodiments of sub-operations 120a and 120b of method 100.
[0155] Each of the components of is a function that quantifies the dependence (as specified by the model) of the corresponding key feature on the value of one or more structural parameters. Thus, for example, the j-th component of is a function that quantifies the dependence of the j-th key feature (e.g., the j-th component of an energy signal) on
[0156] According to some embodiments, in sub-operation 320, Indicates a pair of vectors and The mathematical distance between (the mathematical distance may be a norm or not) (thus yes and Non-limiting examples of distances include L 1 Norm and L 2 norm. (Note that in this case the norm is L 2 ,and exist In some such embodiments, the optimization problem allows for analytical solutions. Double vertical bars indicate vector norms (e.g., the Euclidian norm: More generally, and as described in detail below, the To estimate Where M1 and M2 are matrices with appropriately chosen properties as specified below. In particular, each of M1 and M2 may be a positive definite matrix (optionally diagonal) having a respective minimum eigenvalue greater than a respective predetermined (positive) threshold.
[0157] According to some embodiments, regularization term(s) may be added to the norm (Or more generally, or the first term (depending on the loss function and )) to stabilize the solution or as (multiple) constraints reflecting some prior knowledge about one or more structural parameters.
[0158] According to some embodiments, the extrapolation is a linear function. That is, in is with A is a vector of K corresponding to the values of the key features. f ×K p Matrix. K f is the number of key features, that is, (as well as and The dimension of each of K p yes (and ), i.e., the number of one or more structural parameters of the tested sample (and each of the reference samples) to be estimated. Specified with (i.e., by the nominal value The value of the key feature corresponding to the nominal sample (characterized).
[0159] According to some embodiments, including the GT value obtained by actual measurement, in the same manner as described above for the test sample or reference sample.
[0160] According to some embodiments,
[0161]
[0162] for each 1 ≤ n ≤ N, B is a K f × K p matrix. The double vertical bars indicate the matrix norm (e.g., the Frobenius norm). Optionally, according to some embodiments, (one or more) regularization terms may be added to the matrix norm, in which case the minimization is understood to be performed on the sum of the matrix norm and the (one or more) regularization terms.
[0163] Alternatively, according to some embodiments, similar to matrix A, it is determined by optimization In particular, according to some such embodiments, and both matrix A are obtained as the solution of, where the double vertical bars indicate the matrix norm (e.g., the Frobenius norm). Optionally, according to some embodiments, (one or more) regularization terms may be added to the matrix norm, in which case the minimization is understood to be performed on the sum of the matrix norm and the (one or more) regularization terms.
[0164] In some embodiments, not all of the in the matrix are treated in the same manner. Some non - limiting exemplary embodiments are provided below.
[0165] In some embodiments, where is a predefined function class and where K f is the number of key features, i.e., the dimension of f, and is a general loss function that may or may not depend on i. In some embodiments, all data are treated in the same manner by making D i the same for all i. However, in some embodiments, certain samples are treated in a different manner from other samples by using different D i for different samples or for different groups of samples.
[0166] including non - limiting examples of D that include the L 1 norm and the L 2 norm were discussed above with reference to sub - operation 320, and where and is a pair of vectors (such that is the (mathematical) distance between and M1 and M2 are matrices with appropriately chosen properties as specified in the discussion of sub-operation 320 above. As mentioned above, different D i can be used for different samples or different groups of samples.
[0167] A further non-limiting example is a function constituted by a linear function (an "affine function" which may include a constant term). Here, is of the form (or for example ) of a class of functions. In this case, the optimization becomes and the resulting is
[0168] According to some such embodiments similar to the embodiments mentioned above with respect to sub-operation 320, the function is a linear function with a Frobenius / L2 norm, where D i (v1, v2) = α i ‖v1 - v2‖ 2 . Thus, matrix A is equal to or obtain and both matrix A as 's solution. As mentioned above, different D i are used for different samples or different groups of samples, and here for all data, α i can be the same (resulting in equal treatment for all data), or different for different data. Optionally, according to some embodiments, (a) regularization term(s) can be added to the matrix norm, in which case the minimization is understood to be performed on the sum of the matrix norm and (a) regularization term(s).
[0169] According to some embodiments, the extrapolation can be a non-linear function of. As a non-limiting example, according to some such embodiments, is 's square function. That is, in this embodiment, for each 1 ≤ c ≤ K f , 's c-th component will include (in addition to the linear contribution and the constant) a square contribution given by , where indicates the (a, b)-th component of the K p × K p matrix T (c) .
[0170] Standard local and / or global optimization algorithms can be used to solve the optimization problems specified throughout the present application, such as gradient descent or quasi-Newton methods. According to some embodiments, where the specified optimization problem allows a known analytical solution, the quantity being optimized (e.g., matrix A, ) can be calculated directly from the analytical solution (the function defining the analytical solution). As a non-limiting example, in an embodiment where is linear in (i.e., where is a matrix) and the norm is L 2 , the optimization problem assumes the form such that where the matrix is the Moore-Penrose inverse matrix of . As another example, in an embodiment where allows a (known) analytical solution, such as when the matrix norm is the Frobenius norm, A can be obtained by inserting and into the analytical solution (the function defining the analytical solution). More precisely, the analytical solution is given by A = (F - F0)Q, where Q is the matrix whose columns are composed of and is the Moore-Penrose inverse matrix of , F is the matrix whose columns are composed of , and F0 is the matrix whose columns are each composed of . Finally, in an embodiment where allows a (known) analytical solution, such as when the matrix norm is the Frobenius norm, and can be inserted into the analytical solution (the function defining the analytical solution) to obtain and A. Through appropriate manipulation, the analytical solution can be obtained in substantially the same manner as in the case of determining only A (i.e., when is given).
[0171] Optionally, according to some embodiments, the measurement data analysis operation 300 may further include an (optional) sub-operation (not specified in Figure 3 ) before the sub-operation 320: obtaining by subjecting to a (k = N)-NN classifier with respect to a larger set of N' > N vectors of key features corresponding to N' > N additional reference samples The set of N' > N vectors of key features includes (optionally, relabeled) The N' can be selected as described above in the description of method 100 Complete set.
[0172] Referring again to method 100, according to some embodiments, in sub-operation 120a (and thus also sub-operation 310), in order to derive Fit a corresponding curve to each of the X-ray emission spectra (obtained for each of the e-beams projected in measurement operation 110). According to some embodiments, this is done by Figures 4A to 4E Illustrated by way of example in, in which case, where the energy signal is given by an energy signal associated with a single target substance and a single spectral line (i.e., a single characteristic X-ray spectral line) The more general case will be described below, where the energy signal is associated with multiple target substances, and / or for at least some of the target substances, multiple spectral lines of the target substance are considered.
[0173] In some embodiments, by the same method described herein for deriving Fit a corresponding curve to each of the X-ray emission spectra obtained for the reference sample used for deriving
[0174] Refer to Figure 4A , Figure 4A Depicts the measured (X-ray emission) spectrum 400, which is obtained by implementing measurement operation 110 for the test sample (e.g., sample 20). Also in Figures 4B to 4E In each case, the horizontal axis corresponds to the photon energy ε (or equivalently, frequency) of the emitted X-rays, and the vertical axis corresponds to the intensity I of the emitted X-rays. The scale on each of the horizontal and vertical axes is linearly spaced, where ε i < ε i+1 , and I i < I i+1 . The peak 410 of the measured spectrum 400 is substantially centered on the characteristic X-ray spectral line of the target substance, which is included in the test sample, and an energy signal of the target substance will be obtained. Figure 4B Depicts the optimized curve 450 fitted to the measured spectrum 400. Figure 4C Depicts the optimized curve 450 superimposed on the measured spectrum 400.
[0175] According to some embodiments, fitting to the measured spectrum 400 involves values of one or more adjustable parameters of the optimized curve (also referred to as the "free curve") to obtain the optimized curve 450. The values of the one or more adjustable parameters are fixed by minimizing the distance between the free curve and the measured spectrum (for the one or more adjustable parameters).
[0176] One or more adjustable parameters may include a (first) adjustable parameter, the value of the (first) adjustable parameter indicating the intensity of X-rays emitted for a characteristic X-ray line of a target substance. According to some such embodiments, the adjustable parameter is a multiplicative coefficient of a normalized cap function (e.g., a normalized Gaussian), the multiplicative coefficient being centered on the characteristic X-ray line. According to some embodiments, the one or more adjustable parameters include a plurality of adjustable parameters, which, in addition to the first adjustable parameter, may include an additive bias parameter, at least one parameter controlling the shape of the cap function (e.g., the width of a normalized Gaussian), and / or a (characteristic X-ray) spectral shift parameter controlling the center position of the cap function.
[0177] More generally, according to some embodiments, the free curve may be a sum of at least two adjustable functions: an adjustable cap function centered on the characteristic X-ray line, and an adjustable second function that quantifies the (continuous) spectrum of the bremsstrahlung (i.e., background radiation) component of the corresponding measured X-ray emission spectrum (e.g., the background radiation near the characteristic X-ray line). As a non-limiting example, at least one landing energy includes N E e-beam landing energies such that N E X-ray emission spectra are measured: Here, ε indicates the photon energy of the emitted X-rays, and s i (ε), the i-th measured X-ray emission spectrum, is the measured X-ray emission spectrum caused by projecting an e-beam at a landing energy E i . According to some embodiments, a set of N E free curves may be fitted to the set of measured spectra . According to some embodiments, for each 1 ≤ i ≤ N E , c i (ε) = G i (ε) + b i (ε), where G i (ε) is an adjustable cap function, and b i (ε) is an adjustable second function. G i (ε) = a i ·g i (ε), where g i (ε) is a normalized cap function, and a i is a multiplicative coefficient. According to some embodiments, g i (ε) may be a (normalized) Gaussian, in which case the width and optionally the center of g i (ε) may be adjustable parameters (optimized for the adjustable parameters). According to some alternative embodiments, gi (ε) can be a (normalized) gamma distribution or a generalized Gaussian distribution. According to some embodiments, b i (ε) can be a polynomial with adjustable coefficients (e.g., a first-order polynomial or a second-order polynomial). Alternatively, according to some embodiments, b can be determined from Kramer's rule i (ε).
[0178] Since g i (ε) is normalized, a i is substantially equal to the intensity (or equivalently, the number of photons) of the X-rays emitted due to the transition corresponding to the characteristic X-ray spectrum of the target substance and collected (detected) by the X-ray measurement module.
[0179] Are respectively represented by and the adjustable parameters of g i (ε) and b i (ε). For each 1 ≤ i ≤ N E , the optimized values of the adjustable parameters i (ε), s i (ε)) for a i , and can be obtained by minimizing D(c and D(c i (ε), s i (ε)) is the distance between c i (ε) and s i (ε). More generally, according to some embodiments, the optimized values can be obtained by minimizing a loss function that depends at least on c i (ε) and s i (ε). As a non-limiting example, according to some embodiments, where g i (ε) is Gaussian and b i (ε) is a second-order polynomial: where g i,1 and g i,2 parameterize the width and center of the Gaussian; and where b i,0 , b i,1 and b i,2 are the zero-order, first-order, and second-order coefficients of the polynomial. In particular, according to some embodiments, D(c i (ε), s i (ε)) = ∫dε|c i (ε) - s i (ε)| 2(or a discrete equivalent expression). According to some embodiments, a regularization term can be added to D(c i (ε), s i (ε)) to account for the existing knowledge about any free parameters and / or to stabilize the solution (of the minimization algorithm).
[0180] According to some alternative embodiments, where there is existing knowledge that relates at least some of the free parameters to each other, by jointly optimizing all the adjustable parameters, i.e., subject to the constraints imposed by the aforementioned existing knowledge, to obtain a complete set of optimized values, i.e., (or equivalently where and indicate optimization functions defined by and respectively). More specifically, in such embodiments,
[0181]
[0182] is subject to constraints,
[0183] where is a set of N c constraints. (That is, each of the Q l is an equation or inequality that relates at least some of the free parameters to each other.)
[0184] As a non-limiting example, according to some embodiments depicted in Figures 4B to 4E , the free curve is the sum of three adjustable functions. In addition to g i (ε) which is a Gaussian and b i (ε) which is a second-order polynomial, the sum additionally includes Gaussian Y i (ε). Referring to Figure 4D , curve 460 corresponds to . (which is also a Gaussian) is obtained by optimizing the free parameters for Y i (ε). Centered on the characteristic X-ray spectrum of the target substance. Centered on the characteristic X-ray spectrum of a second (non-target) substance present in the test sample. The characteristic spectrum of the second substance is close to the characteristic spectrum of the target substance, and thus the characteristic spectrum of the second substance is considered to improve the determination of (and thus ) accuracy. Referring to Figure 4E , curve 470 corresponds to . Curve 470 is also plotted in Figure 4C .
[0185] According to some embodiments, an X-ray emission spectrum of a single characteristic X-ray spectral line for a single target substance (included in a sample to be tested) is used to determine The number of components of is equal to the number of e-beam landing energies. According to some embodiments, for each 1 ≤ j ≤ J, is equal to —— the j-th component of the energy signal. More generally, according to some embodiments, for each 1 ≤ j ≤ J, where is and a function of. That is, for each 1 ≤ j ≤ J, the j-th component of the energy signal is and a function of both coefficients of. According to some such embodiments, where q is a function of the coefficient of. As a non-limiting example, according to some embodiments, and where the triangular brackets indicate averaging over an interval equal to the width of centered on .
[0186] According to some embodiments, key features can be derived based on the dependence of the intensity of the X-rays emitted with respect to each of a plurality of different characteristic X-ray spectral lines on the e-beam landing energy. According to some such embodiments, where N L is the number of different characteristic X-ray spectral lines, the key feature is specified by a J = N × N E × N L component vector with components E where 1 ≤ n E ≤ N L ≤ N L (N E is the number of landing energies). The first index indicates the e-beam landing energy, and the second index indicates the characteristic X-ray spectral line. That is, In such an embodiment, in measurement operation 110, for each e-beam landing energy, the X-ray emission spectrum is measured over one photon energy range or a plurality of photon energy ranges including a plurality of characteristic X-ray spectral lines. The components of related to the same characteristic X-ray spectral line (e.g., ) can be obtained in the situation as described above, where N L = 1. According to some embodiments, where at least one target substance includes N sub (N sub ≤ N L ) target substances, N LThe characteristic X-ray spectrum includes characteristic X-ray spectra corresponding to each of N sub target substances, respectively.
[0187] As described above, according to some embodiments, M1 and M2 are matrices with appropriately selected properties (e.g., positive definiteness and symmetry as specified below), and denotes the mathematical distance between and To make the description more specific by way of non-limiting examples, an embodiment is elaborated in detail in which the X-ray emission spectrum of a single characteristic X-ray spectrum of a single target substance is used to determine such that for each 1 ≤ j ≤ N E , is equal to and (i.e., the number of components of E ) is 2N. That is, the first N E components of E are non-normalized energy signal components, and the last N components are bremsstrahlung (i.e., background radiation) components. According to some such embodiments, such that M1 is equal to the identity matrix and M2 = M. M is a diagonal matrix, and the diagonal terms of the diagonal matrix are pairwise equal in the sense that for each 1 ≤ j ≤ N E , where T is a predetermined (positive) threshold. That is, for each 1 ≤ j ≤ N E , the (N E +j)-th component along the diagonal of M is equal to the j-th component along the diagonal. Thus, for each 1 ≤ j ≤ N E , and are weighted by the same corresponding factor. The inclusion of M and the minimization over M can account for potentially different scalings of the corresponding components of and E , and such that for at least some 1 ≤ j ≤ N and the scaling of According to some embodiments, one or more regularization terms can be added to More generally, according to some embodiments, where is a loss function that depends on and and the loss function is equal to (where M has the properties and symmetries specified above, i.e., pairwise equalities). is dependent on and a function of (e.g., the loss function). According to some embodiments, wherein is a regularization function including one or more regularization terms.
[0188] According to some embodiments, wherein a spectrometer is used to obtain an X-ray emission spectrum, the measurement data analysis operation 120 may include an initial preprocessing sub-operation, wherein the X-ray emission spectrum may be preprocessed to remove noise. According to some embodiments, a similar preprocessing sub-operation for a reference sample may be included, wherein the X-ray emission spectrum is preprocessed to remove noise.
[0189] According to some embodiments of method 100, sub-operation 120a includes an initial sub-operation, wherein a difference spectrum for the test sample is obtained by subtracting the reference spectrum of a control sample from the measured spectrum of the test sample obtained in measurement operation 110. The reference spectrum may be that of a gold standard sample, and it is known or assumed that the gold standard sample closely matches the intended design of the test sample in that the deviation of one or more structural parameters (optionally, also structural parameters not estimated by method 100) from the nominal values of the one or more structural parameters does not exceed 1%, 2%, or 5%. Each possibility corresponds to a separate embodiment. In some embodiments, the spectrum of the control sample includes a background radiation intensity similar to the intensity measured for the test sample.
[0190] More specifically, (in sub-operation 120a) the difference spectrum may be processed to extract a vector of key features for the test sample from the difference spectrum substantially as described above in Figures 4A to 4E wherein the difference is that any consideration of bremsstrahlung is eliminated (due to the subtraction).
[0191] In such an embodiment, the method further includes an operation in which a difference spectrum for the reference sample is also obtained by subtracting the reference spectrum of the control sample from the measured spectrum of the reference sample, wherein the measured spectrum of the reference sample is obtained by the same method as described herein for the test sample.
[0192] In some embodiments, instead of obtaining the difference spectra of the test sample and the reference sample by subtracting the spectrum of the control sample (as described above), certain key features of the test sample and the reference sample are obtained by subtracting the corresponding key features obtained for the control sample from the initially obtained key features.
[0193] In this embodiment, the computer algorithm of sub-operation 120b is configured to receive as input and output for each 1 ≤ n ≤ N, in the same manner as obtained for the test sample in the same manner as to obtain including obtaining a differential spectrum by subtracting a reference spectrum from the GT measurement spectrum and processing it to extract a vector of key features
[0194] According to some such embodiments, method 100 may additionally include obtaining a reference spectrum, for example, by performing measurement operation 110 relative to a reference sample.
[0195] It should be understood that the applicability of method 100 is not limited to samples including nominally flat layers (as shown by the non-limiting examples in Figures 2A to 2D ), and more generally, layered samples. Regions with different material compositions can in principle be of any shape. In particular, method 100 can be applied to structures including locally embedded (buried) features, such as nanowires, gate-all-around nanosheets, and more generally, channels. Method 100 can also be applied to samples characterized by a continuously varying density of the substances included in the sample, which varies with the depth coordinate and / or, in three dimensions, with the lateral coordinate. In addition, those skilled in the art will readily realize that method 100 can be applied to samples including cavities and / or holes.
[0196] System
[0197] According to one aspect of some embodiments, a computerized system is provided for non-destructive three-dimensional probing and characterization of a sample (such as, for example, a semiconductor structure included in a patterned wafer) based on X-ray measurements and subsequent analysis of the measurement data obtained using actual measurements of a GT reference sample. Figure 5 Such a system according to some embodiments is schematically depicted, namely computerized system 500. As will be apparent from the description of system 500, system 500 can be used to implement method 100 (including specific embodiments of method 100, which includes measurement data analysis operation 300). In particular, system 500 can be used to estimate the values of one or more structural parameters characterizing the test sample. Non-limiting examples of the structural parameters that can be estimated using system 500 are listed above in the method subsection of the description of method 100.
[0198] System 500 includes an e-beam source 502, an X-ray detector 504 (or more generally, an X-ray sensing assembly including two or more X-ray detectors), processing circuitry 506, and a controller 508. According to some embodiments, system 500 may further include a stage 520 (e.g., an xyz stage) configured to receive a (tested) sample 50. According to some embodiments, the e-beam source 502, the X-ray detector 504, and the controller 508 form part of a scanning electron microscope. According to some embodiments, the sample 50 may be a patterned wafer or a structure (e.g., a semiconductor structure) included in or on a patterned wafer. According to some such embodiments, the sample 50 may be a preliminary structure in one of the manufacturing stages of a patterned wafer, or an auxiliary structure employed in one of the manufacturing stages of a patterned wafer. According to some embodiments, the sample 50 may be or include one or more memory components and / or logic components (such as a gate stack, e.g., a high-k metal gate stack). Note that the sample 50 does not form part of system 500.
[0199] The dashed lines between the elements indicate a functional or communication association between them.
[0200] The e-beam source 502 is configured to generate e-beams at a plurality of e-beam landing energies. In particular, substantially as described above in the description of sub-operation 110a of method 100, the e-beam source 502 is configured to generate e-beams at each of the plurality of landing energies so as to allow probing of the sample 50 to a plurality of depths, respectively.
[0201] The greater the depth of the tested sample to be probed, the greater the maximum e-beam landing energy, and optionally, the greater the number of e-beam landing energies. According to some embodiments, the plurality of e-beam landing energies may include landing energies up to about 5 keV, about 10 keV, about 15 keV, about 20 keV, or even about 30 keV. Each possibility corresponds to a different embodiment. In silicon, an e-beam with a landing energy of about 15 keV can penetrate to a depth of about 3 μm.
[0202] The duration of the projection of the e-beam may be determined by the (i.e., the value of one or more structural parameters characterizing the sample 50) required precision.
[0203] According to some embodiments, an e-beam 505 generated by the e-beam source 502 is shown incident on the sample 50 (on its outer surface 54). As a result of the e-beam 505 impinging on the sample 50 and the e-beam 505 penetrating the sample 50, X-rays, particularly characteristic X-rays, are generated. A portion of these X-rays constituted by the X-rays 515 reaches the X-ray detector 504.
[0204] According to some embodiments, the X-ray detector 504 is sensitive to electromagnetic radiation in an X-ray photon energy range (at least within the characteristic X-ray limit or one or more sub-ranges of the characteristic X-ray limit). According to some embodiments, the X-ray detector 504 can be an EDX spectrometer or a WDX spectrometer. According to some embodiments, an X-ray detector assembly including both an EDX spectrometer and a WDX spectrometer can be used in place of a single X-ray detector. In such an embodiment, both the EDX spectrometer and the WDX spectrometer can be used to obtain an X-ray emission spectrum, where the WDX spectrometer is used to "magnify" at the characteristic X-ray spectral line. In particular, compared to the EDX spectrometer, the greater resolution of the WDX spectrometer (which makes it slower) allows for narrower peaks and valleys to be obtained. According to some embodiments, where the spectrometer is a WDX spectrometer, the X-ray detector 504 can be configured to allow scanning over an extended photon energy range (thus allowing an X-ray emission spectrum to be obtained over an extended photon energy range). The X-ray detector 504 is configured to relay the measurement data thus collected (e.g., the spectrum of the X-rays incident on the X-ray detector) (optionally, via the controller 508) to the processing circuitry 506.
[0205] According to some embodiments, the system 500 can additionally include a window (not shown) positioned between the X-ray detector 504 and the stage 520, the window being configurable to controllably and differentially attenuate the spectrum of the emitted X-rays and / or protect the X-ray sensitive surface of the spectrometer.
[0206] According to some alternative embodiments, the X-ray detector 504 is configured to measure the intensity of electromagnetic X-ray radiation (i.e., electromagnetic radiation in the X-ray photon energy range) at or near the characteristic X-ray spectral line of the (target) substance included in the sample 50, without additionally measuring the intensity of electromagnetic X-ray radiation over an extended photon energy range outside the immediate vicinity of the characteristic X-ray spectral line. According to some such embodiments, the system 500 can additionally include an optical filter (not shown). The optical filter is configured to block electromagnetic radiation having a photon energy outside the immediate vicinity of the characteristic X-ray spectral line from reaching the X-ray detector 504.
[0207] According to some embodiments, system 500 may include additional elements. The additional elements may include electro-optical devices (not shown; e.g., (a)n electrostatic lens(es) and (a)n magnetic deflector(s)) that may be used to direct and manipulate the e-beam generated by e-beam source 502. Additionally or alternatively, the additional elements may include collection optics configured to direct electromagnetic radiation generated due to the e-beam impinging on sample 50 and the e-beam penetrating the sample onto X-ray detector 504. According to some embodiments, the additional elements may include a filter configured to block electromagnetic radiation outside of a characteristic X-ray boundary and / or one or more sub-ranges of the characteristic X-ray boundary.
[0208] According to some embodiments, at least e-beam source 502 and stage 520 may be housed within vacuum chamber 530. While X-ray detector 504 is shown positioned within vacuum chamber 530 in Figure 5 , according to some alternative embodiments, X-ray detector 504 may be positioned outside of vacuum chamber 530.
[0209] Controller 508 may be functionally associated with e-beam source 502 and optionally with stage 520. More specifically, controller 508 is configured to control and synchronize the operation and function of the instruments and components listed above during the probing of the sample under test (e.g., instructing the e-beam source to change the e-beam landing energy).
[0210] Processing circuitry 506 may include one or more processors and optionally include RAM and / or non-volatile memory components (not shown). The one or more processors are configured to execute software instructions, for example, stored in the non-volatile memory components. By executing the software instructions, the measurement data (e.g., obtained by X-ray detector 504) of the sample under test (e.g., sample 50) is processed to estimate Figure 1 and Figure 3 as described above in
[0211] More specifically, processing circuitry 506 is configured to process the measurement data to estimate the value of one or more structural parameters characterizing the sample under test (represented by ), as detailed above in the method subsection. To this end, as detailed above in the description of sub-operation 120b of method 100, processing circuitry 506 is configured to extract from the measurement data a vector specifying the values of key features obtained for the sample under test In particular, according to some embodiments, the key features include a so-called "energy signal" (e.g., constituted by a so-called "energy signal"). According to some embodiments, each component of the energy signal may correspond to the absolute, normalized, or relative intensity of a corresponding characteristic X-ray spectral line.
[0212] According to some embodiments, where the measurement data is X-ray emission spectra, the processing circuitry 506 may be configured to fit a corresponding (free) curve to each of the X-ray emission spectra, thereby obtaining an optimized curve. As described above in the description of sub-operation 120a of method 100, next, According to some such embodiments, the processing circuitry 506 may be configured to perform one or more optimization algorithms (e.g., to solve the optimization problems specified above in Figures 4A to 4E the description). Examples of relevant optimization algorithms include standard iterative optimization algorithms such as gradient descent or Newton's method. According to some embodiments, a customized iterative optimization algorithm may be employed, obtained by "fine-tuning" a standard iterative optimization algorithm to solve constraints and / or ensure global minimization is achieved.
[0213] More specifically, in order to estimate the processing circuitry 506 is configured to additionally consider a reference vector set including the measured key features as defined above for the method aspect. For each 1 ≤ n ≤ N, specify values of one or more structural parameters characterizing the nth reference sample, and each of which may be obtained by actual measurement in the same manner as for the test sample being obtained and as described above in the method subsection. According to some embodiments, and as extended above in the description of method 100, sample a selected hypervolume centered at p in a K dimensional vector space defined by one or more structural parameters, where K p is the number of one or more structural parameters.
[0214] According to some embodiments, the processing circuitry 506 may be configured to calculate where k (k < N) mark the k closest to and the triangular brackets indicate taking the average, optionally weighted, of the k . To this end, according to some embodiments, the processing circuitry 506 may be configured to apply a k-nearest neighbor (k-NN) regression algorithm with respect to towards . According to some embodiments, the processing circuitry 506 may be configured to obtain by calculating the median of the corresponding to the k closest As explained above with respect to the method aspect, weights can be assigned to k-NN or the median to bias towards the GT data points.
[0215] According to some alternative embodiments, minimize a loss function that is a function of at least a key feature and a vector-valued function where the vector-valued function is extrapolated from (Optionally, in addition to the first term that depends on and , the loss function may further include one or more regularization terms.) Thus, the processing circuitry 506 can be configured to: (i) perform an optimization algorithm to minimize the loss function with respect to (e.g., minimize the distance between and ); and (ii) in embodiments where the minimization of the loss function has a known analytical solution, additionally or alternatively, directly estimate from the analytical solution (the function that defines the analytical solution). According to some such embodiments, the processing circuitry 506 can be configured to estimate by (numerically or analytically) solving an optimization problem or more generally, where M is a positive definite matrix (e.g., a diagonal positive definite matrix with pairwise equal diagonal terms as specified above), the minimum fixed value of the positive definite matrix is greater than a predetermined (positive) threshold, and even more generally, where M1 and M2 are appropriately chosen matrices, and is the mathematical distance between . is a vector-valued function of the key feature (extrapolated from ) and models the dependence of the key feature on the values of one or more structural parameters. As explained above with respect to the method aspect, the vector-valued function the optimization and / or loss function may include different treatments for different data points or for different groups of data points.
[0216] According to some embodiments, the processing circuitry 506 can be further configured to perform extrapolation (and thereby obtain ). According to some such embodiments, the extrapolation can be a linear function, i.e., indicating the deviation from the nominal value of one or more structural parameters . is related to A vector of values of corresponding key features. As detailed above in the description of method 300, A is a matrix that can be determined by optimization, taking into account As explained above with respect to the method aspect, extrapolation can include different processing for different data points or for different groups of data points.
[0217] According to some embodiments, the processing circuitry 506 can be configured to, by subjecting to a (k = N)-NN classifier with respect to select from a larger set of reference samples (N′ > N)
[0218] As used herein, when used in the context of curve fitting, the terms "fitted" and "optimized" may be interchangeable.
[0219] According to some embodiments, for example when implemented by a single computer, the processing circuitry 506 and the controller 508 can be housed in a common housing.
[0220] In the specification and claims of this application, the words "comprising" and "having" and their forms are not limited to the members in the lists that may be associated with these words.
[0221] As used herein, according to some embodiments, the term "about" can be used to specify the value of a quantity or parameter (e.g., the length of an element) as being within a continuous range of values close to (and including) a given (prescribed) value. According to some embodiments, "about" can specify that the value of a parameter is between 80% and 120% of a given value. For example, the statement "the length of the element is about equal to 1 m" is equivalent to the statement "the length of the element is between 0.8 m and 1.2 m". According to some embodiments, "about" can specify that the value of a parameter is between 90% and 110% of a given value. According to some embodiments, "about" can specify that the value of a parameter is between 95% and 105% of a given value.
[0222] As used herein, according to some embodiments, the terms "substantially" and "about" may be interchangeable.
[0223] According to some embodiments, an estimated quantity or an estimated parameter may be said to be "about optimized" or "about optimal" when it falls within 5%, 10%, or even 20% of its optimal value. Each possibility corresponds to a separate embodiment. In particular, the expressions "about optimized" and "about optimal" also cover the case where the estimated quantity or parameter is equal to the optimal value of the quantity or parameter. In principle, mathematical optimization software can be used to obtain the optimal value. Thus, for example, when the estimated value is not greater than 101%, 105%, 110%, or 120% (or some other predefined threshold percentage) of the optimal value of the quantity, the estimate (e.g., the estimated residual) may be referred to as "about minimized" or "about minimum / minimal". Each possibility corresponds to a separate embodiment.
[0224] For ease of illustration, a three-dimensional Cartesian coordinate system (with orthogonal x, y, and z axes) is introduced in some of the figures. Note that the orientation of the coordinate system relative to the depicted object may vary from one figure to another. Additionally, the symbol ⊙ is used to represent an axis pointing "out of the page", while the symbol may be used to represent an axis pointing "into the page".
[0225] In flowcharts, optional operations and sub-operations are depicted by dashed lines. Similarly, in block diagrams, optional elements may be depicted by dashed lines. Further, (in block diagrams) dashed lines connecting elements may be used to represent a functional association or at least a one-way or two-way communication association between the connected elements.
[0226] It should be understood that, for clarity, certain features of the present disclosure described in the context of separate embodiments may also be provided in combination in a single embodiment. Conversely, for brevity, the various features of the present disclosure described in the context of a single embodiment may also be provided separately, or in any suitable sub-combination, or appropriately in any other described embodiment of the present disclosure. The features described in the context of an embodiment are not considered to be essential features of the embodiment unless expressly so specified.
[0227] Although the operations of the methods according to some embodiments may be described in a particular order, the methods of the present disclosure may include some or all of the described operations performed in a different order. In particular, it should be understood that, unless the context otherwise clearly indicates, the order of any of the operations and sub-operations of any of the described methods may be reordered, e.g., when a subsequent operation requires the output of a previous operation as input, or when a subsequent operation requires the result of a previous operation. The methods of the present disclosure may include some or all of the described operations. No particular operation in the disclosed methods is considered to be an essential operation of the method unless expressly so specified.
[0228] Although the present disclosure has been described in connection with specific embodiments thereof, it will be apparent that many alternatives, modifications, and variations will be obvious to those skilled in the art. Accordingly, the present disclosure includes all such alternatives, modifications, and variations that fall within the scope of the appended claims. It should be understood that the present disclosure need not be limited to the details of the construction and arrangement of components and / or methods set forth herein. Other embodiments may be practiced and the embodiments may be carried out in various ways.
[0229] The terminology and phraseology used herein is for descriptive purposes only and should not be considered limiting. The citation or identification of any reference in this application should not be construed as an admission that such reference is available as prior art to the present disclosure. The section headings used herein are for ease of understanding the specification and should not be construed as required limitations.
[0230] All technical and / or scientific terms used herein, unless otherwise defined, have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In case of conflict, the patent specification, including definitions, will control. As used herein, the indefinite article "a / an" means "at least one" or "one or more" unless the context clearly dictates otherwise.
[0231] Unless otherwise specifically stated, as will be apparent from the present disclosure, it should be understood that, according to some embodiments, terms such as "processing", "operation", "computing", "determining", "estimating", "evaluating", "measuring", etc. may refer to actions and / or processes of a computer or a computing system or similar electronic computing device that manipulate and / or transform data represented as physical (e.g., electronic) quantities within the registers and / or memories of the computing system into other data similarly represented as physical quantities within the memories, registers, or other such information storage, transmission, or display devices of the computing system.
[0232] Embodiments of the present disclosure may include apparatuses for performing the operations herein. These apparatuses may be specifically constructed for the desired purposes or may include a general-purpose computer selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a computer-readable storage medium, such as but not limited to any type of disk, including floppy disks, optical disks, CD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, solid-state drives (SSD), or any other type of medium suitable for storing electronic instructions and capable of being coupled to a computer system bus.
[0233] The processes and displays presented herein are not inherently related to any particular computer or other apparatus. A variety of general purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the desired method. The following description shows the structures desired for various systems. In addition, embodiments of the present disclosure have been described without reference to any particular programming language. It should be understood that a variety of programming languages may be used to implement the teachings of the present disclosure described herein.
[0234] Aspects of the present disclosure may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Embodiments disclosed herein may also be practiced in a distributed computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including memory storage devices.
Claims
1. A system for non-destructive characterization of a sample, the system comprising: An electron beam (e-beam) source configured to project an e-beam at one or more e-beam landing energies onto a sample being tested; An X-ray detector configured to sense X-rays emitted from the sample being tested; And Processing circuitry configured to: Receive X-ray measurement data related to one or more e-beam landing energies from the X-ray detector; Extract a vector of values specifying key features of the X-ray measurement data from the X-ray measurement data and Based on and a set of vectors to estimate values of one or more structural parameters of the sample under test The set of vectors includes vectors of key features corresponding to ground truth (GT) reference samples, and for each 1 ≤ n ≤ N, specifies values of one or more actually measured structural parameters of the n-th sample and is obtained by an actual measurement of the emission of X-rays from the n-th reference sample The emission of the X-rays from the n-th reference sample is caused by bombarding the n-th reference sample with an e-beam at each of the one or more landing energies.
2. The system according to claim 1, wherein minimizing a loss function, the loss function being at least of the key features and a vectorization function which is extrapolated from 3. The system according to claim 2, wherein the processing circuitry is configured to estimate by calculating the minimum distance between and 4. The system according to claim 1, wherein the processing circuitry is configured to estimate by calculating the distance between and 5. The system according to claim 1, wherein the reference sample includes a sample having the same or a similar intended design as the sample being tested; and / or the reference sample includes a specially prepared sample that exhibits a selected variation relative to the intended design; and / or the reference sample is selected to include an expected variation of the one or more structural parameters.
6. The system according to claim 1, wherein the one or more structural parameters include one or more of the following: the total concentration of at least one material included in the sample being tested; and optionally, when the sample being tested includes a structure embedded in or on the sample being tested, the width of the embedded structure.
7. The system according to claim 1, wherein the sample being tested includes a plurality of layers, and the one or more structural parameters include one or more of the following: (i) at least one thickness of at least one of the layers; (ii) a combined thickness of at least two or more of the layers; (iii) at least one mass density of at least one of the layers; and (vi) at least one relative concentration of at least one material in one or more of the layers.
8. The system according to claim 7, wherein the one or more e-beam landing energies cause the emission of X-rays from at least two of the plurality of layers.
9. The system according to claim 1, wherein the one or more e-beam landing energies cause the emission of X-rays about one or more characteristic X-ray spectral lines related to one or more target substances included in the sample being tested; wherein the X-ray detector is configured to sense at least one measured spectrum of the emitted X-rays in at least one photon energy range that includes at least one of the characteristic X-ray spectral lines; and wherein the X-ray measurement data includes the measured spectrum.
10. The system according to claim 9, wherein the key feature is the intensity of the characteristic X-ray spectral line and / or the intensity of the background radiation, includes the intensity of the characteristic X-ray spectral line and / or the intensity of the background radiation, or is a function of the intensity of the characteristic X-ray spectral line and / or the intensity of the background radiation.
11. The system according to claim 2, wherein where the double vertical bars indicate the vector norm.
12. The system according to claim 2, wherein wherein specify nominal values of the one or more structural parameters, specify a deviation from the nominal values, is a vector of values of the critical feature corresponding to and A is a matrix.
13. The system according to claim 12, wherein said matrix A is equal to where the double vertical bars denote the matrix norm, and for each 1 ≤ n ≤ N, 14. The system according to claim 12, wherein obtaining and the matrix A as a solution of, where the double vertical bars indicate the matrix norm, and for each 1 ≤ n ≤ N, 15. The system according to claim 4, wherein the processing circuitry is configured to, as part of an estimate relative to apply a k-nearest neighbor (k-NN) regression algorithm to determine the k vectors closest to the of 16. The system according to claim 2, wherein the processing circuitry is further configured to subject to a (k = N)-NN classifier to obtain where N' > N, and comprising and obtaining an additional N' - N vectors by actually measuring additional reference samples.
17. The system according to claim 10, wherein, in order to derive the intensity of the characteristic X-ray spectral line, the processing circuitry is configured to fit a free curve to each interval of the measured spectrum, each interval being centered approximately on a corresponding characteristic X-ray spectral line and consisting of a vicinity region of the characteristic X-ray spectral line, so as to obtain a corresponding optimized curve.
18. The system according to claim 17, wherein the free curve is a sum of functions, the sum of functions including a convex function and a second function, the second function being a polynomial; and the processing circuitry is configured to, as part of fitting the free curve, fit the convex function to the peak of the characteristic X-ray spectral line of the corresponding measured spectrum, and fit the second function so as to take into account the background intensity component of the corresponding measured spectrum.
19. The system according to claim 1, wherein optionally, in one of the manufacturing stages of the patterned wafer, the test sample is the patterned wafer or a part of the patterned wafer.
20. A method for non-destructive characterization of a sample, the method comprising: a measurement operation, the measurement operation including sub-operations for each of one or more landing energies: projecting an e-beam onto the test sample; and obtaining measurement data by measuring the intensity of X-rays emitted from the test sample due to the penetration of the e-beam through the test sample; and a measurement data analysis operation, the measurement data analysis operation including sub-operations: Extract a vector of values of specified key features from the measurement data and Based on and a set of vectors to estimate the value of one or more structural parameters of the sample under test The set of vectors includes vectors of key features corresponding to ground truth (GT) samples, and for each 1 ≤ n ≤ N, specifies the value of one or more actually measured structural parameters of the n-th sample, and is obtained by measuring the emission of X-rays from the n-th sample The emission of the X-rays from the n-th sample is caused by bombarding the n-th sample with an e-beam at each of the one or more landing energies.