Computer-implemented method for detection of defects in imaging data of wafers, corresponding computer-readable medium, computer
By using observational characterization and characteristic elements to verify defect standards in wafer imaging datasets, the problem of distinguishing defects and interference is solved, and the accuracy of defect detection is improved.
Patent Information
- Application Number
- CN202380069923.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-28
- Filing Date
- 2023-09-06
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively distinguish defects from interference in wafer imaging data sets, resulting in limited accuracy of defect detection.
By acquiring the imaging dataset of the wafer and the defect-free observation imaging dataset, defect criteria, including observational characterization and characteristic elements, generate defect information to distinguish defects from interference.
It improves the accuracy of defect detection methods, can effectively distinguish actual defects and interference, and is suitable for quality control and quality assurance processes.
Smart Images

Figure CN119998833A_ABST
Abstract
Description
[0001] This application claims priority to German patent application number 10 2022 125 015.6 filed on September 28, 2022, which is hereby incorporated by reference in its entirety. Technical Field
[0002] The present invention relates to systems and methods for quality control and quality assurance of wafers including semiconductor structures, and more particularly to computer-implemented methods, computer-readable media, computer program products, and corresponding systems for defect detection in imaging data sets of wafers. The methods, computer-readable media, computer program products, and systems are based on a computer-implemented method for defect detection, which includes acquiring an imaging data set of a wafer including a semiconductor structure and verification of defect criteria. The methods, computer program products, and systems for semiconductor inspection can be used for quantitative metrology, defect detection, process monitoring, or defect review of integrated circuits within a semiconductor wafer. Background Art
[0003] Semiconductor manufacturing involves very fine and precise manipulations in the nanometer range on materials such as silicon or oxides, such as etching. Therefore, quality management processes including quality assurance and quality control are very important to ensure high quality standards of the manufactured wafers. Quality assurance refers to a set of activities that ensure high-quality products by avoiding any defects that may occur during the development process. Quality control refers to the system that checks the final quality of the product. Quality control is a part of the quality assurance procedure.
[0004] Wafers made of thin sheets of silicon serve as substrates for microelectronic devices containing semiconductor structures built in and on the wafer. Semiconductor structures are built layer by layer using repeated processing steps involving repeated chemical, mechanical, thermal and optical processes. The size, shape and layout of semiconductor structures and patterns are subject to various influences. For example, in the process of 3D storage devices, the key processes currently are etching and deposition. Other related process steps, such as photolithography exposure or implantation, also have an impact on the characteristics of the integrated circuit components. As a result, the manufactured semiconductor structures have rare and different defects. Devices for quantitative metrology, defect detection or defect review are looking for these defects. These devices are not only needed during wafer manufacturing. Due to the complexity and high nonlinearity of the process, it is difficult to optimize the production process parameters. As a remedy, an iterative scheme called process window verification (PWQ) can be applied. In each iteration, test wafers are manufactured according to the current best process parameters, and different bare cores of the wafer are exposed to different manufacturing conditions. By detecting and analyzing the defects in different bare cores based on the quality assurance process, the best manufacturing process parameters can be selected. In this way, the production process parameters can be adjusted to achieve the best state. Hence, high precision quality control processes and apparatus for metrology of semiconductor structures in wafers are necessary.
[0005] Therefore, the identified defects can be used for root cause analysis. It can be used as feedback to improve process parameters of the manufacturing process during quality assurance, such as exposure time, focus variation, etc., or it can be used to ensure the quality of the manufactured wafers during quality control. For example, bridging defects may indicate under-etching, broken lines may indicate over-etching, persistent defects may indicate a defective mask, and missing structures suggest suboptimal material deposition, etc. Other defects are caused by defects or contamination from various sources, such as degradation of the photolithography mask or particle contamination.
[0006] The manufactured semiconductor structure is based on prior knowledge. The semiconductor structure is manufactured from a series of layers parallel to the substrate. For example, in a logic type sample, metal lines extend parallel in the metal layer or a high aspect ratio (HAR) structure, and metal vias extend perpendicular to the metal layer. The angle between the metal lines in different layers is 0° or 90°. On the other hand, for a VNAND type structure, it is known that its cross section is circular on average. In addition, the semiconductor wafer has a diameter of 300 mm and consists of multiple areas (so-called dies), each of which contains at least one integrated circuit pattern, for example for a memory chip or for a processor chip. During manufacturing, the semiconductor wafer undergoes about 1000 process steps and about 100 or more parallel layers are formed within the semiconductor wafer, including transistor layers, layers between lines and interconnect layers, as well as a 3D array of storage cells in a storage device.
[0007] The aspect ratio and number of layers of integrated circuits are increasing, and the structures are moving towards the third (vertical) dimension. The height of memory stacks is now more than a dozen microns. In contrast, the feature size is getting smaller. The minimum feature size or critical dimension is below 10nm, such as 7nm or 5nm, and in the near future the feature size will approach below 3nm. While the complexity and size of semiconductor structures are growing into the third dimension, the lateral dimensions of integrated semiconductor structures are becoming smaller. Therefore, it becomes challenging to measure the shape, size and orientation of 3D features and patterns and their superposition with high accuracy. The lateral measurement resolution of charged particle systems is usually limited by the sampling grating of each image point or the dwell time of each pixel on the sample and the diameter of the charged particle beam. The sampling grating resolution can be set in the imaging system and can be adapted to the charged particle beam diameter on the sample. Typical grating resolution is 2nm or less, but the grating resolution limit can be reduced without physical limitations. The size of the charged particle beam diameter is limited, depending on the operating conditions of the charged particle beam and the lens. The beam resolution is limited to about half the beam diameter. The lateral resolution may be below 2 nm, for example even below 1 nm.
[0008] An important task in semiconductor inspection is to determine a certain set of parameters of the semiconductor object, such as high aspect ratio (HAR) structures within the inspection volume. These parameters are, for example, size, area, shape or other measurement parameters. Typically, the measurement work of the prior art involves multiple computational steps, such as object detection, feature extraction and any type of metrology operation, such as calculating distance, radius or area based on the extracted features. Each of these many steps requires a lot of computational work.
[0009] Generally speaking, semiconductors contain many repeated three-dimensional structures. During process or process development, some selected physical or geometric parameters of representative multiple three-dimensional structures must be measured with high accuracy and high throughput. In order to monitor the process, an inspection volume is defined, which includes the representative multiple three-dimensional structures. The inspection volume is then analyzed, for example, by a slicing and imaging method, thereby generating a 3D volume image of the inspection volume with high resolution obtained by slicing and imaging multiple cross sections within the inspection volume.
[0010] The multiple repeated three-dimensional structures in the examination volume may exceed hundreds or even thousands of individual structures. As a result, a large number of cross-sectional images are generated, for example, at least one hundred three-dimensional structures are studied through one hundred cross-sectional image slices, so the number of measurements to be performed can easily reach ten thousand or more.
[0011] Current technologies such as multi-beam scanning electron microscopy (multi-beam SEM) can be used to image large areas of a wafer surface with high resolution in a short time. To do this, multi-beam SEM uses multiple parallel single beams, each covering a separate part of the surface with pixel sizes as small as 2nm. The resulting data sets are so large that they cannot be analyzed manually.
[0012] To analyze large amounts of data that require a large number of measurements, machine learning methods can be used. These are suitable for analyzing large amounts of data while limiting the interaction with the user to a minimum.
[0013] Machine learning is a field of artificial intelligence. Machine learning methods are usually based on training data consisting of a large number of samples to build a parametric machine learning model. After training, the method is able to generalize the knowledge gained from the training data to new samples that have not been encountered before, and then make predictions on new data. There are many machine learning methods, for example, linear regression, k-means, support vector machines, neural networks or deep learning methods.
[0014] Methods for automatically detecting defects are typically based on a die-to-die or die-to-database principle. The die-to-die principle compares portions of a wafer to other portions of the same wafer to find deviations from a typical or average wafer design. The die-to-database principle compares portions of a wafer to defect-free reference data, such as observed defect-free wafer images or generated wafer images (e.g., simulated images or CAD files), to find deviations from the ideal data. Due to the large differences, unexpected patterns (i.e., anomalies) in the imaging dataset are detected and subsequently analyzed to derive classification criteria, such as thresholds, area coverage, aspect ratios, etc.
[0015] However, not all anomalies are defects: for example, anomalies may also include, for example, imaging artifacts, image acquisition noise, varying imaging conditions, variations in semiconductor structures within standards, rare semiconductor structures or variations due to imperfect lithography, altered manufacturing conditions or altered wafer processing, registration errors, etc. Such anomalies are not defects, but still deviate from the specification for some reason, and therefore, anomalies detected by a certain anomaly detection method are hereinafter referred to as nuisances.
[0016] Therefore, defect detection methods applied to wafer imaging datasets may face the problem of very high noise rate n, which is the inverse of the precision rate p, i.e. n=1-p, because too many and mostly irrelevant deviations are found on the wafer surface. Therefore, defect detection algorithms usually require extensive post-processing to distinguish between noise and real defects.
[0017] State-of-the-art methods typically use a die-to-database approach by registering an observed wafer imaging dataset with a reference imaging dataset and thresholding the differences. Alternatively, die-to-die approaches are also common and are based on machine learning models such as autoencoders. These models learn to reconstruct only defect-free images and can therefore detect defects based on the difference between the observed image and its reconstructed image. However, all of these methods are susceptible to detecting both noise and actual defects, so these methods have limited usability for wafer defect detection.
[0018] For example, US 6678404 B1 discloses a defect detection method for computer vision applications. Defects are detected based on thresholding the difference between a reference image and an input image. In order to improve the defect detection results, an average reference image and a variance reference image are calculated from multiple reference images. The average reference image and the variance reference image contain at each pixel the average or the respective variance of all reference images at the pixel. Then, the deviation of the input image from the reference image at the pixel is weighted based on the average and variance at the pixel of the average and variance images. However, this method is not suitable for distinguishing between interference and actual defects.
[0019] US2022 / 0044391 A1 describes a defect detection method for wafer images that uses a generative adversarial neural network to estimate the design underlying the wafer image. Defects are then detected by comparing the estimated design with the true design behind the wafer image. To improve the estimated design, multiple estimated designs can be averaged. However, this method is not suitable for distinguishing between interference and defects.
[0020] WO2021181749A1 discloses a defect detection method in which a reference image is learned from multiple object images, and the error between the input image and the reference image is reduced by considering the statistics of the input image pixels. Again, this method is not suitable for distinguishing between interference and actual defects.
[0021] US10504692 B2 discloses a defect detection method, wherein defects in a region of an input image are detected by calculating a sparse representation of the region and calculating the number of atoms required for its representation. However, calculating the number of atoms representing the region is not suitable for distinguishing between actual defects and interference.
[0022] Therefore, it is an object of the present invention to provide a wafer inspection method for measuring semiconductor structures in an inspection volume with high accuracy. Another object of the present invention is to improve the accuracy of the defect detection method, in particular to distinguish between actual defects and disturbances. Another object of the present invention is to adapt the defect detection method to an imaging data set of the wafer, in particular for use in quality control or quality assurance processes. An object of the present invention is to provide a generic wafer inspection method for measuring semiconductor structures in an inspection volume, which can be quickly adapted to changes in the measurement work, the measurement system or changes in the semiconductor object of interest. Another object of the present invention is to provide a robust and reliable measurement method for describing a set of parameters of a semiconductor structure in an inspection volume, the method having high accuracy and with reduced measurement artifacts.
[0023] These objects are achieved by the invention specified in the independent claim. Advantageous embodiments and further developments of the invention are specified in the dependent claims. Summary of the invention
[0024] Embodiments of the present invention relate to a computer-implemented method, computer-readable medium, computer program product, and system for performing a defect detection method on an imaging data set of a wafer.
[0025] A first embodiment relates to a computer-implemented method for detecting defects, which includes: acquiring an imaging dataset of a chip, the chip including a semiconductor structure; verifying a defect standard, the defect standard being used to detect defects in a subset of the imaging dataset of the chip, the defect standard including: an observation representation of the subset of the imaging dataset, which is related to a plurality of characteristic elements derived from a reference image of the semiconductor structure, wherein the observation representation and the characteristic elements define a reconstruction of a minimum reconstruction error of the subset of the imaging dataset, and tolerance statistics of a defect-free representation of a subset of defect-free observation imaging datasets of the chip, wherein each of the defect-free representation and the characteristic elements defines a reconstruction of a minimum reconstruction error of the subset of the defect-free imaging dataset; and generating defect information of the subset of the imaging dataset based on the defect standard.
[0026] This method can distinguish defects from disturbances, thereby improving the accuracy of the defect detection method for the following reasons. The characteristic elements are derived from a reference image of the semiconductor structure, i.e., an image without defects, such as a CAD file. Therefore, the characteristic elements mainly represent defect-free structures. Based on the characteristic elements, an observational representation of a subset of observed, possibly defective imaging data sets can be obtained. If the observed imaging data set contains defects or disturbances, these deviations from the ideal or defect-free semiconductor structure encoded by the characteristic elements cannot be represented, thereby generating considerable reconstruction errors. However, if only the reconstruction error is used as an indicator of defects, disturbances and defects will be detected as well. Therefore, the tolerance statistics are learned from the defect-free representation of defect-free observed imaging data sets. These data sets contain unavoidable disturbances, such as line shortening, line thinning or edge roughness, but no defects. Therefore, the tolerance statistics obtained from the defect-free representation of the defect-free observed imaging data sets include deviations due to disturbances, but no deviations due to actual defects. Therefore, this tolerance statistic allows to distinguish between the observational representations of a subset of observed imaging data sets that include defects and those that only include disturbances.
[0027] In the present application, the imaging data set, the defect-free observed imaging data set or the reference image may include the grayscale values of the original image data itself or values derived from these grayscale values, which are obtained by applying certain operations to the imaging data set, the defect-free observed imaging data set or the reference image, such as gradients, derivatives, feature vectors of one or more dimensions, such as filter responses, such as smoothing filters, values obtained by certain preprocessing methods, such as edge detection image values, etc. In this way, the method disclosed herein can be applied to the obtained original image data or any type of preprocessed image data.
[0028] In the whole text of the present invention, a subset of an imaging data set means a part or a whole of an imaging data set. An imaging data set includes one or more images, such as a large number of images. An imaging data set can be obtained by a charged particle beam imaging system.
[0029] The defect information generated for the subset based on the defect criteria may, for example, include an indicator "defective" / "non-defective", a defect probability or a defect segmentation.
[0030] A statistic is any quantity calculated from multiple samples or observations that is considered for a statistical purpose, such as a mean, variance, moment, probability density function. Statistical purposes include, but are not limited to, estimating a population parameter or population, describing a sample, or evaluating a hypothesis.
[0031] The tolerance statistics obtained from the defect-free representation of the defect-free observed imaging data set can be used in different ways, for example, 1) as a direct indicator of defects based on the comparison of the observed representation of a subset of the imaging data set with the tolerance statistical characteristics. According to an example of the first embodiment of the present invention, based on the statistical characteristics of the obtained observed representation with respect to the tolerance statistics, the defect criterion includes detecting defects in the subset. Specifically, the statistical characteristics may include a quantile of the statistic, in particular a threshold, a confidence interval or a moment of the statistic, in particular a mean and / or a variance. Based on the statistical characteristics, the observed representation of the subset of the imaging data set can be directly marked as "defect" or "non-defect", thereby distinguishing interference from defects.
[0032] The tolerance statistics can also be used in different ways, for example, 2) as in the previous optimization problem, to obtain an observational representation of the number of characteristic elements of a subset of the imaging data set. After obtaining the observational representation as a solution to the optimization problem, defects can be detected based on the reconstruction error associated with the obtained observational representation. Therefore, in an example of the first embodiment of the present invention, the observational representation of the subset is obtained by solving an optimization problem including a priori tolerance statistics including a reconstruction error and a defect-free representation. The defect criterion may include detecting defects in the subset of the acquired imaging data set based on the reconstruction error of the solution to the optimization problem. By using the tolerance statistics as a priori in the optimization problem used to calculate the observational representation of the subset, the tolerance statistics directly affects the observational representation of the subset by preventing low-likelihood observational representations. This will increase the reconstruction error of the defect, but will not increase the reconstruction error of the interference, which, according to the tolerance statistics, is more likely to refer to interference. Therefore, defects can be detected based on the reconstruction error represented by the acquired observations.
[0033] Throughout the present invention, the "defect-free" property of a data set refers to a predominantly defect-free data set, i.e., less than 10% of the data set, preferably less than 5% of the data set, more preferably less than 2% of the data set, and most preferably less than 1% of the data set contains defects.
[0034] The term "characteristic element" derived from a reference image of a semiconductor structure may refer to a feature, such as a set of feature vectors or images that represent a feature of a reference image. For example, a feature may include a subset of a reference image or a processed subset of a reference image (e.g., by modifying contrast, brightness, intensity, color, or by applying a filter such as an edge detector, shape detector, etc.). For example, a feature may include any type of subspace or basis derived from a reference image defined by a set of feature vectors or images (e.g., by using subspace methods such as principal component analysis or independent component analysis, dictionary methods, cluster methods, wavelet or Fourier basis, etc.). The term "characteristic element" may also refer to a set of parameters of a function that maps a subset of an imaging data set to an observed representation of a subset of the imaging data set, where the set of parameters is derived from a reference image, for example, the term "characteristic element" may refer to parameters of a machine learning model learned from a reference image, particularly parameters of a neural network.
[0035] In an example of the first embodiment of the present invention, the defect criterion further includes modifying the defect detection result or the intermediate result of the defect detection method by a trained machine learning model. In this way, the accuracy of the defect detection method can be improved by using a second source of information. The trained machine learning model can be applied to a subset of the imaging data set of the wafer, and / or to the difference between the subset of the imaging data set of the wafer and the aligned reference image (particularly the simulated aligned reference image), and / or to the reconstruction error of the observational representation of the subset of the imaging data set of the wafer. The trained machine learning model can use a region of interest including a subset of the imaging data set of the wafer and / or a region of interest including the difference between the subset of the imaging data set of the wafer and the aligned reference image (particularly the simulated aligned reference image), and / or a region of interest including the reconstruction error of the observational representation of the subset of the imaging data set of the wafer as input. The trained machine learning model may include an autoencoder or a segmentation model.
[0036] A second embodiment of the present invention relates to a computer-implemented method for obtaining tolerance statistics on defect-free characterizations of a subset of defect-free observed imaging data sets of a wafer, the method comprising the following steps: obtaining a defect-free observed imaging data set of a wafer including a semiconductor structure; generating a defect-free characterization of a subset of the defect-free observed imaging data set of the wafer, which is related to a plurality of characteristic elements derived from a reference image of the semiconductor structure, wherein the defect-free characterization and each of the characteristic elements define a reconstruction of a minimum reconstruction error of the subset of the defect-free observed imaging data set; obtaining tolerance statistics on the defect-free characterization. This method allows deriving tolerance statistics of characteristics of a defect-free observed imaging data set that contains interference but does not contain defects. The tolerance statistics then allow distinguishing between observation characterizations of a defect subset of the imaging data set that has a lower probability and observation characterizations of a subset of the imaging data set that has a higher probability of containing only interference.
[0037] An example of the second embodiment of the present invention may also include, before generating the defect-free representation, obtaining a plurality of characteristic elements from a reference image of the semiconductor structure by solving an optimization problem for a minimum reconstruction error including a reconstruction of the reference image, the reconstruction being defined by the reference representation and the characteristic elements.
[0038] By taking the characteristic elements from the reference image, while the tolerance statistics are taken from the defect-free observed imaging dataset, the tolerance statistics are able to model the difference between disturbances and defects. The tolerance statistics include the deviation from the reference image due to disturbances, because disturbances also occur in the defect-free observed imaging dataset, but the tolerance statistics do not include the deviation due to defects. In this way, disturbances can be distinguished from defects.
[0039] According to an example of the second embodiment of the present invention, the optimization problem includes at least one constraint or prior of the characteristic element. In this way, a characteristic element that meets certain requirements can be calculated, or a plurality of solutions with only significantly different characteristics can be avoided, or a set of solutions to the optimization problem can be appropriately constrained to simplify the optimization. For example, the constraint or prior can involve the Lp norm (Lp-norm) of the characteristic element, or the sparsity of the characteristic element, in particular the L0 norm (L0-norm) or L1 norm of the characteristic element.
[0040] According to an example of the second embodiment of the invention, the optimization problem includes at least one constraint or prior on the reference representation. In this way, a reference representation that meets certain requirements, such as sparsity or smoothness requirements, can be calculated, thereby obtaining a result with higher accuracy. For example, the constraint or prior may involve the Lp norm, in particular the L2 norm or the L1 norm, of the reference representation of a neighboring subset of the reference image or the gradient of the reference representation. The constraint or prior is a measure of the sparsity of the reference representation, in particular the L0 norm or the L1 norm or the kurtosis of the reference representation.
[0041] A sparse representation is one that contains only a few non-zero elements. This improves the accuracy of the method because a sparse representation consists of a subset of only a few characteristic elements. This prevents defects from being approximately represented by a combination of many different characteristic elements, which can lead to low reconstruction errors that are not detected as defects. Kurtosis measures the degree of normality of the distribution of the elements of the representation. Therefore, sparse representations have a lower kurtosis.
[0042] According to an example of the first or second embodiment of the present invention, the tolerance statistics include a probability density function obtained from a defect-free characterization of a defect-free observed imaging data set by a density estimation technique. This is beneficial because the estimated probability density function can better estimate the true underlying probability density function than sample or relative frequency statistics, for example, the probability density function can be unbiased and continuous. Therefore, the accuracy of this method is improved.
[0043] According to an example of the first or second embodiment of the present invention, the tolerance statistics include a joint probability density function f(S, R) or a conditional probability density function f(S|R) obtained by a density estimation technique, wherein S includes an observed representation of a subset of an observed imaging data set and / or a defect-free representation of a subset of defect-free observed imaging data sets, and wherein R includes a reference representation of a subset of reference images of multiple characteristic elements. By using a joint or conditional probability density function, rare semiconductor structures or rare interferences can be modeled by a probability density function whose probability is not close to 0, thereby improving the accuracy of the method.
[0044] The characterization S may be based on the same number of characteristic elements as the characterization in R, or on an additional number of characteristic elements, e.g. derived from the observed characterizations of the subset of imaging datasets and / or the defect-free characterizations of the subset of defect-free observed imaging datasets. According to examples of the first or second embodiment of the present invention, a number of additional characteristic elements may be derived from the observed imaging datasets, and the observed characterizations and / or the defect-free characterizations S may be based on the additional characteristic elements, while the reference characterization R of the reference image may be based on the characteristic elements derived from the reference image.
[0045] The probability density function of the tolerance statistics, in particular the probability density function of a Gaussian or Gaussian mixture model, can be obtained by means of parameter density estimation techniques. The advantage of this is that only a few parameters of the predefined probability density function need to be estimated from the defect-free representations, so a small number of defect-free representations can still produce satisfactory results. In addition, some probability density functions are particularly simple to handle, such as the Gaussian probability density function.
[0046] Alternatively, the probability density function of the tolerance statistic can be obtained by nonparametric density estimation techniques, in particular the Parzen density estimator. These methods have the advantage of higher accuracy because for an infinite number of samples, the estimated probability density function converges to the true underlying probability density function.
[0047] In an example, tolerance statistics can also include a machine learning model trained on defect-free representations, particularly a one-class SVM or support vector data description (SVDD). For example, a one-class SVM is trained on a defect-free observed imaging dataset and is able to identify outliers, i.e., defects, based on distance measurements.
[0048] The tolerance statistics may include only a subset of dimensions that are not defective. This can save computation time and limit the tolerance statistics to important dimensions. The tolerance statistics may also include separate tolerance statistics for each dimension of the subset of dimensions that are not defective. This may simplify the computation of tolerance statistics or the application of tolerance statistics in defect detection, as single-dimensional statistics are easier to handle than multivariate statistics, for example, for computing quantiles or confidence intervals.
[0049] According to an example of the first or second embodiment of the present invention, the observed representation of the subset of the imaging data set includes a registration vector indicating an offset between the subset of the imaging data set and a characteristic element in the form of a corresponding subset of the reference image, so that the corresponding subset of the reference image is registered with the subset of the imaging data set by the registration vector, and wherein the defect-free representation of the subset of the defect-free observed imaging data set includes a registration vector indicating an offset between the subset of the defect-free observed imaging data set and the characteristic element in the form of the corresponding subset of the reference image, so that the corresponding subset of the reference image is registered with the subset of the defect-free observed imaging data set by the registration vector. In this way, the registration vector between the characteristic element in the form of the reference image subset and the corresponding subset of the imaging data set of the wafer, or vice versa, is calculated, thereby minimizing the reconstruction error.
[0050] The reconstruction error of the subset of the imaging data set may include a warping error between the subset of the imaging data set and the corresponding subset of the reference image, or vice versa, and the reconstruction error of the defect-free representation of the subset of the defect-free observed imaging data set may include a warping error between the subset of the defect-free observed imaging data set and the corresponding subset of the reference image, or vice versa. The warping error between the first and second images includes the deviation of the first image from the registered second image, that is, the second image warped according to the associated registration vector. For example, the warping error can be measured by the deviation of the subset of the reference image warped according to the registration vector and the subset of the imaging data set of the wafer, or vice versa. For example, the deviation can be measured by the sum of the squares of the differences in grayscale values. For a subset of the imaging data set containing a plurality of pixels, the registration vector field can be calculated by optimizing a registration optimization problem known to those skilled in the art, such as an optimization problem including the warping error and the gradient norm of adjacent registration vectors. Based on the tolerance statistics of the registration vectors, the possibility of a defect can be assigned to each registration vector of the registration vector field. Using the registration vectors as an observed representation of a subset of the imaging data set and / or as a defect-free representation of a subset of the defect-free observed imaging data set improves the accuracy of the defect detection method.
[0051] According to an example of the first or second embodiment of the present invention, the plurality of characteristic elements may include a machine learning model trained on a reference image of a semiconductor structure, in particular a neural network in the form of an autoencoder, and the observed representation of a subset of the imaging data set may include the output of the machine learning model when applied to the subset of the imaging data set, and the defect-free representation of a subset of defect-free observed imaging data sets may include the output of the machine learning model when applied to the subset of defect-free observed imaging data sets. For example, restrictions may be imposed on the parameters of the machine learning model, such as restrictions on the size or number of layers of the neural network. For example, the machine learning model may learn to reduce the dimensionality of the input data. Using the output of the machine learning model as the observed representation of the subset of the imaging data set and / or as the defect-free representation of the subset of defect-free observed imaging data sets improves the accuracy of the defect detection method.
[0052] According to an example of the first or second embodiment of the present invention, the observed representation of a subset of the imaging data set includes coefficients of a decomposition of the subset of the imaging data set, which are related to multiple characteristic elements, and the defect-free representation of the subset of the defect-free observed imaging data set includes coefficients of a decomposition of the subset of the defect-free observed imaging data set, which are related to multiple characteristic elements. In this way, characteristic elements can be learned from reference images that meet the requirements of any suitable defect detection task. For example, a small number of orthogonal characteristic elements generate low-dimensional representations, such as subspace learning techniques, thereby reducing computing time and workload. In contrast, a large number of characteristic elements that represent the typical structure of the input data, as in sparse coding techniques, leads to a high-precision defect detection method.
[0053] Instead of using a subset of imaging data sets, a subset of defect-free imaging data sets, or a subset of reference images, the method disclosed herein may also be applied to difference images, for example, a difference image applied to a subset of imaging data sets aligned with a corresponding subset of a reference image, a difference image applied to a subset of defect-free imaging data sets aligned with a corresponding subset of a reference image, a difference image applied to a subset of imaging data sets and an aligned subset of defect-free observed imaging data sets, etc.
[0054] Characteristic elements and tolerance statistics can also be derived from the difference image of the subset of defect-free observed imaging data set and the aligned subset of reference images, rather than from the reference image, and the observed representation and defect-free representation of the subset can include coefficients of the decomposition of the difference image of the subset and the aligned reference image with respect to the number of characteristic elements. In this way, the observed representation and defect-free representation only encode the difference between the images, rather than the information contained in the images themselves, which reduces the complexity of the characteristic elements and the observed representation and defect-free representation, thereby improving the accuracy of the defect detection method.
[0055] The decomposition of the subsets can be nonlinear or linear. The characteristic elements can include basis elements, such as wavelet basis, Fourier basis, or elements of a principal component basis obtained by principal component analysis. The characteristic elements can also include elements of an overcomplete frame. An overcomplete frame refers to a set of vectors that can be linearly related and based on which each subset of the reference image can be arbitrarily well approximated in norm by a finite combination of vectors. The characteristic elements can include dictionary elements obtained by dictionary learning or multiple independent components obtained by independent component analysis. The characteristic elements can also include multiple image blocks obtained by unsupervised clustering methods. The specific selection of characteristic elements can adapt the defect detection method to many different use cases, making it universal, thereby improving the accuracy of the defect detection method.
[0056] In an example of the first or second embodiment of the present invention, a reference image of a semiconductor structure includes a subset of a defect-free observed imaging data set of the semiconductor structure or a subset of a defect-free generated image of the semiconductor structure, in particular a composite image of a defect-free semiconductor structure. The defect-free generated image of the semiconductor structure may include a plurality of polygons representing the semiconductor structure, or an image generated by a defect-free CAD model of a wafer, or a defect-free image generated by a neural network. The reference image may also include a defect-free generated image of the semiconductor structure and a defect-free observed image of the semiconductor structure. The reference image may be aligned to improve accuracy, for example, multiple reference images may be aligned with respect to each other. In this way, characteristic elements, such as typical structures of defect-free images, may be learned from a large number of aligned reference images. The reference image may also be aligned with the observed imaging data set of the wafer. In this way, the structure of the observed imaging data set and the reference image may be compared, for example, the corresponding structure of the observed structure and the reference image, for example, by a difference image. In this way, defects may be detected.
[0057] By simulating the image acquisition process and the lithography process, the generated semiconductor structure image can be simulated to have a similar appearance to the observed imaging data set of the wafer to improve accuracy.
[0058] To simulate the image, a physically inspired forward simulation of the imaging process of a given charged particle beam imaging system can be applied to the image. The simulation typically involves scaling the image according to the pixel raster of the chosen scanning method. The spatially resolved image contrast is determined by the material contrast of the material present in the image. After scaling and applying the material contrast value, a convolution of the image with a convolution kernel is performed according to the point spread function of the imaging system. The point spread function can be determined according to the expected interaction volume generated by the primary charged particle imaging beam at a cross section through the wafer. The interaction volume typically depends on the electron energy. A noise level can be added according to the dwell time at each raster position. Thus, the finite detection count of the chosen detector geometry is also taken into account. The imaging parameters can also depend on the material composition within the inspection volume of the wafer, and the imaging parameters can include the curtain effect of the milling operation according to the material composition in the cross section to be milled. The curtain effect can be explained by a simple model of the milling operation and can therefore also be taken into account. For example, the milling effect generates additional topographic contrast superimposed on the material contrast. The physical simulation can further take into account additional structures within the inspection area, such as word lines.
[0059] Alternatively, to simulate images, a machine learning model can be trained based on reference images (e.g., layout archives) as input and corresponding defect-free observed imaging datasets as output. In this way, the model learns to simulate image acquisition and optical lithography processes.
[0060] The observed representation of a subset of the imaging data set may include spatial information about the position of the subset in the imaging data set, and / or the defect-free representation of a subset of the defect-free observed imaging data set may include spatial information about the position of the subset in the defect-free observed imaging data set, and / or the representation of a subset of the reference image may include spatial information about the position of the subset in the reference image. For example, the spatial information may include a position encoding, in particular Fourier functions of different frequencies. In this way, spatial information may be taken into account in the defect detection method, thereby obtaining results with higher accuracy. For example, if typical types of defects mainly occur in specific areas, or if areas have a higher probability of defects, such as boundary areas, the positions of the subsets may include valuable information, or about the correlation of subsets from different imaging data sets or reference images.
[0061] In an example of the first or second embodiment of the present invention, the subset includes a single pixel. In this way, a small portion of the wafer imaging data set can be checked for defects. Alternatively, the subset can include multiple pixels that are checked for defects together. An observed representation of the subset of imaging data sets can be obtained from a region of interest of the subset including the imaging data sets, and a non-defective representation of the subset of non-defective observed imaging data sets can be obtained from a region of interest of the subset including the non-defective observed imaging data sets. This allows the configuration of the respective subsets to be taken into account during defect detection, thereby improving the accuracy of the method.
[0062] In an example of the first or second embodiment of the present invention, a machine learning model is trained to assign a defect type from a predefined set of defect types to an observed representation of a subset of the wafer imaging data set, the observed representation being based on the number of characteristic elements. In this way, the detected defects can also be classified, which can then be fed back directly to the user or the specific hardware unit responsible for the detected defect type.
[0063] The imaging data set of the wafer can be acquired by a charged particle beam system. Charged particle beam systems include, but are not limited to, scanning electron microscopes (SEMs), focused ion beam microscopes, such as helium ion microscopes. Another example of a charged particle beam system is a corrected electron scanning microscope, which includes correction means for correcting chromatic aberration and spherical aberration.
[0064] Examples of the first or second embodiment of the present invention further include directing observed representations of a subset of the imaging data set of the wafer, and / or defect-free representations of a subset of defect-free observed imaging data sets, and / or reference representations of reference images and / or characteristic elements, and / or detected defects in the imaging data set of the wafer to a display device or dashboard for visualization. Examples of the first or second embodiment of the present invention further include directing detected defects in the imaging data set of the wafer to a display device or dashboard for visualization, wherein the detected defects are highlighted or marked according to the defect type. This improves usability.
[0065] In an example of the first or second embodiment of the present invention, reference images, characteristic elements and / or tolerance statistics are provided by replaceable hardware, so that the data can be reused in other applications and can be easily exchanged, thereby improving the usability of the method.
[0066] Examples of the first or second embodiments of the present invention further include determining one or more measurement values of identified defects in a subset of the imaging data set of the chip, particularly size, area, dimensions, shape parameters, distance, radius, aspect ratio, type, number of defects, density, spatial distribution of defects, presence of any defects, and the like.
[0067] Based on these measurements, this example may further include evaluating the quality of the wafer based on the one or more measurements and at least one quality assessment rule, or this example may further include controlling at least one wafer manufacturing process parameter based on one or more measurements of defects identified in the imaging data set of the wafer. Wafer manufacturing process parameters include, but are not limited to, parameters of exposure time, etching, deposition, implantation, thermal treatment, and other processes related during manufacturing.
[0068] The present invention also relates to a computer-readable medium on which a computer program executable by a computer device is stored, the computer program comprising a program code for executing the method according to any one of the embodiments of the present invention.
[0069] The invention also relates to a computer program product comprising instructions which, when a computer executes the program, cause the computer to perform a method according to any one of the embodiments of the invention.
[0070] The present invention also relates to a system for controlling the quality of wafers produced in a semiconductor manufacturing plant, the system comprising: an imaging device, which is suitable for providing an imaging data set of the wafer; one or more processing devices; one or more machine-readable hardware storage devices, which contain instructions executable by the one or more processing devices to perform operations including a method for evaluating the quality of the wafer.
[0071] The present invention also relates to a system for controlling the production of chips in a semiconductor manufacturing plant, the system comprising: a component for producing chips controlled by at least one manufacturing process parameter; an imaging device suitable for providing an imaging data set of the chip; one or more processing devices; one or more machine-readable hardware storage devices comprising instructions executable by one or more processing devices to perform operations including a method for controlling at least one chip manufacturing process parameter.
[0072] Any of the above systems may include a display device and / or a user interface.
[0073] Although examples and embodiments of the present invention are described with respect to semiconductor wafers, it should be understood that the present invention is not limited to semiconductor wafers, but can also be applied to semiconductor fabrication reticles or masks or other fabricated objects.
[0074] The present invention described using the examples and embodiments is not limited to the embodiments and examples, but can be implemented by those skilled in the art through various combinations or modifications of the embodiments and examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 shows the photolithography process used to manufacture wafers;
[0076] Figure 2 shows the inspection process for quality control in wafers;
[0077] Figure 3 a to 3d show typical disturbances and defects in the imaging data set of a wafer;
[0078] Figure 4 a and 4b show the difference between interference and defect;
[0079] Figure 5 a-5d illustrate the inadequacy of the die-to-die approach for defect detection in a wafer imaging dataset;
[0080] Figure 6 A flow chart illustrating the steps of a standard die-to-database method is shown;
[0081] Figure 7 A flow chart illustrating the steps of a first embodiment of the present invention is shown;
[0082] Figure 8 A flow chart illustrating the steps of a second embodiment of the present invention is shown;
[0083] Fig. 9 An example of a first embodiment of the present invention is shown;
[0084] Fig.10 An example of a first embodiment of the present invention is shown;
[0085] Fig.11 An example of a first embodiment of the present invention is shown;
[0086] Fig.12 A system is schematically shown which may be used to control the quality of wafers produced in a semiconductor manufacturing plant;
[0087] Fig.13 A system is schematically shown that may be used to control the production of wafers in a semiconductor manufacturing plant. DETAILED DESCRIPTION
[0088] Advantageous exemplary embodiments of the present invention will be described below and schematically illustrated in the accompanying drawings. In all drawings and descriptions, the same reference numerals are used to describe the same features or components. Dashed lines represent optional elements.
[0089] Semiconductor manufacturing enables 3D template design on physical materials at the sub-nanometer scale. Imprinting is done layer by layer, where each iteration contains lithography-based manufacturing steps and quality control steps.
[0090] Figure 1 A layered photolithography process 10 is shown for fabricating a wafer 24. At each iteration, a photoresist 14 is deposited on a substrate 12. A mask 16, which includes a template of the desired semiconductor pattern, is used to selectively expose the photoresist 14 to destructive radiation 15. The substrate 12 beneath these areas is then removed by etching 18. Finally, the remaining photoresist 14 is removed by cleaning 20, thereby achieving the pattern of the semiconductor structure specified by the mask 16.
[0091] Figure 2 An inspection process 22 is shown for quality control of each layer in a wafer 24. During fabrication, the wafer 24 is imaged at a suitable resolution using a microscope 26. The resulting image data set 28 of the layer is inspected for defects in the inspection process 22. The lithography and inspection process continues until all layers of the design are satisfactorily imprinted on the physical substrate.
[0092] The complex process of deposition, mask exposure and etching may lead to many anomalies, resulting in a significant decrease in yield. Therefore, it is important to detect defects in the imaging data set 28 of the wafer 24 in order to perform root cause analysis and attribute the detected defects to specific steps of the manufacturing process. In this way, a quality assurance and / or quality control mechanism can be established. Quality assurance ensures that the approaches, technologies, methods and processes of wafer manufacturing are implemented as required. It aims to improve the parameters or conditions of the wafer 24 production process in the laboratory, such as deposition, exposure and etching processes. To this end, known and unknown defects must be identified and analyzed. In contrast, quality control aims to ensure the quality of the final manufactured product in the online manufacturing process. To this end, known defects must be identified and analyzed.
[0093] Figure 3 a to 3d show typical disturbances 34 and defects 39 in the imaging data set 28 of the wafer 24 . Figure 3 a shows a mask 16 or layout image comprising an ideal semiconductor structure of layers in the form of polygons to be imprinted on substrate 12 during wafer fabrication. Figure 3 3b and 3c show an imaging data set 28 of a wafer 24 covered by a mask 16 generated during the inspection process 22. Figure 3 In b, the manufacturing process is defect-free, resulting in a defect-free observed imaging dataset 30. However, small deviations from the ideal layout are unavoidable, such as line shortening 36, line thinning 37, and edge roughness 38. These random deviations do not affect the functionality of the wafer 24 and therefore fall into the category of disturbance 34. Figure 3 c, the manufacturing process is prone to error, generating an error-prone imaging data set 32, which includes Figure 3 d, such as line thinning defects 40, bridge defects 42, long bridge defects 44, intrusion defects 46, broken line defects 48, excursion defects 50, and line pullback defects 52. These defects 39 indicate problems in the manufacturing process, for example, consistent line pullback defects 52 at the same structure of each bare die indicate defects in the mask 16, bridge defects 42 indicate underexposure, and line thinning defects 40 indicate overetching 18. Other defects are caused by defects or contamination from various sources, such as degradation of the photolithography mask 16 or particle contamination. Such defects 39 are useful for root cause analysis of quality assurance or quality control processes.
[0094] Figure 4 4a and 4b illustrate the difference between disturbance 34 and defect 39. Figure 4 a shows a portion of the mask 16 and the implemented design exhibits edge roughness 38. This deviation belongs to the category of disturbance 34, because the deviation from the ideal design is unavoidable and does not affect the function of the wafer 24. On the contrary, Figure 4 b shows the same part of the mask 16 and the realized design, but with a larger structure outside the mask 16. This deviation is a defect 39 called an offset defect 50. In order to be able to make meaningful statements about the quality of the wafer 24, it is important to distinguish between disturbances 34 and defects 39.
[0095] To detect defects 39 in an imaging dataset 28 of a wafer 24, a popular approach is based on a die-to-die analysis approach. The die-to-die approach compares portions of a wafer 24 to other portions of the same wafer 24 to find deviations from a typical or average wafer design. Such an approach allows for the distinction between disturbances 34 and real defects 39 if trained using a defect-free observed imaging dataset 30. However, a limitation of this approach is that defects 39 are consistent across multiple dies, e.g., mask-related defects or defects that occur frequently cannot be found. However, the detection of such defects 39 is particularly important. In addition, it is not possible to infer defects 39 based solely on the limited spatial context in the imaging dataset 28 as in the die-to-die approach. Figure 5 The inadequacy of a die-to-die approach for defect detection in the imaging data set 28 of the wafer 24 is illustrated.
[0096] Figure 5 a shows a mask 16 or layout image comprising an ideal semiconductor structure of layers to be imprinted on a substrate during wafer fabrication. Figure 5 5b and 5c show an imaging data set 28 of a wafer 24 covered by a mask 16 generated during the inspection process 22. Figure 5 b, the manufacturing process is defect-free, generating a defect-free observed imaging dataset 30. However, Figure 5 As shown in Figure 3, minor deviations from the ideal layout cannot be avoided, such as line shortening 36, line thinning 37 at dense areas due to mask optical interactions during exposure, and edge roughness 38 due to the statistical nature of exposure and etching 18. Other differences may occur due to complex registration problems between the reference image 66 and the observed image, for example due to nonlinearities in the imaging process. Such random deviations do not affect the functionality of the wafer 24 and therefore fall into the category of disturbances 34. Figure 5 d shows the problem of inferring a defect 39 based only on the local environment 76, 80 of the semiconductor structure. The die-to-die approach derives its knowledge from similar structures in different locations of the same imaging data set 28. Therefore, a defect 39 that looks like the correct structure in the imaging data set cannot be detected. The same applies to defects 39 that appear in multiple locations of the imaging data set 28. Marks 74, 78 represent portions of defect-free semiconductor structures. However, if these structures are inspected by a die-to-die approach based only on the local environment 76, 80, it is impossible to determine whether the structure is correct because the local environment of the structure looks exactly the same. Therefore, the die-to-die approach is not suitable for reliable defect detection in semiconductor structure inspection.
[0097] Instead, a die-to-database approach may be used that compares portions of the wafer 24 to a defect-free reference image 66, such as a defect-free observed imaging dataset 30 of the wafer 24 or a generated image of the wafer 24, such as a simulated image or CAD file, to identify deviations from the ideal data.
[0098] Figure 6 is a flow chart showing the steps of a standard die to database method 56. The inputs to this method are an observed imaging dataset 28 and a reference image 66 of a wafer 24 to be inspected. The reference image 66 contains information about polygons representing ideal semiconductor structures and their spatial positions, which can be compared with the observed imaging dataset 28. These polygons are rasterized onto an image grid in a rasterization step 58. In a subsequent anchor point step 60, corresponding feature points are detected in the rasterized layout image and the imaging dataset 28. These anchor points are used to align the layout image and the imaging dataset 28 in an alignment step 62. Image alignment can also be performed or supported by manual intervention. During a simulation step 64, the rasterized layout image can be optionally textured to make it look like the observed image by simulating an image acquisition process, for example, by a multi-beam electron microscope and a lithography process 10. This simulated alignment image is used as a simulated alignment reference image 67, i.e., a model indicating the ideal semiconductor structure that should be imprinted on the wafer 24. This simulated alignment reference image 67 is then compared with the observed imaging dataset 28 in a differencing step 68. The differences between the observed and reference images reveal anomalies, i.e. defects 39 and disturbances 34, etc. In order to reduce disturbances 34 in the detection, the anomalies are post-processed in a post-processing step 70. For example, the anomalies can be presented to an expert who labels them as disturbances 34 of no interest or defects 39 of interest, thereby generating a number of defect suggestions 72. This information can then be used for root cause analysis.
[0099] Although bare-die-to-database approaches are more robust for defect detection, they cannot distinguish between interference34 and real defects39 and therefore require extensive post-processing.
[0100] Therefore, one object of the present invention is to propose a defect detection method which is both robust to defect detection and capable of distinguishing disturbances 34 from real defects 39 .
[0101] Figure 7A flow chart illustrating the steps of a first embodiment of the present invention is shown. The first embodiment relates to a computer-implemented method 82 for defect detection, comprising the steps of acquiring an imaging data set 28 of a wafer 24 including a semiconductor structure in an imaging step 84; verifying a defect standard for defect detection in a subset of the imaging data set 28 of the wafer 24 in a defect standard verification step 86, the defect standard comprising an observed representation 88 of the subset of the imaging data set 28 with respect to a plurality of characteristic elements 90 derived from a reference image 66 of the semiconductor structure, wherein the observed representation and the characteristic elements 90 define a reconstruction of a minimum reconstruction error of the subset of the imaging data set; and a tolerance statistic 92 of defect-free representations 94 of a subset of defect-free observed imaging data sets 30 of the wafer 24, wherein each defect-free representation and the characteristic element 90 define a reconstruction of a minimum reconstruction error of the subset of the defect-free imaging data set. The method generates defect information for the subset of the imaging data set 28 based on the defect standard. The detected defects 39 may be used for quality assurance or quality control of the wafer.
[0102] If the reference image 66 is as Figure 6 The described alignment reference image or simulated alignment reference image 67 is advantageous if the reference image 66 and the defect-free observed imaging data set 30 comprise the same semiconductor structure as the imaging data set 28 of the wafer 24 to be inspected.
[0103] The observational representation 88 of a subset of the imaging dataset 28 may be obtained by solving the following general optimization problem:
[0104]
[0105] in Observation representation 88 representing a subset x of imaging data set 28 that minimizes the reconstruction error E with respect to a plurality of characteristic elements C C .
[0106] The optimization problem may include at least one constraint or prior on the observation representation 88. For example, the constraint or prior may be a measure of the sparsity of the observation representation 88, in particular the L0 norm or L1 norm or the kurtosis of the observation representation 88.
[0107] For example, equation (I) can be further constrained as follows:
[0108]
[0109] where q is a prior on r, such as the L1 norm, and λ is a weighting factor.
[0110] For example, equation (I) can also be constrained as follows:
[0111]
[0112] And for example, Where v is a predefined value.
[0113] The defect criteria can use the tolerance statistics 92 in different ways, for example as a direct indication of defect information, such as defect probability, or as a prior in the optimization problem (I) described above. Other uses of the tolerance statistics 92 are also contemplated.
[0114] The first option uses the tolerance statistics 92 as a direct indicator of defects 39 . In an example of the first embodiment of the present invention, the defect criteria includes detecting defects 39 in a subset of the imaging data set 28 based on statistical properties of the acquired observation representations 88 with respect to the tolerance statistics 92 .
[0115] For example, let P be indicative of a probability distribution estimated from a sample of tolerance statistics 92, i.e., a subset of defect-free observed imaging data sets 30 with respect to the number of characteristic elements 94. The probability distribution P can then be used to assign defect probabilities to the observed characteristics of the subset of imaging data sets 28.
[0116] Right now
[0117] in
[0118] The tolerance statistic 92 may also be used directly without first estimating a probability distribution from the sample, for example by computing the relative frequencies of the observed representations 88 based on a histogram.
[0119] Instead of deriving a probability of defect, a binary decision of "defective" / "non-defective" may also be obtained. To this end, the statistical properties may for example comprise quantiles of said tolerance statistics 92, in particular threshold values.
[0120] Let F be the cumulative distribution function of the probability density function estimated from the plurality of defect-free representations 94 of the defect-free observed imaging data set 30 according to the tolerance statistics 92, then x p is the p-quantile, if
[0121] F(x p )≥p and The empirical p-quantile can also be used to characterize multiple defect-free values based on tolerance statistics 94x1, ..., x n to estimate without first estimating the cumulative distribution function. (p) is the empirical p-quantile if for at least p·n there are no defective characterizations
[0122] x i ≤x(p)
[0123] Established, and at least (1-p)·n defect-free characterization
[0124] x i ≥x( p)
[0125] Established.
[0126] According to the tolerance statistic 92, the quantile x p or x (p) can be used as a threshold to separate the fraction p of representations with high probability from the fraction (1-p) of representations with low probability (i.e., outliers). For a known threshold x p , the corresponding p-value can also be determined.
[0127] The quantile may be determined for only a subset of the dimensions of the defect-free representation 94, in particular for a single dimension, or for each dimension individually. Thus, the threshold may be a threshold vector for a subset of the dimensions of the defect-free representation 94, in particular for only a single value for a single dimension. Based on the quantile or threshold, the observed representation 88 of the subset of the imaging data set 28 may be marked as a defect 39, for example, if the corresponding value of the observed representation 88 exceeds the quantile x p or threshold x p :
[0128] in
[0129] In an example of a first embodiment of the present invention, the statistical properties of the tolerance statistics 92 include confidence intervals. The confidence intervals may be determined by fitting a probability density function to the defect-free representation 94. If the observed representation 88 of a subset of the observed imaging data set 28 is outside the confidence interval, the observed representation 88 and the corresponding subset of the imaging data set 28 are outliers with respect to the tolerance statistics 92 and are marked as defects 39. Again, confidence intervals may be determined for subsets of dimensions of the defect-free representation 94, particularly for a single dimension, or for each dimension separately. l ,b u ] is set to represent the confidence interval, with the lower limit being b l , the upper limit is b u , if the characterization of the imaging dataset 28 If it is outside the confidence interval, the label “defect” is assigned to the subset of imaging dataset 28:
[0130] in
[0131] In an example of a first embodiment of the present invention, the statistical properties include moments of the tolerance statistics 92, in particular the mean μ and / or the variance σ. For example, a Gaussian distribution may be fitted to the defect-free representation 94, and a confidence interval may be obtained based on the statistical mean and variance of the defect-free representation 94, such as b l =μ-3σ,b u =μ+3σ.
[0132] Alternatively, the distance of the observed representation 88 of a subset of the observed imaging data set 28 from the mean value may be used as a defect indicator.
[0133] The second option uses the tolerance statistics 92 as a prior in the optimization problem to obtain the observed representations 88 of the subset of the imaging data set 28. Therefore, the optimization problem contains a prior that includes the tolerance statistics on the defect-free representations 94. For example, the optimization problem can be formulated as follows based on the optimization problem (I) above:
[0134]
[0135] Where γ is a weighting factor. In this case, the tolerance statistic 92P is formulated as a logarithmic probability function. Other terms of the optimization problem that include the tolerance statistic 92 are conceivable, such as the formulation for a single-class SVM described below. The tolerance statistic 92 on the defect-free representation 94 of the defect-free observed image dataset 30 can also be defined or modified by a human.
[0136] In an example of a first embodiment of the invention, the defect criterion comprises detecting defects 39 in a subset of the acquired imaging data set 28 based on a reconstruction error of a solution to an optimization problem. Since the characteristic elements 90 are derived from a reference image 66 without the defect 39, they cannot adequately represent the defect 39. Therefore, in case the subset contains defects that cannot be characterized by the characteristic element 90, the observed representation 88 of the imaging data set 28 with respect to the characteristic element 90 deviates from the original subset x. Therefore, the reconstruction error of the observed representation 88 with respect to the characteristic element 90 is used as defect information in the form of a defect indicator D:
[0137] in By applying a threshold to the reconstruction error, D can also be a binary function.
[0138] The reconstruction error can be measured, for example, by an Lp norm, such as the L1 or L2 norm, or a weighted Lp norm:
[0139]
[0140] Where R C(x) is the reconstruction of x with respect to characteristic element C, and w(x) represents a weight or a weight vector. Various reconstructions of x with respect to multiple characteristic elements C are described below.
[0141] The first embodiment of the present invention can be used in conjunction with other methods for defect detection (e.g. Figure 6 To this end, the defect criteria may also include modifying the defect detection results by training the machine learning model 95, such as Figure 7 For example, the defect information generated according to the defect criteria can be combined with the output of the machine learning model, for example, the defect probabilities can be multiplied, or if both methods assign the label "defect" to the subset, then only the subset is labeled as defective, or, in order to obtain a more sensitive method, if one of the methods assigns the label "defect" to the subset, then the subset is labeled with the label "defect".
[0142] The machine learning model 95 can also be used to modify intermediate results of the defect detection method according to the exemplary first embodiment of the present invention before post-processing the intermediate results. The modified intermediate results can then be further processed by the defect detection method. For example, the observed representations 88 of a subset of the imaging data set 28 of the wafer 24 can be modified by a machine learning model trained to suppress interference 34, for example, by reducing the length of the registration vector or reducing the difference between grayscale values.
[0143] The trained machine learning model 95 may include, for example, defect detection, anomaly detection, defect segmentation, or anomaly segmentation methods. Anomalies refer to deviations between semiconductor structures and predefined specifications. They include defects 39 and interferences 34, among others.
[0144] The trained machine learning model 95 may be applied, for example, to a subset of the imaging dataset 28 of the wafer 24, and / or differences between the subset of the imaging dataset 28 of the wafer 24 and the alignment reference image 67, in particular the simulated alignment reference image 67, and / or reconstruction errors of the observed representations 88 of the subset of the imaging dataset 28 of the wafer. The machine learning model 95 may also learn to render the reference image 66 to fit the image distribution of the imaging dataset 28 instead of the alignment reference image 66. The trained machine learning model 95 may also use, as input, a region of interest containing a subset of the imaging dataset 28 of the wafer 24, and / or a region of interest including differences between the subset and the reference image 66, and / or a region of interest including reconstruction errors of the observed representations 88 of the subset of the imaging dataset 28 of the wafer 24. In this way, the machine learning model 95 learns to suppress the disturbances 34 while preserving the defects 39.
[0145] To train the machine learning model 95, multiple user-annotated samples of defects 39, interferences 34, and non-defective data may be presented to the machine learning model 95. The samples do not have to cover all types of defects 39 or interferences 34, as the machine learning model 95 generalizes to unknown defects 39 and interferences 34.
[0146] The trained machine learning model may include an autoencoder. The autoencoder learns the expected statistical variation 30 of the defect-free observed imaging data set. The autoencoder may be trained using a subset of the defect-free observed imaging data set 30 and a corresponding subset of the reference image 66 and / or differences or reconstruction errors thereof with respect to characteristic elements 90 or regions of interest including the input data. The autoencoder learns a compressed representation of the input data. If a subset of the observed imaging data set 28 of the wafer 24 does not have defects 39, the subset is reconstructed with high fidelity by the autoencoder. However, if the subset contains defects 39, the corresponding spatial region is reconstructed with reduced fidelity. Defects 39 may then be detected by calculating the reconstruction error of the output of the autoencoder with respect to the input of the autoencoder.
[0147] According to one aspect of the example of the first embodiment of the present invention, the trained machine learning model 95 may include a segmentation model. Defects 39 acquired by the method according to the example of the first embodiment of the present invention may be compared with the results of the segmentation model and marked based on a combination of the two results.
[0148] Figure 8 A flow chart illustrating the steps of a second embodiment of the present invention is shown. The second embodiment relates to a computer-implemented method 96 for obtaining tolerance statistics 92 for defect-free representations 94 of a subset of defect-free observed imaging data sets 30 of a wafer 24, based on a plurality of characteristic elements 90 derived from a reference image 66 of a semiconductor structure, the method comprising the following steps: obtaining a defect-free observed imaging data set 30 of a wafer 24 including a semiconductor structure in an imaging step 98; generating a defect-free representation 94 of a subset of defect-free observed imaging data sets 30 of the wafer 24 with respect to a plurality of characteristic elements 90 derived from the reference image 66 of the semiconductor structure, wherein each of the defect-free representations 94 and the characteristic elements 90 defines a reconstruction of a minimum reconstruction error of the subset of defect-free observed imaging data sets 30 in a representation generating step 100, and obtaining tolerance statistics 92 for the defect-free representations 94 in a tolerance statistics step 102. The tolerance statistics 92 may be used for defect detection of the wafer, for example, to perform the method of any example of the first embodiment of the present invention, or for quality assurance or quality control of the wafer. It is advantageous if the defect-free observed imaging data set 30 comprises the same semiconductor structure as the imaging data set 28 of the wafer 24 to be inspected.
[0149] An example of the second embodiment of the present invention may also include a characteristic element step 99 before the representation generation step 100, wherein the number of characteristic elements 90 is obtained from a reference image 66 of the semiconductor structure, for example, from a simulated registered reference image 67 including the same semiconductor structure, and the reconstruction is defined by the reference representation 104 and the characteristic elements 90 by solving an optimization problem for the minimum reconstruction error of the reconstruction including the reference image 66. Alternatively, the characteristic elements 90 may be obtained from another source, for example from a previous use case or from a database, or they may be loaded from a memory.
[0150] It is advantageous to improve the accuracy of the method if the reference image 66 is aligned before deriving the characteristic elements 90. Figure 6 The alignment process described, for example, uses rasterization, anchoring, alignment or registration techniques or manual intervention. If the reference image 66 is simulated, that is, the textures of the reference image 66 are adapted so that they look similar to those of the reference image 66 as described above, Figure 6 The described observed imaging dataset 28 is also suitable. It is also advantageous if the reference image 66 comprises the same semiconductor structure as the imaging dataset 28 of the wafer 24 to be inspected.
[0151] The optimization problem of obtaining the characteristic element C can generally be expressed as:
[0152]
[0153] According to an example of the second embodiment, the optimization problem comprises at least one constraint or prior on the characteristic element 90. For example, the constraint or prior may relate to the Lp norm of the characteristic element 90, in particular the L0 norm or the L1 norm of the characteristic element 90.
[0154] The optimization problem including constraints on the characteristic element 90 can be based on the set The statement is as follows:
[0155]
[0156] For example Where v is a predetermined value, such as 1.
[0157] The optimization problem involving the prior q on the characteristic element 90 with adjustable weight λ is expressed as follows
[0158]
[0159] Here, the prior plays a regularizing role on the characteristic element 90, e.g.
[0160] According to an example of the first or second embodiment of the invention, the optimization problem comprises at least one constraint or prior on the reference representation 104. For example, the constraint or prior may relate to the Lp norm of the reference representation 104, in particular the L2 norm or the L1 norm or the L0 norm or the kurtosis.
[0161] For example, equation (I) can be further constrained a priori as follows:
[0162]
[0163] where q is a prior on r, such as the L1 norm, and λ is a weighting factor.
[0164] For example, equation (I) can be further restricted by the following constraints:
[0165]
[0166] For example Where v is a predefined value.
[0167] For example, the optimization problem in equation (II) for obtaining the characteristic element 90 can also be constrained by constraining the reference representation 104 as follows:
[0168]
[0169] For example or For a specific value v.
[0170] For example, the optimization problem in equation (II) for obtaining the characteristic element 90 can be further constrained by the prior q and the weight factor λ on the reference representation 104 as follows:
[0171]
[0172] The prior here plays a regularizing role on the reference representation 104 and can be formulated as: or q(r i )=‖r i ‖1 or q(r i )=‖r i ‖0.
[0173] In an example of the first or second embodiment of the invention, the constraint or prior relates to the Lp norm, in particular the L2 norm or the L1 norm, of the gradient of the reference representation 104 of a neighboring subset of the reference image, e.g.
[0174] or
[0175] For example, equation (I) can be further constrained a priori as follows:
[0176]
[0177] In an example of the first or second embodiment of the invention, the constraint or prior is a measure of the sparsity of the reference representation 104 , in particular the L0 norm or the L1 norm or the kurtosis of the reference representation 104 .
[0178] In one example, the optimization problem can take the following form:
[0179]
[0180] According to an example of the first or second embodiment of the present invention, the tolerance statistics 92 include a probability density function obtained from the defect-free representation 94 of the defect-free observed imaging data set 30 by a density estimation technique.
[0181] The tolerance statistics may include a joint probability density function f(S, R) or a conditional probability density function f(S|R) obtained by density estimation techniques, where S includes observed representations 88 of a subset of the observed imaging data set 28 and / or defect-free representations 94 of a subset of defect-free observed imaging data sets 30, and where R includes reference representations 104 of a subset of reference images 66 with respect to a plurality of characteristic elements 90, for example, the same number of characteristic elements 90 or an additional number of characteristic elements 91. In this way, rare semiconductor structures or rare interferences 34 may be modeled by the probability density function without having a probability close to zero, thereby improving the accuracy of the method. Thus, representations 88, 94, 104 may be derived based on different sets of characteristic elements. For example, a plurality of additional characteristic elements 91 may be derived from the observed imaging data set 28 and / or the defect-free observed imaging data set 30. Then, the reference representation 104R of the reference image 66 may be based on the characteristic elements 90 and the representation S comprising the observed representation 88 of the observed imaging dataset 28 and / or the defect-free representation 94 of the defect-free observed imaging dataset 30 may be based on the additional characteristic elements 91 , or vice versa.
[0182] For density estimation, parametric or non-parametric methods may be used. For example, the probability density function of the tolerance statistics may be obtained by parametric density estimation techniques, in particular the probability density function of a Gaussian or Gaussian mixture model 92. Alternatively, the probability density function of the tolerance statistics 92 may be obtained by non-parametric density estimation techniques, in particular the Parzen density estimator. The tolerance statistics 92 may also include a machine learning model trained on the defect-free representation 94, in particular a single-class SVM or SVDD.
[0183] According to one aspect of this example, the tolerance statistics 92 include only a subset of the dimensions of the defect-free characterization 94. Specifically, the tolerance statistics 92 may include only a single dimension of the defect-free characterization 94. The tolerance statistics 92 may also include separate tolerance statistics 92 for each dimension of the subset of dimensions of the defect-free characterization 94.
[0184] In an example of the first or second embodiment of the invention, the observed representation 88 of the subset of the imaging data set 28 comprises a registration vector indicating an offset between the subset of the imaging data set 28 and a corresponding subset of the reference image 66 in the form of a characteristic element 90, such that the corresponding subset of the reference image 66 is registered with the subset of the imaging data set 28 by the registration vector, and wherein the defect-free representation 94 of the subset of the defect-free observed imaging data set 30 comprises a registration vector indicating an offset between the subset of the defect-free observed imaging data set 30 and a corresponding subset of the reference image 66 in the form of a characteristic element 90, such that the corresponding subset of the reference image 66 is registered with the subset of the defect-free observed imaging data set 30 by the registration vector. In this example, the characteristic element 90 can be understood as a corresponding part of the reference image 66, which is registered with the subset. Based on the registration vectors and the tolerance statistics 92 of these registration vectors, defects 39 can be detected.
[0185] Fig. 9 An example of a first embodiment of the invention is shown. In an imaging step 84, a subset of the imaging data set 28 of the wafer 24 is obtained, for which defects 39 are to be detected. The imaging data set 28 contains line thinning defects 40 and false structural defects 54. First, in a simulation step 64, a corresponding reference image 66 containing the correct semiconductor layout is simulated to produce a simulated reference image 106. The simulation is optional 116. In a registration step 108, the simulated reference image 106 or the corresponding reference image 66 without simulation is registered with the subset of the imaging data set 28, for example by a machine learning registration method or based on an optimization problem for minimizing the reconstruction error. To this end, the reconstruction error may include a distortion error between the subset and a corresponding subset of the reference image 66, or vice versa, for example,
[0186] Among them I O (a) = x, c = I ref (a+r), where I O (a) represents a subset of the imaging dataset observed at position a, and I ref (a+r) represents the reference image 66 at position a+r. For example, the observation representation 88 including the registration vector (i.e., the registration vector field) can be obtained by solving an optimization problem including the warping error and the regularization on the registration vector field:
[0187]
[0188] A tolerance statistic 92 of the registration vector is obtained from the registration vector of the defect-free observed image dataset 30. Based on this tolerance statistic 92, defect information is generated for the calculated registration vector in the defect criteria verification step 86. For example, a zero offset registration vector 112 indicates no defect or a low defect probability, while a non-zero offset registration vector 114 indicates a defect 39 or a high defect probability. The registration vectors may be guided from a subset in the reference image 66 to a corresponding subset of the observed imaging dataset 28, or vice versa, from a subset of the observed imaging dataset 28 to a corresponding subset of the reference image 66. Instead of calculating the tolerance statistic 92 based on the registration vector of the defect-free observed image dataset 30, the tolerance statistic 92 may be defined by a human, for example, a defect probability may be assigned based on the length of the registration vector, or a minimum length may be defined as a threshold, for example, a zero offset registration vector 112 indicates no defect, while a non-zero offset registration vector 114 exceeding the minimum length indicates a defect 39.
[0189] In an example of the first or second embodiment of the present invention, the plurality of characteristic elements 90 comprises a machine learning model, in particular a neural network 118 comprising an autoencoder, which is trained on the reference images 66 of the semiconductor structure, and the observed representations 88 of the subset of the imaging datasets 28 comprise the output of the machine learning model applied to the subset of the imaging datasets 28. Each defect-free representation 94 of the subset of the defect-free observed imaging datasets 30 comprises the output of the machine learning model applied to the subset of the defect-free observed imaging datasets 30. Fig.10 An example of a first embodiment of the present invention is shown. A machine learning model, such as a neural network 118, decodes a subset of the imaging data set 28 of the wafer 24 obtained in the imaging step 84 to obtain an observation representation 88 of the subset, which has, for example, the following reconstruction error:
[0190] Where r = C(x).
[0191] C(x) is the output of the machine learning model applied to the input x, i.e., a subset of the imaging data set 28. The machine learning model is trained by solving an optimization problem, e.g., training the neural network 118 to minimize the reconstruction error of a subset of the reference images 66. Further restrictions may be imposed on the machine learning model. A tolerance statistic 92 may be derived from the defect-free characterization 94, i.e., from the output of the machine learning model applied to a subset of the defect-free observed imaging data set 30. Based on this tolerance statistic 92, a label "defect" or "defect-free" or a defect probability may be assigned to a subset of the observed imaging data set 28 in the defect criteria verification step 86, thereby generating defect information.
[0192] The neural network 118 may include, for example, an autoencoder. The autoencoder learns a compressed internal representation of the defect-free reference image 66. As a result, the model is able to perfectly reconstruct the defect-free reference image 66. In contrast, the defect-free observed imaging data set 30 including the disturbance 34 and the defective subset of the observed imaging data set 28 are not completely reconstructed. However, the disturbance 34 can be distinguished from the defect 39 based on the tolerance statistics 92.
[0193] In an example of the first or second embodiment of the invention, the observed characterization 88 of the subset of the imaging data set 28 comprises coefficients of a decomposition 120 of the subset of the imaging data set 28 with respect to the number of characteristic elements 90, and wherein the defect-free characterization 94 of the subset of the defect-free observed imaging data set 30 comprises coefficients of a decomposition 120 of the subset with respect to the number of characteristic elements 90. Let Denote a vectorized subset of the imaging data set 28 of width w and height h, and let is a matrix containing n feature components of size w·h, and let The reconstruction error then measures the deviation between the subset and its decomposition 120, for example:
[0194]
[0195] The observed representation of the subset can be obtained by solving, for example, the following optimization problem 88:
[0196]
[0197] Fig.11 An example of a first embodiment of the present invention is shown. The feature elements include a dictionary 121 obtained by dictionary learning. For an observed representation 88 of a subset of the imaging data set 28 acquired in the imaging step 84, defect information can be generated. To this end, the subset is decomposed with respect to the dictionary elements, resulting in an observed representation 88 including coefficients of the decomposition 120. Based on the tolerance statistics 92 obtained from the defect-free representation 94 of the defect-free observed imaging data set 30, defects 39 can be detected in the defect standard verification step 86, for example by using the tolerance statistics P as a direct defect indicator
[0198] in or by using the tolerance statistic P as a prior, e.g. by
[0199]
[0200] And use the reconstruction error as a defect indicator.
[0201] In an example of the first or second embodiment of the invention, instead of using the observed imaging dataset itself, the method can directly operate on the difference between a subset of the observed imaging dataset and a corresponding subset of the reference image. In this way, only the difference image needs to be processed, which contains much less information than the observed imaging dataset. This reduces the complexity of the model, i.e., the decomposition, characteristic elements and representation, thereby improving the accuracy of the method. In an example of the first or second embodiment of the invention, therefore, the number of characteristic elements 90 and the tolerance statistics 92 are derived from the difference image of the subset of the defect-free observed imaging dataset 30 and the aligned subset of the reference image 66, while the observed representation 88 of the subset of the imaging dataset 28 includes the coefficients of the decomposition 120 of the difference image of the subset and the aligned reference image with respect to a plurality of characteristic elements 90, as well as the reconstruction error measure of the deviation between the subset and its decomposition 120.
[0202] The decomposition 120 may be linear or nonlinear. For example, the characteristic elements 90 may include elements of a basis, such as a wavelet basis or a Fourier basis. The characteristic elements 90 may also include a plurality of principal components obtained by principal component analysis, such as a subset of principal components. The characteristic elements 90 may include an overcomplete framework. The characteristic elements 90 may include a dictionary 121 obtained by dictionary learning.
[0203] The reference images 66x1, ..., x can be obtained by solving the following optimization problem, for example k A subset of (particularly the simulated aligned reference image 67) learns the dictionary C 121
[0204]
[0205] The L1 norm of the reference representation 104 enforces a sparse reconstruction of the subset x with respect to the dictionary elements (called atoms), that is, a reconstruction with a small number of non-zero elements, i.e., linear combinations of only a few atoms. This prevents the subsets with defects 39 from being reconstructed almost entirely based on a combination of many different atoms, thereby ensuring that these subsets still have large reconstruction errors, and therefore these subsets are marked as "defective". In this way, the accuracy of the defect detection method is improved.
[0206] The elements of dictionary 121 are constrained to have unit norm so that any scaling is contained in the reference representation 104. Without this constraint, additional regularization of the reference representation norm would be ineffective, since scaling can be contained in dictionary 121 and the reference representation norm can become arbitrarily small.
[0207] Since the dictionary 121 and the reference representation r are optimized simultaneously i104 is a non-convex problem, so an alternating optimization technique can be used. To this end, the problem is divided into a) updating the dictionary 121 of the fixed reference representation 104 and b) refining the reference representation 104 given the updated dictionary 121. Both problems are solved in an alternating manner using the alternating direction method of multipliers (ADMM). In the case of optimizing the dictionary elements, a constrained version of ADMM can be used to handle the constraints on the dictionary elements (unit or bounded norm), which is equivalent to projecting to the feasible set between every two iterations.
[0208] To compute the observation representation 88 of a subset of the imaging data set 28 based on the known dictionary 121 , an optimization technique depends on an optimization problem.
[0209] If the tolerance statistic P is used as a direct defect indicator:
[0210] in ADMM can be used for optimization.
[0211] In the case where the tolerance statistic P is used as a prior in the optimization problem and the prior contains a single-class SVM, a gradient descent step needs to be performed on the single-class SVM. For this purpose, the generalized proximal gradient method can be used, which is a combination of the generalized forward-backward splitting and the Chambolle-Pock optimization algorithm.
[0212] Rather than obtaining the tolerance statistics 92 solely from the defect-free representations 94, the tolerance statistics 92 may include a joint probability density function f(S, R) or a conditional probability density function f(S|R) obtained by density estimation techniques, where S includes the observed representations 88 of a subset of the observed imaging dataset 28 and / or the defect-free representations 94 of a subset of the defect-free observed imaging dataset 30, and where R includes the reference representations 104 of a subset of the reference images 66 with respect to a plurality of characteristic elements 90. The corresponding probability distributions P(S, R) or P(S|R) model the joint distribution or the conditional distribution, respectively, and thereby assign likelihoods to pairs of the reference images 66 and the subset of the observed images 28. The observed representations 88 of the subset of the observed imaging dataset 28 and the defect-free representations 94 of the subset of the defect-free observed imaging dataset 30 may be obtained based on a plurality of additional characteristic elements 91, which may be derived from the observed imaging dataset 28 and / or the defect-free observed imaging dataset 30. For example, let x1, ..., x2, ..., x3, ..., x4, ..., x5, ..., x6, ..., x7, ..., x8, ..., x9, ..., x1 ...1, ..., x2, ..., x3, ..., x4, ..., x5, ..., x6, ..., x8, ..., x9, ..., x1, ..., x1, ..., x1, ..., x1, ..., x1, ..., x1, ..., x1, ..., x1, ..., x1, ..., x1, ..., x1, ..., x1, ..., x1, k represents a subset of the reference image 66, and y1, ..., y m represents a subset of the observed imaging data set 28 and / or the defect-free observed imaging data set 30, the characteristic element 90C can be obtained by solving the following optimization problem:
[0213]
[0214] The additional characteristic element 91A can be obtained by solving the following optimization problem:
[0215]
[0216] For a known pair including a subset of the observed imaging dataset 28 and / or a subset of the defect-free observed imaging dataset 30x and a corresponding subset of the reference image 66y, the reference characterization 104 regarding the feature element 90C and the observed characterization 88 and / or the defect-free characterization 94 regarding the number of additional feature elements 91A may be obtained as follows:
[0217]
[0218] Then we can base it on the joint probability Or based on conditional probability and characteristic properties of the distribution, such as thresholds: or To detect defects39.
[0219] Alternatively, the distribution may be used as a prior in an optimization problem involving reconstruction errors for a subset x of the imaging dataset 28, where the corresponding subset of reference images 66y has the reference representation
[0220]
[0221] Then the reconstruction error can be To detect defects 39, for example, for a threshold τ:
[0222] If the reference image 66 includes the same or nearly the same semiconductor structure as the imaging datasets 28, 30, then the reference image 66 is a corresponding reference image 66 for the imaging datasets 28, 30. A subset of the reference images 66 is a corresponding subset of the reference images 66 for the subset of the imaging datasets 28, 30 if it includes the same or nearly the same semiconductor structure as the subset of the imaging datasets 28, 30.
[0223] According to one aspect of the example, the characteristic element 90 includes independent components obtained by independent component analysis. The characteristic element 90 may also include multiple image blocks obtained by an unsupervised clustering method (eg, by k-means, agglomerative clustering, or perceptually driven clustering, etc.).
[0224] In an example of the first or second embodiment of the present invention, the reference image 66 of the semiconductor structure includes a subset of the defect-free observed imaging data set 30 of the semiconductor structure. The reference image 66 of the semiconductor structure may also include a subset of defect-free generated images of the semiconductor structure, such as a composite image of the defect-free semiconductor structure. The generated defect-free image of the semiconductor structure may include a plurality of polygons representing the semiconductor structure, such as Figure 3 a and 5a. The defect-free generated image of the semiconductor structure may include an image generated from a defect-free CAD model of the wafer. In this case, a mask may be applied to the CAD model to ignore irrelevant portions of the CAD model, such as irrelevant portions, defective portions, or portions containing insufficient information.
[0225] By simulating the image acquisition process and the lithography process 10, the generated image can be simulated as an observed imaging data set 28 similar in appearance to the wafer 24. The simulated image can be calculated by a machine learning model. The reference image 66 can include a defect-free generated image of the semiconductor structure and a defect-free observed image of the semiconductor structure.
[0226] In an example of the first or second embodiment of the invention, the observed representation 88 of the subset of the observed imaging dataset 28 comprises spatial information about the position of the subset within the imaging dataset 28, and / or the defect-free representation 94 of the subset of the defect-free observed imaging dataset 30 comprises spatial information about the position of the subset within the defect-free observed imaging dataset 30, and / or the reference representation 104 of the subset of the reference image 66 comprises spatial information about the position of the subset within the reference image 66. For example, the pixel positions may be encoded in this way. To this end, the spatial information may comprise a position encoding, in particular a Fourier function of different frequencies.
[0227] Position encoding, also known as "Fourier features", is a popular technique for encoding spatial coordinates by generating position features as a set of sine and cosine waves with different frequencies. For example, the feature of a one-dimensional position x can be represented by the following vector:
[0228] (sin(x·π),cos(x·μ),sin(x·μ / 2),cos(x·μ / 2),sin(x·μ / 4),cos(x·μ / 4),…) T
[0229] The position encoding vector may include the same number of dimensions as the representation vector. The two vectors may be concatenated to form a single representation. For example, tolerance statistics 92 may be derived from these defect-free representations 94 including spatial information.
[0230] In an example of the first or second embodiment of the present invention, the subset includes a single pixel. The subset of the imaging data set 28 may also include a portion of the observed imaging data set 28. The subset of the defect-free observed imaging data set 30 may also include a portion of the defect-free observed imaging data set 30. The subset of the reference image 66 may include a portion of the reference image 66. An observed representation 88 of the subset of the imaging data set 28 may be obtained from a region of interest including the subset of the imaging data set 28. A defect-free representation 94 of the subset of the defect-free observed imaging data set 30 may be obtained from a region of interest including the subset of the defect-free observed imaging data set 30. A reference representation 104 of the subset of the reference image 66 may be obtained from a region of interest including the subset of the reference image 66.
[0231] The detected defects 39 may be classified according to the type of defect 39. In an example of the first or second embodiment of the present invention, therefore, the machine learning model is trained to assign a defect type from a predetermined set of defect types to an observed representation 88 of a subset of the imaging data set 28 of the wafer 24, the observed representation 88 being based on the number of characteristic elements 90. The machine learning model may be trained to assign a defect type (e.g. Figure 3 d) is associated with defect 39. Additionally or alternatively, defects may be labeled by a human using their defect type. Additionally or alternatively, defects may be labeled by a rule-based algorithm that applies predefined rules to known defects 39 to infer the type of defect 39. Depending on the type of defect 39, the defect 39 may be located directly to the corresponding hardware or system part, such as a bridge defect 42 or a line thinning defect 40 of an etching unit, a missing structure defect of an illumination unit, etc.
[0232] According to one aspect of the examples of the first or second embodiments of the present invention, information about the computer-implemented method, such as characteristic elements 90, tolerance statistics 92, reference images 66, defect-free observed imaging data set 30, or any other parameters of the defect detection method for future use cases or for analysis of defect standards or learning processes, can be stored. Intermediate results, such as characteristic elements, difference images, or reconstruction errors, can also be provided as input data to other methods. Fixed inputs such as reference images 66, characteristic elements 90, or tolerance statistics 92 can be provided by exchangeable hardware.
[0233] In any example of the first or second embodiment of the present invention, an imaging data set of a wafer can be obtained by a charged particle beam system. Charged particle beam systems include, but are not limited to, scanning electron microscopes (SEMs), focused ion beam microscopes, such as helium ion microscopes. Another example of a charged particle beam system is a corrected electron scanning microscope, which includes a correction device for correcting chromatic aberration and spherical aberration.
[0234] In order to present the input data, intermediate or final results to the user, a visualization device may be used. According to an example of the first or second embodiment of the present invention, the observed representation 88 of a subset of the imaging data set 28 of the wafer 24 and / or the defects 39 detected in the imaging data set 28 of the wafer 24 and / or the characteristic elements 90 are directed to a display device 136 or a dashboard for visualization. Characteristic elements 90 such as dictionaries including a plurality of atoms can be visualized, for example, by a heat map. The same is true for defect probabilities. In order to obtain an overview of the results, it is advantageous to visualize the inspection subset of the imaging data set 28 of the wafer 24 together with the characteristic elements 90, the observed representation 88 of the subset and the detected defects 39. In this way, the detected defects 39 can be monitored in real time. Alternatively, the data can be stored in a long-term memory for further analysis, for example, for generating statistics of the defects 39. In a further example, the identified defects 39 can be cached in memory for a specified time span, for example, 48 hours, so that the detected defects 39 can be further analyzed without requiring a large amount of memory.
[0235] Examples of the first or second embodiment of the present invention also include directing the defects 39 detected in the imaging data set 28 of the wafer 24 to a display device 136 or dashboard for visualization, wherein the detected defects 39 are highlighted or marked according to the type of defect 39. For example, a specific type of defect 39, such as a bridging defect 42, may be marked with a specific color or corresponding text.
[0236] For quality assurance or quality control processes, it is important to obtain further information about the detected defects. Therefore, examples of the first or second embodiment of the present invention may also include determining one or more measurements of the identified defects 39 in a subset of the imaging data set 28 of the wafer 24, in particular the size, area, dimension, shape parameter, distance, radius, aspect ratio, type, number of defects, density, spatial distribution of defects, whether there are any defects, etc. Such measurements may be obtained only for specific types of defects or only in specific areas of the imaging data set 28 that may be marked by a mask, for example, a boundary or a die area or a user-defined area.
[0237] For quality control, examples of the first or second embodiment of the present invention may also include evaluating the quality of the wafer 24 based on one or more measurements and at least one quality evaluation rule, for example, according to the DIN-ISO quality specification, which defines an upper limit for acceptable non-ideal wafers. For example, the density of a specific defect type at the nucleus should be less than 10 / nm 2 .
[0238] According to any of the various embodiments of the present invention, at least one wafer fabrication process parameter may be controlled based on one or more measurements of defects identified in an imaging dataset of a wafer.
[0239] The present invention also relates to a computer-readable medium having a computer program executable by a computer device stored thereon, the computer program comprising a program code for executing the method according to any embodiment of the present invention.
[0240] The invention also relates to a computer program product comprising instructions which, when executed by a computer, cause the computer to perform a method according to any one of the embodiments of the invention.
[0241] Fig.12 A system 122 is schematically shown that can be used to control the quality of wafers 24 produced in a semiconductor manufacturing plant. The system 122 includes an imaging device 124 and a processing device 126. The imaging device 124 is coupled to the processing device 126. The imaging device 124 is configured to acquire an imaging data set 28 of the wafer 24. The wafer 24 may include semiconductor structures, such as transistors such as field effect transistors, storage cells, etc. Exemplary embodiments of the imaging device 124 may be a SEM or a multi-beam SEM, a helium ion microscope (HIM), or a cross-beam device including a FIB and a SEM, or any charged particle imaging device.
[0242] The imaging device 124 may provide the imaging data set 28 to the processing device 126. The processing device 126 includes a processor, for example implemented as a CPU 128 or a GPU. The processor may receive the imaging data set 28 via an interface 130. The processor may load program code from a memory 132. The processor may execute the program code. When executing the program code, the processor performs techniques such as those described herein according to the first or second embodiment of the present invention, for example, defect detection, measuring detected defects, calculating observation representations 88 of a subset of the imaging data set 28 with respect to a plurality of characteristic elements 90, calculating tolerance statistics 92 from defect-free representations 94 of defect-free observed image data sets 30, and calculating characteristic elements 90 from reference images 66. For example, the processor may execute when loading the program code from the memory 132. Figure 7 or the computer-implemented method shown in 8. The processing device 126 may optionally include a user interface 134 for receiving user input, such as defect measurement types, quality assessment rules, parameters of machine learning models, simulation parameters, parameters for aligning the imaging data set 28 and the reference image 66, etc. The processing device 126 may optionally include a display device 136 for displaying defect detection results, input data, or intermediate results to a user, such as in real time or buffered.
[0243] Fig.13A system 140 is schematically shown that can be used to control the production of wafers 24 in a semiconductor manufacturing plant. The system 122 includes Fig.12 , and the above also applies to the various components here. In addition, the system 122 has a device 138 for producing a wafer 24 controlled by at least one wafer manufacturing process parameter. To this end, the imaging data set 28 is provided to the processing device 126 by the imaging device 124. The processor of the processing device 126 is configured to perform one of the disclosed methods, including controlling at least one wafer manufacturing process parameter based on one or more measured characteristics of the defects 39 identified in the imaging data set 28 of the wafer 24. For example, a detected bridge defect 42 indicates insufficient etching, so the etching amount is increased; a detected broken line defect 48 indicates excessive etching, so the etching amount is reduced; a persistent anomaly or defect 39 indicates that the mask 16 is defective, so the mask 16 must be checked; and a defect 39 caused by a missing structure suggests that the material deposition is not ideal, so the material deposition is modified.
[0244] Embodiments, examples and aspects of the invention may be described by the following clauses:
[0245] 1. A computer-implemented method 82 for detecting defects, comprising:
[0246] - acquiring an imaging data set 28 of a wafer 24, the wafer comprising a semiconductor structure;
[0247] - verifying defect criteria for defect detection in a subset of the imaging data set 28 of the wafer 24, the defect criteria comprising:
[0248] i. an observed representation 88 of the subset of the imaging data set 28 in relation to a plurality of characteristic elements 90 derived from the reference image 66 of the semiconductor structure, wherein the observed representation and said characteristic elements 90 define a reconstruction of a minimum reconstruction error of the subset of the imaging data set 28, and
[0249] ii. tolerance statistics 92 of a defect-free characterization 94 of a subset of the defect-free observed imaging data set 30 of the wafer 24, wherein each of the defect-free characterization and the characteristic elements 90 defines a reconstruction of a minimum reconstruction error of the subset of the defect-free imaging data set 30;
[0250] - generating defect information for the subset of the imaging data set 28 based on the defect criterion.
[0251] 2. A method as described in clause 1, wherein the observed representation 88 of the subset of the imaging data set 28 includes coefficients of a decomposition 120 of the subset of the imaging data set 28, which are related to the multiple characteristic elements 90, and wherein the defect-free representation 94 of the subset of the defect-free observed imaging data set 30 includes coefficients of a decomposition 120 of the subset of the defect-free observed imaging data set 30, which are related to the multiple characteristic elements 90.
[0252] 3. The method of clause 2, wherein the decomposition 120 is a linear decomposition 120.
[0253] 4. A method as described in clause 2 or 3, wherein the characteristic element 90 comprises an element of a base.
[0254] 5. The method of clause 4, wherein the characteristic elements 90 comprise elements of a wavelet basis.
[0255] 6. The method of clause 4 or 5, wherein the characteristic elements 90 comprise elements of a Fourier basis.
[0256] 7. The method according to any one of clauses 4 to 6, wherein the characteristic element 90 comprises a plurality of principal components obtained by principal component analysis.
[0257] 8. A method as described in any of clauses 2 to 7, wherein the property elements 90 include elements of an overcomplete framework.
[0258] 9. The method of any one of clauses 2 to 8, wherein the characteristic elements 90 comprise elements of a dictionary 121 acquired by dictionary learning.
[0259] 10. The method according to any one of clauses 2 to 9, wherein the characteristic element 90 comprises a plurality of independent components obtained by independent component analysis.
[0260] 11. The method according to any one of clauses 2 to 10, wherein the characteristic element 90 comprises a plurality of image-patches obtained by an unsupervised clustering method.
[0261] 12. A method as described in clause 1, wherein the observed representation 88 of the subset of the imaging dataset 28 includes a registration vector indicating an offset between the subset of the imaging dataset 28 and a characteristic element 90 in the form of a corresponding subset of the reference image 66, so that the corresponding subset of the reference image 66 is aligned with the subset of the imaging dataset 28 through the registration vector, and wherein the defect-free representation 94 of the subset of the defect-free observed imaging dataset 30 includes a registration vector indicating an offset between the subset of the defect-free observed imaging dataset 30 and the characteristic element 90 in the form of a corresponding subset of the reference image 66, so that the corresponding subset of the reference image 66 is aligned with the subset of the defect-free observed imaging dataset 30 through the registration vector.
[0262] 13. A method as described in claim 12, wherein the reconstruction error of a subset of the imaging data set 28 includes a distortion error between the subset of the imaging data set 28 and the corresponding subset of the reference image 66, and wherein the reconstruction error of the defect-free representation 94 of the subset of the defect-free observation imaging data set 30 includes a distortion error between the subset of the defect-free observation imaging data set 30 and the corresponding subset of the reference image 66.
[0263] 14. A method as described in claim 1, wherein the plurality of characteristic elements 90 include a machine learning model trained on the reference image 66 of the semiconductor structure, and wherein the observed representation 88 of the subset of the imaging data set 28 includes the output of the machine learning model when applied to the subset of the imaging data set 28, and wherein the defect-free representation 94 of the subset of the defect-free observed imaging data set 30 includes the output of the machine learning model when applied to the subset of the defect-free observed imaging data set 30.
[0264] 15. A method as described in clause 14, wherein the machine learning model includes a neural network 118.
[0265] 16. The method of any of the preceding clauses, wherein the defect criterion comprises detecting defects 39 in the subset of the imaging data set 28 based on a statistical property of the obtained observation representation 88 in relation to the tolerance statistic 92.
[0266] 17. A method as described in clause 16, wherein the statistical property comprises a quantile of the tolerance statistic 92, in particular a threshold value.
[0267] 18. A method as described in clause 16 or 17, wherein the statistical property includes a confidence interval.
[0268] 19. The method according to any one of clauses 16 to 18, wherein the statistical property comprises a moment of the tolerance statistic 92, in particular a mean value and / or a variance.
[0269] 20 . The method of any of the preceding clauses, wherein the observed representation 88 of the subset of the imaging data set 28 is obtained by solving an optimization problem including the reconstruction error and including a priori tolerance statistics 92 on the defect-free representation 94 .
[0270] 21. The method of clause 20, wherein the defect criterion comprises detecting defects 39 in the subset of the acquired imaging data sets 28 based on the reconstruction error of the solution to the optimization problem.
[0271] 22. The method of any of the preceding clauses, wherein the tolerance statistics 92 comprise a probability density function obtained from the defect-free representation 94 of the defect-free observed imaging dataset 30 by a density estimation technique.
[0272] 23. The method of clause 22, wherein the probability density function of the tolerance statistic 92, in particular the probability density function of a Gaussian model or a Gaussian mixture model, is obtained by a parameter density estimation technique.
[0273] 24. The method of clause 22, wherein the probability density function of the tolerance statistic 92 is obtained by a non-parametric density estimation technique, in particular a Parzen density estimator.
[0274] 25. A method as described in any of the preceding clauses, wherein the tolerance statistics 92 include a machine learning model trained on the defect-free representations 94 of the subset of the defect-free observed imaging data set 30, in particular a single-class SVM or support vector data description.
[0275] 26 . The method of any of the preceding clauses, wherein the tolerance statistics 92 include only a subset of dimensions of the defect-free representations 94 of the subset of the defect-free observed imaging data sets 30 .
[0276] 27. The method of any of the preceding clauses, wherein the tolerance statistics 92 comprise separate tolerance statistics for each dimension in a subset of dimensions of the defect-free representation 94 of the subset of the defect-free observed imaging data sets 30.
[0277] 28. The method of any of the preceding clauses, wherein the reference image 66 of the semiconductor structure comprises a subset of the defect-free observed imaging data set 30 of the semiconductor structure.
[0278] 29. The method of any of the preceding clauses, wherein the reference image 66 of the semiconductor structure comprises a subset of defect-free generated images of the semiconductor structure.
[0279] 30. The method of clause 29, wherein the defect-free generated image of the semiconductor structure comprises a composite image of the defect-free semiconductor structure.
[0280] 31. The method of clause 29 or 30, wherein the defect-free generated image of the semiconductor structure comprises a plurality of polygons representing the semiconductor structure.
[0281] 32. The method of any one of clauses 29 to 31, wherein the defect-free generated image of the semiconductor structure comprises an image generated from a defect-free CAD model of the wafer.
[0282] 33. The method of any one of clauses 29 to 32, wherein the generated image is simulated by simulating an image acquisition process and a lithography process to have a similar appearance to the observed imaging dataset 28 of the wafer 24.
[0283] 34. The method of any of the preceding clauses, wherein the reference image 66 comprises a defect-free generated image of a semiconductor structure and a defect-free observed image of the semiconductor structure.
[0284] 35. A method as described in any of the preceding clauses, wherein the reference image 66 is aligned.
[0285] 36. A method as described in any of the preceding clauses, wherein the observed representation 88 of the subset of the observed imaging dataset 28 includes spatial information about the position of the subset in the imaging dataset 28, and wherein the defect-free representation 94 of the subset of the defect-free observed imaging dataset 30 includes spatial information about the position of the subset in the defect-free observed imaging dataset 30.
[0286] 37. The method of clause 36, wherein the spatial information comprises a position code comprising Fourier functions of different frequencies.
[0287] 38. A method as described in any of the preceding clauses, wherein the subset comprises a single pixel.
[0288] 39. A method as described in any of the preceding clauses, wherein the observation representation 88 of the subset of the imaging dataset 28 is obtained from a region of interest of the subset of the imaging dataset 28, and wherein the defect-free representation 94 of the subset of the defect-free observation imaging dataset 30 is obtained from a region of interest of the subset of the defect-free observation imaging dataset 30.
[0289] 40. A method as described in any of the preceding clauses, wherein the defect standard further includes modifying the defect detection result through a trained machine learning model 95.
[0290] 41. The method of any of the preceding clauses, further comprising modifying intermediate results of the computer-implemented method for defect detection 82 via a trained machine learning model 95.
[0291] 42. A method as described in claim 41, wherein the trained machine learning model 95 is applied to the subset of the imaging dataset 28, and / or to the difference between the subset of the imaging dataset 28 and the aligned reference image 66, and / or to the reconstruction error of the observation representation 88 of the subset of the imaging dataset 28, wherein the aligned reference image 66 is particularly a simulated aligned reference image 67.
[0292] 43. A method as described in any of clauses 40 to 42, wherein the trained machine learning model 95 comprises an autoencoder.
[0293] 44. A method as described in any of clauses 40 to 43, wherein the trained machine learning model 95 includes a segmentation model.
[0294] 45. A method as described in any of the preceding clauses, wherein a machine learning model is trained to assign a defect type from a predefined set of defect types to a subset of the imaging data set 28 of the chip 24, and to communicate the defect 39 to the specific hardware unit responsible for the defect.
[0295] 46. The method according to any of the preceding clauses, wherein the imaging dataset 28 is acquired by means of a charged particle beam system, in particular a multi-beam scanning electron microscope.
[0296] 47. The method as described in any of the preceding clauses further includes directing the observed representations 88 and / or characteristic elements 90 of a subset of the imaging data set 28 of the chip 24 and / or the detected defects 39 in the imaging data set 28 of the chip 24 to a display device 136 or a dashboard for visualization.
[0297] 48. The method as described in any of the preceding clauses, further comprising directing the detected defects 39 in the imaging data set 28 of the wafer 24 to a display device 136 or dashboard for visualization, wherein the detected defects 39 are highlighted or marked according to the defect type.
[0298] 49. The method of any of the preceding clauses, wherein the reference image 66, the characteristic elements 90 and / or the tolerance statistics 92 are provided by replaceable hardware.
[0299] 50. A computer-implemented method 96 for obtaining tolerance statistics 92 on defect-free representations 94 of a subset of defect-free observed imaging data sets 30 of a wafer 24, the method comprising the steps of:
[0300] i. Acquire a defect-free observation imaging data set 30 of the wafer 24 including the semiconductor structure;
[0301] ii. generating a defect-free representation 94 of a subset of the defect-free observed imaging data set 30 of the wafer 24, which is related to a plurality of characteristic elements 90 derived from the reference image 66 of the semiconductor structure, wherein the defect-free representation 94 and each of the characteristic elements 90 define a reconstruction of a minimum reconstruction error of the subset of the defect-free observed imaging data set 30;
[0302] iii. Obtaining tolerance statistics 92 on the defect-free representation 94 .
[0303] 51. The method of clause 50, further comprising, before step ii, obtaining characteristic elements 90 from a reference image 66 of the semiconductor structure by solving an optimization problem for a minimum reconstruction error of a reconstruction comprising the reference image 66, said reconstruction being defined by a reference representation 104 and said characteristic elements 90.
[0304] 52. The method of clause 51, wherein the optimization problem includes at least one constraint or prior on the property element 90.
[0305] 53. A method as described in clause 52, wherein the constraint or prior relates to the sparsity of the characteristic element 90, in particular the L0 norm or L1 norm or the kurtosis of the characteristic element 90.
[0306] 54. A method as described in any of clauses 51 to 53, wherein the optimization problem includes at least one constraint or prior on the reference representation 104.
[0307] 55. A method as described in clause 54, wherein the constraint or prior is a measure of the sparsity of the reference representation 104, in particular the L0 norm or L1 norm or the kurtosis of the reference representation 104.
[0308] 56. The method as described in any of the preceding clauses further includes determining one or more measurement results of the identified defects (39) in a subset of the imaging data set 28, in particular the size, area, dimension, shape parameter, distance, radius, aspect ratio, type, number of defects, density, spatial distribution of defects, existence of defects, etc.
[0309] 57. The method of clause 56, further comprising assessing a quality of the wafer based on the one or more measurement results and at least one quality assessment rule.
[0310] 58. The method of clause 56, further comprising controlling at least one wafer fabrication process parameter based on one or more measurements of defects identified in the imaging dataset 28.
[0311] 59. A computer readable medium having stored thereon a computer program executable by a computing device, the computer program comprising code for performing the method according to any one of clauses 1 to 58.
[0312] 60. A computer program product comprising instructions which, when executed by a computer, cause the computer to perform the method according to any one of clauses 1 to 58.
[0313] 61. A system 122 for controlling the quality of wafers 24 produced in a semiconductor manufacturing plant, the system 122 comprising:
[0314] An imaging device 124 adapted to provide an imaging data set 28 of the wafer 24;
[0315] - one or more processing devices 126;
[0316] - One or more machine-readable hardware storage devices comprising instructions executable by one or more processing devices 126 to perform operations including the method described in clause 57.
[0317] 62. A system 140 for controlling the production of wafers 24 in a semiconductor manufacturing plant, the system 140 comprising:
[0318] - means (138) for producing wafers 24, which are controlled by at least one manufacturing process parameter;
[0319] An imaging device 124 adapted to provide an imaging data set 28 of the wafer 24;
[0320] - one or more processing devices 126;
[0321] - One or more machine-readable hardware storage devices comprising instructions executable by one or more processing devices 126 to perform operations including the method described in clause 58.
[0322] 63. The system 122 , 140 of clause 61 or 62 , further comprising a display device 136 .
[0323] 64. The system 122 , 140 of any one of clauses 61 to 63 , further comprising a user interface 134 .
[0324] Reference numerals list
[0325] 10 Photolithography process
[0326] 12 substrates
[0327] 14 Photoresist
[0328] 15. Radiation
[0329] 16 Mask
[0330] 18 Etching
[0331] 20. Cleaning
[0332] 22 Inspection process
[0333] 24 chips
[0334] 26 Microscope
[0335] 28 Imaging Datasets
[0336] 30 defect-free observation imaging datasets
[0337] 32 Misleading Imaging Datasets
[0338] 34 Interference
[0339] 36 lines shortened
[0340] 37 Line Thinning
[0341] 38 Edge Roughness
[0342] 39 Defects
[0343] 40 line thinning defect
[0344] 42 Bridging Defects
[0345] 44 Long bridging defect
[0346] 46 Intrusion Defects
[0347] 48 wire breakage defect
[0348] 50 offset defect
[0349] 52 line pullback defect
[0350] 54 Stray structural defects
[0351] 56 Bare core to database workflow
[0352] 58 Rasterization Steps
[0353] 60 Anchor Point Steps
[0354] 62 Alignment Steps
[0355] 64 simulation steps
[0356] 66 reference images
[0357] 67 Aligned Simulation Reference Images
[0358] 68 Differentiation Steps
[0359] 70 Post-processing steps
[0360] 72 Defect suggestions
[0361] 74 Mark
[0362] 76 Local Environment
[0363] 78 Mark
[0364] 80 Local Environment
[0365] 82 Computer-implemented methods
[0366] 84 Imaging Steps
[0367] 86 Defect Standard Verification Steps
[0368] 88 Observational Characterization
[0369] 90 characteristic elements
[0370] 91 Additional characteristic elements
[0371] 92 Tolerance Statistics
[0372] 94 Defect-free characterization
[0373] 95 Machine Learning Models
[0374] 96Computer-implemented method
[0375] 98 Imaging Steps
[0376] 99 Characteristic Elements Steps
[0377] 100 Characterization Generation Steps
[0378] 102 Tolerance Statistics Steps
[0379] 104 Reference Characterization
[0380] 106 simulation reference images
[0381] 108 Registration Steps
[0382] 112 Zero offset registration vector
[0383] 114 Non-zero offset registration vector
[0384] 116 Optional
[0385] 118 Neural Network
[0386] 120 Decomposition
[0387] 122 System
[0388] 124 Imaging device
[0389] 126 Processing Device
[0390] 128CPU
[0391] 130 interface
[0392] 132 Memory
[0393] 134 User Interface
[0394] 136 Display devices
[0395] 138 devices
Claims
1. A computer-implemented method (82) for detecting defects, comprising: - acquiring an imaging data set (28) of a wafer (24) comprising a semiconductor structure; - verifying defect criteria for defect detection in a subset of the imaging data set (28) of the wafer (24), the defect criteria comprising: i. an observed representation (88) of the subset of the imaging data set (28) in relation to a characteristic element (90) derived from a reference image (66) of the semiconductor structure, wherein the observed representation and the characteristic element (90) define a reconstruction of the subset of the imaging data set (28) with a minimum reconstruction error, and ii. tolerance statistics (92) of defect-free characterizations (94) of a subset of defect-free observed imaging data sets (30) for a wafer (24), wherein each of the defect-free characterizations and the characteristic elements (90) defines a reconstruction of a minimum reconstruction error of the subset of defect-free imaging data sets (30); - generating defect information for the subset of the imaging data set (28) based on the defect criterion.
2. The method of claim 1, wherein the observed representation (88) of the subset of the imaging data set (28) comprises coefficients of a decomposition (120) of the subset of the imaging data set (28) relating to a plurality of characteristic elements (90), and wherein the defect-free representation (94) of the subset of the defect-free observed imaging data set (30) comprises coefficients of a decomposition (120) of the subset of the defect-free observed imaging data set (30) relating to a plurality of characteristic elements (90).
3. The method of claim 2, wherein the decomposition (120) is a linear decomposition (120).
4. A method as claimed in claim 2 or 3, wherein the characteristic elements (90) comprise elements of a basis.
5. The method of claim 4, wherein the characteristic elements (90) comprise elements of a wavelet basis.
6. A method as claimed in claim 4 or 5, wherein the characteristic elements (90) comprise elements of a Fourier basis.
7. The method according to any one of claims 4 to 6, wherein the characteristic elements (90) include a plurality of principal components obtained by principal component analysis.
8. The method of any one of claims 2 to 7, wherein the property elements (90) comprise elements of an overcomplete framework.
9. The method according to any one of claims 2 to 8, wherein the characteristic elements (90) include elements of a dictionary (121) acquired by dictionary learning.
10. The method according to any one of claims 2 to 9, wherein the characteristic element (90) comprises a plurality of independent components obtained by independent component analysis.
11. The method according to any one of claims 2 to 10, wherein the characteristic element (90) comprises a plurality of image patches obtained by an unsupervised clustering method.
12. The method of claim 1, wherein the observed representation (88) of the subset of the imaging data set (28) includes a registration vector indicating an offset between the subset of the imaging data set (28) and a characteristic element (90) in the form of a corresponding subset of a reference image (66), such that the corresponding subset of the reference image (66) is registered with the subset of the imaging data set (28) by the registration vector, and wherein the defect-free representation (94) of the subset of the defect-free observed imaging data set (30) includes a registration vector indicating an offset between the subset of the defect-free observed imaging data set (30) and a characteristic element (90) in the form of a corresponding subset of the reference image (66), such that the corresponding subset of the reference image (66) is registered with the subset of the defect-free observed imaging data set (30) by the registration vector.
13. The method of claim 12, wherein a reconstruction error of a subset of the imaging data set (28) comprises a distortion error between the subset of the imaging data set (28) and a corresponding subset of the reference image (66), and wherein a reconstruction error of the defect-free representations (94) of a subset of the defect-free observed imaging data set (30) comprises a distortion error between the subset of the defect-free observed imaging data set (30) and the corresponding subset of the reference image (66).
14. The method of claim 1, wherein the plurality of characteristic elements (90) comprises a machine learning model trained on the reference image (66) of the semiconductor structure, and wherein the observed representation (88) of the subset of the imaging data set (28) comprises an output of the machine learning model when applied to the subset of the imaging data set (28), and wherein the defect-free representation (94) of the subset of the defect-free observed imaging data set (30) comprises an output of the machine learning model when applied to the subset of the defect-free observed imaging data set (30).
15. The method of claim 14, wherein the machine learning model comprises a neural network (118).
16. A method as claimed in any preceding claim, wherein the defect criteria comprises detecting defects (39) in the subset of the imaging data set (28) based on a statistical property of the obtained observation representation (88) with respect to the tolerance statistic (92).
17. The method as claimed in claim 16, wherein the statistical property comprises a quantile of the tolerance statistic (92), in particular a threshold value.
18. The method of claim 16 or 17, wherein the statistical property comprises a confidence interval.
19. The method as claimed in any one of claims 16 to 18, wherein the statistical property comprises a moment of the tolerance statistic (92), in particular a mean value and / or a variance.
20. A method as claimed in any one of the preceding claims, wherein the observed representation (88) of the subset of the imaging data set (28) is obtained by solving an optimization problem including the reconstruction error and including a priori the tolerance statistics (92) on the defect-free representation (94).
21. The method of claim 20, wherein the defect criterion comprises detecting defects (39) in the subset of the acquired imaging data sets (28) based on a reconstruction error of a solution to the optimization problem.
22. The method of any preceding claim, wherein the tolerance statistics (92) comprises a probability density function derived from the defect-free representation (94) of the defect-free observed imaging data set (30) by a density estimation technique.
23. The method as claimed in claim 22, wherein a probability density function of the tolerance statistic (92), in particular a probability density function of a Gaussian model or a Gaussian mixture model, is obtained by means of a parameter density estimation technique.
24. The method of claim 22, wherein the probability density function of the tolerance statistic (92) is obtained by a non-parametric density estimation technique, in particular a Parzen density estimator.
25. A method as described in any of the preceding claims, wherein the tolerance statistics (92) include a machine learning model trained on the defect-free representation (94) of the subset of the defect-free observation imaging data set (30), in particular a single-class SVM or a support vector data description.
26. The method of any preceding claim, wherein the tolerance statistics (92) comprises only a subset of dimensions of the defect-free representations (94) of the subset of the defect-free observed imaging data sets (30).
27. The method of any preceding claim, wherein the tolerance statistics (92) comprise separate tolerance statistics for each dimension in a subset of dimensions of the defect-free representation (94) of the subset of the defect-free observed imaging data set (30).
28. The method of any preceding claim, wherein the reference image (66) of the semiconductor structure comprises a subset of the defect-free observed imaging data set (30) of the semiconductor structure.
29. The method of any of the preceding claims, wherein the reference image (66) of the semiconductor structure comprises a subset of defect-free generated images of the semiconductor structure.
30. The method of claim 29, wherein the defect-free generated image of the semiconductor structure comprises a composite image of the defect-free semiconductor structure.
31. The method of claim 29 or 30, wherein the defect-free generated image of the semiconductor structure comprises a plurality of polygons representing the semiconductor structure.
32. The method of any one of claims 29 to 31, wherein the defect-free generated image of the semiconductor structure comprises an image generated from a defect-free CAD model of the wafer.
33. A method as claimed in any one of claims 29 to 32, wherein the generated image is simulated by simulating an image acquisition process and a lithography process to have a similar appearance to an observed imaging dataset (28) of the wafer (24).
34. The method of any of the preceding claims, wherein the reference image (66) comprises a defect-free generated image of a semiconductor structure and a defect-free observed image of the semiconductor structure.
35. A method as claimed in any preceding claim, wherein the reference images (66) are aligned.
36. A method as described in any of the preceding claims, wherein the observed representation (88) of the subset of the observed imaging data set (28) includes spatial information about the position of the subset in the imaging data set (28), and wherein the defect-free representation (94) of the subset of the defect-free observed imaging data set (30) includes spatial information about the position of the subset in the defect-free observed imaging data set (30).
37. The method of claim 36, wherein the spatial information comprises a position code comprising Fourier functions of different frequencies.
38. A method as claimed in any preceding claim, wherein the subset comprises a single pixel.
39. A method as described in any of the preceding claims, wherein the observed representation (88) of the subset of the imaging data set (28) is obtained from a region of interest of the subset of the imaging data set (28), and wherein the defect-free representation (94) of the subset of the defect-free observation imaging data set (30) is obtained from a region of interest of the subset of the defect-free observation imaging data set (30).
40. The method of any preceding claim, wherein the defect criteria further comprises modifying the defect detection results via a trained machine learning model (95).
41. The method of any of the preceding claims, further comprising modifying intermediate results of the computer-implemented method for defect detection (82) via a trained machine learning model (95).
42. A method as claimed in claim 41, wherein the trained machine learning model (95) is applied to the subset of the imaging dataset (28), and / or to the difference between the subset of the imaging dataset (28) and an aligned reference image (66), and / or to the reconstruction error of the observed representation (88) of the subset of the imaging dataset (28), wherein the aligned reference image (66) is in particular a simulated aligned reference image (67).
43. A method as claimed in any one of claims 40 to 42, wherein the trained machine learning model (95) comprises an autoencoder.
44. A method as claimed in any one of claims 40 to 43, wherein the trained machine learning model (95) includes a segmentation model.
45. A method as described in any of the preceding claims, wherein a machine learning model is trained to assign a defect type from a predefined set of defect types to a subset of an imaging data set (28) of a wafer (24), and to communicate the defect (39) to a specific hardware unit responsible for the defect.
46. The method as claimed in any of the preceding claims, wherein the imaging data set (28) is acquired by means of a charged particle beam system, in particular a multi-beam scanning electron microscope.
47. The method as claimed in any of the preceding claims further comprises directing the observed representations (88) and / or characteristic elements (90) of a subset of the imaging data set (28) of the chip (24) and / or the detected defects (39) in the imaging data set (28) of the chip (24) to a display device (136) or a dashboard for visualization.
48. The method as claimed in any of the preceding claims, further comprising directing the detected defects (39) in the imaging data set (28) of the chip (24) to a display device (136) or a dashboard for visualization, wherein the detected defects (39) are highlighted or marked according to the defect type.
49. A method as claimed in any preceding claim, wherein the reference image (66), the characteristic element (90) and / or the tolerance statistic (92) are provided by means of replaceable hardware.
50. A computer-implemented method (96) for obtaining tolerance statistics (92) on defect-free representations (94) of a subset of a defect-free observed imaging data set (30) of a wafer (24), the method comprising the steps of: i. Acquiring a defect-free observation imaging data set (30) of a wafer (24) including a semiconductor structure; ii. generating a defect-free characterization (94) of a subset of the defect-free observed imaging data set (30) of the wafer (24) in relation to characteristic elements (90) derived from a reference image (66) of the semiconductor structure, wherein each of the defect-free characterization (94) and the characteristic elements (90) defines a reconstruction of a minimum reconstruction error of the subset of the defect-free observed imaging data set (30); as well as iii. Obtaining tolerance statistics (92) on the defect-free representation (94).
51. The method of claim 50, further comprising, before step ii, obtaining characteristic elements (90) from a reference image (66) of the semiconductor structure by solving an optimization problem for a minimum reconstruction error of a reconstruction including the reference image (66), wherein the reconstruction is defined by a reference representation (104) and the characteristic elements (90).
52. The method of claim 51, wherein the optimization problem includes at least one constraint or prior on a property element (90).
53. A method as claimed in claim 52, wherein the constraint or prior relates to the sparsity of the characteristic element (90), in particular the L0 norm or the L1 norm or the kurtosis of the characteristic element (90).
54. The method of any one of claims 51 to 53, wherein the optimization problem comprises at least one constraint or prior on a reference representation (104).
55. A method as claimed in claim 54, wherein the constraint or prior is a measure of the sparsity of the reference representation (104), in particular the L0 norm or the L1 norm or the kurtosis of the reference representation (104).
56. The method as claimed in any of the preceding claims further comprises determining one or more measurement results of the identified defects (39) in a subset of the imaging data set (28), in particular size, area, dimension, shape parameter, distance, radius, aspect ratio, type, number of defects, density, spatial distribution of defects, presence of defects, etc.
57. The method of claim 56, further comprising assessing the quality of the wafer based on the one or more measurement results and at least one quality assessment rule.
58. The method of claim 56, further comprising controlling at least one wafer fabrication process parameter based on one or more measurements of defects identified in the imaging data set (28).
59. A computer readable medium having stored thereon a computer program executable by a computing device, the computer program comprising codes for executing the method according to any one of claims 1 to 58.
60. A computer program product comprising instructions which, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 58.
61. A system (122) for controlling the quality of wafers (24) produced in a semiconductor manufacturing plant, the system (122) comprising: - an imaging device (124) adapted to provide an imaging data set (28) of the wafer (24); - one or more processing devices (126); - One or more machine-readable hardware storage devices comprising instructions executable by one or more processing devices (126) to perform operations including the method of claim 57.
62. A system (140) for controlling production of wafers (24) in a semiconductor manufacturing plant, the system (104) comprising: - an apparatus (138) for producing a wafer (24), which is controlled by at least one manufacturing process parameter; - an imaging device (124) adapted to provide an imaging data set (28) of the wafer (24); - one or more processing devices (126); - One or more machine-readable hardware storage devices comprising instructions executable by one or more processing devices (126) to perform operations including the method of claim 58.
63. The system of claim 61 or 62, further comprising a display device (136).
64. The system of any one of claims 61 to 63, further comprising a user interface (134).
Citation Information
Patent Citations
Method and system for generating a synthetic image of a region of an object
US10504692B2
Deep learning based defect detection
US20220044391A1
Automatic referencing for computer vision applications
US6678404B1
Learning device, image inspection device, learned parameter, learning method, and image inspection method
WO2021181749A1
Cited By
Wafer quality analysis method based on image recognition
CN120894675A