Training and using neural networks for defect detection

By training the neural network model, the semiconductor sample images are transformed from the input space to the latent space, solving the problem of high precision and high uniformity monitoring of sub-micron feature structures in semiconductor manufacturing, and improving defect detection accuracy and yield.

CN120471131APending Publication Date: 2025-08-12APPL MATERIALS ISRAEL LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510156426.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-12
Filing Date
2025-02-12
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently monitor the high precision and uniformity of submicron feature structures during semiconductor manufacturing, resulting in insufficient equipment reliability and yield.

Method used

By training the neural network model, the semiconductor sample image elements in the input space are transformed into the latent space, and the expected probability function and training loss value L are used to optimize element allocation to achieve efficient defect detection and classification.

Benefits of technology

It improves the accuracy and efficiency of defect detection in the semiconductor manufacturing process, ensures high accuracy and uniformity of the equipment characteristic structure, and improves the yield rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471131A_ABST
    Figure CN120471131A_ABST
Patent Text Reader

Abstract

A system is provided for training a model representing elements in an input space, the elements in the input space each having N dimensions and being associated with an image of a semiconductor sample, as a potential space representing an equal number of elements, the elements in the potential space each having M (M < = N) dimensions. The system includes a processor configured to obtain a desired probability function for transforming elements in an input space into clusters of elements in a potential space. Then, the desired probability function is used to repeatedly transform elements in the input space into equal elements in the potential space that conform to an actual probability function indicating an actual allocation of elements to the cluster (s) until specified criteria are met. Finally, a training loss value L associated with an element in the potential space is determined and a test is made as to whether the training loss value L satisfies a specified criterion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The presently disclosed subject matter relates generally to the field of training and using neural networks for defect detection. Background Art

[0002] The current demand for high density and high performance associated with ultra-large-scale integration of manufactured devices requires submicron features, increased transistor and circuit speeds, and enhanced reliability. As semiconductor processes advance, pattern sizes such as line widths and other types of critical dimensions continue to shrink. This demand necessitates the formation of device features with high precision and uniformity, which in turn necessitates careful monitoring of the manufacturing process, including automated inspection of the devices while they are still in the form of semiconductor wafers.

[0003] During or after fabrication of a sample to be inspected, inspection can be performed using non-destructive inspection tools. Inspection generally involves generating some output (e.g., an image, a signal, etc.) of the sample by directing light or electrons to and detecting the light or electrons from the wafer. Examples of various non-destructive inspection tools include, by way of non-limiting example, scanning electron microscopes, atomic force microscopes, optical inspection tools, and the like.

[0004] The inspection process may include multiple inspection steps. The semiconductor device manufacturing process may include various processes such as etching, deposition, planarization, growth such as epitaxial growth, implantation, etc. The inspection step may be performed multiple times, for example, after certain process steps and / or after fabrication of certain layers. Additionally or alternatively, each inspection step may be repeated multiple times, for example, for different wafer locations, or for the same wafer location with different inspection settings.

[0005] Inspection processes are used at various steps during semiconductor manufacturing to perform, for example, defect-related operations. The efficiency of inspection can be improved by automating certain processes, such as defect detection, automatic defect classification (ADC), automatic defect review (ADR), image segmentation, and / or other operations. Automated inspection systems ensure that manufactured parts meet expected quality standards and provide useful information about possible adjustments to manufacturing tools, equipment, and / or components based on the identified error types, to promote higher yields. Summary of the Invention

[0006] According to one aspect of the present invention, a system is provided for training a model representing a plurality of elements in an input space into a latent space representing an equal plurality of elements, wherein the plurality of elements in the input space each have N dimensions and are associated with at least one image of a semiconductor sample, and the plurality of elements in the latent space each have M (M≤N) dimensions, the system comprising a processor and memory circuitry (PMC), the PMC configured to:

[0007] a) obtaining a desired probability function for transforming an element in the input space into one or more corresponding clusters of elements in the latent space;

[0008] b) using the desired probability function to repeatedly transform elements in the input space into an equal plurality of elements in the latent space that conform to the actual probability function until a specified criterion is satisfied, the actual probability function indicating an actual assignment of the elements to one or more corresponding clusters;

[0009] c) determining a training loss value L associated with the element in the latent space and testing whether the training loss value L satisfies a specified criterion; determining the training loss value L based on at least:

[0010] a. The first item L Rec , where the first term indicates the distance between an element in the input space and an element in the output space reconstructed from an element in the latent space; and

[0011] b. The second item L Prob , the second term indicates the statistical distance between the expected probability function and the actual probability function.

[0012] In addition to the above features, the system according to the aspects of the presently disclosed subject matter may include one or more of the features (i) to (xi) listed below in any desired combination or arrangement that is technically possible:

[0013] (i) wherein the training loss value L satisfies the following equation:

[0014] α*L Rec +(1-α)*L Prob .

[0015] (ii) wherein the system facilitates more efficient analysis of elements associated with actual probability functions that are sufficiently similar to desired probability functions in the latent space, rather than performing hypothetical analysis of elements in the input space that inherently do not conform to the specified desired probability function.

[0016] (iii) wherein the training is semi-supervised learning or fully supervised learning, such that at least some of the plurality of elements in the input space are labeled with corresponding classes from at least two classes for transforming the elements in the input space into elements in one or more corresponding clusters of elements in the latent space.

[0017] (iv) wherein elements in the input space are assigned to a certain number (I ≥ 2) of mutually discriminable clusters in the latent space, and the (PMC) is further configured as:

[0018] a. Obtain data indicating a certain number (K ≥ I) of classes, each of which is associated with each

[0019] The corresponding clusters of the clusters are associated, thereby generating I clusters;

[0020] b. labeling each element of at least a subset of the input space with a selected class from the K classes;

[0021] c. The training loss value L is also determined based on the following terms:

[0022] The third item L Sup , the third item indicates the degree of allocation of the transformed elements labeled with a class to the corresponding cluster in I clusters, each of the classes being associated with a corresponding group in I class groups, so that the fewer the transformed elements are allocated to clusters other than the corresponding cluster in the I clusters, L Sup The lower the value.

[0023] (v) wherein the training loss value L satisfies the following equation: α*L Rec +β*L Prob +(1-α-β)*L Sup .

[0024] (vi) wherein elements in the input space are assigned to a certain number (I ≥ 2) of mutually discriminable clusters in the latent space, and the (PMC) is further configured as:

[0025] d. obtain data indicating a certain number (k < I) of classes;

[0026] e. labeling each element of at least one subset of the input space with a selected class from the K classes;

[0027] f. The training loss value L is also determined based on the following terms:

[0028] The third item L Sup , the third item indicates the degree of allocation of the transformed elements labeled with k classes to the corresponding clusters in I clusters, so that the fewer the transformed elements allocated to clusters other than the corresponding cluster in I clusters, L Sup The lower the value.

[0029] (vii) wherein the model is a neural network comprising an encoder and a decoder.

[0030] (viii) where the statistical distance between the expected probability function and the actual probability function is calculated using Jenson-Shannon or Kullback-Leibler divergence.

[0031] (ix) wherein at least one of the clusters is characterized by a known statistical distribution.

[0032] (x) Wherein at least two of the clusters are characterized by Gaussian mixture modeling (GMM) or Gaussian.

[0033] (xi) wherein the N-dimensional input space associated with at least one image provides information of pixel values and / or at least two of the following: average intensity level, deviation from average pixel value, defect size, SNR (signal-to-noise ratio), correlation with a predefined template, image moments.

[0034] According to other aspects of the disclosed subject matter, a system is provided for analyzing elements in a latent space using a trained model; the latent space represents a plurality of elements each having M dimensions, the plurality of elements transformed from elements in an input space having N (M≤N) dimensions and associated with at least one image of a semiconductor sample; the transformed elements conforming to a probability function; the system comprising a processor and memory circuitry (PMC) configured to:

[0035] a) obtaining at least one element in an input space associated with an image of a semiconductor sample;

[0036] b) utilizing the trained model to transform at least one element into an equal number of elements in the latent space;

[0037] c) For each transformed element, determining a distance between the element and a reference to the probability function, wherein the element is checked based on the determined distance.

[0038] The described aspects of the presently disclosed subject matter may comprise, mutatis mutandis, one or more of the features (i) to (v) listed below with respect to the described system in any desired combination or permutation that is technically possible.

[0039] (i) wherein said reference to the probability function is the center of the probability function.

[0040] (ii) wherein the model is trained by PMC, and the training comprises:

[0041] a) obtaining a desired probability function for transforming an element in the input space into one or more corresponding clusters of elements in the latent space;

[0042] b) using the desired probability function to repeatedly transform elements in the input space into an equal plurality of elements in the latent space that conform to the actual probability function until a specified criterion is satisfied, the actual probability function indicating an actual assignment of the elements to one or more corresponding clusters;

[0043] c) determining a training loss value L associated with the element in the latent space and testing whether the training loss value L satisfies a specified criterion; determining the training loss value L based on at least:

[0044] a. The first item L Rec , where the first term indicates the distance between an element in the input space and an element in the output space reconstructed from an element in the latent space; and

[0045] b. The second item L Prob , the second term indicates the statistical distance between the expected probability function and the actual probability function.

[0046] (iii) wherein the analyzing comprises determining anomalies of the transformed elements.

[0047] (iv) wherein the analyzing comprises determining an association of each transformed element with one of the clusters.

[0048] (v) wherein the analyzing comprises generating at least one new element in the latent space, the at least one new element being reconstructible into a corresponding at least one output element in the output space, wherein each of the at least one output element constitutes a new synthetic input element for training the model.

[0049] According to other aspects of the disclosed subject matter, there is provided a method for training a model representing a plurality of elements in an input space to a latent space representing an equal plurality of elements, the plurality of elements in the input space each having N dimensions and associated with at least one image of a semiconductor sample, the plurality of elements in the latent space each having M (M≤N) dimensions, the method comprising, by a processor and memory circuitry (PMC):

[0050] a) obtaining a desired probability function for transforming an element in the input space into one or more corresponding clusters of elements in the latent space;

[0051] b) using the desired probability function to repeatedly transform elements in the input space into an equal plurality of elements in the latent space that conform to the actual probability function until a specified criterion is satisfied, the actual probability function indicating an actual assignment of the elements to one or more corresponding clusters;

[0052] c) determining a training loss value L associated with the element in the latent space and testing whether the training loss value L satisfies a specified criterion; determining the training loss value L based on at least:

[0053] a. The first item L Rec , where the first term indicates the distance between an element in the input space and an element in the output space reconstructed from an element in the latent space; and

[0054] b. The second item L Prob , the second term indicates the statistical distance between the expected probability function and the actual probability function.

[0055] The described aspects of the presently disclosed subject matter may comprise, mutatis mutandis, one or more of the features (i) to (xi) listed above with respect to the described system in any desired combination or permutation that is technically possible.

[0056] According to other aspects of the disclosed subject matter, a method for analyzing elements in a latent space using a trained model is provided; the latent space represents a plurality of elements each having M dimensions, the plurality of elements being transformed from elements in an input space having N (M≤N) dimensions and associated with at least one image of a semiconductor sample; the transformed elements conforming to a probability function; the method comprising, by a processor and memory circuitry (PMC):

[0057] a. obtaining at least one element associated with an image of a semiconductor sample in an input space;

[0058] b. Using the trained model to transform at least one element to be equal in the latent space

[0059] number of elements;

[0060] c. For each transformed element, determine a distance between the element and a reference to the probability function, wherein the element is checked based on the determined distance.

[0061] The described aspects of the presently disclosed subject matter may comprise, mutatis mutandis, one or more of the features (i) to (v) listed above with respect to the described system in any desired combination or permutation technically possible.

[0062] According to other aspects of the disclosed subject matter, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium tangibly embodying instructions that, when executed by a computer, cause the computer to perform a method for training a model representing a plurality of elements in an input space to a latent space representing an equal plurality of elements, the plurality of elements in the input space each having N dimensions and associated with at least one image of a semiconductor sample, the plurality of elements in the latent space each having M (M≤N) dimensions, the method comprising, by a processor and memory circuitry (PMC):

[0063] a) obtaining a desired probability function for transforming an element in the input space into one or more corresponding clusters of elements in the latent space;

[0064] b) using the desired probability function to repeatedly transform elements in the input space into an equal plurality of elements in the latent space that conform to the actual probability function until a specified criterion is satisfied, the actual probability function indicating an actual assignment of the elements to one or more corresponding clusters;

[0065] c) determining a training loss value L associated with an element in the latent space and testing whether the training loss value L satisfies a specified criterion; determining the training loss value L based on at least the following:

[0066] Loss of value L:

[0067] a. The first item L Rec , where the first term indicates the distance between an element in the input space and an element in the output space reconstructed from an element in the latent space; and

[0068] b. The second item L Prob , the second term indicates the statistical distance between the expected probability function and the actual probability function.

[0069] The described aspects of the presently disclosed subject matter may comprise, mutatis mutandis, one or more of the features (i) to (xi) listed above with respect to the described system in any desired combination or permutation that is technically possible.

[0070] According to other aspects of the disclosed subject matter, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium tangibly embodying instructions that, when executed by a computer, cause the computer to perform a method for analyzing elements in a latent space using a trained model; the latent space represents a plurality of elements each having M dimensions, the plurality of elements transformed from elements in an input space having N (M≤N) dimensions and associated with at least one image of a semiconductor sample; the transformed elements conform to a probability function; the method comprising, by a processor and memory circuitry (PMC):

[0071] a. obtaining at least one element associated with an image of a semiconductor sample in an input space;

[0072] b. Using the trained model to transform at least one element to be equal in the latent space

[0073] number of elements;

[0074] c. For each transformed element, determine a distance between the element and a reference to the probability function, wherein the element is checked based on the determined distance.

[0075] The described aspects of the presently disclosed subject matter may comprise, mutatis mutandis, one or more of the features (i) to (v) listed above with respect to the described system in any desired combination or permutation technically possible. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to understand the present disclosure and to see how it may be carried out in practice, embodiments will now be described, by way of non-limiting example only, with reference to the accompanying drawings, in which:

[0077] Figure 1 depicts a generalized block diagram of a training system according to certain embodiments of the presently disclosed subject matter;

[0078] Figure 2 depicts a generalized block diagram of a sequence of operations for training a model according to certain embodiments of the presently disclosed subject matter;

[0079] Figure 3A schematically depicts elements in a latent space transformed into a single cluster according to a desired probability function, according to certain embodiments of the presently disclosed subject matter;

[0080] Figure 3B schematically depicts elements in a latent space transformed into two clusters according to a desired probability function according to certain embodiments of the presently disclosed subject matter;

[0081] Figures 3C to 3E schematically depicting three respective examples of transforming elements from an input space into elements in a latent space according to certain embodiments of the disclosed subject matter;

[0082] Figure 4 depicts a generalized block diagram of an inference system according to certain embodiments of the disclosed subject matter; and

[0083] Figure 5 Depicted is a generalized block diagram of a sequence of operations performed in an inference system according to certain embodiments of the disclosed subject matter. DETAILED DESCRIPTION

[0084] In the field of analyzing input elements of semiconductor samples (such as wafers), for example, processing numerous input elements each associated with data indicating many dimensions (see below) can hinder successful analysis of such input elements (e.g., determining whether an input element is a defect, detecting anomalies, determining whether an element falls into any specified class, and so on).

[0085] Intuitively, according to certain embodiments, a system is provided for training a model to reduce an input space representing a plurality of elements to a latent space representing an equal plurality of elements, wherein the plurality of elements in the input space each have N dimensions (e.g., average intensity level, deviation from an average pixel value of a pixel matrix, etc.) and are associated with (a plurality of) images of a semiconductor sample, and the plurality of elements in the latent space each have M (M≤N) dimensions. The process may include:

[0086] a) obtaining a desired probability function (e.g., Gaussian) for assigning elements in the input space to one or more corresponding clusters of elements in the latent space;

[0087] b) using the desired probability function to repeatedly transform elements in the input space into an equal plurality of elements in the latent space that conform to the actual probability function until a specified criterion is satisfied, the actual probability function indicating an actual assignment of the elements to one or more corresponding clusters;

[0088] c) determining a training loss value L associated with an element in the latent space and testing whether the training loss value L satisfies a specified criterion; determining the training loss value L based on at least:

[0089] a. The first item L Rec , the first term indicates the distance between an element in the input space and a reconstructed element in a reconstructed space that can be reconstructed from an element in the latent space; and

[0090] b. The second item L Prob , the second term indicates the statistical distance between the expected probability function and the actual probability function.

[0091] Elements in the context of the present invention should be interpreted as indicating a portion (e.g., an image portion) of a semiconductor sample (e.g., a die or wafer), including but not limited to normal portions, defects of interest (DOIs) (occasionally also referred to as "defects"), interference that may indicate abnormal portions that are not of interest and / or not considered defects, etc. It is well known that defects in semiconductor dies can occur at any stage of the manufacturing process, from the initial growth of silicon wafers to the final packaging of ICs. These defects can range in size from microscopic flaws to large cracks or fragments. There are many different types of semiconductor die defects, such as particle contamination, structural defects, process defects, etc.

[0092] Dimensions in the context of the present invention should be broadly interpreted to include data associated with the element. For example, a dimension may provide information about pixel values. As an example, consider that the element provides information about a 32×32 pixel matrix (extracted from the die), and the N=1024 dimensions are the grayscale pixel values of the pixel matrix. As another example, the N dimensions may indicate metadata associated with the element. For example, consider the previous example where the element provides information about a 32×32 pixel matrix (e.g., as extracted from the die), and the N dimensions may be, for example, at least two of the following: average intensity level, deviation from the average pixel value defect size, SNR (signal-to-noise ratio), correlation with a predefined template, image moments, and / or other, depending on the specific application. Note that the present invention is not constrained by these examples.

[0093] Based on this, please note Figure 1 , which depicts a functional block diagram of an inspection system according to certain embodiments of the presently disclosed subject matter.

[0094] exist Figure 1The inspection system 100 shown in FIG. 1 can be used to inspect elements in a sample (e.g., a semiconductor wafer, die, and / or portion of a sample) as part of a sample manufacturing process. References to inspection herein may be interpreted as encompassing any type of operation on the sample, including defect inspection / detection, defect classification, segmentation, operations such as, for example, critical dimension (CD) measurement, overlay, and the like. The system 100 includes one or more inspection tools 120 configured to scan the sample and capture images of the sample for further processing for various inspection applications.

[0095] Without limiting the scope of the present disclosure, it is also noted that the inspection tool 120 can be implemented as various types of inspection machines, such as an optical inspection machine, an electron beam inspection machine (e.g., a scanning electron microscope (SEM) [e.g., defect review], an atomic force microscope (AFM), or a transmission electron microscope (TEM), etc.), etc. In some cases, the same inspection tool can provide low-resolution image data and high-resolution image data. The resulting image data (low-resolution image data and / or high-resolution image data) can be transmitted to the system 101 directly or via one or more intermediate systems. The present disclosure is not limited to any particular type of inspection tool and / or the resolution of the image data generated by the inspection tool.

[0096] In some embodiments, at least one of the inspection tools 120 may be configured to capture an image and perform an operation on the captured image.

[0097] According to certain embodiments, the inspection tool may be an electron beam tool, such as, for example, a scanning electron microscope (SEM). An SEM is an electron microscope that generates an image of a sample by scanning the sample with a focused electron beam. The electrons interact with atoms in the sample, generating various signals that contain information about the sample's surface topography and / or composition.

[0098] According to certain embodiments of the presently disclosed subject matter, the inspection system 100 includes a computer-based system 101 operably connected to an inspection tool 120, which includes, but is not limited to, online operation, wherein images obtained by the inspection tool are processed by various modules of the PMC 102, or according to other non-limiting embodiments, images obtained by the inspection tool 120 are received via an I / O module 126 and stored in a storage module 122 for later offline processing by the PMC 102, all of which will be explained in greater detail below.

[0099] Specifically, the system 101 includes a processor and memory circuitry (PMC) 102 operatively connected to a hardware-based I / O interface 126. Figure 2As further described in detail in FIG3 , PMC 102 is configured to provide the processing required by the operating system and includes one or more processors (not separately shown) operably connected to a memory (not separately shown). The processor(s) of PMC 102 can be configured to execute several functional modules according to computer-readable instructions implemented on a non-transitory computer-readable memory included in the PMC. Such functional modules are hereinafter referred to as being included in the PMC.

[0100] Functional modules included in the PMC 102 of the system 101 may include, for example, a training module 104 , which in turn may include a transformation module 105 , a reconstruction module 107 , and a criterion testing module 106 .

[0101] PMC 102 may be configured to obtain data indicative of a desired probability function for assigning elements in the input space to one or more corresponding clusters of elements in the latent space via I / O interface 126 and from inspection tool 120 , all of which will be explained in more detail below.

[0102] Will refer to Figure 2 Figure 3 further describes in detail the operation of systems 100, 101, 102 and their (multiple) PMCs and functional modules therein for training a model for reducing an input space representing multiple elements to a latent space representing an equal multiple elements, wherein the multiple elements in the input space each have N dimensions and are associated with at least one image of a semiconductor sample, and the multiple elements in the latent space each have M (M≤N) dimensions.

[0103] In some cases, the inspection system 100 may include one or more inspection modules, such as, for example, a defect detection module and / or an automatic defect review module (ADR) and / or an automatic defect classification module (ADC) and / or other inspection modules, in addition to the system 101. The one or more inspection modules may be implemented as standalone computers, or the functionality (or at least a portion of the functionality) of the one or more inspection modules may be integrated with the inspection tool 120. In some cases, the output of the system 101 (such as, for example, an image associated with data indicating the outline of the bottom of the hole) may be provided to the one or more inspection modules for further processing.

[0104] According to certain embodiments, the system 101 may include a storage unit 122. The storage module 122 may be configured to store any data required to operate the system 101, such as data relating to the input and output of the system 101, as well as intermediate processing results generated by the system 101. As an example, the storage module 122 may be configured to store images of samples and / or derivatives of the samples generated by the inspection tool 120. Thus, images may be retrieved from the storage module 122 and provided to the PMC 102 for further processing. The output of the system 101 may be sent to the storage module 122 for storage. As an example, a designated storage unit may further store a desired probability function, a training criterion, a training loss value L, etc., all of which will be explained in more detail below.

[0105] In some embodiments, the system 100 may optionally include a computer-based graphical user interface (GUI) 124 configured to enable user-specified input related to the system 101. For example, a visual representation of a sample, including image data of the sample, may be presented to the user (e.g., via a display forming part of the GUI 124). The GUI may provide the user with options for defining certain operational parameters. The user may also annotate reference images via the GUI. The user may also view operational results on the GUI.

[0106] In some cases, the system 101 may be further configured to send the output data to one or more of the inspection tools 120 and / or one or more inspection modules as described above for further processing via the I / O interface 126. In some cases, the system 101 may be further configured to send certain output data to a storage unit 122 and / or an external system (e.g., a yield management system (YMS) of a fabrication facility (semiconductor foundry)).

[0107] Those skilled in the art will readily appreciate that the teachings of the disclosed subject matter are not limited to Figure 1 The system depicted is not limited, and in particular is not limited to any specific modules 104, 105, 106 and 107 and / or operations performed thereby, as described below with reference to Figure 2 As described to Figure 3. Equivalent and / or modified functionality may be combined or divided in another manner and may be implemented in any suitable combination of software and firmware and / or hardware.

[0108] Note that this can be implemented in a distributed computing environment Figure 1 The system shown, wherein Figure 1 The aforementioned components and functional modules shown may be distributed across several local and / or remote devices and may be connected via a communication network. For example, the inspection tool 120 and the system 101 may be located at the same entity (in some cases hosted by the same device) or distributed across different entities.

[0109] It should also be noted that in some embodiments, at least some of the inspection tools 120, storage module 122, and / or GUI 124 may be external to the inspection system 100 and operate in data communication with the systems 100 and 101 via the I / O interface 126. As described above, the system 101 may be implemented as a standalone computer(s) used in conjunction with the inspection tools and / or additional inspection modules. Alternatively, the corresponding functionality of the system 101 may be at least partially integrated with one or more inspection tools 120, thereby improving and enhancing the functionality of the inspection tools 120 in inspection-related processes.

[0110] Despite Figure 1 1 , but in some cases, the functionality of system 110 may be at least partially integrated with system 100. As an example, the functional modules of system 110 may be incorporated into PMC 102 in system 101.

[0111] Although not necessarily so, the operation of systems 101 and 100 may be similar to that described with respect to Figure 2 3 corresponds to some or all of the stages of the method described in FIG. Figure 2 3 and their possible implementations can be implemented by systems 101 and 100 using modules 104, 105, 106 and 107. Figure 2 The embodiments discussed up to FIG. 3 may also be implemented mutatis mutandis as various embodiments of systems 101 and 100 , and vice versa.

[0112] Now pay attention Figure 2 , which depicts a generalized block diagram of a sequence of operations for training a model 200 according to certain embodiments of the presently disclosed subject matter. Note that model training may be performed in a training module 104, which may utilize "black box" modules (105 and 107) for training and reconstruction operations.

[0113] As shown, the input (designated as x 201), the so-called input space, represents a plurality of elements, each of which has N dimensions and is associated with at least one image of a semiconductor sample. In many applications, there may be numerous elements and a large number (N) of dimensions.

[0114] While utilizing the transformation model 200, the elements may be transformed (in a manner described in detail below) into a so-called latent space representing a corresponding number of elements, each of which has M (M≤N) dimensions. In some embodiments, M<N results in a smaller amount of data in the latent space.

[0115] In addition to the specified elements, another input fed to the model 200 may be a desired probability function (designated p(z) 203) that specifies the desired assignment of elements in the input space to one or more corresponding clusters of elements in the latent space, all of which will be explained in more detail below. The clusters may be characterized by statistical distributions, all of which will be described in more detail below.

[0116] Thus, the inputs fed to the model may include a specified input element, the associated N-dimensional data, the desired probability function p(z), and a desired number M that provides information about the number of dimensions in the latent space. Note that the N-dimensional data includes a specific specification of each dimension (e.g., grayscale value or average brightness value, etc.), while the M dimensions are not specified (except that the number M provides information about the number of dimensions in the latent space) because their details are determined by the model, all of which are known per se.

[0117] The specified inputs to the model may be retrieved, for example, from the storage module 122. Data indicative of elements associated with the image of the semiconductor sample and their associated N dimensions may be received, for example, from the inspection tool 120 and fed to the storage module via the I / O module.

[0118] Intuitively, the goal of the transformation is to transform elements from the (N-dimensional) input space to the M-dimensional latent space to comply with a specified desired probability function 203 and some other conditions, as will be explained in more detail below.

[0119] The transformation model that can be used is, for example, a neural network module known per se (using, for example, an autoencoder). Model training can be performed in an iterative manner, repeatedly transforming using a desired probability function until a specified criterion is met, such as a loss function that satisfies a certain criterion, all of which will be explained in more detail below. During the training phase, the elements in the input space will be transformed into an equal number of elements in the latent space that conform to the actual probability function, which indicates the actual assignment of the elements to one or more corresponding clusters.

[0120] More specifically and by way of example, in each iteration, the elements in the input space are transformed (eg, in the transformation module 105) into a corresponding number of elements that conform to the desired probability function. The transformed elements conform to the actual probability function q θ (z)(see Figure 2 204 in the ), the actual probability function should ideally match the expected probability function p(z). Note that z represents the latent space and θ represents the parameters of the encoder 202. In a simple, non-limiting case, all elements in the latent space are assigned to a single cluster 3000, as will be referenced below. Figure 3AAs shown. Naturally, in the first iteration, the "similarity" between the expected probability function and the actual probability function is not optimal and therefore additional iterations are required in order to improve the "similarity". It should be noted that the expected probability function and the actual probability function can be estimated by utilizing, for example, the KDE algorithm known per se.

[0121] This can be done, for example, by determining the statistical difference between the expected probability function and the actual probability function (see below with reference to L Prob The latter is just a component that specifies how many iterations will be performed until "success" is achieved, i.e., a specified criterion is met.

[0122] Note that the model also includes: a reconstruction module using a decoder which may itself be known (e.g. Figure 1 107) reconstructing 206 the M-dimensional transformed elements (in the latent space) into reconstructed output elements 207, all of which are explained in more detail below.

[0123] According to some embodiments, at each iteration of training the transformation model 202, a training loss value L associated with an element in the latent space is determined and tested (e.g., in the criterion testing module 106 - see Figure 1 , relative to a specified criterion, and if the latter is satisfied, the training is complete. According to some embodiments, the training loss value L is determined based at least on:

[0124] a. The first item L Rec , where the first term indicates the distance between an element in the input space and an element in the output space that can be reconstructed from an element in the latent space; and

[0125] b. The second item L Prob , the second term indicates the statistical distance between the expected probability function 203 and the actual probability function 204.

[0126] According to some embodiments, L is calculated Prob The specified statistical distance of the components may utilize the Jenson-Shannon (JS) or Kullback-Leibler (KL) divergence which is known per se.

[0127] In addition to the calculated L Prob In addition to (indicating the statistical distance between probability functions), another component of the loss function L can be L Recterm. Intuitively, the latter indicates how "similar" the so-called output elements (reconstructed from the elements in the latent space) are compared to the input elements in the input space. In this context, it is noted that an element in the latent space "sufficiently corresponds" to an element in the input space. This can be achieved by, for example, reconstructing 206 the output space from the transformed elements in the latent space (208) using a decoder (206) that operates, for example, in the reverse manner to the encoder 202 discussed above. This is achieved by the output element in (207).

[0128] Once the output element is reconstructed, the reconstructed output element is calculated 207 and the “similarity” between the input element x (e.g., L Rec - indicates the distance between the input element and the reconstructed output element). As an example, L rec The following equation is met:

[0129]

[0130] Note that the specified equation corresponds to the distance between one input element and one output element (or vice versa). In the case of computing distances between multiple elements (for example, in batch mode discussed below), the specified L between each pair of corresponding input and output elements in the batch is rec The values can be averaged, for example, over all pairs in the batch, yielding the combined L rec Value, combined L rec The value provides information about the distance between a batch of input elements and a batch of output elements. The present invention is not restricted to the latter example.

[0131] It is also noted that the decoding operation of the decoder 206 (as part of the neural network model (200)) as described above is also trained simultaneously with the encoder 202, all of which will be explained in more detail below.

[0132] With this in mind, according to some embodiments, based at least on the specified L Prob and L Rec term to determine the training loss value L, (e.g., L = α * L Rec +(1-α)*L Prob ), where 0<α<1, and is tested against a specified criterion, for example, L should fall below a given threshold. Intuitively, the statistical difference L Prob The smaller the term, the more "similar" the actual probability function is to the expected probability function, and L Rec The smaller the term, the more "similar" the reconstructed output element is to the input element. Note that the value of the coefficient α can be determined based on the specific application.

[0133] Therefore, it is important to note that the reconstruction component of the model (e.g., decoder) is fully trained until the N-dimensional reconstruction element 207 is "sufficiently similar" to the input element x. This can be determined in a manner known per se The distance between and x is achieved. Note that the distance can be determined element-by-element or batch-by-batch, etc. For example, consider the non-limiting example where the model (during the training phase) is fed input elements serially one after another. Each of the input elements x is characterized by N dimensions (e.g., N=n1×n1 pixel values associated with the input image). Reconstruct There may be corresponding elements with N dimensions (N = n1×n1 reconstructed pixel values). The distance between the n1×n1 pixel values associated with the input image and the n1×n1 reconstructed pixel values may be calculated. The procedure is repeated for the next input element (fed to the model) relative to its corresponding reconstructed element, and so on.

[0134] According to another non-limiting example, the model is fed a batch of B elements one after another. By way of example, the distances between the elements of a batch are calculated. Thus, for example, consider a batch of B elements fed to the model, where each element in the batch is characterized by, for example, N dimensions (all constituting x). Reconstructing the features of the batch of B elements is done by N dimensions (constituting ). Through the examples, the distances between a batch of input elements and a batch of reconstructed elements can be calculated, all as discussed above. It should be noted that the present invention is not limited to these examples.

[0135] Once specified criteria are met, training can be terminated, and elements in the latent space can be processed and analyzed, all of which will be exemplified in more detail below.

[0136] Notice Figure 3A , which schematically illustrates elements in a latent space transformed into a single cluster according to a desired probability function according to certain embodiments of the disclosed subject matter. Figure 3A In the example, the input element ( Figure 3A , each feature in latent space 3000 (not shown) has N = 2 dimensions, and each element in latent space 3000 has two dimensions (M = 2), such that each element can be represented as a value at (z1, z2). The desired probability function is, for example, a Gaussian. Model training proceeds through several iterations until the model's transformation and reconstruction functions produce a loss function (e.g., one of the above) that satisfies a specified criterion, such as L falling below a predefined threshold.

[0137] As described above, model (eg, neural network 200) training involves simultaneously training the transformation and reconstruction functions (eg, decoder 202 and decoder 206).

[0138] Please note that reference Figure 1 Description and reference of the system architecture Figure 2 The operational sequence of FIG3 and FIG4 relates to a training phase of the system. According to another aspect of the present invention, a system is provided that is configured to operate in an inference phase. During the inference phase, a trained model (e.g., by following the above reference Figure 2 3 ) to analyze elements in a latent space. As can be recalled, the latent space represents a plurality of elements, each having M dimensions, transformed from elements in an input space having N (M≤N) dimensions and associated with at least one image of a semiconductor sample. The transformed elements conform to a given probability function. This aspect will now be described in greater detail.

[0139] Therefore, please note Figure 4 , which depicts a generalized block diagram of an inference system according to certain embodiments of the disclosed subject matter. System 401, PMC 402, and inspection tool 420, storage module 422, GUI 424, and I / O 426 apply mutatis mutandis to the corresponding system 101, PMC 102, and inspection tool 120, storage module 122, GUI 124, and I / O 416. Note that although reference is made to Figure 4 The reasoning system is described as a separate system, but according to some embodiments, the reasoning system may be partially or fully integrated with the reference system. Figure 1 Thus, as an example, according to certain embodiments, any one of a designated PMC, an inspection tool, a storage module, a GUI, and an I / O module may be shared by a training system and a reasoning system.

[0140] Specifically, the system 401 may include a processor and memory circuitry (PMC) 402 operatively connected to a hardware-based I / O interface 426. Figure 5 As further described in detail, PMC 402 is configured to provide the processing required by the operating system and includes one or more processors (not separately shown) operatively connected to a memory (not separately shown). The processor(s) of PMC 402 can be configured to execute several functional modules according to computer-readable instructions implemented on a non-transitory computer-readable memory included in the PMC. Such functional modules are hereinafter referred to as being included in the PMC.

[0141] Functional modules included in the PMC 402 of the system 401 may include, for example, a transformation module 405 and an analysis module 406 .

[0142] Will refer to Figure 5The operations of systems 100, 101, 102 and their (multiple) PMCs and functional modules therein are further described in detail for inference operations using a trained model in the latent space.

[0143] Thus, reference is also made below Figure 5 to an inference sequence describing operations in accordance with certain embodiments of the present invention. For better understanding and for illustrative purposes only, the description will occasionally refer to Figure 3A an example in which the figure schematically depicts the results of the training phase, showing elements falling within a single cluster, such as 3000 (characterized by, for example, a Gaussian distribution), in two dimensions in the latent space, where each element is represented in two dimensions (M = 2), and these elements are transformed from elements in an input space (where N = 2 dimensions) that conforms to a desired probability function, all in accordance with certain embodiments of the subject matter of the present disclosure. The elements in the input space are associated with at least one image of a semiconductor sample. As an example, during the training phase, elements in the input space that are transformed into elements falling within cluster 3000 are a priori selected to represent interference samples.

[0144] Intuitively, once the system is sufficiently trained and a newly fed input element is fed into the trained system (e.g., obtained from a newly inspected sample), if the newly fed element is obtained from an interference sample, it will be transformed and fall within the designated cluster (since the elements in the latent space corresponding to the interference sample (by the example) conform to the designated desired probability function, i.e., a Gaussian distribution). Thus, in the latter example, all elements in the latent space originating from elements in the (input space) associated with the input image will conform to a Gaussian distribution and will fall within the cluster (e.g.) 3000. Thus, if in the following inference phase, the newly fed input element represents, for example, a defect, it is likely that its characteristics do not conform to the designated desired probability function, and thus it will be transformed into an element that does not fall within the designated cluster 3000 but is away from the center of the cluster (e.g., element 3001 drawn in an enlarged shape for clarity). Note that throughout the specification, the term center (in the context of the distance from the cluster / probability function) is an example of "distance from the reference cluster" and means that the distance is not necessarily measured with respect to the center but may be measured with respect to one or more other points of interest associated with the cluster (probability function).

[0145] As an example, consider another embodiment in which the model is trained with a mixture of, for example, a majority of I interference input elements and a minority of j defect input elements (j << i), but still with a constraint such as Figure 3ABy way of example, there may be a higher probability that the elements in the input space representing interference samples will be transformed into corresponding elements in the latent space that are closer to the center of the cluster 3000 (e.g., 3002), while the elements in the input space representing defective samples will be transformed into corresponding elements in the latent space that are farther away from the center of the cluster 3000 (e.g., 3003). Although not necessarily so, the operation process of the systems 401 and 400 may be similar to that of the system 401 and 400. Figure 5 Some or all of the stages of the described method may correspond. Figure 5 The described method and its possible implementations can be implemented by systems 401 and 400 using modules 405 and 406. Therefore, it should be noted that Figure 5 The discussed embodiments may also be implemented mutatis mutandis as various embodiments of the systems 401 and 400 , and vice versa.

[0146] With this in mind, please note Figure 5 Thus, at the start 501, at least one element in the input space associated with an image of a semiconductor sample is obtained. Then, at 502, the at least one element is transformed into an equal number of elements in the latent space using a trained model (e.g., performed in the transformation module 405 - see Figure 4 ). Then, for each transformed element, the distance between the element and the probability function (e.g., the center of the probability function) is determined 503. Finally, at 504, an inspection of the element can be performed, which is based on the determined distance (e.g., in the analysis module 406).

[0147] Consider the example of training a model using only elements that represent noise samples. For example, if the distance of the transformed element from the center of the cluster falls below a given threshold, then inspection of the element will produce a noise sample. On the other hand, if it exceeds a specified threshold (i.e., deviates from the cluster), this may indicate an anomaly, possibly a defect, and can be further processed in a manner known per se.

[0148] Other implementations of the inference phase are possible, all of which are described in more detail below.

[0149] Having described an exemplary inference sequence of operation according to certain embodiments of the present invention, attention is now turned to the training phase (cf. Figure 2and FIG3 ) for illustrating additional non-limiting examples. As can be recalled, the transformation to the latent space is not limited to only one cluster, but according to certain embodiments is limited to two or more (I ≥ 2) mutually distinguishable clusters. By way of the described embodiment, the input fed to the system further comprises data indicating a certain number (K ≥ I) of classes, each of which is associated with a respective class group of each cluster, thereby generating I class groups, and further, each element of at least a subset of the input space is labeled with a selected class from the K classes.

[0150] To simplify the explanation, consider a non-limiting example of two clusters and two classes, with input elements labeled with either class. As a non-limiting example, the first class indicates the portion of interference on the wafer(s) or die(s) derived from, for example, a first process variation, and the other class indicates the portion of interference on the wafer(s) or die(s) derived from, for example, a second process variation. By way of this example, the desired probability function, comprising two different clusters, aims to transform elements derived from the first process variation (labeled class A) into elements that will fall into the first cluster, and to transform elements derived from the second process variation (labeled class B) into elements that will fall into the second cluster. For simplicity, it is further assumed that each of the clusters is characterized by a different Gaussian distribution.

[0151] The basic assumption in the latter simplified example is that the images of disturbances derived from class A should conform (after sufficient training of the transformation function) to a common distribution (e.g. with the first cluster (e.g. Figure 3B centered at c1 in ), and the images of interfering wafers of class B originating from the interference should conform (after sufficient training of the transformation function) to a common but still distinct distribution (e.g. centered at c2 in the second cluster (e.g. Figure 3B Note that the training of the model using the examples is supervised learning.

[0152] To better understand the latter example, note that Figure 3B , which schematically illustrates elements in the latent space transformed into two clusters according to a desired probability function according to certain embodiments of the presently disclosed subject matter. For simplicity, Figure 3B The examples in assume M=2 dimensions of transformed elements in a latent space drawn as a two-dimensional representation, where any transformed element can be represented as a (z1, z2) value.

[0153] Figure 3B The following diagram shows the use of Figure 2The result achieved after the neural network 200 repeatedly trains the transformation model, wherein each input element is labeled with, for example, the first class or the second class discussed above. Therefore, after appropriate repeated training of the model, the input elements labeled with the first class should preferably be transformed into elements falling into the first cluster 311 in the latent space (labeled with a hashed circle and representing, for example, a first Gaussian distribution), and the input elements labeled with the second class should preferably be transformed into elements falling into the second cluster 312 in the latent space (labeled with a hashed circle and representing, for example, a second Gaussian distribution). By way of example, clusters 311 and 312 are included in the probability function 300, and the characteristics of the clusters can be Gaussian mixture modeling (GMM).

[0154] continue Figure 3B , the dark gray area 313 shows a plurality of “points”, each “point” representing an element in the latent space that falls into the first cluster 311, and the light gray area 314 shows a plurality of “points”, each “point” representing an element in the latent space that falls into the second cluster 312.

[0155] Note in passing that for simplicity, each cluster indicates a known statistical distribution.

[0156] Now back Figure 3B , as mentioned above, the figure represents the final stage after repeated training of the transformation function.

[0157] In order to achieve a specified result that an element will fall into a specified cluster, the loss function L may be modified. As may be recalled, according to certain embodiments (discussed above), the loss function L conforms to the following equation L = α * L Rec +(1-α)*L Prob , and if L is tested and satisfies the specified criteria, then the training is successful, intuitively meaning that the transformed elements in the latent space will be "sufficiently similar" to the input elements when reconstructed (e.g., decoded into the output space), and additionally, the actual probability function (204) will be "sufficiently similar" to the expected probability function (203). As an example, consider Figure 3B , and assuming that the clusters associated with the desired probability function both represent Gaussian distributions, then the actual clusters 311 and 312 of the actual probability function 300 represent the desired corresponding distributions.

[0158] Therefore, according to certain embodiments, the loss function L is modified to determine the training loss value L also based on:

[0159] The third item L Sup , the third item indicates the degree of allocation of the transformed elements marked with the class to the corresponding cluster in the I clusters, so that the fewer the transformed elements allocated to clusters other than the corresponding cluster in the I clusters, the SupThe lower the value.

[0160] Therefore, in Figure 3B In the example above, if fewer elements are assigned to the “wrong cluster”, then L Sup Thus, by way of example, the fewer “bright” elements 314 are assigned to cluster 311 and / or the fewer “dark” elements 313 are assigned to cluster 312, the lower L Sup The lower the value.

[0161] Please note that L Sup The equation is only a non-limiting example. As an example, in the case where the clusters do not form a circle, it is clear that the r value is irrelevant.

[0162] Therefore, according to one embodiment, the training loss value L satisfies the following equation:

[0163] α*L Rec +β*L Prob +(1-α-β)*L Sup , where 0<α+β<1, and L Sup It should be noted that the values of coefficients α and / or β can be determined according to the specific application.

[0164] For simplicity, the above description relates to the case of M=2 (two dimensions in the latent space) and two classes. Of course, the present invention is not restricted to the example described.

[0165] It should be noted that the supervised learning described above is illustrated with respect to two classes, one class representing elements originating from a first type of interference (e.g., providing information about a first process variation) and the other class representing elements originating from a second type of interference (e.g., providing information about a second process variation). Of course, the present invention is not limited by the illustrated examples of elements and classes that can be fed into the model for training. Thus, by way of another example, two classes representing elements corresponding to different types of interference elements can be used, or by way of yet another non-limiting example, two different classes representing defects of interest (DOIs) corresponding to different types can be used, such as a first type of DOI originating from bridges and a second type of DOI originating from particles.

[0166] By way of yet another example, two classes may be assigned to elements providing information about interference (labeled as class A), and another class may be assigned to elements providing information about DOI (labeled as class B). Note that the present invention is not limited to these examples. Obviously, only two classes are discussed for illustrative purposes, and by other embodiments, elements labeled with more than two classes may be fed into the model.

[0167] According to certain other embodiments, the input data 201 includes data indicating a certain number (K ≥ I) of classes, each of which is associated with a corresponding cluster group for each cluster, thereby generating I cluster groups. Each cluster group corresponds to a cluster. To better understand the foregoing, consider, for example, Figure 3C , the figure shows input data 330 associated with two groups (in the example, I=2 groups) labeled 331 and 332 respectively. Group 331 is associated with two classes A and B, and group 332 is in turn associated with classes C and D. In other words, by way of example, during the training phase, and as mentioned above with reference to Figure 2 and Figure 3B As described in detail, each (or a subset) of the input elements is labeled with any of classes A-D.

[0168] continue Figure 3C In the example of , the distribution function is associated with two clusters (333 and 334 in the example, corresponding to the number of groups (two in the example). Figure 3C As can be easily shown in the example of , it is expected that elements labeled with class A or B (and belonging to group 331) will be transformed into cluster 333 in the latent space, and elements labeled with class C or D (and belonging to group 332) will be transferred to cluster 334 in the latent space. Therefore, with this particular example, there are K=4 classes, I=2 groups, and corresponding I=2 clusters.

[0169] According to the above description, if the loss function L (such as α*L Rec +β*L Prob +(1-α-β)*L Sup ) satisfies the specified criteria, the input elements will be transformed and assigned to the desired two clusters (as determined by L Prob Items tested), will fully fall into the corresponding cluster (333 or 334) according to their labeled classes (as determined by L Sup Item test), and will be reconstructed in the output space 335 (from elements in the latent space) with “sufficient similarity” to the elements in the input space (as determined by L Rec Please note that Figure 3C In the specific example of , the clusters in the latent space 333 and 334 conform to the actual probability function q, which in turn is sufficiently similar to the expected probability function p (as schematically depicted by the two clusters in 336).

[0170] Figure 3D Draws the Figure 3C A similar transformation is trained in

[15] , except that the tested loss function L does not take into account the fact that the elements are assigned to clusters according to their classes. Thus, by the example described, the loss function L can be based on α*L Rec +(1-α)*LProb And L is required to meet the specified criteria. Note that L is eliminated in the calculation Sup , so the specified loss function L can satisfy the specified criterion even if elements are assigned to clusters other than their specified classes. Thus, as an example, however and as Figure 3C As shown, in Figure 3D In the example, elements marked with class A or B are ideally assigned to cluster 333, and elements marked with class C or D are ideally assigned to cluster 334. Sup Due to the fact that cluster 333 may also include elements marked C, although the latter preferably fall into cluster 334 (see Figure 3C ).

[0171] Note that, as a non-limiting example, Figure 3C and Figure 3D Each of the clusters of probability functions depicted in may be characterized by a known distribution (eg, a Gaussian distribution).

[0172] Figure 3E Also shown with reference Figure 3C The scenario described is similar to the scenario described in Figure 3E , the statistical function 350 may be characterized by a Gaussian distribution and composed of two clusters 351 and 352, respectively, divided by a boundary line 353, such that elements labeled with A or B are transformed into elements falling into cluster 351 (i.e., they are "above" the boundary line 353), and elements labeled with C or D are transformed into elements falling into cluster 352 (i.e., they are "below" the boundary line 353). Note also that the loss function L that aims to separate the transformed elements "above" and "below" the boundary line 353 has a value L. SUP Components differ from reference Figure 3B Discussed L SUP (Aiming to identify preferably (but not necessarily) non-overlapping clusters).

[0173] Of course, the present invention is not limited to the specified L SUP Example constraints.

[0174] Provide reference Figures 3C to 3E The specified example of is for illustrative purposes only and is in no way limiting. For example (and for graphical presentation) assume that there are only M=2 dimensions in the latent space and only N=2 dimensions in the input space.

[0175] Note that while the above discussion illustrates supervised learning and unsupervised learning (the latter not utilizing class labels), according to certain embodiments, the present invention also encompasses semi-supervised learning, where, mutatis mutandis, a portion of the elements fed to the model are labeled with appropriate classes, while other elements are not labeled with appropriate classes.

[0176] Also note that although reference Figures 3C to 3E The description involves more classes than clusters, but according to other embodiments, elements with I classes may be transformed into J clusters (I<J) in the latent space, e.g., elements labeled with one of I possible classes may be fed to a model with the constraint that the latent space should conform to J>I.

[0177] Having described various embodiments of the training aspects of the present invention, it is again noted that for example, reference is made above to Figure 5 The interference aspect described. There may be applications other than anomaly detection as described above. For example, according to certain embodiments of the present invention, the interference phase may be directed to determine whether the tested element falls into any given cluster. Thus, for example, consider Figure 3B After the encoder is fully trained to transform the input elements to fall into the latent space into clusters 311 or 312, depending on the class labels of the training set assigned to the elements, in the subsequent inference phase, the tested elements can be fed into the system and transformed (using the trained encoder 202 - see Figure 2 ) to position 315 in latent space. The transformed elements can now be tested to determine which cluster they belong to. Given that it is closer to cluster 311 than 313 (e.g., shorter distance to the center), it is determined to belong to the former, not the latter. This can represent the following real-life scenario.

[0178] According to yet another non-limiting example of an application of reasoning, the system may be used to generate synthetic examples. For example, referring to e.g. Figure 3C Consider the following scenario. The system has been fully trained to transform elements labeled as class A into regions A that fall into cluster 333. As described above, all transformed elements in the latent space originate from corresponding elements in the input space (class A elements of input space 331). For the sake of discussion, assume that additional input elements (of class A) need to be obtained. Therefore, it is desirable to generate qualitative "synthetic examples," that is, to generate high-quality synthetic input elements that will resemble the real input elements. Given that the model is fully trained, it is guaranteed that the actual probability function q is sufficiently similar to the expected probability function p, and that the reconstructed output elements are sufficiently similar to the input elements. Therefore, according to certain embodiments, sample points are generated from the expected probability function p(z) and then processed by the model's decoder to obtain new reconstructed output element samples. Given that the reconstructed output elements can be used as input elements, the latter (output elements) can be used as synthetic input elements, for example, to train other models. In the case where a newly trained model requires a large amount of data during the training phase, the latter procedure can be used to generate as many new qualitative (synthetic) input elements as are needed to train the new model.

[0179] According to certain embodiments, at least one of the following advantages is achieved:

[0180] (i) It is more efficient to perform analysis on elements associated with actual probability functions that are sufficiently similar to the expected probability function in the latent space, rather than performing hypothetical analysis on elements in the input space that inherently do not conform to the specified expected probability function.

[0181] (ii) in the case of M < N dimensions, it is more efficient to analyze the elements in the latent space (characterized by M < N dimensions) rather than to perform a hypothetical analysis of the elements in the input space (characterized by N > M dimensions). The term "more efficient" is herein meant to be more efficient in terms of computational complexity and / or smaller required computer storage space.

[0182] Note that the present invention is not limited to the specific examples of utilizing the trained system in the inference phase; these examples are provided for illustration purposes only.

[0183] It should be noted that the examples and values illustrated in this disclosure (such as, for example, specifying N and M dimensions, statistical distances and / or distance criteria, etc.) are illustrated for illustrative purposes and should not be considered to limit this disclosure in any way. In addition to or in place of the above, other appropriate examples / implementations may be used.

[0184] It should also be noted that any mathematical terms used herein should be interpreted as also including equivalents of the terms.

[0185] In the detailed description, numerous specific details are set forth in order to fully understand the present disclosure. However, it will be understood by those skilled in the art that the present disclosure can be practiced without these specific details. In other cases, well-known methods, procedures, components, and circuits are not described in detail in order to avoid obscuring the present disclosure.

[0186] Unless otherwise specifically stated, as will be apparent from the discussion, it should be understood that throughout this specification, discussions utilizing terms such as monitor, contain, determine, represent, analyze, and include refer to the action(s) and / or processing(s) of a computer that manipulates data and / or transforms data into other data, the data being represented as physical quantities (such as electronic quantities) and / or the data representing physical objects. The term "computer" should be broadly interpreted to encompass any type of hardware-based electronic device with data processing capabilities, such as, for example, a computer system. Figure 1 or Figure 4 described.

[0187] The processor referred to in this disclosure may represent one or more general-purpose processing devices, such as a microprocessor, a central processing unit, and the like. More particularly, the processor may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor that implements other instruction sets, or a processor that implements a combination of instruction sets. The processor may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, and the like. The processor is configured to execute instructions for performing the operations and steps described herein.

[0188] The memory mentioned in this document may include main memory (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) (such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), etc.) and static memory (e.g., flash memory, static random access memory (SRAM), etc.).

[0189] As used herein, the terms "non-transitory memory" and "non-transitory storage medium" should be broadly interpreted to encompass any volatile or non-volatile computer memory suitable for the subject matter of the present disclosure. The terms should be deemed to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more instruction sets. The terms should also be understood to include any medium capable of storing or encoding an instruction set executed by a computer and causing the computer to perform any one or more of the methods of the present disclosure. Thus, the terms should be deemed to include, but are not limited to, read-only memory ("ROM"), random access memory ("RAM"), magnetic disk storage media, optical storage media, flash memory devices, and the like.

[0190] The term "sample" as used in this specification should be broadly interpreted to encompass any type of semiconductor sample, such as wafers, masks, reticles, and other structures, and combinations and / or portions thereof, that can be used, for example, to manufacture semiconductor integrated circuits, magnetic heads, flat panel displays, and other semiconductor products. The sample is also exemplified herein as a semiconductor sample, and the sample can be produced by manufacturing equipment that performs the corresponding manufacturing process.

[0191] The term "inspection" as used in this specification should be broadly interpreted to cover any type of operation involving defect detection, defect review and / or various types of defect classification, segmentation, and / or other operations during and / or after the sample manufacturing process. Inspection is performed during or after the manufacture of the sample to be inspected by using a non-destructive inspection tool. As a non-limiting example, the inspection process may include runtime scanning (in a single scan or multiple scans), imaging, sampling, detection, review, measurement (including, for example, measurement of sample holes and bottom characteristics of the holes), classification and / or other operations performed on the sample or part of the sample using the same or different inspection tools. Similarly, inspection may be performed before the sample to be inspected is manufactured, and the inspection may include, for example, generating (multiple) inspection recipes and / or other setup operations. It should be noted that, unless otherwise specifically stated, the term "inspection" or derivatives of "inspection" used in this specification are not limited to the resolution or size of the inspection area. As a non-limiting example, various non-destructive inspection tools include scanning electron microscopes (SEMs), atomic force microscopes (AFMs), optical inspection tools, etc.

[0192] The term "inspection tool(s)" as used herein should be broadly interpreted to encompass any tool that may be used in an inspection-related process, including, by way of non-limiting example, scanning (in a single or multiple scans), imaging, sampling, reviewing, measuring, sorting, and / or other processes performed on a sample or portion thereof.

[0193] It is noted that the term "image(s)" as used herein may refer to original images of a sample captured by an inspection tool during the manufacturing process, derivatives of captured images obtained through various pre-processing stages, and / or computer-generated images based on design data. It is noted that, in some cases, the images referred to herein may include image data (e.g., captured images, processed images, etc.) and associated numerical data (e.g., metadata, handcrafted attributes, etc.). It is also noted that the image data may include data related to one or more layers of interest of the sample.

[0194] According to certain embodiments, the terms "similar" or "sufficiently similar", "distance" and "statistical distance" used in this specification should be broadly interpreted to cover any kind of well-known techniques, such as measuring distance (e.g., L1 norm, L2 norm) and measuring statistical distance (e.g., KL divergence, JS divergence).

[0195] It will be understood that, unless specifically stated otherwise, certain features of the disclosed subject matter described in the context of separate embodiments may also be provided in combination in a single embodiment. Conversely, various features of the disclosed subject matter described in the context of a single embodiment may also be provided individually or in any suitable subcombination. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the methods and apparatus.

[0196] It should be noted that, according to certain embodiments, the order of the computational stages described herein with reference to the accompanying drawings is not necessarily binding. For example, the order of the steps may be changed, steps may be modified or deleted, and / or other steps may be added to replace or supplement the steps disclosed herein.

[0197] It is to be understood that the disclosure is not limited in its application to the details set forth in the description contained herein or illustrated in the drawings.

[0198] It should also be understood that the system according to the present disclosure can be implemented at least in part on a suitably programmed computer. Likewise, the present disclosure contemplates a computer program readable by a computer for performing the method of the present disclosure. The present disclosure further contemplates a non-transitory computer-readable memory tangibly embodying a program of instructions executable by a computer for performing the method of the present disclosure.

[0199] The present disclosure is capable of other embodiments and can be practiced and carried out in various ways. Therefore, it should be understood that the phraseology and terminology used herein are for descriptive purposes only and should not be considered restrictive. Therefore, those skilled in the art will understand that the concepts upon which the present disclosure is based can be readily used as a basis for designing other structures, methods, and systems for achieving the several purposes of the subject matter of the present disclosure.

[0200] Those skilled in the art will readily appreciate that various modifications and changes can be made to the embodiments of the present disclosure described above without departing from the scope of the present disclosure as defined by the appended claims.

Claims

1. A system for training a model representing a plurality of elements in an input space to a latent space representing an equal plurality of elements, wherein the plurality of elements in the input space each have N dimensions and are associated with at least one image of a semiconductor sample, wherein the plurality of elements in the latent space each have M (M≤N) dimensions, the system comprising a processor and memory circuitry (PMC), the PMC configured to: a) obtaining a desired probability function for transforming the elements in the input space into one or more corresponding clusters of elements in the latent space; b) using the expected probability function to repeatedly transform elements in the input space into the plurality of equal elements in the latent space that conform to an actual probability function, the actual probability function indicating an actual assignment of the elements to the one or more corresponding clusters, until a specified criterion is satisfied; c) determining a training loss value L associated with the element in the latent space and testing whether the training loss value L satisfies the specified criterion; The training loss value L is determined based on at least the following items: a. The first item L Rec , the first term indicates the distance between the element in the input space and the element in the output space reconstructed from the element in the latent space; as well as b. The second item L Prob , the second term indicates the statistical distance between the expected probability function and the actual probability function.

2. The system according to claim 1, wherein the training loss value L satisfies the following equation: a*L Rec +(1-a)*L Prob 。 3. The system of claim 1 , wherein the system facilitates more efficient analysis of elements associated with the actual probability function that are sufficiently similar to the expected probability function in the latent space, rather than performing hypothetical analysis of elements in the input space that inherently do not conform to a specified expected probability function.

4. A system according to claim 1, wherein the training is semi-supervised learning or fully supervised learning, so that at least some of the multiple elements in the input space are labeled with corresponding classes of at least two classes for transforming the elements in the input space into elements in one or more corresponding element clusters in the latent space.

5. The system of claim 4, wherein: Elements in the input space are assigned to a certain number (I ≥ 2) of mutually discriminable clusters in the latent space, and the (PMC) is further configured as follows: a) obtaining data indicating a certain number (K≥I) of classes, each of which is associated with a corresponding cluster group for each cluster, thereby producing I cluster groups; b) labeling each element of at least a subset of the input space with a selected class from the K classes; c) further determining the training loss value L based on the following items: The third item L Sup , the third item indicates the degree of allocation of the transformed elements labeled with a class to the corresponding cluster in the I clusters, each of the classes being associated with a corresponding group in the I class groups, so that the fewer the transformed elements allocated to clusters other than the corresponding cluster in the I clusters, the lower the L Sup The lower the value.

6. The system according to claim 4, wherein the training loss value L satisfies the following equation: a*L Rec +β*L Prob +(1-a-b)*L Sup 。 7. The system of claim 3, wherein: Elements in the input space are assigned to a certain number (I ≥ 2) of mutually discriminable clusters in the latent space, and the (PMC) is further configured as follows: a) obtaining data indicating a certain number (k < I) of classes; b) labeling each element of at least a subset of the input space with a selected class from the K classes; c) further determining the training loss value L based on the following items: The third item L Sup , the third item indicates the degree of allocation of the transformed elements labeled with k classes to the corresponding clusters in the I clusters, so that the fewer the transformed elements are allocated to clusters other than the corresponding cluster in the I clusters, the smaller the L Sup The lower the value.

8. The system of claim 1, wherein the model is a neural network comprising an encoder and a decoder.

9. The system of claim 1, wherein the statistical distance between the expected probability function and the actual probability function is calculated using Jenson Shannon or Kullback-Leibler divergence.

10. The system of claim 1, wherein at least one of the clusters is characterized by a known statistical distribution.

11. The system of claim 4, wherein at least two of the clusters are characterized by Gaussian mixture modeling (GMM) or Gaussian.

12. The system of claim 1 , wherein the N-dimensional input space associated with at least one image provides information of pixel values and / or at least two of the following: average intensity level, deviation from average pixel value, defect size, SNR (signal-to-noise ratio), correlation with a predefined template, image moments.

13. A system for analyzing elements in a latent space using a trained model; the latent space represents a plurality of elements each having M dimensions, the plurality of elements transformed from elements in an input space having N (M≤N) dimensions and associated with at least one image of a semiconductor sample; the transformed elements conforming to a probability function; the system comprising a processor and memory circuitry (PMC), the PMC configured to: a) obtaining at least one element in the input space associated with an image of a semiconductor sample; b) utilizing the trained model to transform the at least one element into an equal number of elements in the latent space; c) For each transformed element, determining a distance between said element and a reference to said probability function, wherein said element is checked based on the determined distance. The system of claim 13 , wherein the reference to the probability function is a center of the probability function.

15. The system of claim 13, wherein the model is trained by PMC, the training comprising: a) obtaining a desired probability function for transforming the elements in the input space into one or more corresponding clusters of elements in the latent space; b) using the expected probability function to repeatedly transform elements in the input space into an equal plurality of elements in the latent space that conform to an actual probability function, the actual probability function indicating an actual assignment of the elements to the one or more corresponding clusters, until a specified criterion is satisfied; c) determining a training loss value L associated with the element in the latent space and testing whether the training loss value L satisfies the specified criterion; The training loss value L is determined based on at least the following items: a. The first item L Rec , the first term indicates the distance between the element in the input space and the element in the output space reconstructed from the element in the latent space; as well as b. The second item L Prob , the second term indicates the statistical distance between the expected probability function and the actual probability function. The system of claim 13 , wherein the analyzing comprises determining anomalies of the transformed elements.

17. The system of claim 13, wherein the analyzing comprises determining an association of each transformed element with a cluster of the clusters.

18. The system of claim 13, wherein the analyzing comprises generating at least one new element in the latent space, the at least one new element being reconstructible into a corresponding at least one output element in the output space, wherein Each of the at least one output element constitutes a new synthetic input element for training the model.

19. A method for training a model representing a plurality of elements in an input space to a latent space representing an equal plurality of elements, wherein the plurality of elements in the input space each have N dimensions and are associated with at least one image of a semiconductor sample, wherein the plurality of elements in the latent space each have M (M≤N) dimensions, the method comprising: a) obtaining a desired probability function for transforming the elements in the input space into one or more corresponding clusters of elements in the latent space; b) using the expected probability function to repeatedly transform elements in the input space into the plurality of equal elements in the latent space that conform to an actual probability function, the actual probability function indicating an actual assignment of the elements to the one or more corresponding clusters, until a specified criterion is satisfied; c) determining a training loss value L associated with the element in the latent space and testing whether the training loss value L satisfies the specified criterion; The training loss value L is determined based on at least the following items: a. The first item L Rec , the first term indicates the distance between the element in the input space and the element in the output space reconstructed from the element in the latent space; as well as b. The second item L Prob , the second term indicates the statistical distance between the expected probability function and the actual probability function.

20. A method for analyzing elements in a latent space using a trained model; the latent space represents a plurality of elements each having M dimensions, the plurality of elements transformed from elements in an input space having N (M≤N) dimensions and associated with at least one image of a semiconductor sample; the transformed elements conforming to a probability function; the method comprising, by a processor and memory circuitry (PMC): a) obtaining at least one element in the input space associated with an image of a semiconductor sample; b) utilizing the trained model to transform the at least one element into an equal number of elements in the latent space; c) For each transformed element, determining a distance between said element and a reference to said probability function, wherein said element is checked based on the determined distance.

21. A non-transitory computer-readable storage medium tangibly embodying instructions that, when executed by a computer, cause the computer to perform a method for training a model representing a plurality of elements in an input space to a latent space representing an equal plurality of elements, the plurality of elements in the input space each having N dimensions and associated with at least one image of a semiconductor sample, the plurality of elements in the latent space each having M (M≤N) dimensions, the method comprising, by a processor and memory circuitry (PMC): a) obtaining a desired probability function for transforming the elements in the input space into one or more corresponding clusters of elements in the latent space; b) using the expected probability function to repeatedly transform elements in the input space into the plurality of equal elements in the latent space that conform to an actual probability function, the actual probability function indicating an actual assignment of the elements to the one or more corresponding clusters, until a specified criterion is satisfied; c) determining a training loss value L associated with the element in the latent space and testing whether the training loss value L satisfies the specified criterion; The training loss value L is determined based on at least the following items: a. The first item L Rec , the first term indicates the distance between the element in the input space and the element in the output space reconstructed from the element in the latent space; as well as b. The second item L Prob , the second term indicates the statistical distance between the expected probability function and the actual probability function.

22. A non-transitory computer-readable storage medium tangibly embodying instructions that, when executed by a computer, cause the computer to perform a method for analyzing elements in a latent space using a trained model; the latent space represents a plurality of elements each having M dimensions, the plurality of elements transformed from elements in an input space having N (M≤N) dimensions and associated with at least one image of a semiconductor sample; The transformed elements conform to a probability function; the method comprising, by a processor and memory circuitry (PMC): a) obtaining at least one element in the input space associated with an image of a semiconductor sample; b) utilizing the trained model to transform the at least one element into an equal number of elements in the latent space; c) For each transformed element, determining a distance between said element and a reference to said probability function, wherein said element is checked based on the determined distance.