Detecting Defects in a Semiconductor Specimen Using Weak Markers
Through machine learning models combined with high and low resolution images, efficient automated defect detection and classification of semiconductor samples' patterns of interest is achieved, solving the problems of low detection efficiency and high cost in the prior art, and improving detection accuracy and automation.
Patent Information
- Application Number
- CN202410416642.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-03
- Filing Date
- 2021-04-26
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-04-26
AI Technical Summary
The prior art is difficult to efficiently and automatically detect and classify defects in submicron features during semiconductor manufacturing, especially when performing detailed analysis of potential defects at high resolution, there are problems of low efficiency and high cost.
Using machine learning models, the neural network is trained using high-resolution training images and low-resolution label derivatives to generate defect possibility scores per pixel block, and combined with low-resolution optical inspection tools and high-resolution scanning tools to achieve efficient classification of patterns of interest in semiconductor samples.
It improves the automation degree and accuracy of semiconductor sample defect detection, reduces the dependence of manual annotation, and improves detection efficiency and accuracy.
Smart Images

Figure CN118297906B_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application for invention, with the application date of April 26, 2021, the application number of "202110455269.9", and the invention title of "Detecting Defects in Semiconductor Specimens Using Weak Labels". Technical Field
[0002] The presently disclosed subject matter generally relates to the field of wafer specimen inspection and, more particularly, to detecting defects in a specimen. Background Art
[0003] Current demands for high density and performance related to the very large scale integration of fabricated devices require sub-micron features, increased transistor and circuit speeds, and improved reliability. These demands require the formation of device features with high precision and uniformity, which in turn requires careful monitoring of the manufacturing process, including automatically inspecting the device while it is still in the form of a semiconductor wafer.
[0004] As a non-limiting example, in-run inspection can employ a two-stage procedure, such as viewing a specimen and then examining sampled locations of potential defects. During the first stage, the surface of the specimen is viewed at high speed and relatively low resolution. During the first stage, a defect map is generated to show suspicious locations on the specimen having a high probability of defects. During the second stage, at least some of the suspicious locations are analyzed more thoroughly at a relatively high resolution. In some cases, the two stages can be implemented by the same viewing tool, while in some other cases, the two stages are implemented by different viewing tools.
[0005] Inspection processes are used at various steps during semiconductor manufacturing to detect and classify defects on a specimen. The effectiveness of inspection can be improved through the automation of one or more processes, such as automatic defect classification (ADC), automatic defect review (ADR), etc. Summary of the Invention
[0006] According to one aspect of the presently disclosed subject matter, there is provided a system for classifying a pattern of interest (POI) on a semiconductor specimen, the system including a processor and a memory circuit (PMC), the processor and the memory circuit being configured to:
[0007] Obtain data providing information of a high-resolution image of the POI on the specimen;
[0008] Generate data usable for classifying the POI according to defect-related classifications,
[0009] wherein the generating utilizes a machine learning model that has been trained according to at least a plurality of training samples, each training sample including:
[0010] High-resolution training images, which are captured by scanning corresponding training patterns on a specimen, and the corresponding training patterns are similar to the POI,
[0011] wherein the corresponding training pattern is associated with a label derivative of a low-resolution inspection of the corresponding training pattern.
[0012] In addition to the above features, the system according to this aspect of the currently disclosed subject matter may include one or more of the following features (i) to (xiii) in any desired combination or arrangement that is technically possible:
[0013] (i) The high-resolution training image of each training sample in the plurality of training samples is a scanning electron microscope (SEM) image
[0014] (ii) The label of each training sample in the plurality of training samples is a derivative of an optical inspection of the corresponding training pattern
[0015] (iii) The machine learning model includes a neural network, and the neural network includes:
[0016] A first series of one or more neural network layers configured to output a feature map given input data providing information about a high-resolution image of a POI, and
[0017] A second series of neural network layers configured to output data indicating at least one per-pixel block classification score given the input feature map, where each per-pixel block classification score belongs to a region of the POI represented by pixels of a corresponding pixel block of the high-resolution image, and the per-pixel block classification score indicates the likelihood of a defect in the corresponding region, thereby generating data that can be used to classify the POI according to a defect-related classification;
[0018] And wherein the machine learning model has been trained according to:
[0019] a) Applying the first series of neural network layers to a first high-resolution image of a first POI of a first training sample in the plurality of training samples, the first POI being associated with a label indicating a defect, thereby generating a suspicious feature map according to the current training state of the neural network;
[0020] b) Applying the second series of neural network layers to the suspicious feature map, thereby generating data indicating at least one suspicious per-pixel block score according to the current training state of the neural network,
[0021] Each of the suspicious per-pixel block scores belongs to the region of the first POI represented by the pixels of the corresponding pixel block of the first high-resolution image, and the suspicious per-pixel block score indicates the likelihood of a defect in the corresponding region;
[0022] c) Applying the first series of neural network layers to a second high-resolution image of a second POI of a second training sample among the plurality of training samples, the second POI being associated with a label indicating no defect, thereby generating a reference feature map according to the current training state of the neural network;
[0023] d) Applying the second series of neural network layers to the reference feature map, thereby generating data indicating at least one reference per-pixel block score according to the current training state of the neural network,
[0024] where each reference per-pixel block score belongs to the region of the second POI represented by the pixels of the corresponding pixel block of the second high-resolution image, and the reference per-pixel block score indicates the likelihood of a defect in the corresponding region;
[0025] e) Adjusting at least one weight of at least one of the first series of neural network layers and the second series of neural network layers according to a loss function, the loss function utilizing at least the following:
[0026] The distance metric derivative of the suspicious feature map and the reference feature map, the at least one suspicious per-pixel block score, and the at least one reference per-pixel block score; and
[0027] f) Repeating a)-e) for one or more additional first and second training samples among the plurality of training samples.
[0028] (iv) The neural network layer is a convolutional layer
[0029] (v) The distance metric is based on the Euclidean difference between the suspicious feature map and the reference feature map
[0030] (vi) The distance metric is based on the cosine similarity between the suspicious feature map and the reference feature map
[0031] (vii) The additional second training samples are the same training samples
[0032] (viii) The loss function further utilizes:
[0033] Annotation data associated with a set of pixels of the first high-resolution image, the annotation data indicating a defect in the region of the first POI represented by the set of pixels.
[0034] (ix) The annotation data is a derivative of human annotation of the high-resolution image
[0035] (x) The processor compares each of the at least one per-pixel block classification scores with a defect threshold to produce an indication of whether the POI is defective.
[0036] (xi) The processor warns an operator based on the indication of whether the POI is defective.
[0037] (xii) The processor determines a defect bounding box based on at least one per-pixel block classification score.
[0038] (xiii) The system further includes:
[0039] A low-resolution inspection tool configured to capture a low-resolution image of the POI and classify the POI according to a defect-related classification using an optical inspection of the low-resolution image; and
[0040] A high-resolution inspection tool configured to capture a high-resolution image of the POI in response to the low-resolution tool classifying the POI as defective.
[0041] According to another aspect of the presently disclosed subject matter, there is provided a method of classifying a pattern on a semiconductor specimen based on a high-resolution image of a pattern of interest (POI) on the specimen, the method including:
[0042] The processor receives data providing information of a high-resolution image of the POI on the specimen; and the processor generates data usable to classify the POI according to a defect-related classification, wherein the generation utilizes a machine learning model that has been trained based on at least a plurality of training samples, each training sample including:
[0043] A high-resolution training image captured by scanning a corresponding training pattern on the specimen, the training pattern being similar to the POI,
[0044] wherein the corresponding training pattern is associated with a label derivative of a low-resolution inspection of the corresponding training pattern.
[0045] This aspect of the disclosed subject matter may optionally include, in any desired combination or arrangement that is technically possible with necessary modifications in details, one or more of the features (i) to (xii) described above with respect to the system.
[0046] According to another aspect of the presently disclosed subject matter, there is provided a non-transitory program storage device readable by processing and memory circuitry, the non-transitory program storage device tangibly embodying computer-readable instructions executable by the processing and memory circuitry to perform a method of classifying a pattern according to a high-resolution image of a pattern of interest (POI) on a semiconductor specimen, the method comprising:
[0047] Receiving data providing information of a high-resolution image of the POI on the specimen; and generating data useable to classify the POI according to defect-related classifications,
[0048] wherein the generating utilizes a machine learning model that has been trained according to at least a plurality of training samples, each training sample comprising:
[0049] A high-resolution training image captured by scanning a corresponding training pattern on a specimen, the training pattern being similar to the POI,
[0050] wherein the corresponding training pattern is associated with a label derivative of a low-resolution inspection of the corresponding training pattern.
[0051] This aspect of the disclosed subject matter may optionally include, in any desired combination or arrangement, with necessary modifications in details, one or more of the features (i) to (xii) of the system listed above.
[0052] According to another aspect of the presently disclosed subject matter, there is provided a non-transitory program storage device readable by processing and memory circuitry, the non-transitory program storage device tangibly embodying computer-readable instructions executable by the processing and memory circuitry to perform a method of training a neural network to generate data indicative of at least one per-pixel block classification score in the case of input data providing information of a high-resolution image of a POI, wherein each per-pixel block classification score belongs to a region of the POI represented by pixels of a corresponding pixel block of the high-resolution image, the per-pixel block classification score indicating a defect likelihood of the corresponding region, thereby producing data useable to classify the POI according to defect-related classifications, wherein the training utilizes at least a plurality of training samples, each training sample comprising:
[0053] A high-resolution training image captured by scanning a corresponding training pattern on a specimen, the training pattern being similar to the POI,
[0054] wherein the corresponding training pattern is associated with a label derivative of a low-resolution inspection of the corresponding training pattern,
[0055] The method comprising:
[0056] a) Apply a first series of neural network layers to a first high - resolution image of a first POI of a first training sample among the plurality of training samples, where the first POI is associated with a label indicating a defect, so as to generate a suspicious feature map according to the current training state of the neural network;
[0057] b) Apply the second series of neural network layers to the suspicious feature map, so as to generate data indicating at least one suspicious per - pixel block score according to the current training state of the neural network,
[0058] where each suspicious per - pixel block score belongs to a region of the first POI represented by the pixels of the corresponding pixel block of the first high - resolution image, and the suspicious per - pixel block score indicates the defect likelihood of the corresponding region;
[0059] c) Apply the first series of neural network layers to a second high - resolution image of a second POI of a second training sample among the plurality of training samples, where the second POI is associated with a label indicating no defect, so as to generate a reference feature map according to the current training state of the neural network;
[0060] d) Apply the second series of neural network layers to the reference feature map, so as to generate data indicating at least one reference per - pixel block score according to the current training state of the neural network,
[0061] where each reference per - pixel block score belongs to a region of the second POI represented by the pixels of the corresponding pixel block of the second high - resolution image, and the reference per - pixel block score indicates the defect likelihood of the corresponding region;
[0062] e) Adjust at least one weight of at least one of the first series of neural network layers and the second series of neural network layers according to a loss function, where the loss function utilizes at least the following:
[0063] the distance - metric derivative of the suspicious feature map and the reference feature map, the at least one suspicious per - pixel block score, and the at least one reference per - pixel block score; and
[0064] f) Repeat a) - e) for one or more additional first and second training samples among the plurality of training samples.
[0065] This aspect of the disclosed subject matter may optionally include, in any desired combination or arrangement technically possible with necessary modifications in details, one or more of the features (i) to (ix) of the system listed above. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] To understand the present invention and to see how it may be practiced in practice, embodiments will be described by way of non-limiting examples with reference to the accompanying drawings, in which:
[0067] Figure 1 A general block diagram of an inspection system in accordance with certain embodiments of the subject matter disclosed herein is shown;
[0068] Figure 2 A flowchart depicting an example method of classifying a pattern of interest (POI) on a semiconductor specimen based on weak label derivatives of optical inspection in accordance with certain embodiments of the currently disclosed subject matter is shown.
[0069] Figure 3 An exemplary layer of a machine learning model in accordance with certain embodiments of the currently disclosed subject matter is shown, which exemplary layer can be used to receive data providing information of a SEM image (or other high-resolution image) of a POI and generate data indicative of per-pixel block scores that can be used to classify the POI using defect-related classifications.
[0070] Figure 4 An exemplary method of training a machine learning model such that the model can receive an input SEM image of a POI and then generate data indicative of per-pixel block scores that can be used to classify the POI in accordance with some embodiments of the currently disclosed subject matter is shown.
[0071] Figure 5 An exemplary machine learning model and training data flow in accordance with some embodiments of the currently disclosed subject matter is shown. Detailed Description
[0072] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, those skilled in the art will understand that the presently disclosed subject matter may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the presently disclosed subject matter.
[0073] Unless otherwise expressly stated, as will be apparent from the following discussion, it should be understood that throughout the description of the specification, terms such as "train", "obtain", "generate", "calculate", "utilize", "feed", "provide", "register", "apply", "adjust", etc. refer to actions and / or processes in which a processor manipulates and / or converts data into other data, where the data is represented physically, such as electronically, numerically, and / or the data represents physical objects. The term "processor" encompasses any computing unit or electronic unit having data processing circuitry that can perform tasks based on instructions stored in a memory (such as a computer, server, chip, hardware processor, etc.). It encompasses a single processor or multiple processors, which may be located in the same geographical area or may be at least partially located in different areas and may be capable of communicating with each other.
[0074] As used herein, the terms "non-transitory memory" and "non-transitory medium" should be construed broadly to cover any volatile or non-volatile computer memory applicable to the presently disclosed subject matter.
[0075] As used in this specification, the term "defect" should be construed broadly to cover any kind of anomaly or undesired feature formed on or within a specimen.
[0076] As used in this specification, the term "specimen" should be construed broadly to cover any kind of wafer, mask, and other structures, their combinations and / or components used in the manufacture of semiconductor integrated circuits, magnetic heads, flat panel displays, and other semiconductor products.
[0077] The term "inspection" as used in this specification should be construed broadly to cover any type of metrology-related operation, as well as operations related to the detection and / or classification of defects in the specimen during specimen manufacture. Inspection is provided by using non-destructive inspection tools during or after the manufacture of the specimen to be inspected. As a non-limiting example, the inspection process may include runtime scans (in single or multiple scans), sampling, review, measurement, classification, and / or other operations on the specimen or a portion thereof provided by the same or different inspection tools. Similarly, inspection may be provided before the manufacture of the specimen to be inspected and may include, for example, generating inspection recipes and / or other setup operations. It should be noted that unless otherwise specifically stated, the term "inspection" or its derivatives as used in this specification is not limited in terms of the resolution or size of the inspection area. As non-limiting examples, various non-destructive inspection tools include scanning electron microscopes, atomic force microscopes, optical inspection tools, etc.
[0078] Embodiments of the presently disclosed subject matter are not described with reference to any particular programming language. It should be understood that multiple programming languages may be used to implement the teachings of the presently disclosed subject matter as described herein.
[0079] The present invention contemplates a computer program that can be read by a computer to execute one or more methods of the present invention. The present invention also contemplates a machine-readable memory tangibly embodying a program of instructions executable by a computer for performing one or more methods of the present invention.
[0080] It should be borne in mind that Figure 1 , Figure 1 FIG. shows a functional block diagram of an inspection system in accordance with certain embodiments of the presently disclosed subject matter. Figure 1 The inspection system 100 shown in FIG. can be used to inspect a specimen (e.g., a semiconductor specimen such as a wafer and / or a portion thereof), which is part of a specimen manufacturing process. The inspection system 100 shown includes a computer-based system 103 capable of automatically determining metrology-related and / or defect-related information using images of one or more specimens. The system 103 can be operatively connected to one or more low-resolution inspection tools 101 and / or one or more high-resolution inspection tools 102 and / or other inspection tools. The inspection tools are configured to capture images of the specimen and / or review the captured images and / or implement or provide measurements related to the captured images. The system 103 can be further operatively connected to a data repository 109. The data repository 109 can be operatively connected to a manual annotation input device 110 and can receive manual annotation data from the manual annotation input device 110.
[0081] The system 103 includes a processor and memory circuit (PMC) 104. The PMC 104 is configured to provide the processing required by the operating system 103, as further described in detail in various embodiments below; and includes a processor (not shown separately) and a memory (not shown separately). In Figure 1 FIG., the PMC 104 is operatively connected to a hardware-based input interface 105 and a hardware-based output interface 106.
[0082] The processor of the PMC 104 can be configured to execute a number of functional modules according to computer-readable instructions implemented on a non-transitory computer-readable memory included in the PMC. Such functional modules are hereinafter referred to as being included in the PMC. The functional modules included in the PMC 104 can include a machine learning unit (ML unit) 112. The ML unit 112 can be configured to implement data processing using a machine learning model / machine learning algorithm to output application-related data based on an image of the specimen.
[0083] The ML unit 112 may include a supervised or unsupervised machine learning model (for example, for a deep neural network (DNN)). The machine learning model of the ML unit 112 may include layers organized according to a corresponding DNN architecture. As a non-limiting example, the DNN layers may be organized according to a convolutional neural network (CNN) architecture, a recurrent neural network architecture, a recursive neural network architecture, a generative adversarial network (GAN) architecture, etc. Optionally, at least some of the layers may be organized in multiple DNN sub-networks. Each layer of the machine learning model may include a plurality of basic computational elements (CEs), which are commonly referred to in the art as dimensions, neurons, or nodes. In some embodiments, the machine learning model may be a neural network, where each layer is a neural network layer. In some embodiments, the machine learning model may be a convolutional neural network, where each layer is a convolutional layer.
[0084] Generally, the computational elements of a given layer may be connected to the CEs of the previous layer and / or the next layer. Each connection between the CEs of the previous layer and the CEs of the next layer is associated with a weighted value. A given CE may receive an input from the CEs of the previous layer via the corresponding connection, and each given connection is associated with a weighted value that can be applied to the input of the given connection. The weighted value may determine the relative strength of the connection and thus the relative impact of the corresponding input on the output of the given CE. A given CE may be configured to calculate an activation value (for example, a weighted sum of the inputs), and further derive an output by applying an activation function to the calculated activation. The activation function may be, for example, an identity function, a deterministic function (such as a linear function, a sigmoid function, a threshold function, etc.), a random function, or other suitable functions.
[0085] The output from a given CE may be transmitted via the corresponding connection to the CEs of the next layer. Similarly, as described above, each connection at the output of the CE may be associated with a weighted value that can be applied to the output of the CE before being accepted as an input to the CEs of the next layer. In addition to the weighted values, there may also be thresholds (including limiting functions) associated with the connections and CEs.
[0086] The weighted values and / or thresholds of a machine learning model can be initially selected before training and can be further iteratively adjusted or modified during training to achieve an optimal set of weighted values and / or thresholds in the trained ML network. After each iteration, the difference (also known as the loss function) between the actual output produced by the machine learning model and the target output associated with the corresponding training set of data can be determined. This difference can be referred to as the error value. Training can be determined to be complete when the cost or loss function indicating the error value is less than a predetermined value or when additional iterations only result in limited performance improvement. Optionally, at least some of the training machine learning model sub-networks (if any) in the machine learning model network can be trained individually before training the entire machine learning model network.
[0087] The set of machine learning model input data for adjusting the weights / thresholds of a neural network is herein referred to as the training set.
[0088] System 103 can be configured to receive input data via input interface 105, where the input data can include data generated by an inspection tool (and / or derivatives of the data and / or metadata associated with the data) and / or data generated and / or stored in one or more data repositories 109 and / or another associated data repository. It should be noted that the input data can include images (e.g., captured images, images derived from captured images, simulated images, synthetic images, etc.) and associated scalar data (e.g., metadata, manual / image annotations, automatic attributes, etc.). It should also be noted that the image data can include data related to the layer of interest of the specimen and / or one or more other layers of the specimen.
[0089] When processing input data (e.g., low-resolution image data and / or high-resolution image data, optionally together with other data, such as design data, synthetic data, inspection data, etc.), system 103 can send the results (e.g., data related to instructions) to any of the inspection tools in the inspection tool via output interface 106, store the results (e.g., defect attributes, defect classifications, etc.) in storage system 107, present the results via GUI 108, and / or send the results to an external system (e.g., to the yield management system (YMS) of the FAB). GUI 108 can be further configured to implement user-specified inputs related to system 103.
[0090] As a non-limiting example, one or more low-resolution inspection tools 101 (e.g., optical inspection systems, low-resolution SEMs, etc.) may inspect a specimen. The low-resolution inspection tool 101 may transmit the resulting low-resolution image data, which may provide information about a low-resolution image of the specimen, (either directly or via one or more intermediate systems) to the system 103. Alternatively or additionally, a high-resolution tool 102 may inspect the specimen (e.g., a subset of potential defect locations selected for inspection based on the low-resolution image may subsequently be reviewed by a scanning electron microscope (SEM) or an atomic force microscope (AFM)). The high-resolution tool 102 may transmit the resulting high-resolution image data, which may provide information about a high-resolution image of the specimen, (either directly or via one or more intermediate systems) to the system 103.
[0091] It should be noted that the image data may be received and processed together with metadata associated therewith (e.g., pixel size, textual description of the defect type, parameters of the image capture process, etc.).
[0092] In some embodiments, the image data may be received and processed together with annotation data. As a non-limiting example, a human reviewer may select a region (e.g., a hand-marked oval region) and label it as defective or mark it with a label indicating a defect. As described below, manual or other annotation data may be utilized during training.
[0093] Those skilled in the art will readily understand that the teachings of the presently disclosed subject matter are not limited by Figure 1 the system shown; equivalent and / or modified functionality may be combined or partitioned in another manner and may be implemented in any suitable combination of software with firmware and / or hardware.
[0094] Without in any way limiting the scope of the present disclosure, it should also be noted that the inspection tool may be implemented as various types of inspection machines, such as optical imaging machines, electron beam inspection machines, etc. In some cases, the same inspection tool may provide both low-resolution image data and high-resolution image data. In some cases, at least one inspection tool may have metrology capabilities.
[0095] It should be noted that Figure 1 the inspection system shown may be implemented in a distributed computing environment, where Figure 1The above-described functional modules may be distributed across several local and / or remote devices and may be linked via a communication network. It should also be noted that in another embodiment, at least some of inspection tools 101 and / or 102, data warehouse 109, manual annotation input 110, storage system 107, and / or GUI 108 may be external to inspection system 100 and may operate to communicate data with system 103 via input interface 105 and output interface 106. System 103 may be implemented as a stand-alone computer for use in conjunction with an inspection tool. Alternatively, the corresponding functions of the system may be at least partially integrated with one or more inspection tools.
[0096] Now turning to Figure 2 , which shows a flowchart in accordance with certain embodiments of the presently disclosed subject matter that depicts an exemplary processor-based method for classifying a pattern of interest (POI) on a semiconductor specimen based on a high-resolution image of the pattern.
[0097] In the monitoring of semiconductor manufacturing processes, it may be desirable to determine various data and metrics related to the manufactured specimens, and more particularly, data and metrics related to patterns on the specimens. Such data / metrics may include: the defect status of the pattern (e.g., defective / non-defective), the defect region (e.g., identification of a set of pixels in the high-resolution image that represent the defective pattern), and the degree of certainty regarding the accuracy of the defect determination.
[0098] Such data may be determined, for example, by human annotation of high-resolution images. However, such annotation is costly and time-consuming.
[0099] In some embodiments of the presently disclosed subject matter, a machine learning model is used to determine defects in regions of a high-resolution image, the machine learning model being trained using a training set that includes “weakly labeled” training samples, where a “weakly labeled” training sample is an image that has an associated label that applies as a whole to (e.g.) the image, and the image having the associated label is, for example, derived from a low-resolution classification method such as optical inspection. In the context of the present disclosure, “training” may include any suitable method of configuring a machine learning model, the method including: training methods such as the methods described below with reference to Figure 4 and setting the parameters of the model based on settings of another model trained using a training set that utilizes weakly labeled samples, etc.
[0100] Advantages of some of these methods are that they can provide an accurate assessment of pixel region defects based on samples that are readily available, without the need for annotated data.
[0101] System 103 (e.g., PMC 104) may receive (200) data providing information of a high-resolution image (such as a scanning electron microscope (SEM) image) of a POI on a specimen.
[0102] Inspection system 100 may be configured to enable system 103 (e.g., PMC 104) to receive an image. In some embodiments: The low-resolution inspection tool 101 may be configured to capture a set of one or more low-resolution images of a specimen (e.g., a wafer or a die) or portions of the specimen, and in so doing capture a low-resolution image of a pattern of interest on the specimen. The low-resolution inspection tool 101 may be further configured to classify the pattern of the specimen according to a defect-related classification (e.g., defective / non-defective) using optical inspection of the low-resolution image (e.g., as known in the art). The high-resolution tool 102 may be configured to capture a high-resolution image of the pattern of interest (POI) (e.g., using SEM) in response to the POI being classified as defective by one or more low-resolution inspection tools 101. The high-resolution tool 102 may be further configured to provide the captured high-resolution image of the POI to system 103.
[0103] The computer-based system 103 (e.g., PMC 104) may then generate (210) data indicating a score for each pixel block in at least one pixel block of the high-resolution image, and the score may be used to classify the POI according to a defect-related classification (such as a certain region of the POI or the whole POI being defective / non-defective).
[0104] More specifically, in some embodiments, e.g., using the method described below with reference to Figures 3 to 5 the computer-based system 103 (e.g., PMC 104) may generate data that can be used to classify the POI according to a defect-related classification. In some embodiments, the generated data provides information of one or more scores (e.g., a score matrix), where each score is derived from a pixel block of the high-resolution image, and the generated data indicates the defect likelihood (e.g., an estimate of the defect likelihood) of the region of the POI represented by the pixels of the pixel block. Such a matrix is hereinafter referred to as a "hierarchy map".
[0105] In some embodiments, the ranking map generated by system 103 (e.g., ML model 112) has the same dimensions as the input high-resolution image of the POI, and the entries of the matrix are scalar values (e.g., between 0 and 1), where the scalar value indicates the likelihood that the corresponding pixel of the input image of the POI corresponds to a defective area of the POI. In some other embodiments, the generated ranking map is smaller than the image of the POI (e.g., a 512×512 matrix can be generated from a 1024×1024 image, where each matrix entry contains the score for a corresponding 2×2 pixel block), and in such cases, the scalar value of the matrix indicates the likelihood that the corresponding pixel block (e.g., 2×2 or 4×4 block or pixel block of another dimension) of the POI image corresponds to a defective area of the POI. It should be noted that the term "pixel block" in this specification can include a single pixel as well as horizontal and / or vertical groups of pixels of various dimensions.
[0106] In some embodiments, system 103 (e.g., ML unit 112) generates a ranking map by utilizing a machine learning model that has been trained based on a set of "weakly labeled" training examples (e.g., images with an associated label applied as a whole to the image). More specifically, in some such embodiments, each training example is a high-resolution image (or data providing information about the high-resolution image) captured by scanning a training pattern similar to the POI. In some embodiments, the label associated with the training pattern is derived from an optical inspection (or other low-resolution inspection tool) of the corresponding training pattern. For example, an optical inspection tool can inspect the pattern, compare the inspected pattern with a reference pattern, and label the inspected pattern as "defective" or "non-defective". An exemplary method for training the machine learning model is described below with reference to Figure 4 the description of the exemplary method for training the machine learning model.
[0107] A wafer or die can be manufactured in such a way that multiple instances of a pattern are repeated on the wafer or die, and then the wafer or die is divided into many device instances. In the context of this disclosure, the term "similar" in the context of one POI being similar to another POI should be interpreted broadly to include multiple instances of a pattern on a single wafer die, as well as multiple instances of a pattern on multiple instances of the same wafer or die, etc.
[0108] In some embodiments, system 103 (e.g., ML unit 112) may generate a data structure in addition to the ranking map, which still indicates the per-pixel (or per-pixel block) score that can be used to classify the POI according to defect-related classifications.
[0109] Optionally: System 103 (e.g., PMC 104) may then compare (220) each value in the rank map (or corresponding alternative data representation) with a defect threshold (e.g.,.5 on a scale from 0 to 1), thereby generating an indication of whether the POI is defective and also classifying the POI according to a defect-related classification (e.g., defective / non-defective).
[0110] Optionally, if there is a score in the rank map that meets the defect threshold, system 103 (e.g., PMC 104) may take an action. Optionally, the action may include warning (230) an operator based on an indication of whether the POI is defective (or whether a series or plurality of POIs are defective, etc.).
[0111] In some embodiments, the rank map may be used to classify POIs according to other defect-related classifications. For example, system 103 (e.g., PMC 104) may determine a bounding box of a defect based on the per-pixel block score of the output.
[0112] It should be noted that the teachings of the presently disclosed subject matter are not limited by Figure 2 the flowcharts shown. It should also be noted that although the flowcharts are described with reference to the elements of the Figure 1 or Figure 3 system, this is in no way a limitation, and the operations may be performed by elements other than those described herein.
[0113] Now note Figure 3 , Figure 3 illustrates an exemplary layer of a machine learning model according to certain embodiments of the presently disclosed subject matter, which exemplary layer may be used, for example, by PMC 104 (more specifically, e.g., by ML unit 112) to receive data providing information of a SEM image (or other high-resolution image) of a POI and generate data indicating a score for each pixel block in at least one pixel block of the image, and the score may be used to classify the POI using a defect-related classification.
[0114] The machine learning model may include a first series of machine learning network layers 320 (e.g., neural network layers or convolutional neural network layers), the first series of machine learning network layers being configured to receive a SEM image (or another type of high-resolution image) 310 of a POI as an input. Then, the first series of machine learning network layers 320 may generate a feature map 330 based on the SEM image 310.
[0115] The feature map 330 may be an intermediate output of the machine learning network. Specifically, as described above, the machine learning network may include multiple layers L1 to L N , and the feature map 330 may be used as layer L jThe output is obtained, where 1 < j < N (in some embodiments, the intermediate layers from layer L1 to layer L j can form a convolutional neural network, where j < N). As described above, in a machine learning model, each layer L j provides an intermediate output, which is fed to the next layer L j+1 , until the last layer L N provides the final output. Assume, for example, that the SEM image 310 has dimensions X1, Y1, Z1, where:
[0116] - X1 corresponds to the dimension of the SEM image 310 along the first axis X;
[0117] - Y1 corresponds to the dimension of the SEM image 310 along the second axis Y; and
[0118] - Z1 corresponds to the number of values associated with each pixel, where Z1 > 1. For example, if a representation using three colors (red, green, and blue) is used, then Z1 = 3. This is not restrictive, and other representations can be used (e.g., in an SEM microscope, electrons from each pixel are collected by multiple different collectors, each collector having a different position, and each channel corresponds to one dimension, so Z1 is the total number of channels).
[0119] In some embodiments, the feature map 330 has dimensions X2, Y2, Z2, where X2 < X1, Y2 < Y1, and Z2 > Z1. Z2 can depend on the number of filters present in layer L j .
[0120] The machine learning model can include a second series of ML network layers 340 (e.g., neural network layers, or convolutional neural network layers), which are configured to receive the feature map 330 as an input. Then, the second series of ML network layers 340 can generate, for example, a rank map 350 that conforms to the reference Figure 2 described above based on the feature map 330.
[0121] It should be noted that the teachings of the presently disclosed subject matter are not limited by the machine learning model layers Figure 3 described in the reference. Equivalent and / or modified functions can be combined or divided in another way, and can be implemented in any suitable combination of software with firmware and / or hardware and executed on a suitable device.
[0122] Now note Figure 4 , Figure 4 illustrates a machine learning model (e.g., including having the above reference Figure 3The neural network of the described structure) is trained so that the machine learning model can receive an input high-resolution (e.g., SEM) image of a POI and then generate data indicating a per-pixel-block classification score for each pixel block in at least one pixel block of the image, where the per-pixel-block classification score indicates the likelihood of a defect in the region of the POI represented by the pixels of the pixel block. According to some embodiments of the presently disclosed subject matter, the per-pixel-block classification score can be used to classify the POI according to defect-related classifications. The description of Figure 4 the method shown refers to Figure 5 the training data stream shown.
[0123] In some embodiments, the PMC 104 (e.g., the ML unit 112) uses paired weakly labeled (e.g., image-level labeled) training samples from a training set to train the machine learning model, where one training sample (referred to as the suspect training sample) includes a high-resolution image (e.g., SEM scan) of a POI that has previously been labeled as defective (or suspected of being defective), and the second sample (referred to as the reference training sample) includes a high-resolution image (e.g., SEM scan) of a reference POI (i.e., a POI that has previously been labeled as non-defective (e.g., by an earlier classification)).
[0124] The labeling of the high-resolution image can be derived from, for example, a low-resolution (e.g., optical) inspection of the corresponding POI. Alternatively, the labeling of the high-resolution image can be derived from human inspection or another suitable method.
[0125] In some embodiments, the training set consists entirely of training samples where the image-level labels are derived from optical inspection or other low-resolution inspection tools. In some embodiments, the training set consists of such training samples as well as other training samples.
[0126] In some embodiments, some or all of the image-level labeled training samples can be associated with annotation data. The annotation data can be, for example, a derivative of human annotation (e.g., a human marking an ellipse around a region of a high-resolution image).
[0127] The annotation data can include data indicating the defectiveness (and, in some embodiments, the type of defect) of the portion of the specimen pattern represented by the annotated pixel group of the high-resolution image. In this context, the term "pixel group" can refer to a single pixel or the entire image. In this context, if, for example, a portion of the specimen of the pattern is substantially or entirely depicted by a particular pixel group, then that portion of the specimen can be considered to be represented by the pixel group.
[0128] As will be described below, in some embodiments, Figure 4The training method shown adjusts the weights of the machine learning model layers according to a loss function calculated from the distance metric derivatives of a suspicious feature map (i.e., a feature map derived from an image of a POI labeled as defective or suspicious) and a reference feature map (i.e., a feature map derived from an image of a POI labeled as non-defective).
[0129] In some embodiments, the loss function may seek to maximize the difference between two feature maps in a region of the feature map derived from pixels representing defective portions of the POI, and seek to minimize the difference between two feature maps in a region of the feature map derived from pixels representing non-defective portions of the POI.
[0130] In some embodiments, the loss function thus constitutes an attention mechanism. A first series of machine learning model layers may generate a feature map capable of identifying semantic regions, and a second series of machine learning model layers may score the defect likelihood of the regions.
[0131] PMC 104 (e.g., ML unit 112) may apply a first series of neural network layers 320 to (400) a high-resolution image 510 of a POI, the high-resolution image being associated with a label indicating defectiveness (e.g., determined or suspected defectiveness). In some embodiments, the label indicating defectiveness is associated with the image because optical inspection indicates the defectiveness of the POI. Then, the first series of neural network layers 320 may generate a feature map 530 from the high-resolution image. The feature map is generated according to the current training state of the neural network, and thus, as training progresses, the resulting feature map will change according to the progress of the machine learning model training. The feature map generated from an image associated with a label indicating a defect is referred to herein as a "suspicious feature map".
[0132] PMC 104 (e.g., ML unit 112) may then apply a second series of neural network layers 340 to (410) the suspicious feature map 530. The second series of neural network layers 340 may then generate a rank map 560 (e.g., a score for each pixel block in one or more pixel blocks of the high-resolution image 510, where each score indicates the defect likelihood of the region of the first POI represented by the pixels of the pixel block). The score output for a pixel block of the image may be referred to as a per-pixel-block score, and the rank map may thus be referred to as a set of per-pixel-block scores. The rank map is generated according to the current training state of the neural network, and thus, as training progresses, the resulting rank map will change according to the progress of the machine learning model training. The rank map generated from the suspicious feature map is referred to herein as a "suspicious rank map".
[0133] PMC 104 (e.g., ML unit 112) may then apply a first series of neural network layers 320 to (420) the high-resolution image 520 of the POI associated with the label indicating no defect. In some embodiments, the label is associated with the image because the image is determined (e.g., by optical inspection) to be defect-free, or is otherwise determined or assumed to be defect-free. Then, the first series of neural network layers 320 may generate a feature map 540 (calculated according to the current training state of the machine learning model) from the high-resolution image 520. The feature map generated from the image associated with the label indicating no defect is referred to herein as a "reference feature map".
[0134] PMC 104 (e.g., ML unit 112) may then apply a second series of neural network layers 340 to (430) the reference feature map 540. The second series of neural network layers 340 may then generate a reference score map 570 (e.g., a score for each pixel block in one or more pixel blocks of the reference image 520, where each score indicates the likelihood of a defect in the region of the reference POI represented by the pixels of the pixel block), calculated according to the current training state of the machine learning model.
[0135] PMC 104 (e.g., ML unit 112) may then adjust the weights of the first series of neural network layers 320 and the second series of neural network layers 340 according to a loss function (e.g., at least one weight of at least one layer, or e.g., all weights of all layers). For example, PMC 104 (e.g., ML unit 112) may calculate (440) the loss function 590 and employ, for example, backpropagation to adjust the weights of the first series of neural network layers 320 and the second series of neural network layers 340.
[0136] In some embodiments, the loss function 590 utilizes at least a distance metric (e.g., a value or set of values representing the difference between the reference feature map 540 and the suspicious feature map 530), the suspicious score map 560, and the reference score map 570. In some embodiments, as described below, the distance metric may be a differential feature map based on, for example, Euclidean distance or cosine similarity.
[0137] In some embodiments, the differential feature map 550 may be based on the Euclidean distance between the reference feature map 540 and the suspicious feature map 530. For example: The differential feature map 550 may be calculated by computing the Euclidean distance between the reference feature map 540 and the suspicious feature map 530 (i.e., in this case, the differential feature map 550 is a matrix where each entry is the arithmetic difference between the corresponding entries in the two feature maps).
[0138] In other embodiments, the differential feature map 550 can be based on the cosine similarity between a value in the reference feature map 540 and the corresponding value in the suspicious feature map 530. For example, the differential feature map 550 can be a matrix where each entry is calculated by computing the cosine similarity between a value in the reference feature map 540 and the corresponding value in the suspicious feature map 530.
[0139] In other embodiments, the differential feature map 550 can be a different representation of the difference between the suspicious feature map 530 and the reference feature map 540. In other embodiments, the loss function 590 can use different distance metrics representing the difference between the suspicious feature map 530 and the reference feature map 540.
[0140] Optionally, as described above, in some embodiments, annotation data 580 may be available. The annotation data can include data indicating a specific group of pixels in the suspicious image 510 that corresponds to a defective region for a corresponding POI. In such embodiments, the loss function 590 can utilize the annotation data 580 as well as the differential feature map 550, the reference rank map 570, and the suspicious rank map 560.
[0141] The PMC 104 (e.g., the ML unit 112) can repeat (450) applying two series of neural network layers 320, 340 to additional pairs of training samples among the plurality of training samples, and can adjust the weights of the neural network layers for each pair of samples according to a loss function that utilizes a distance metric such as the differential feature map 550, the suspicious rank map 560, and the reference rank map 570. For example, the PMC 104 (e.g., the ML unit 112) can be trained using all available suspicious training samples, and in conjunction with each suspicious training sample, the PMC 104 (e.g., the ML unit 112) can utilize a reference training sample from the set of training samples. In some embodiments, the PMC 104 (e.g., the ML unit 112) uses the same reference training sample in each training iteration.
[0142] It should be noted that the teachings of the presently disclosed subject matter are not limited by Figure 4 the flowchart shown in Figure 1 and the operations shown may not occur in the order shown. For example, the operations 400 and 420 shown consecutively can be performed substantially simultaneously or in the reverse order. It should also be noted that although the flowchart is described with reference to Figure 3 the elements of the Figure 5 system,
[0143] the neural network layers of Figure 5Constraints on the described data flow. Equivalent and / or modified functionality may be combined or divided in another way and may be implemented in any suitable combination of software with firmware and / or hardware and executed on a suitable device.
[0144] As described above, training a machine learning model in this way (e.g., using a suspicious image and a reference image and leveraging an attention mechanism based on a differential feature map) enables fast training and provides high classification accuracy.
[0145] The update of the weights may use techniques such as the feedforward / backpropagation method and may rely on any optimizer (e.g., stochastic gradient descent - SGD, ADAM, etc.).
[0146] It should be noted that the reference image 520 is an image of a reference area of a specimen (e.g., a die, a cell, etc.), where the corresponding image data represents a reference area without defects. The reference image may be an image captured from a reference (golden) die, a reference cell, or other areas verified to be without defects. Alternatively or additionally, the reference image may be simulated using CAD data and / or may be enhanced after capture to exclude defects (if any) in the reference area.
[0147] It should also be noted that in some embodiments, the suspicious image 510 is comparable to the reference image 520 (e.g., die to die, cell to cell, die to bin, etc.) and provides information about a first area of a semiconductor specimen. The first image is assumed to provide information about multiple defects associated with the first area. The first image may be an image captured from the first area. Optionally, the first image may be further enhanced and / or may include synthetic defects introduced after capture. The first area is configured to meet a similarity criterion with respect to the reference area and may belong to the same or different semiconductor specimens. The similarity criterion may be defined, for example, as the first area and the reference area corresponding to the same physical component or corresponding to similar areas of a semiconductor specimen (e.g., similar dies, cells, etc.).
[0148] It should be noted that in some embodiments, to ensure compatibility between the images, at least one of the reference image 520 and the first image 510 in the training samples must undergo a registration procedure. It should also be noted that at least a portion of different training samples may include the same reference image.
[0149] It should be noted that the various features described in the various embodiments can be combined according to all possible technical combinations. It should be understood that the application of the present invention is not limited to the details set forth in the description contained herein or shown in the drawings. The present invention is capable of having other embodiments and can be practiced and carried out in various ways. Therefore, it should be understood that the wording and terminology used herein are for the purpose of description and should not be regarded as restrictive. Thus, those skilled in the art should recognize that the concepts on which this disclosure is based can readily be used as a basis for other structures, methods, and systems designed for several purposes of implementing the subject matter disclosed herein. Those skilled in the art will readily appreciate that various modifications and changes can be applied to the embodiments of the present invention as described above without departing from the scope of the present invention as defined in the appended claims and by the appended claims.
Claims
1. A system for classifying a pattern of interest on a semiconductor specimen, the system comprising a processor and a memory circuit, the processor and the memory circuit being configured to: obtain data providing information of a high-resolution image of the pattern of interest on the specimen; and generate data usable for classifying the pattern of interest according to a defect-related classification, wherein the generating utilizes a machine learning model that has been trained according to at least a plurality of training samples, each training sample being obtained by: capturing a high-resolution training image by using a high-resolution inspection tool to scan a corresponding training pattern on the specimen, the corresponding training pattern being similar to the pattern of interest, and associating a label with the high-resolution training image, the label being a derivative of a low-resolution inspection of the corresponding training pattern.
2. The system according to claim 1, wherein the high-resolution inspection tool is a scanning electron microscope.
3. The system according to claim 1, wherein the corresponding label associated with each of the plurality of training samples is a derivative of an optical inspection of the corresponding training pattern.
4. The system according to claim 1, wherein the machine learning model comprises a neural network, the neural network comprising: a first series of one or more neural network layers configured to output a feature map given input data providing information of the high-resolution image of the pattern of interest, and a second series of one or more neural network layers configured to generate data indicating at least one per-pixel block classification score given an input of the feature map, wherein each per-pixel block classification score belongs to a region of the pattern of interest represented by the pixels of a corresponding pixel block of the high-resolution image, the per-pixel block classification score indicating the defect likelihood of the corresponding region, thereby generating data usable for classifying the pattern of interest according to the defect-related classification.
5. The system according to claim 4, wherein the machine learning model has been trained according to: a) applying the first series of one or more neural network layers to a first high-resolution training image of a first pattern of interest of a first training sample among the plurality of training samples, the first pattern of interest being associated with a label indicating a defect, to generate a suspicious feature map according to a current training state of the neural network; b) applying the second series of one or more neural network layers to the suspicious feature map to generate data indicating at least one suspicious per-pixel block classification score according to the current training state of the neural network, wherein each suspicious per-pixel block classification score belongs to a region of the first pattern of interest represented by the pixels of a corresponding pixel block of the first high-resolution training image, the suspicious per-pixel block classification score indicating the defect likelihood of the corresponding region; c) Apply one or more neural network layers of the first series to a second high-resolution training image of a second pattern of interest of a second training sample among the plurality of training samples, the second pattern of interest being associated with a label indicating no defect, to generate a reference feature map according to the current training state of the neural network; d) Apply one or more neural network layers of the second series to the reference feature map to generate data indicating at least one reference per-pixel block classification score according to the current training state of the neural network, wherein each reference per-pixel block classification score belongs to a region of the second pattern of interest represented by the pixels of a corresponding pixel block of the second high-resolution training image, and the reference per-pixel block classification score indicates the likelihood of a defect in the corresponding region; e) Adjust at least one weight of at least one of the one or more neural network layers of the first series and the one or more neural network layers of the second series according to a loss function, the loss function being based on at least the following: The distance metric derivative of the suspicious feature map and the reference feature map, The at least one suspicious per-pixel block classification score, and The at least one reference per-pixel block classification score; And f) Repeat a)-e) for one or more additional first training samples and additional second training samples among the plurality of training samples.
6. The system according to claim 4, wherein the one or more neural network layers of the first series and the one or more neural network layers of the second series are convolutional layers.
7. The system according to claim 4, wherein the processor and the memory circuit are further configured to: compare each per-pixel block classification score among the at least one per-pixel block classification score with a defect threshold, thereby providing an indication of whether the pattern of interest is defective.
8. The system according to claim 7, wherein the processor and the memory circuit are further configured to: warn an operator according to the indication of whether the pattern of interest is defective.
9. A method for classifying a pattern of interest based on a high-resolution image of the pattern of interest on a semiconductor specimen, the method being executed by a processor and a memory circuit, the method comprising: Receiving data providing information of a high-resolution image of the pattern of interest on the specimen; And Generating data usable for classifying the pattern of interest according to a defect-related classification, wherein the generating utilizes a machine learning model that has been trained according to at least a plurality of training samples, and each training sample is obtained by: Capturing a high-resolution training image by using a high-resolution inspection tool to scan a corresponding training pattern on the specimen, the training pattern being similar to the pattern of interest, and Associating the corresponding training pattern with a label that is a derivative of a low-resolution inspection of the corresponding training pattern.
10. The method according to claim 9, wherein the machine learning model includes a neural network, and the neural network includes: A first series of one or more neural network layers configured to output a feature map given input data providing information of the high-resolution image of the pattern of interest, and A second series of neural network layers configured to generate data indicative of at least one per-pixel block classification score given an input of the feature map, wherein each per-pixel block classification score pertains to a region of the pattern of interest represented by the pixels of the corresponding pixel block of the high-resolution image, and the per-pixel block classification score indicates the likelihood of a defect in the corresponding region, thereby producing data usable to classify the pattern of interest according to the defect-related classification; The method further includes: comparing each of the at least one per-pixel block classification scores with a defect threshold, thereby providing an indication of whether the pattern of interest is defective.
11. The method according to claim 10, further comprising: Warning an operator based on the indication of whether the pattern of interest is defective.
12. The method according to claim 9, wherein the machine learning model includes a neural network, and the neural network includes: A first series of one or more neural network layers configured to output a feature map given input data providing information of the high-resolution image of the pattern of interest, and A second series of neural network layers configured to generate data indicative of at least one per-pixel block classification score given an input of the feature map, wherein each per-pixel block classification score pertains to a region of the pattern of interest represented by the pixels of the corresponding pixel block of the high-resolution image, and the per-pixel block classification score indicates the likelihood of a defect in the corresponding region, thereby producing data usable to classify the pattern of interest according to the defect-related classification; The method further includes determining a defect bounding box based on the at least one per-pixel block classification score.
13. A computer program product comprising program instructions that, when read by a processor, cause the processor to execute a method of training a machine learning model to generate data usable to classify a pattern of interest according to a defect-related classification given input data providing information of a high-resolution image of the pattern of interest, the method including: a) Obtaining a plurality of training samples, each training sample obtained by: Capturing a high-resolution training image by using a high-resolution inspection tool to scan a corresponding training pattern on a specimen, the training pattern being similar to the pattern of interest, and Associating the corresponding training pattern with a label that is a derivative of a low-resolution inspection of the corresponding training pattern; And b) Training the machine learning model according to the plurality of training samples.
14. The computer program product according to claim 13, wherein the machine learning model includes a neural network, and the neural network includes: A first series of one or more neural network layers configured to output a feature map given input data providing information of the high-resolution image of the pattern of interest, and A second series of neural network layers configured to generate data indicative of at least one per-pixel block classification score given an input of the feature map, wherein each per-pixel block classification score pertains to a region of the pattern of interest represented by the pixels of a corresponding pixel block of the high-resolution image, and the per-pixel block classification score indicates the likelihood of a defect in the corresponding region, thereby producing data that can be used to classify the pattern of interest according to defect-related classifications.
15. The computer program product according to claim 14, wherein training the machine learning model comprises: a) applying a first series of neural network layers to a first high-resolution training image of a first pattern of interest of a first training sample among the plurality of training samples, the first pattern of interest being associated with a label indicative of a defect, to produce a suspicious feature map according to a current training state of the neural network; b) applying the second series of neural network layers to the suspicious feature map to produce data indicative of at least one suspicious per-pixel block classification score according to the current training state of the neural network, wherein each suspicious per-pixel block classification score pertains to a region of the first pattern of interest represented by the pixels of a corresponding pixel block of the first high-resolution training image, and the suspicious per-pixel block classification score indicates the likelihood of a defect in the corresponding region; c) applying the first series of neural network layers to a second high-resolution training image of a second pattern of interest of a second training sample among the plurality of training samples, the second pattern of interest being associated with a label indicative of no defect, to produce a reference feature map according to the current training state of the neural network; d) applying the second series of neural network layers to the reference feature map to produce data indicative of at least one reference per-pixel block classification score according to the current training state of the neural network, wherein each reference per-pixel block classification score pertains to a region of the second pattern of interest represented by the pixels of a corresponding pixel block of the second high-resolution training image, and the reference per-pixel block classification score indicates the likelihood of a defect in the corresponding region; e) adjusting at least one weight of at least one of the first series of neural network layers and the second series of neural network layers according to a loss function that utilizes at least the following: a distance metric derivative of the suspicious feature map and the reference feature map, the at least one suspicious per-pixel block classification score, and the at least one reference per-pixel block classification score; and f) repeating a)-e) for one or more additional first training samples and additional second training samples among the plurality of training samples.
16. The computer program product according to claim 15, wherein the distance metric is a difference map calculated according to the Euclidean difference between the suspicious feature map and the reference feature map or according to the cosine similarity between the suspicious feature map and the reference feature map.
17. The computer program product according to claim 15, wherein the loss function further utilizes: Annotation data associated with a set of pixels of the first high-resolution training image, the annotation data indicating a defect in a region of the first pattern of interest represented by the set of pixels.
Citation Information
Patent Citations
Method of generating a training set usable for examination of a semiconductor specimen and system thereof
CN110945528A
Training a neural network for defect detection in low resolution images
US20190303717A1