Method and system for classifying defects in a wafer using wafer defect images based on deep learning
By utilizing a mixture of synergies and patterns among multiple modes of wafer defect images in a deep learning network, combined with reference images, the problem of difficulty in accurately determining wafer defects in the prior art is solved, and efficient defect detection and classification are achieved.
Patent Information
- Application Number
- CN202010831725.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-18
- Filing Date
- 2020-08-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-08-18
AI Technical Summary
The prior art is difficult to accurately determine defects in wafers by considering different aspects/modes of wafer defect images.
A deep learning network-based method is adopted to make classification decisions using the coordination between multiple modes of wafer defect images, and a reference image is combined with reference images to improve classification accuracy by adding a mixture of patterns, such as color images, internal crack imaging images, black and white images, etc.
Accurate detection and classification of wafer defects is achieved, significantly reducing the number of labeled images and training time of deep learning models, and improving the defect detection efficiency in the manufacturing process.
Smart Images

Figure CN113627457B_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims the benefit and priority of U.S. Provisional Application No. 63 / 015,101, filed Apr. 24, 2020. Field of the Invention
[0003] This application generally relates to neural networks for semiconductor applications. Specifically, but not limited to, the present invention relates to a method and system for classifying defects in a wafer using wafer defect images based on a deep learning network. Background of the Invention
[0004] Generally, semiconductor substrate (i.e., wafer) manufacturing techniques have been continuously improved to incorporate an increasing number of features of semiconductor elements and multiple layers into a smaller surface area of a semiconductor wafer. Thus, the photolithography / lithography process can be used in semiconductor wafer manufacturing, which is more complex, thereby allowing more and more features to be incorporated into a smaller area of the semiconductor wafer (i.e., for achieving higher performance of the semiconductor wafer). As a result, the size of potential defects on a semiconductor wafer can be in the micron to sub-micron range due to the addition of more and more features. In addition, defects in a wafer can be, for example, defects caused by real and physical phenomena of the wafer, as well as false events / nuisance defects (i.e., a nuisance can be an irregular or false defect on the wafer but not a defect of interest).
[0005] Typically, defects in a semiconductor wafer can be detected based on obtaining a higher resolution image of the wafer using at least one of a high magnification optical system or a scanning electron microscope (SEM). To determine parameters such as the thickness, roughness, size, etc. of a defect, a high resolution wafer defect image can be generated. In addition, conventional systems disclose an imaging system that can be configured to scan a multimode energy source (e.g., light or electrons) on a physical version of a wafer and thereby generate an actual image for the physical version of the wafer. Further, a defect region can be determined by comparing the defect image with a reference image for anomaly detection and defect classification. Conventional systems can use a single deep learning model to detect and classify defects in a wafer. However, conventional systems may not be able to accurately determine defects in a wafer by considering different aspects / modes of the defect image corresponding to the wafer. Summary of the Invention
[0006] The present invention provides a method and system for classifying defects in a wafer using wafer defect images based on a deep learning network. Embodiments herein use the synergy between multiple modalities of wafer defect images to make classification decisions. Additionally, by adding a mixture of modalities, information can be obtained from different sources, such as: color images, internal crack imaging (ICI) images, black and white images, etc., to classify defect images. In addition to the mixture of modalities, reference images (e.g., a golden die image) can be used for each modality. The advantage of providing a reference image to each modality image is that it focuses on the defect itself rather than the underlying lithography associated with the defect image. Further, the reference image can be provided to the training process of the deep learning model, which can significantly reduce the number of labeled images and the training time required for the deep learning model to converge (i.e., when the entire dataset is passed forward and backward through the deep learning neural network).
[0007] Embodiments herein can use a Directed Acyclic Graph (DAG) as a combination of deep learning models, and each deep learning can use defect wafer images to handle different aspects of the problem or different forms of defects in the wafer. Additionally, the created DAG can have any number of models and multiple different images for each deep learning model. Further, a post-processing decision module can be configured to combine parameters, such as, two aspects of a defect inspection image and the result label of the defect inspection image, values from each deep learning model of the DAG, and metrology information (metadata) of the defect or parameters previously collected in a scanner machine. Based on the deep learning network, the DAG including the deep learning models can be used to accurately classify wafer defects using wafer defect images.
[0008] The features disclosed in the present invention help to accurately detect defects and classify defects in a wafer by analyzing multiple modalities of wafer defect images during manufacturing.
[0009] In one aspect, a computer-implemented method for classifying and inspecting defects in a semiconductor wafer includes: providing one or more imaging units; providing a computing unit; receiving a plurality of images taken of one or more wafers on the semiconductor wafer inspected by the one or more imaging units, wherein the plurality of images are captured using a plurality of imaging modes; providing one or more machine learning (ML) models, one of the plurality of ML models being associated with at least one computer processor, a database, and a memory associated with the computing unit; identifying and classifying one or more defects present in the semiconductor wafer into one or more defect categories from the plurality of ML models, the computer processor, wherein the plurality of ML models are configured in a directed acyclic graph (DAG) architecture, wherein each node in the DAG architecture represents an ML model, and wherein one or more of the ML models are configured as root nodes in the DAG architecture; the plurality of ML models being configured to classify one or more defects on one or more wafers in the semiconductor wafer, wherein the training includes: providing a plurality of labeled images and a plurality of reference images of the semiconductor wafer stored in the database to the one or more ML models from the plurality of ML models; configuring each ML model from the plurality of ML models to classify the plurality of labeled images into one or more defect categories using a corresponding reference image from the plurality of reference images; storing the one or more defect categories; inspecting one or more wafers on the semiconductor wafer for the presence of defects by imaging the one or more wafers; attempting to match an image of the one or more wafers with any one or more of the one or more defect categories; if there is a match between the one or more wafers and the one or more defect categories, classifying the one or more matching wafers as defective and transmitting the identification of the one or more defective wafers and rejecting them as defective.
[0010] In another aspect, the one or more ML models have a plurality of images from the plurality of imaging modes and a plurality of labeled images belonging to the imaging modes. In addition, each of the plurality of ML models is one of a supervised model, a semi-supervised model, and an unsupervised model.
[0011] In a further aspect, the plurality of modes includes at least one of: X-ray imaging, internal crack imaging (ICI), grayscale imaging, black and white imaging, and color imaging. In addition, the plurality of ML models are deep learning models.
[0012] In another aspect, the plurality of labeled images include labels associated with the one or more defect categories, wherein the plurality of labeled images are generated using a label model.
[0013] In an additional aspect, the computing unit according to the present invention includes one or more processors and a memory configured to execute the above method steps.
[0014] In one aspect, a method for classifying defects in a semiconductor wafer includes: capturing a plurality of images of the semiconductor wafer detected by one or more imaging units, wherein the plurality of images are captured using a plurality of imaging modes; providing the plurality of images from the plurality of ML models to one or more machine learning (ML) models to identify one or more defects in the semiconductor wafer and classify them into one or more defect categories, wherein the plurality of ML models are configured in a directed acyclic graph (DAG) architecture, wherein each node in the DAG architecture represents an ML model, and wherein one or more ML models are configured as root nodes in the DAG architecture; wherein the plurality of ML models are trained to classify one or more defects in the semiconductor wafer, and wherein the training includes: providing a plurality of labeled images and a plurality of reference images of the semiconductor wafer stored in a database from the plurality of ML models to one or more ML models; and configuring each ML model from the plurality of ML models to classify the plurality of labeled images into one or more defect categories using a corresponding reference image from the plurality of reference images.
[0015] In another aspect, the one or more ML models have a plurality of images from a plurality of imaging modes and a plurality of labeled images belonging to one imaging mode. In addition, each of the plurality of ML models is one of a supervised model, a semi-supervised model, and an unsupervised model. The plurality of modes include at least one of: X-ray imaging, internal crack imaging (ICI), grayscale imaging, black and white imaging, and color imaging. Each of the plurality of ML models is a deep learning model.
[0016] In another aspect, the plurality of labeled images include labels associated with the one or more defect categories, wherein the plurality of labeled images are generated using historical images of the semiconductor wafer. Features extracted from the plurality of modes are combined using one of late fusion techniques, early fusion techniques, or hybrid fusion techniques. Post-processing is further included, wherein the post-processing includes accurately classifying the plurality of images into one or more defect categories using classification information from each of the plurality of ML models.
[0017] In one aspect, a system for classifying and inspecting defects in a semiconductor wafer includes: one or more imaging units configured to capture multiple images of one or more wafers on a semiconductor wafer inspected by the one or more imaging units, wherein the multiple images are captured using multiple imaging modes; a computing unit including at least a computer processor, a database, and a memory, and configured to: provide the multiple images from multiple ML models to one or more machine learning (ML) models to identify and classify more defects in one or more wafers on the one or more machine learning (ML) semiconductor wafers to form one or more defect categories, wherein the multiple ML models are configured in a directed acyclic graph (DAG) architecture, wherein each node in the DAG architecture represents an ML model, wherein the one or more ML models are configured as root nodes in the DAG architecture, and the multiple ML models are configured to be trained to classify one or more defects on one or more wafers in a semiconductor wafer, wherein the computing unit is configured to: provide multiple labeled images and multiple reference images of the semiconductor wafer stored in the database to one or more ML models from multiple ML models;
[0018] Configure each ML model from the multiple ML models to classify the multiple labeled images into one or more defect categories using a corresponding reference image from the multiple reference images; then, store the defect categories and accept or reject one or more wafers on the premise of inspection, respectively, according to the presence of the match between the one or more wafers and the one or more defect categories.
[0019] In another aspect, the one or more imaging units include an automatic optical inspection (AOI) device, an automatic X-ray inspection (AXI) device, a joint test action group (JTAG) device, and an in-circuit test (ICT) device. In addition, the computing unit receives multiple labeled images including labels related to one or more defect categories from a label model, wherein the label model generates the multiple labeled images using historical images of the semiconductor wafer.
[0020] In yet another aspect, one of a late fusion technique, an early fusion technique, or a hybrid fusion technique is used to combine features extracted from multiple modes. The computing unit is further configured to post-process the outputs of the multiple ML models, wherein the computing unit precisely classifies the multiple images into one or more defect categories using classification information from each of the multiple ML models. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The embodiments of the present invention itself, as well as its preferred usage patterns, its further objectives and advantages, will be best understood by referring to the following detailed description of illustrative embodiments when read in conjunction with the accompanying drawings. One or more embodiments are now described with reference to the drawings, wherein:
[0022] Figure 1 Shows a block diagram of a system for classifying defects in a wafer using a wafer defect image based on a deep learning network, according to some embodiments of the present invention;
[0023] Figure 2 Shows a block diagram of a multi-modal late fusion deep learning model according to some embodiments of the present invention, which can be used as one of the deep learning models for classifying defects in a wafer using a wafer defect image;
[0024] Figure 3 Shows a block diagram of a multi-modal hybrid fusion deep learning model according to some embodiments of the present invention, which can be used as one of the deep learning models for classifying defects in a wafer using a wafer defect image;
[0025] Figure 4 Shows a block diagram of a multi-modal early fusion deep learning model according to some embodiments of the present invention, which can be used as one of the deep learning models for classifying defects in a wafer using a wafer defect image;
[0026] Figure 5a Shows a schematic diagram of a DAG topology using a series of deep learning models according to some embodiments of the present invention;
[0027] Figure 5b Shows a schematic diagram of exemplary result labels for each deep learning model that customizes the flow path in a DAG, according to some embodiments of the present application;
[0028] Figure 6a Is a flow chart describing a method for classifying defects in a wafer using a wafer defect image based on a deep learning network, according to some embodiments of the present invention; and
[0029] Figure 6b Is a flow chart that describes a method for calculating features representing defect metadata according to some embodiments of the present invention if the defect metadata of a wafer defect image is not stored in an electronic device.
[0030] The accompanying drawings depict embodiments of the present invention for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods shown herein can be used without departing from the principles of the disclosure described herein. Detailed Description
[0031] The foregoing generally outlines the features and technical advantages of the present invention, so as to better understand the following detailed description of the present invention. Those skilled in the art should understand that the disclosed concepts and specific embodiments can be readily used as a basis for modifying or designing other structures for achieving the same purpose of the present invention.
[0032] The technical features of the present invention, which are considered novel features, including its organization and method of operation, as well as further objects and advantages, will be better understood in the following consideration in conjunction with the accompanying drawings. However, it should be clearly understood that each of the provided drawings is for illustrative and descriptive purposes only and should not be regarded as a limiting definition of the present invention.
[0033] Figure 1 A block diagram of a system 100 for classifying defects in a wafer based on a deep learning network using wafer defect images according to some embodiments of the present invention is shown.
[0034] Throughout the invention, the term "wafer" generally refers to a substrate formed of semiconductor or non-semiconductor materials. For example, semiconductor or non-semiconductor materials may include, but are not limited to, single crystal silicon, gallium arsenide, indium phosphide, etc. The wafer may include one or more layers, and the layers may include, for example, but are not limited to, resist, node materials, conductive materials, semiconductor materials, etc. For example, one or more layers formed on the wafer may be patterned or unpatterned. For example, the wafer may include multiple wafers, each wafer having repeatable pattern features. The formation and processing of these material layers may result in a complete device. Additionally, as used herein, the term "surface defect" or "defect" refers to a defect (e.g., a particle) that is entirely above the upper surface of the wafer and a defect that is partially below or entirely below the upper surface of the wafer. Thus, the classification of defects is particularly useful for semiconductor materials such as wafers and the materials formed on the wafers. Additionally, for bare silicon wafers, silicon-on-insulator (SOI) films, strained silicon films, and dielectric films, it may be particularly important to distinguish between surface and subsurface defects. The embodiments herein can be used to inspect wafers containing silicon or silicon-containing layers formed thereon, such as, for example, silicon carbide, carbon-doped silicon dioxide, silicon-on-insulator (SOI), strained silicon, silicon-containing dielectric films, etc.
[0035] In Figure 1In an embodiment, the system 100 includes an imaging device 102 and an electronic device 104. The imaging device 102 is associated with the electronic device 104 via a communication network 106. The communication network 106 can be a wired network or a wireless network. In one embodiment, the imaging device 102 can be, but is not limited to, at least one of an automatic optical inspection (AOI) device, an automatic X-ray inspection (AXI) device, a joint test action group (JTAG) device, an in-circuit test (ICT) device, etc. The imaging device 102 includes at least one of, but is not limited to, a light source 108, a camera lens 110, a defect detection module 112, and an imaging storage unit 126. For example, the defect detection module 112 associated with the imaging device 102 can detect multiple surface feature defects of a wafer, such as, but not limited to, at least one of silicon nodes (i.e., bumps), scratches, stains, dimensional defects (such as open circuits, short circuits, and solder thinning, etc.). In addition, the defect detection module 112 can also detect incorrect components, missing components, and mis-placed components, because the imaging device 102 is capable of performing all visual inspections.
[0036] In addition, the electronic device 104 can be, but is not limited to, at least one of a mobile phone, a smart phone, a tablet, a handheld device, a phablet, a laptop computer, a computer, a personal digital assistant (PDA), a wearable computing device, a virtual / augmented reality device, an Internet of Things device (IoT device), etc. The electronic device 104 further includes a storage unit 116, a processor 118, and an input / output (I / O) interface 120. In addition, the electronic device 104 includes a deep learning module 122. The deep learning module 122 enables the electronic device 104 to classify defects in the wafer using the wafer defect images obtained from the imaging device 102. The electronic device 104 can also include an application management framework for classifying defects in the wafer using a deep learning network. The application management framework can include different modules and sub-modules to perform operations of classifying wafer defects using wafer defect images based on the deep learning network. Additionally, the module and sub-module can include at least one or both of a software module or a hardware module.
[0037] Therefore, the embodiments described herein are configured for image-based wafer process control and yield improvement. For example, one embodiment herein relates to a system and method for classifying defects in a wafer using wafer defect images based on a deep learning network.
[0038] In one embodiment, the imaging device 102 may be configured to capture images of wafers placed therein. For example, the images may include, for example, at least one of inspection images, optical or electron beam images, wafer inspection images, optical and SEM-based defect inspection images, simulation images, clips from design layouts, and the like. Additionally, the imaging device 102 may subsequently be configured to store the captured images in an image storage unit 126 associated with the imaging device 102. In one embodiment, an electronic device 104 communicatively coupled to the imaging device 102 may be configured to retrieve images stored in the image storage unit 126 associated with the imaging device 102. For example, the images may include black and white images, color images, internal crack imaging (ICI) images, images previously scanned using the imaging device 102 (e.g., AOI machine, etc.), images from the imaging storage unit 126 or several storage units (not shown), images obtained in real-time from the imaging device 102 (e.g., AOI machine, etc.), and the like. Then, the electronic device 104 is configured to load from an external database (not shown) or from a storage unit 116 associated with the electronic device 104 at least one reference image corresponding to at least one of a black and white reference image, a color reference image, and an ICI reference image representing the same scan for a wafer area of a defect-free inspection image of a wafer. Further, the electronic device 104 is configured to provide the reference images and the wafer images having associated patterns to a deep learning module 122. In one aspect of the present invention, multiple deep learning models or deep learning classifiers may be trained using different types of classifications of defects in wafers. The multiple deep learning models or deep learning classifiers may be, but are not limited to, at least one of a convolutional neural network (CNN) (e.g., LeNet, AlexNet, VGGNet, GoogleNet, ResNet, etc.), a recurrent neural network (RNN), a generative adversarial network (GAN), a random forest algorithm, an autoencoder, and the like. The purpose of training several deep learning models is that each model can be created to handle the synergy of concentrated patterns of defects. Thus, several deep learning models can be established to hierarchically classify all defects based on similarities and differences and distribute the training classification process. In another embodiment, in order to shorten the training process and the number of classified images, the reference images of each pattern image may be added to the architecture of each deep learning model. The reference images during the training process may enable faster adjustment or training of the deep learning internal parameters by providing information about the internal relationship between the inspected images and the reference images to the deep learning module 122. Additionally, if defects appear in different wafers, the deep learning module 122 trained to classify specific defects in a single wafer may also dynamically classify the trained defects. Thus, the training process may discard common events, such as underlying lithography, etc., and may focus on actual defects.
[0039] In one embodiment, the deep learning model can be connected in a parallel architecture or a serial architecture. Further, the electronic device 104 can be configured to use a directed acyclic graph (DAG) architecture of multiple deep learning models to generate a classification determination of the wafer image. For example, multiple trained deep learning models can be called from the deep learning module 122 and subsequently connected in a directed acyclic graph (DAG) architecture for classification processing of wafer defects. In addition, the electronic device 104 can be configured to save the classified wafer image including relevant directories and metadata results (i.e., defect metadata) in an external database or storage unit 116 associated with the electronic device 104.
[0040] In another embodiment, the electronic device 104 can be configured to load previously calculated and stored defect metadata from an external database or storage unit 116 associated with the electronic device 104. For example, the metadata includes different features, but is not limited to, the size of the defect, the histogram of the defect, the maximum color or grayscale value of the defect, the minimum color or grayscale value of the defect, etc. If the defect metadata is not stored, then the electronic device 104 can be configured to calculate the features representing the defect metadata. Then, the electronic device 104 can be configured to provide the inspected image, the reference image, and the defect metadata (i.e., the metadata features of the defect) to the trained deep learning model. Therefore, the electronic device 104 can be configured to use a directed acyclic graph (DAG) architecture of multiple deep learning models to generate a classification determination of the wafer image. In addition, the electronic device 104 can be configured to store the classified wafer image including the associated metadata results (i.e., defect metadata) in the storage unit 116 or external database associated with the electronic device 104.
[0041] In addition, the image and defect metadata can also be stored in an external database (not shown). For example, an external memory can be used for the training process of the deep learning model / classifier. As an example, the images stored in the external database (or the image storage unit 126) can be black-and-white images, color images, ICI images, images previously scanned by the AOI device, images containing wafer defects, false events, and nuisance defects, etc. The images can be labeled before being stored in the external database (or the image storage unit 126). For each defect found in the images stored in the external database (or the image storage unit 126), a set of metadata features extracted from the defect image is stored in the external database (or the image storage unit 126). The metadata defect features can be provided by the user or by the AOI scanner results (or the metadata defect features can be created for data retrieval of the deep learning classifier). In addition, reference images including color reference images, black-and-white reference images, and / or ICI reference images (e.g., gold wafer images) can also be stored in the external database. The reference images are images of the same wafer. The external database can also be used to perform the training process of the deep learning model.
[0042] The embodiments herein use the synergy between the centralized patterns of wafer defect images for classification decisions. In addition, by adding a mixture of patterns, information can be obtained from different sources such as color images, ICI, black-and-white images, etc. to classify the defect images. In addition to the mixture of patterns, reference images (e.g., gold wafer images) can be used for each pattern. The advantage of providing the reference images to each pattern image is to focus on the defect itself rather than the underlying lithography associated with the defect image. This method saves processing power, memory utilization, and time. In addition, the reference images provided to the training process of the deep learning model can significantly reduce the number of labeled images and the training time (i.e., when a complete data set is passed forward and backward through the deep learning neural network) required for the fusion of the deep learning model.
[0043] Figure 2 A block diagram of a multi-modal late fusion deep learning model according to some embodiments of the present invention is shown, which can be used as one of the deep learning models for classifying defects in a wafer using wafer defect images.
[0044] In one embodiment, the electronic device 104 includes a multi-modal convolutional neural network (CNN) configured to integrate images acquired by different image sensors in a single forward pass. A deep learning model such as a multi-modal late fusion deep learning model can reason about two sensor images, for example, by using the ICI image of the first deep learning model and the color image of the second deep learning model. In addition, as Figure 2As shown, the multi-modal CNN model includes a CNN model for encoding a color image and an ICI image respectively and making a decision to combine the two. The trained multi-modal post-fusion deep learning model can be used to process each mode to allow separate decisions for each mode. Finally, a central classification layer can provide a joint decision based on different modes.
[0045] Figure 3 FIG. shows a block diagram of a multi-modal hybrid fusion deep learning model according to some embodiments of the present invention, which can be used as one of the deep learning models for classifying defects in a wafer using wafer defect images.
[0046] A multi-modal CNN model, such as a multi-modal hybrid fusion deep learning model, may include a first CNN model for encoding a color image, a second CNN model for encoding an ICI image, and a third CNN model for jointly representing the color and ICI defect images. The third / last CNN model can learn the inter-model relationship between the color image and the ICI image before making a classification decision.
[0047] Figure 4 FIG. shows a block diagram of a multi-modal early fusion deep learning model according to some embodiments of the present invention, which can be used as one of the deep learning models for classifying defects in a wafer using wafer defect images.
[0048] The multi-modal early fusion deep learning model may include a CNN model for jointly representing a color defect image and an ICI defect image by simultaneously processing joint feature points in a single multi-modal image.
[0049] Figure 5a FIG. shows a schematic diagram of a DAG topology using a series of deep learning models according to some embodiments of the present invention.
[0050] As Figure 5a shown, multiple deep learning models can be connected as a polytree, and the polytree can be a directed acyclic graph (DAG) of deep learning models, the underlying undirected graph of which can be a tree, as Figure 5a shown, a multi-modal hybrid fusion deep learning model, a multi-modal early fusion deep learning model, a deep learning model with a single input image, an autoencoder and / or a generative adversarial network (GAN) deep learning model with one or two input images. The DAG may include a unique topological order, and each deep learning model can be located at a node of the DAG. In addition, each node can be directly connected to one or more previous nodes and then to one or more nodes. Moreover, the result label of each deep learning model defines the flow path in the DAG. For example, as Figure 5aAs shown, the result image "Label 1" in "Model 1" will continue to be evaluated in "Model 3".
[0051] As an example, considering that the result labels of each model are described in Figure 5b . The result label "Label 1: A" of "Model 1" may have a probability value of 0.9, while the probability value of "Label 1: B" may be 0.1. Similarly, the probability value of the result label "Label 2: B" of "Model 3" may be 0.2, the probability value of "Label 3: B" may be 0.7, and the probability value of "Label 3: C" may be 0.1. In addition, the probability value of the result label "Label 5: A" of "Model 5" is 0.1, the probability value of "Label 5: B" is 0.1, the probability value of "Label 5: C" is 0.2, and the probability value of "Label 5: D" is 0.6. Each deep learning model in the DAG can be unique and can be designed to handle a specific part of the classification problem. For example, one deep learning model in the DAG can be a ResNet model, another can be a GoogleNet model, and still another can be a multi-modal deep learning model. At the end of the DAG path, each image (as Figure 5a shown) can be evaluated in a post-processing module, where a decision can be made based on the results of the deep learning models that interact with the image.
[0052] Figure 6a Described is a flowchart of a method 600a for classifying defects in a wafer using a wafer defect image based on a deep learning network according to some embodiments of the present application.
[0053] In block 601, an image of the wafer is captured by imaging device 102. In block 602, the captured image is stored by the imaging device 102 ( Figure 1 ) in an image storage unit 126 ( Figure 1 ) associated with the imaging device 102. In block 603, the image stored in the image storage unit 126 associated with the imaging device 102 is processed by an electronic device 104 ( Figure 1)The retrieval is performed. At block 604, the electronic device 104 receives at least one reference image corresponding to at least one black and white reference image, a color reference image, and an ICI reference image, where the reference image represents the same area of a wafer scanned with an inspected image without defects in the wafer. At block 605, the electronic device 104 uses a plurality of trained deep learning models / classifiers with associated expected pattern images from the deep learning module 122 of the electronic device 104. At block 606, the plurality of trained deep learning models are connected by the electronic device 104 in a directed acyclic graph (DAG) architecture for the classification process of wafer image defects. At block 607, the electronic device 104 uses the directed acyclic graph (DAG) architecture of the plurality of deep learning models to generate a classification determination of the wafer image. Finally, at block 608, the electronic device 104 stores the classified wafer image including associated metadata results (i.e., defect metadata) in an external database or storage unit 116 of the electronic device 104.
[0054] Figure 6b Described is a flow of method 600b for calculating features representing defect metadata in the case where the electronic device 104 does not store the defect metadata of the wafer defect image according to some embodiments of the present application.
[0055] At block 611, the electronic device 104 receives previously calculated and stored defect metadata from an external database or storage unit 116 of the electronic device 104. For example, the metadata includes different features of the defect, but is not limited to, the size of the defect, the histogram of the defect, the maximum color or grayscale value of the defect, the minimum color or grayscale value of the defect, etc. At block 612, if the defect metadata is not stored, the electronic device 104 calculates the features representing the defect metadata.
[0056] Embodiments herein may utilize a directed acyclic graph (DAG) as a combination of deep learning models, and each deep learning can use defect wafer images to handle different aspects of the problem or different forms of defects in the wafer. In addition, the DAG can create models with any number of multiple different images (e.g., six images) for each deep learning model. In addition, the post-processing decision module can be configured to combine parameters, e.g., two aspects of the defect inspection image and the result label of the defect inspection image, values from each deep learning model of the DAG, and metrology information (metadata) of the defect or defects previously collected in the scanner. Based on the deep learning network, the DAG including the deep learning models can be used to accurately classify wafer defects using wafer defect images.
[0057] For essentially any use of plural and / or singular terms herein, those skilled in the art can convert from plural to singular and / or from singular to plural according to the context and / or application. For the sake of clarity, various singular / plural permutations can be explicitly set forth herein.
[0058] Those skilled in the art will understand that, generally speaking, the terms used herein are usually "open-ended" terms (e.g., the term "including" should be interpreted as "including but not limited to", the term "having" should be interpreted as "having at least", the term "includes" should be interpreted as "including but not limited to", etc.). Those skilled in the art should further understand that the specific number of the recited claims is intentional. For example, as an aid to understanding, the detailed description may include the use of introductory phrases "at least one" and "one or more" to introduce the claims. However, the use of such phrases should not be construed as limiting any particular claim that includes the introduced claim statement by the indefinite article "a" or "an" to an invention that only includes one such recitation, even if the same claim includes the introductory phrase "one or more" or "at least one" and the indefinite article, such as "a" or "an" (e.g., "a" or "an" should generally be interpreted as "one or more" or "at least one"); the same applies to the use of the definite article to introduce the claims. In addition, even if a specific number of the claims is explicitly recited, those skilled in the art will also recognize that such a recitation should generally be interpreted as meaning at least the recited number (e.g., "two recitations" without any other modifiers generally means at least two recitations, or two or more recitations).
[0059] Although aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for illustrative purposes and not for limitation, and the true scope and spirit are represented by the following detailed description.
[0060] Reference numeral
[0061] Reference numeral Detailed description 100 System 102 Imaging device 104 Electronic device 106 Communication network 108 Light source 110 Camera lens 112 Defect detection module 116 Storage unit 118 Processor 120 I / O interface 122 Deep learning module 126 Image storage unit
Claims
1. A method for classifying defects in a semiconductor wafer, the method comprising: providing one or more imaging units configured to capture images in different imaging modes; providing a computing unit; providing two or more machine learning models from a plurality of machine learning models, the two or more machine learning models being at least associated with a computer processor, a database, and a memory of the computing unit, wherein the two or more machine learning models are trained to classify one or more defects on one or more wafers in one or more semiconductor wafers, wherein the training of the two or more machine learning models comprises: providing the two or more machine learning models with a plurality of labeled images of the one or more semiconductor wafers stored in the database, and configuring each of the two or more machine learning models to classify the plurality of labeled images into at least one defect category; connecting the two or more trained machine learning models in a directed acyclic graph architecture, wherein the directed acyclic graph architecture includes non-leaf nodes and terminal nodes, wherein at least some of the non-leaf nodes in the directed acyclic graph architecture represent respective machine learning models, wherein the terminal nodes represent post-processing procedures, wherein the non-leaf nodes include one or more root nodes, wherein each of the one or more root nodes does not have any incoming edges reaching it, wherein the directed acyclic graph architecture includes at least a first path and a second path, the first path starting from a root node and ending at the terminal node, the first path including a first number of non-leaf nodes, wherein the second path starts from the root node and ends at the terminal node, the second path including a second number of non-leaf nodes, the first number being greater than the second number, wherein the directed acyclic graph architecture includes nodes representing respective machine learning models, wherein the directed acyclic graph architecture includes: a first edge and a second edge, wherein the first edge is associated with a first predicted label of a first defect category, the first edge emanating from the node and entering another non-leaf node, wherein the second edge is associated with a second predicted label of a second defect category, the second edge emanating from the node and entering the terminal node; and detecting a defect located on a wafer included in a semiconductor wafer by: receiving at least one image from the one or more imaging units, the at least one image representing the one defect captured from the wafer in the semiconductor wafer, determining whether at least one machine learning model in the directed acyclic graph architecture is skipped based on an output of a previous machine learning model in the directed acyclic graph architecture, wherein if the at least one machine learning model is skipped, the at least one image is not input to the at least one machine learning model determined to be skipped, According to the directed acyclic graph architecture, the one defect is classified by an unskipped machine learning model, where the unskipped machine learning model excludes the at least one skipped machine learning model, thereby obtaining a classification decision of the unskipped machine learning model. A classification decision of the one defect is generated by performing post-processing associated with the terminal node, where the post-processing is performed based on the classification decision of the unskipped machine learning model, and The generated classification decision is output.
2. The method according to claim 1, wherein a plurality of images belonging to different imaging modalities are provided to the two or more machine learning models.
3. The method according to claim 1, wherein each of the two or more machine learning models is one of the following: a supervised model, a semi-supervised model, and an unsupervised model.
4. The method according to claim 1, wherein the imaging modality includes at least one of the following: X-ray imaging, grayscale imaging, black-and-white imaging, and color imaging.
5. The method according to claim 1, wherein the imaging modality includes internal crack imaging.
6. The method according to claim 1, wherein the training of the two or more machine learning models comprises: configuring a first machine learning model among the two or more machine learning models to classify the plurality of labeled images into a first set of defect categories from the at least one defect category, and configuring a second machine learning model among the two or more machine learning models to classify the plurality of labeled images into a second set of defect categories from the at least one defect category, where the first set of defect categories is different from the second set of defect categories.
7. The method according to claim 6, wherein the number of defect categories in the first set of defect categories is greater than the number of defect categories in the second set of defect categories.
8. The method according to claim 1, wherein the training of the two or more machine learning models comprises: providing a first set of images among the plurality of labeled images to a first machine learning model among the two or more machine learning models and configuring the first machine learning model to classify the first set of images into the at least one defect category, and providing a second set of images among the plurality of labeled images to a second machine learning model among the two or more machine learning models and configuring the second machine learning model to classify the second set of images into the at least one defect category, where the first set of images is different from the second set of images.
9. The method according to claim 8, wherein the first set of images has a first imaging modality, while the second set of images has a second imaging modality, where the first imaging modality is different from the second imaging modality.
10. The method according to claim 1, wherein the classification of the one defect is performed by using the unskipped machine learning model included in the first path, where the method further comprises: Generate a classification decision for a second defect on the wafer included on the semiconductor, wherein generating the classification decision for the second defect includes classifying the second defect by utilizing a second set of machine learning models included in the second path, wherein the classification decision for the one defect is performed by utilizing a larger number of machine learning models as compared to the number of machine learning models for the classification decision for the second defect.
11. The method according to claim 1, wherein each of the two or more machine learning models from the plurality of machine learning models is a deep learning model.
12. The method according to claim 1, wherein the plurality of labeled images includes labels associated with the at least one defect category, wherein the plurality of labeled images is generated using historical images of the one or more semiconductor wafers.
13. The method according to claim 1, wherein the post-processing comprises: Utilizing metrology information together with the classification decisions of the non-skipped machine learning models to determine the classification decision for the one defect.
14. The method according to claim 13, wherein the metrology information includes a size measurement of the one defect.
15. A system for classifying defects in a semiconductor wafer, the system comprises: One or more imaging units of different imaging modes; Two or more machine learning models, wherein the two or more machine learning models are trained to classify one or more defects on one or more wafers in one or more semiconductor wafers, wherein the training of the two or more machine learning models includes: Providing the two or more machine learning models with a plurality of labeled images of the one or more semiconductor wafers stored in a database, and Configuring each of the two or more machine learning models to classify the plurality of labeled images into at least one defect category; Wherein the two or more trained machine learning models are connected in a directed acyclic graph architecture, wherein the directed acyclic graph architecture includes non-leaf nodes and terminal nodes, wherein at least some of the non-leaf nodes in the directed acyclic graph architecture represent respective machine learning models, wherein the terminal nodes represent a post-processing process, wherein the non-leaf nodes include one or more root nodes, wherein each of the one or more root nodes does not have any incoming edges reaching it, Wherein the directed acyclic graph architecture includes at least a first path and a second path, the first path starting from a root node and ending at the terminal node, the first path including a first number of non-leaf nodes, wherein the second path starts from the root node and ends at the terminal node, the second path including a second number of non-leaf nodes, the first number being greater than the second number, The directed acyclic graph architecture includes nodes representing respective machine learning models, and the directed acyclic graph architecture includes: a first edge and a second edge, where the first edge is associated with a first prediction label of a first defect category, the first edge emanates from the node and enters another non-leaf node, and the second edge is associated with a second prediction label of a second defect category, the second edge emanates from the node and enters the terminal node; and a computing unit, the computing unit includes at least a computer processor, a database, and a memory, and is configured to detect a defect on a wafer included in a semiconductor wafer by: Receiving at least one image from the one or more imaging units, the at least one image representing the one defect captured from the wafer in the semiconductor wafer, Based on the output of the previous machine learning model in the directed acyclic graph architecture, determining whether at least one machine learning model in the directed acyclic graph architecture is skipped, where If the at least one machine learning model is skipped, the at least one image is not input to the at least one machine learning model determined to be skipped, Classifying the one defect by the non-skipped machine learning models according to the directed acyclic graph architecture, where the non-skipped machine learning models exclude the at least one skipped machine learning model, thereby obtaining a classification decision of the non-skipped machine learning models, Generating a classification decision of the one defect by performing post-processing, the post-processing being associated with the terminal node and being performed based on the classification decision of the non-skipped machine learning models, and Outputting the generated classification decision.
16. The system according to claim 15, wherein the one or more imaging units include at least one of the following: an automated optical inspection device, an automated X-ray inspection device, a joint test action group device, and an in-line test device.
17. The system according to claim 15, wherein the plurality of labeled images include labels related to the at least one defect category, and the plurality of labeled images are generated using historical images of the one or more semiconductor wafers.
18. The system according to claim 15, wherein the post-processing includes: Utilizing metrology information together with the classification decision of the non-skipped machine learning models to determine the classification decision of the one defect.
19. The system according to claim 15, wherein the system is configured to generate a classification decision for the one defect by following the first path in the directed acyclic graph architecture, and the system is configured to generate a second classification decision for a second defect by following the second path in the directed acyclic graph architecture, and the classification decision for the one defect is performed using a larger number of machine learning models compared to the number of machine learning models for the classification decision of the second defect.
Citation Information
Patent Citations
Method of deep learning-based examination of a semiconductor specimen and system thereof
TW201935590A
Cited By
Chip test failure mode analysis and classification method based on artificial intelligence
CN122336447A
An AI-based method for analyzing and classifying chip test failure modes
CN122336447B