Distributed neural network for table data classification
By adopting the new neural network classifier model (TNNC), the problem of inaccuracy and efficiency in tabular data processing and semiconductor sample defect detection in the prior art is solved, and higher classification accuracy and shorter processing time are achieved.
Patent Information
- Application Number
- CN202510076977.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-21
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, when processing table data, especially in semiconductor sample defect detection, there are problems of accuracy and efficiency, and it is difficult to effectively distinguish between defects of interest and noise.
A new neural network classifier model (TNNC) is adopted, which includes a set input layer, a set integral layer, and multiple neural network threads. Each NN thread is designed to include at least two hidden layers and random inactivation operations are applied on the input feature set to improve classification accuracy and efficiency.
TNNC significantly improves the accuracy and efficiency of tabular data classification, exhibits better true to false positive ratios, and has a shorter processing time.
Smart Images

Figure CN120031106A_ABST
Abstract
Description
Technical Field
[0001] The disclosed subject matter relates to a computer system and method for processing tabular data. The disclosed subject matter also relates to defect detection in semiconductor samples. Background Art
[0002] Tabular data refers to information organized in a structure including rows and columns, such as a table. Tabular data is commonly used in databases, spreadsheets, and datasets and is a common representation used in machine learning (ML) analysis.
[0003] During processing, various types of data are sometimes transformed into tabular data. In image processing, this approach can include extracting features from an image using image processing techniques, and then representing these features as numerical values in a tabular format.
[0004] A wafer is a thin, usually circular slice of semiconductor material, often made of silicon, that is used as a substrate for manufacturing integrated circuits. A die (also referred to as a "semiconductor die") refers to an independent and discrete component of an integrated circuit. Each die contains a set of specific electronic components, all of which are manufactured together on the same wafer. Typically, during the manufacturing process, multiple dies are formed on a single wafer, each of which is a copy of the same integrated circuit, effectively producing identical copies of the integrated circuit design.
[0005] The current demand for high density and high performance associated with ultra-large scale integration of manufactured devices requires sub-micron features, increased transistor and circuit speeds, and increased reliability. As semiconductor processes advance, pattern sizes (such as line widths) and other types of critical dimensions continue to shrink. Such demands require the formation of device features with high precision and high uniformity, which in turn requires careful monitoring of the manufacturing process, including automated inspection of the device while it is still in the form of a semiconductor wafer.
[0006] Runtime inspection typically includes generating images for semiconductor samples (e.g., wafers or dies), and applying defect detection on the imaged output to detect defects of interest (DOIs) in the die, such as impurities or irregularities, which may affect the functionality and / or performance of the die. A defect map may be generated to show suspected locations on the sample that have a high probability of being a defect. Figure 6 Images of a die (A and B) are shown showing defects of interest marked by white circles. Summary of the invention
[0007] Some DOI detection methods involve machine learning techniques used in automated inspection processes, which are particularly helpful in promoting higher yields. For example, supervised machine learning can be used for this purpose. In some cases, as part of this process, the image output of the die is transformed into tabular data, and a machine learning classification algorithm is applied on the tabular data.
[0008] Tabular data used as input for machine learning typically has the appearance of a structured data set organized in rows and columns. Each row (or record) represents a separate sample (or observation), while each column corresponds to a feature associated with the sample.
[0009] Models commonly used to process tabular data include Random Forests and Gradient Boosting Machines, such as XGBoost and catboost. At the same time, neural networks in general and deep learning algorithms in particular are known to have insufficient performance when applied to tabular data. One reason for the insufficient performance is related to the fact that such models, which typically involve neural networks with multiple layers, are designed to learn hierarchical features from unstructured data types (such as images, text, and sequences) and are not well suited to process tabular data.
[0010] Factors that can be used to determine the quality of machine learning classification output of tabular data include accuracy and efficiency. The accuracy of the classification output can be defined based on the ratio between true positive (TP) detections and false positive (FP) detections. TP detection refers to the correct detection of an object of interest in a sample and the classification of the sample as a sample including at least one object of interest. FP detection refers to the incorrect detection of noise in the data as an object of interest, and therefore the incorrect classification of the sample. In the context of semiconductor manufacturing, TP refers to the correct detection of defects of interest (DOI) in a die or wafer, and FP refers to the incorrect detection of defects in a die or wafer. A high TP to FP ratio is desired. The efficiency of the classification can be defined based on the processing time, which includes both the training time and the inference time of the machine learning model.
[0011] Unlike DOI, which refers to actual defects, noise refers to any random or undesired interference that may obscure the true signal or introduce errors in measurement or inspection, resulting in incorrect detection of defects. In the context of semiconductor manufacturing, noise can be generated from various sources, including equipment, environmental factors, or signal processing. Noise also includes defects that have no technical impact on the performance of the sample (e.g., die). During defect detection, it is desirable to increase the detection of DOI (TP detection) while reducing the detection of noise (FP detection).
[0012] The subject matter of the present disclosure includes a novel computer-implemented method and computer system for classifying tabular data using a new neural network classifier model (also referred to herein as a "tabular neural network classifier" or TNNC). The disclosed method and system are characterized by improved accuracy and efficiency compared to other existing tabular data classification techniques such as random forests, XGBoost, etc. The inventors have found that the TNNC generally exhibits a better TP to FP ratio in the classification output and exhibits shorter processing time than existing tabular data classification techniques.
[0013] According to a first aspect of the disclosed subject matter, there is provided a computer system comprising at least one processing circuit system configured to classify tabular data using a machine learning model, the processing circuit system being configured to:
[0014] obtaining tabular data comprising one or more records, each record corresponding to a respective sample and comprising a plurality of features characterizing the sample;
[0015] Utilizing a machine learning (ML) model on the tabular data, wherein the ML model includes an aggregate input layer, an aggregate integration layer, and a plurality of neural network (NN) threads, each NN thread being connected between the aggregate input layer and the aggregate integration layer and being designed as a separate neural network (NN) including at least two hidden layers, and an output layer;
[0016] wherein the processing circuitry is configured to utilize the ML model to:
[0017] Provide the ML model with a set of input features extracted from the tabular data;
[0018] For each NN thread:
[0019] Applying a random dropout operation on the set of input features to thereby obtain a corresponding feature subset selected from the set of input features;
[0020] applying a neural network algorithm on the respective feature subsets to thereby obtain respective NN thread outputs at respective output layers; and
[0021] Provide the corresponding NN thread output to the collective integration layer;
[0022] An ensemble output of the ML model is calculated at an ensemble integration layer based on the plurality of NN outputs, the ensemble output indicating classification of the sample into a class selected from the at least two classes.
[0023] In addition to the above features, the method according to this aspect of the disclosed subject matter may optionally include one or more of the following features (i) to (xix) in any technically possible and technically feasible combination or arrangement:
[0024] i. wherein the processing circuit system is configured to apply the same neural network algorithm in all NN threads.
[0025] ii. wherein the processing circuit system is configured to apply the neural network algorithm in all NN threads concurrently.
[0026] iii. Wherein the sample is a defect in a semiconductor specimen (eg, a die or wafer), and the feature is a property that characterizes the defect; and wherein the aggregate output indicates whether the sample is classified as a defect of interest (DOI).
[0027] iv. wherein the processing circuitry is operatively connected to an inspection system dedicated to detecting defects of interest during operation as part of a semiconductor manufacturing process and is configured to determine whether a suspected defect detected by the inspection system is a DOI or noise.
[0028] v. Wherein the tabular data is generated based on the image data, wherein each record in the tabular data corresponds to a respective sample (eg, a defect) identified in an image (eg, of a semiconductor sample), and the features in the record are features that characterize the sample.
[0029] vi. wherein the processing circuit system is further configured to obtain tabular data to: obtain one or more images; process the one or more images and identify at least one sample and a plurality of corresponding features characterizing the sample; and generate a record in a tabular data format corresponding to the at least one sample and add the features to the record.
[0030] vii. Wherein each NN thread is designed as a separate neural network (NN) comprising exactly two hidden layers.
[0031] viii. wherein each NN thread is designed as a separate neural network (NN), and at least two NN threads include different numbers of hidden layers.
[0032] ix. wherein the processing circuit system is configured to apply at least two different activation functions such that at least two different NN threads each apply a different activation function.
[0033] x. Wherein the processing circuit system is configured to apply the same activation function (e.g., Sigmoid function) on all corresponding output layers in each NN thread.
[0034] xi. Wherein the processing circuit system is configured to apply a unique activation function at each NN thread.
[0035] xii. wherein the processing circuitry is configured to apply dropout on the set of input features received in the set input layer such that at least two different subsets each include a different number of features.
[0036] xiii. wherein the corresponding classification of one or more samples in the tabular data is binary classification.
[0037] xiv. wherein the corresponding classification of one or more samples in the tabular data is multi-class classification.
[0038] xv. wherein the output layer is part of an NN thread.
[0039] xvi. wherein the processing circuitry is configured to apply regularization in at least one of two or more hidden layers to thereby prevent overfitting of the ML model.
[0040] xvii. wherein the corresponding classification of one or more samples in the tabular data is binary classification.
[0041] xviii. wherein the processing circuitry is configured to apply regularization in at least one of two or more hidden layers to thereby prevent overfitting of the ML model.
[0042] xix. wherein the processing circuitry is configured to calculate an ensemble output of the ML model based on an average of a plurality of NN outputs.
[0043] According to a second aspect of the subject matter of the present disclosure, there is provided a computer-implemented method of classifying tabular data using a machine learning (ML) model, the method comprising:
[0044] obtaining tabular data including one or more records, each record corresponding to a respective sample and including a plurality of features characterizing the sample;
[0045] utilizing a machine learning (ML) model on the tabular data, wherein the ML model includes a set input layer, a set integration layer, and a plurality of neural network (NN) threads, each NN thread being connected between the set input layer and the set integration layer and being designed as a separate neural network (NN) including at least two hidden layers, and an output layer;
[0046] wherein utilizing the ML model includes:
[0047] providing to the ML model a set of input features extracted from the tabular data;
[0048] for each NN thread:
[0049] Applying a random dropout operation on the set of input features to thereby obtain a corresponding subset of features selected from the set of input features;
[0050] applying a neural network algorithm on the respective feature subsets to thereby obtain respective NN thread outputs at respective output layers; and
[0051] Provide the corresponding NN thread output to the collective integration layer;
[0052] An ensemble output of the ML model is calculated at an ensemble integration layer based on the plurality of NN outputs, the ensemble output indicating classification of the sample into a class selected from the at least two classes.
[0053] According to a third aspect of the disclosed subject matter, there is provided a non-transitory computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform a computerized method of classifying tabular data using a machine learning (ML) model, the method comprising:
[0054] obtaining tabular data comprising one or more records, each record corresponding to a respective sample and comprising a plurality of features characterizing the sample;
[0055] Utilizing a machine learning (ML) model on the tabular data, wherein the ML model includes an aggregate input layer, an aggregate integration layer, and a plurality of neural network (NN) threads, each NN thread being connected between the aggregate input layer and the aggregate integration layer and being designed as a separate neural network (NN) including at least two hidden layers, and an output layer;
[0056] The ML models used include:
[0057] Provide the ML model with a set of input features extracted from the tabular data;
[0058] For each NN thread:
[0059] Applying a random dropout operation on the set of input features to thereby obtain a corresponding subset of features selected from the set of input features;
[0060] applying a neural network algorithm on the respective feature subsets to thereby obtain respective NN thread outputs at respective output layers; and
[0061] Provide the corresponding NN thread output to the collective integration layer;
[0062] An ensemble output of the ML model is calculated at an ensemble integration layer based on the plurality of NN outputs, the ensemble output indicating classification of the sample into a class selected from the at least two classes.
[0063] According to a fourth aspect of the disclosed subject matter, there is provided an inspection system dedicated to detecting defects of interest at runtime as part of a semiconductor manufacturing process, the system comprising or otherwise operatively connected to at least one processing circuitry configured to classify tabular data using a machine learning model, the processing circuitry configured to:
[0064] obtaining tabular data comprising one or more records, each record corresponding to a respective sample and comprising a plurality of features characterizing the sample;
[0065] Utilizing a machine learning (ML) model on the tabular data, wherein the ML model includes an aggregate input layer, an aggregate integration layer, and a plurality of neural network (NN) threads, each NN thread being connected between the aggregate input layer and the aggregate integration layer and being designed as a separate neural network (NN) including at least two hidden layers, and an output layer;
[0066] wherein the processing circuitry is configured to utilize the ML model to:
[0067] Provide the ML model with a set of input features extracted from the tabular data;
[0068] For each NN thread:
[0069] Applying a random dropout operation on the set of input features to thereby obtain a corresponding subset of features selected from the set of input features;
[0070] applying a neural network algorithm on the respective feature subsets to thereby obtain respective NN thread outputs at respective output layers; and
[0071] Provide the corresponding NN thread output to the collective integration layer;
[0072] An ensemble output of the ML model is calculated at an ensemble integration layer based on the plurality of NN outputs, the ensemble output indicating classification of the sample into a class selected from the at least two classes.
[0073] The methods, systems and non-transitory program storage devices disclosed with reference to the second, third and fourth aspects may mutatis mutandis optionally include one or more of the features (i) to (xix) listed above in any technically possible combination or arrangement.
[0074] The presently disclosed subject matter further contemplates computer systems, methods, and program storage devices having instructions each specifically configured to perform training of a tabular neural network classifier as described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to understand the subject matter of the present disclosure and to see how it may be carried out in practice, the subject matter will now be described, by way of non-limiting example only, with reference to the accompanying drawings, in which:
[0076] Figure 1A schematically illustrates a generalized block diagram of a computer system dedicated to classifying tabular data according to certain examples of the presently disclosed subject matter;
[0077] Figure 1B schematically illustrates a block diagram of a semiconductor inspection system incorporating the system according to claim 1 according to some examples of the disclosed subject matter;
[0078] Figure 2 is a high-level flow chart illustrating operations performed as part of training a machine learning model according to an example of the presently disclosed subject matter;
[0079] Figure 3 is a high-level flow chart illustrating operations performed as part of executing a trained machine learning model according to an example of the presently disclosed subject matter;
[0080] Figure 4 is a diagram schematically illustrating the architecture of a novel tabular neural network classifier (TNNC) according to an example of the presently disclosed subject matter;
[0081] Figure 5 is a flow chart illustrating operations performed by a TNNC model according to an example of the presently disclosed subject matter; and
[0082] Figure 6 Examples of images of a die are shown, with the DOI in each image indicated by a white circle. DETAILED DESCRIPTION
[0083] Although the specific principles and operation of the subject matter of the present disclosure are described herein in the context of DOI detection in semiconductor wafers, it should be noted that this is done for clarity purposes only and as a non-limiting example. The present disclosure contemplates utilizing the novel machine learning models described herein in a wide range of applications, covering the processing of various forms of tabular data obtained from a variety of resources, including image data, text data (e.g., language), audio data, etc.
[0084] The terms "tabular data" and "table" as used herein should be interpreted broadly to include any structured data equivalent to tabular data in addition to the tabular data itself, for example, a hash table in which each key is connected to an array of values. In the context of the subject matter disclosed herein, each key may represent an identifier of a sample, and each value in the array of values may represent a corresponding feature.
[0085] It is worth noting that although the present specification is primarily concerned with the detection of DOIs in semiconductor samples, this is provided as a non-limiting example. The present disclosure also encompasses the classification of other types of data. For example, the sample may include human tissue (e.g., liver tissue), and the objects of interest include cells having some distinguishable characteristics. The TNNC may be used, for example, to classify cells as malignant or benign. In other examples, the sample may include a leaf, and the objects of interest include distinguishable spots on the surface of the leaf. The TNNC may be used, for example, to classify the spots as healthy or disease-related.
[0086] In light of the above, please note that Figure 1A , which is a generalized illustration of a computer system 101 configured with tabular data classification techniques according to an example of the disclosed subject matter. Figure 1A is a general example demonstrating various principles of the subject matter of the present disclosure. As a non-limiting example, system 101 is shown to include more than one computer processing circuit system, the more than one computer processing circuit system including: a machine learning training computer processing circuit system 10, the machine learning training computer processing circuit system is configured to receive a training data set and generate a tabular neural network classifier (TNNC 21); and a machine learning execution computer processing circuit system 12, the machine learning execution computer processing circuit system is configured to utilize the trained TNNC to classify tabular data, such as for the purpose of anomaly detection. According to the illustrated example, the training data set is stored in a computer data store 14, the computer processing circuit system 10 can access the computer data store, and the machine learning output 21 generated by the computer processing circuit system 10 is stored in a computer data store 16. By way of further example, input data (the TNNC is applied to the input data) is stored in a computer data store 18, and the classification output (the classification output is the result (product) of the TNNC) can be stored in a computer data store 20 and / or provided to a user, for example at a user terminal 22. It is worth noting that it can be implemented in a single data storage device Figure 1A Two or more of the computer data storage devices shown.
[0087] Figure 1B A block diagram of an inspection system 100 according to some examples of the disclosed subject matter is illustrated. Inspection system 100 includes or is otherwise operatively connected to computer system 101. In some non-limiting examples, system 101 is implemented as dedicated processing circuitry integrated within system 100.
[0088] Semiconductor manufacturing processes typically require multiple sequential processing steps and / or layers, each of which may produce errors that may result in yield loss. Examples of various processing steps may include lithography, etching, deposition, planarization, growth (e.g., epitaxial growth), and implantation, etc. Various inspection operations such as defect-related inspections (e.g., defect detection, defect review, and defect classification, etc.) and / or metrology-related inspections may be performed at different processing steps / layers during the manufacturing process to monitor and control the process. Inspection operations may be performed multiple times, for example, after certain processing steps and / or after manufacturing certain layers, etc.
[0089] With the continuous advancement of semiconductor manufacturing technology, semiconductor devices are developed into more and more complex structures, and the feature sizes of semiconductor devices are gradually reduced, which makes it more difficult for conventional methods to meet the inspection performance.
[0090] exist Figure 1B The inspection system 100 illustrated in FIG. 1 may be used to inspect semiconductor samples (e.g., wafers, dies, or portions of semiconductor samples) as part of a sample manufacturing process. Reference to inspection herein may be interpreted as covering any type of operation on the sample, such as defect inspection / detection, defect classification, segmentation, and / or metrology operations. The system 100 includes one or more inspection tools 120 configured to scan the sample and capture images of the sample to be further processed for various inspection applications.
[0091] The term "inspection tool" as used herein should be broadly interpreted to encompass any tool that can be used in inspection-related processes, including, as non-limiting examples, scanning (in a single or multiple scans), imaging, sampling, reviewing, measuring, sorting, and / or other processes performed on a sample or portion of a sample. Without limiting the scope of the present disclosure in any way, it should also be noted that the inspection tool 120 can be implemented as various types of inspection machines, such as an optical inspection machine, an electron beam inspection machine (e.g., a scanning electron microscope (SEM), an atomic force microscope (AFM), or a transmission electron microscope (TEM), etc.), etc.
[0092] The one or more inspection tools 120 may include, for example, one or more inspection tools and / or one or more review tools. In some cases, at least one of the inspection tools 120 may be an inspection tool configured to scan a sample (e.g., an entire wafer, an entire die, or a portion of a sample) to capture an inspection image (typically at a relatively high speed and / or low resolution) to detect potential defects (i.e., defect candidates). During inspection, during exposure, the wafer may be moved in steps relative to a detector of the inspection tool (or the wafer and the tool may be moved in opposite directions relative to each other), and the wafer may be scanned stepwise along a swath of the wafer by the inspection tool, wherein the inspection tool images a portion / potion of the sample (within the swath) at a time. As an example, the inspection tool may be an optical inspection tool. At each step, light may be detected from a rectangular portion of the wafer, and such detected light may be converted into a plurality of intensity values at a plurality of points in the portion, thereby forming an image corresponding to a portion / portion of the wafer. For example, in optical inspection, an array of parallel laser beams may scan the surface of the wafer along a swath. The strips are laid out in parallel rows / columns that follow one another to build up an image of the surface of the wafer one strip at a time. For example, the tool may scan a wafer along a strip from top to bottom, then switch to the next strip and scan the wafer from bottom to top, and so on, until the entire wafer has been scanned and an inspection image of the wafer has been collected.
[0093] In some cases, at least one of the inspection tools 120 may be a review tool configured to capture review images of at least some of the defect candidates detected by the inspection tool to determine whether the defect candidates are indeed defects of interest (DOIs). Such review tools are typically configured to inspect fragments of a sample one at a time (typically at a relatively low speed and / or high resolution). As an example, the review tool may be an electron beam tool, such as a scanning electron microscope (SEM), etc. An SEM is an electron microscope that generates an image of a sample by scanning the sample with a focused electron beam. The electrons interact with the atoms in the sample, thereby generating various signals containing information about the surface morphology and / or composition of the sample. An SEM is capable of accurately inspecting and measuring features during the manufacture of semiconductor wafers.
[0094] The inspection tool and the review tool can be different tools located at the same or different locations, or a single tool operating in two different modes. In some cases, the same inspection tool can provide low-resolution image data and high-resolution image data. The resulting image data (low-resolution image data and / or high-resolution image data) can be transmitted to the system 101 directly or via one or more intermediate systems. The present disclosure is not limited to any particular type of inspection tool and / or the resolution of the image data generated by the inspection tool. In some cases, at least one of the inspection tools 120 has a metrology capability and can be configured to capture images and perform metrology operations on the captured images. Such inspection tools are also referred to as metrology tools.
[0095] According to some examples of the disclosed subject matter, inspection system 100 includes computer system 101 operatively connected to inspection tool 120 and configured to receive images output from the inspection tool and perform defect inspection on semiconductor samples at runtime based on runtime images obtained during sample fabrication. System 101 is also referred to as a defect inspection system.
[0096] As an example, system 101 shows the Figure 1A The computer processing circuit system 10 and the computer processing circuit system 12 shown in FIG. Figures 2 to 5 As further described in detail, processing circuit systems 10 and 12 are configured to provide processing required by the operating system. Processing circuit systems 10 and 12 may be configured to execute several functional modules according to computer-readable instructions implemented on a non-transitory computer-readable memory included in the processing circuit system. Such functional modules are hereinafter referred to as being included in the processing circuit system.
[0097] One or more functional modules included in the processing circuit system 12 may include a pre-processing module 108 and a defect inspection module 110. The pre-processing module 108 is configured to receive an image of a sample (e.g., a die) and transform the image into tabular data. For example, an inspection image and / or a review image generated by scanning a semiconductor sample may be received from the inspection tool 120. The anomaly detection module 110 is configured to apply the previously trained TNNC model 106 to process the tabular data and detect DOI. More specifically, the anomaly detection module 110 is configured to process tabular data including information about one or more defect candidates and classify the defect candidates as DOI or noise.
[0098] According to some embodiments of the disclosed subject matter, system 101 may be a runtime defect inspection system configured to perform defect inspection operations using trained TNNC model 106 applied to runtime images obtained during sample fabrication.
[0099] More specifically, the processing circuit system 12 may be configured to: obtain (e.g., via the I / O interface 126) one or more runtime images of a semiconductor sample acquired by an inspection tool; transform the runtime images into corresponding tabular data representations; and provide the tabular data as input to the anomaly detection module 110, which applies the ML model 106 to the tabular data to detect DOI.
[0100] The disclosed subject matter further contemplates an inspection tool that integrates the functionality of system 101, for example, as an add-on to an inspection tool. According to this example, a component integrated into the inspection tool receives an inspection tool output (e.g., an image) generated by the inspection tool and classifies candidate defects detected by the inspection tool into one of at least two classes selected from a DOI class and a noise class, thereby enhancing the inspection tool output.
[0101] In some examples, system 101 may be further configured as a training system capable of training a TNNC model using a specific training set during a training phase. In this case, one or more functional modules included in processing circuit system 10 of system 101 may include training module 104 and TNNC model 106 to be trained.
[0102] The training module 104 may be configured to obtain a training data set including one or more tables, each table including a plurality of records, each record corresponding to a defect candidate identified in an image of a sample (such as a die or a wafer) and annotated according to a corresponding classification, the corresponding classification including at least one class corresponding to DOI and another class corresponding to noise.
[0103] Training module 104 may be configured to train TNNC model 106 using a training set. Once trained, the TNNC model may be provided to processing circuit system 12. As mentioned above, in other examples, the training of the ML model is performed by a processing circuit system in a computer system other than system 101. Once TNNC 106 has been trained, the trained model is stored in a computer data store accessible to computer system 101.
[0104] According to some embodiments, the system 100 may include a data storage unit 122. The storage unit 122 may be configured to store any data required by the operating system 101, such as computer software loaded during the execution of any of the modules described above, intermediate processing results generated by the system 101, training data sets, trained TNNC models, and / or outputs of TNNC modules.
[0105] In some embodiments, the system 100 may optionally include a computer-based graphical user interface (GUI) 124 that is configured to enable user-specified input related to the system 101. For example, a visual representation of the sample, including an image of the sample, etc., may be presented to the user (e.g., on a display operatively connected to the system 100). The user may be provided with the option of defining certain operations and / or parameters through the GUI. The user may also view the results of the operation or intermediate processing results within the GUI, such as, for example, defect inspection output, some graphical simulation of the results, etc.
[0106] It should also be noted that in some embodiments, at least some of inspection tool 120 , storage unit 122 , and / or GUI 124 may be external to inspection system 100 and operate in data communication with systems 100 and 101 , for example via I / O interface 126 .
[0107] As described above, the system 101 may be implemented as a (multiple) stand-alone computer used in conjunction with an inspection tool and / or additional inspection modules. Alternatively, the corresponding functions of the system 101 may be at least partially integrated with one or more inspection tools 120, thereby improving and enhancing the functionality of the inspection tool 120 in the inspection-related process.
[0108] refer to Figure 2 , which illustrates operations performed as part of training a tabular neural network classifier that can be used to classify tabular data, according to certain examples of the presently disclosed subject matter.
[0109] A training data set comprising at least one table may be obtained (201). The table comprises a plurality of records (rows), each record corresponding to a respective sample (or observation) and further comprising a plurality of columns. Each column in the record corresponds to a feature or attribute characterizing the sample and comprises a respective value. The number of features may vary depending, inter alia, on the type and availability of the data. In some examples, the number of features is 25 or greater.
[0110] In the context of semiconductor inspection, each table corresponds to one or more corresponding semiconductor samples. In some examples, each table may correspond to a corresponding die or multiple dies in a wafer. A single table may include multiple records corresponding to the dies of many wafers. Each record in the table corresponds to a corresponding defect candidate, and each column corresponds to a feature characterizing the defect. In the training set, in addition to the feature value, each record is also labeled according to the corresponding class of the defect (including at least two classes, namely, real defects of interest (DOI) and noise).
[0111] In some examples, tabular data is generated by applying computer processing to an image of a semiconductor sample (e.g., by a preprocessing module 108) and identifying defect candidates in the image. As mentioned above, the image may include an inspection image and / or a review image generated by scanning a semiconductor wafer (e.g., generated as part of a semiconductor manufacturing process). Various image processing techniques may be applied to the image to detect defect candidates in each wafer and in the corresponding die and to determine feature values for each defect candidate. These techniques include, for example, edge detection, color segmentation, and object detection. In some examples, data labeling is performed by manually reviewing the die and adding appropriate annotations to the tabular data using appropriate software tools. In other examples, automatic labeling is applied. The TNNC model is trained (203) using a training data set, and a trained model (21) is generated.
[0112] As is known in the art, the training phase (particularly within the domain of neural networks) includes iterative execution of forward and backward passes. During the forward pass of the neural network, the model generates predictions based on the input data and then evaluates these predictions relative to the actual target values, resulting in the calculation of an error or loss metric that quantifies the performance of the model. Subsequently, in the backward pass (commonly referred to as backpropagation), this error is propagated back through the layers of the network. As part of this process, the gradient associated with each parameter (including weights and biases) is calculated. These gradients are used to adjust the parameters of the model to minimize the error through optimization techniques (such as gradient descent). This iterative process continues until the neural network converges, enabling the neural network to improve its predictive ability and generalize the knowledge it acquires to make accurate predictions on new, previously unseen data.
[0113] Figure 3 A generalized flow diagram of operations performed during execution of a trained machine learning model that can be used to classify tabular data is illustrated in accordance with certain examples of the presently disclosed subject matter.
[0114] Real-world test data is obtained (301). The test data includes at least one table, where each table includes at least one row, but generally includes multiple rows, each row corresponding to a corresponding sample, as in training. The table can be generated by processing one or more samples, extracting feature values that characterize the samples, and inserting each value into a corresponding column in a corresponding record.
[0115] As explained above, the TNNC disclosed herein may be used to inspect semiconductor samples, for example, as part of a sample manufacturing process. As explained above with respect to block 201, an image of a semiconductor sample may be received (e.g., from an inspection tool 120) and the image may be transformed into a tabular data representation. In some examples, the image is received during runtime and the image is transformed into tabular data, which is then processed by the TNNC to detect DOI in the sample.
[0116] The tabular data is constructed such that each record corresponds to a corresponding defect candidate and each column corresponds to a corresponding feature characterizing the defect. During the generation of the tabular data, each defect candidate identified in the image of the sample is analyzed to determine the corresponding feature value of the defect. For example, a feature named "circular" characterizes the circularity of a defect candidate in a die, where 1 is assigned to a perfect circular shape and 0 is assigned to a longitudinal shape. Another feature named "volume" can be determined by analyzing the volume of the defect candidate identified in the die.
[0117] The trained TNNC model is applied (303) on the test data, which classifies the records (and corresponding defect candidates) into one of the classes (block 305).As mentioned above, in some examples, classification is used for anomaly detection, ie, determining whether a defect candidate is a DOI or noise.
[0118] Dies identified as containing a DOI to noise ratio less than a certain threshold are recorded for quality control (block 307). In some examples, the system 100 is configured to make decisions based on the TNNC model output regarding whether to accept, reject, rework a die (microchip), and / or stop production based on detected defects. This process is critical to maintaining the quality of microchips used in electronic devices and includes a feedback loop to improve the manufacturing process.
[0119] A new machine learning (ML) model with a new neural network (NN) model architecture specifically designed for processing tabular data is disclosed herein. The new machine learning model (also referred to herein as a "tabular neural network classifier" or abbreviated as TNNC) features a new model architecture that can efficiently and accurately classify data in a tabular format.
[0120] Figure 4 is a schematic illustration of an example TNNC architecture according to the presently disclosed subject matter. Figure 5 1 is a flow chart of operations performed during TNNC use according to an example of the subject matter of the present disclosure. As described above, the training of the TNNC may be performed, for example, by the training module 104, and the execution and inference of the TNNC module may be performed, for example, by the anomaly detection module 110. Figure 2 Describe the principles of supervised learning to perform training.
[0121] like Figure 4 As illustrated, the model includes a single input layer ("aggregate input layer" 41) for receiving (501) features for each record, such as each defect.
[0122] Dropout (43) is applied (503) to the input layer multiple times. Dropout operates by randomly removing certain features from the input layer, thereby retaining only a subset of the features and creating multiple (n) subsets, each of which includes a subset of features randomly selected from the total set of features inserted into the input layer.
[0123] In some examples, all subsets include an equal number of features. In other examples, different subsets may include different numbers of features. Random deactivation may be applied n times concurrently to thereby obtain n subsets substantially simultaneously. In some examples, random deactivation is applied with replacement.
[0124] It is worth noting that random dropout is usually used to prevent or reduce overfitting by randomly deactivating segments of neurons in the hidden layer of a neural network for each forward and backward pass. Differently, in this paper it is proposed to apply random dropout to the input layer used as a pass-through layer, and unlike the hidden layer, usually no processing is applied to the data.
[0125] Furthermore, dropout is not used for the general purpose of dropout to prevent overfitting, but is used to mimic bagging in neural networks. Bagging is commonly used to reduce variance in machine learning models (such as decision trees, random forests, GBMs, etc.) and is not commonly used in neural networks and particularly not in deep neural networks. Neural networks integrate various regularization methods to reduce overfitting. Furthermore, neural networks are inherently complex, and using bagging will further complicate the model and increase computational requirements.
[0126] A neural network (NN) is applied to each of the subsets of features (45; 505). The TNNC employs a distributed neural network in which each subset of features is processed independently and concurrently by a different set of neural network layers. Each separate set of neural network layers is referred to herein as a "NN thread". In some examples, random dropout may be implemented as an additional layer ("random dropout layer") before the hidden layer in each NN thread.
[0127] As an example, Figure 4 The use of four respective NN threads is shown, each NN thread being used to process a respective feature subset. It is noteworthy that the use of four NN threads is a non-binding example, because a different number (e.g., three, five, six, etc.) of NN threads may be used instead. Each NN thread includes two or more of the hidden layers. In some examples, the TNNC is implemented as a deep neural network model, in which some or all of the NN threads each include many hidden layers, typically 3 or more hidden layers. In some examples, the hidden layers implement regularization methods for reducing overfitting.
[0128] In some examples, the same neural network is utilized on all threads, while in other examples, the neural networks in different threads exhibit variations of each other to thereby increase diversity in the outputs of the different threads. One type of variation involves the number of features processed by each NN thread as mentioned above. Another type of variation involves the number of hidden layers, where in some examples all NN threads include the same number of hidden layers, while in other examples different NN threads include different numbers of hidden layers exhibiting different NN depths.
[0129] In some examples, to further increase the diversity in the outputs of different threads, each NN thread applies a different activation function. For example, for n=4 NN threads, the following activation functions may be used in different threads: GELU, SELU, Swish, and ReLU. If a larger number of NN threads are used, additional activation functions may be assigned to the additional NN threads.
[0130] Each NN thread provides a respective output ("neural network output") to a respective output layer (47; 507). The final output is calculated (509) by an output layer common to all NN threads ("collective integration layer" 49). In the collective integration layer, the final output is calculated based on a compilation of outputs from different NN threads. According to one example, the final output is calculated as an average of the output values received from the different NN threads.
[0131] In some examples, the same activation function is used in all output layers (47) of all NN threads to obtain a common output format that can be easily integrated. In one example, where binary classification is required (e.g., classification between DOI and noise), the Sigmoid activation function is used because the Sigmoid activation function provides a binary output of 0 or 1. In other examples, multi-class classification is applied to, for example, classify different types of DOI. In this case, the SoftMax activation function can be used in all output layers (47) of all NN threads.
[0132] As mentioned above, the TNNC disclosed herein has been shown to be more accurate and more efficient than other existing tabular data machine learning classification techniques. In addition to these advantages, it is also noted that the TNNC architecture retains the ability to achieve improved accuracy and efficiency when a small number of hidden layers, even as few as two, are employed. This is one of the reasons why TNNC requires shorter computation time.
[0133] In addition, the TNNC architecture can be implemented using particularly compact computer code compared to other state-of-the-art neural network classifiers with similar functionality. Take ResNet50 as an example, which is a state-of-the-art neural network widely used in computer vision tasks such as image classification, object detection, and image segmentation. A comparison of the trainable parameters between ResNet50 and TNNC disclosed herein shows that while ResNet50 has more than 23 million trainable parameters, TNNC has as few as 480 trainable parameters. It is known that other tree-based classifiers that operate by transforming image data into a tabular data format outperform ResNet50. TNNC, which is superior to such tree-based classifiers (including random forests and XGboost), is implemented as a neural network with a particularly small amount of computer code.
[0134] Unless otherwise specifically stated, as will be apparent from the following discussion, it should be understood that throughout this specification, discussions utilizing terms such as "obtain," "utilize," "provide," "apply," "receive," "compute," etc., include actions and / or processes of a computer that manipulate data and / or transform data into other data, where the data is represented as physical quantities, such as electronic quantities, and / or where the data represents physical objects.
[0135] The terms "computer," "computer system," "computer device," "computerized device," and the like as used herein should be interpreted broadly to include any kind of hardware-based electronic device having one or more data processing circuit systems. Each processing circuit system may include, for example, one or more processors operatively connected to a computer memory, the one or more processors being loaded with executable instructions for performing operations as further described above.
[0136] The one or more processors referred to herein may represent, for example, one or more general-purpose processing devices, such as microprocessors, central processing units, etc. More particularly, a given processor may be one of the following: a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. One or more processors may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a graphics processing unit (GPU), a network processor, etc.
[0137] The memory referred to in this article may include, for example, one or more of the following items: internal memory (such as processor registers and cache), main memory (such as read-only memory (ROM)), flash memory, dynamic random access memory (DRAM) (such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), etc.
[0138] As will be described in further detail with reference to the above figures, the processing circuit system may be configured to execute several functional modules according to computer readable instructions implemented on a non-transitory computer readable storage medium. Such functional modules are hereinafter referred to as being included in the processing circuit system.
[0139] As used herein, the phrases "for example," "such as," "for instance," and variations thereof describe non-limiting embodiments of the subject matter of the present disclosure. References in the specification to "one instance," "some instances," "other instances," or variations thereof mean that a particular feature, structure, or characteristic described in conjunction with the embodiment is included in at least one embodiment of the subject matter of the present disclosure. Therefore, the appearance of the phrases "one instance," "some instances," "other instances," or variations thereof do not necessarily refer to the same embodiment.
[0140] It should be understood that, for the sake of clarity, certain features of the subject matter of the present disclosure described in the context of separate embodiments may also be provided in combination in a single embodiment. Conversely, for the sake of brevity, various features of the subject matter of the present disclosure described in the context of a single embodiment may also be provided separately or in any suitable subcombination.
[0141] In an embodiment of the presently disclosed subject matter, the following may be performed: Figures 2 to 5 In embodiments of the presently disclosed subject matter, one or more steps illustrated in the accompanying drawings may be performed in a different order, and / or one or more groups of steps may be performed simultaneously.
[0142] Figure 1A and Figure 1B A general schematic diagram of a computer system architecture according to some examples of the presently disclosed subject matter is illustrated. Figure 1A and Figure 1B The elements in may be comprised of any combination of software and hardware and / or firmware that performs the functions as defined and explained herein. Figure 1A and Figure 1B The components of the system may be concentrated in one location or distributed in more than one location. For example, each of the processing circuit systems 10 and 12 may be implemented as part of a different computer device, each of which is located in a different geographical location. In addition, in some examples, the system may include more than Figure 1A and Figure 1B There may be fewer, more, and / or different elements than those shown. For example, Figure 1A and Figure 1B Two separate processing circuit systems are shown, each dedicated to performing certain functions of the system, however, it will be clear to those skilled in the art that the functionality may be divided in other ways. For example, operations related to training a machine learning model and the execution of these operations assigned to processing circuit system 10 and processing circuit system 12, respectively, may otherwise be implemented in a single processing circuit.
[0143] Figure 1A and Figure 1B Each component in may represent multiple specific components that are suitable for operating independently and / or collaboratively to process various data and electrical inputs and for implementing operations associated with the systems disclosed herein. In some cases, multiple instances of a component may be utilized for performance, redundancy, and / or availability reasons. Similarly, in some cases, multiple instances of a component may be utilized for functionality or application reasons. For example, different portions of a specific functionality may be placed in different instances of a component.
[0144] In some examples, certain components utilize cloud implementations, such as those implemented in a private or public cloud. Where the various components of the inspection system are not entirely located in one location or one physical entity, communication between the various components of the inspection system may be implemented through any signaling system or communication components, modules, protocols, software languages, and drive signals, and may be wired and / or wireless where appropriate. For example, in some examples, training operations of a tabular neural network classifier (TNNC) may be performed on a cloud computing infrastructure.
[0145] It will also be understood that a system according to the subject matter of the present disclosure may be a suitably programmed computer. Likewise, the subject matter of the present disclosure contemplates a computer program readable by a computer to perform the method of the subject matter of the present disclosure. The subject matter of the present disclosure further contemplates a machine-readable non-transitory memory tangibly embodying a program of instructions executable by a machine to perform the method of the subject matter of the present disclosure.
[0146] It should be understood that the subject matter of the present disclosure is not limited to the details set forth in the specification contained herein or shown in the accompanying drawings in its application to the present disclosure. The subject matter of the present disclosure can have other embodiments and can be practiced and executed in various ways. Therefore, it should be understood that the wording and terminology used herein are for descriptive purposes and should not be considered as limiting. Therefore, it will be understood by those skilled in the art that the concepts on which the present disclosure is based can be easily used as the basis for other structures, methods and systems designed to achieve several purposes of the subject matter of the present disclosure.
Claims
1. A computer system comprising at least one processing circuit system configured to classify tabular data using a machine learning model, the processing circuit system configured to: obtaining tabular data comprising one or more records, each record corresponding to a respective sample and comprising a plurality of features characterizing the sample; Utilizing a machine learning (ML) model on the tabular data, wherein the ML model comprises an aggregate input layer, an aggregate integration layer, and a plurality of neural network (NN) threads, each NN thread being connected between the aggregate input layer and the aggregate integration layer and being designed as a separate neural network (NN) comprising at least two hidden layers, and an output layer; wherein the processing circuitry is configured to utilize the ML model to: providing the ML model with a set of input features extracted from the tabular data; For each NN thread: applying a random dropout operation on the set of input features to thereby obtain a corresponding feature subset selected from the set of input features; applying a neural network algorithm on the respective feature subsets to thereby obtain respective NN thread outputs at respective output layers; as well as providing the corresponding NN thread output to the collective integration layer; An ensemble output of the ML model is calculated at the ensemble integration layer based on a plurality of NN outputs, the ensemble output indicating a classification of the sample to a class selected from at least two classes.
2. The system of claim 1 , wherein the processing circuit system is configured to apply the same neural network algorithm in all NN threads.
3. The system of claim 1, wherein the sample is a defect in a semiconductor specimen and the feature is a property characterizing the defect; and wherein the aggregate output indicates whether the sample is classified as a defect of interest (DOI).
4. The system of claim 3, wherein the processing circuit system is operatively connected to an inspection system dedicated to detecting defects of interest at runtime as part of a semiconductor manufacturing process, and is configured to determine whether a suspected defect detected by the inspection system is a DOI or noise.
5. The system of claim 1, wherein the tabular data is generated based on image data, wherein each record in the tabular data corresponds to a respective sample identified in the image, and the features in the record are features that characterize the sample.
6. The system of claim 5, wherein the processing circuitry is further configured to obtain the tabular data to: obtaining one or more images; processing the one or more images and identifying at least one sample and a plurality of corresponding features characterizing the sample; and A record in a tabular data format corresponding to the at least one sample is generated and the feature is added to the record.
7. The system of claim 1, wherein each NN thread is designed as a separate neural network (NN) comprising exactly two hidden layers.
8. The system of claim 1, wherein each NN thread is designed as a separate neural network (NN), and at least two NN threads include different numbers of hidden layers.
9. The system of claim 1, wherein the processing circuit system is configured to apply a unique activation function at each NN thread.
10. The system of claim 1, wherein the processing circuitry is configured to apply the dropout on the set of input features received in the set input layer such that at least two different subsets each include a different number of features.
11. The system of claim 1, wherein the corresponding classification of one or more samples in the tabular data is a binary classification.
12. A computer-implemented method for classifying tabular data using a deep learning ML model, the method comprising: obtaining tabular data comprising one or more records, each record corresponding to a respective sample and comprising a plurality of features characterizing the sample; Utilizing a machine learning (ML) model on the tabular data, wherein the ML model comprises an aggregate input layer, an aggregate integration layer, and a plurality of neural network (NN) threads, each NN thread being connected between the aggregate input layer and the aggregate integration layer and being designed as a separate neural network (NN) comprising at least two hidden layers, and an output layer; Wherein utilizing the ML model comprises: providing the ML model with a set of input features extracted from the tabular data; For each NN thread: applying a random dropout operation on the set of input features to thereby obtain a corresponding feature subset selected from the set of input features; applying a neural network algorithm on the respective feature subset to thereby obtain a respective NN thread output at a respective output layer; and providing the corresponding NN thread output to the collective integration layer; An ensemble output of the ML model is calculated at the ensemble integration layer based on a plurality of NN outputs, the ensemble output indicating a classification of the sample to a class selected from at least two classes.
13. The method of claim 12, wherein the same neural network algorithm is applied in all NN threads.
14. The method according to claim 12, further comprising: The tabular data is generated based on the image data, wherein each record in the tabular data corresponds to a respective sample identified in the image, and the features in the records are features that characterize the sample.
15. The method of claim 12, wherein the sample is a defect in a semiconductor specimen and the feature is a property characterizing the defect; and wherein the aggregate output indicates whether the sample is classified as a defect of interest (DOI).
16. The method of claim 12, wherein the method is implemented as part of an inspection process, the portion of the inspection process being dedicated to detecting defects of interest at runtime as part of a semiconductor manufacturing process.
17. The method of claim 12, wherein each NN thread is designed as a separate neural network (NN) comprising exactly two hidden layers.
18. A non-transitory computer readable medium comprising instructions that, when executed by a computer, cause the computer to perform a computerized method of classifying tabular data using a machine learning (ML) model, the method comprising: obtaining tabular data comprising one or more records, each record corresponding to a respective sample and comprising a plurality of features characterizing the sample; Utilizing a machine learning (ML) model on the tabular data, wherein the ML model comprises an aggregate input layer, an aggregate integration layer, and a plurality of neural network (NN) threads, each NN thread being connected between the aggregate input layer and the aggregate integration layer and being designed as a separate neural network (NN) comprising at least two hidden layers, and an output layer; Wherein utilizing the ML model comprises: providing the ML model with a set of input features extracted from the tabular data; For each NN thread: applying a random dropout operation on the set of input features to thereby obtain a corresponding feature subset selected from the set of input features; applying a neural network algorithm on the respective feature subset to thereby obtain a respective NN thread output at a respective output layer; and providing the corresponding NN thread output to the collective integration layer; An ensemble output of the ML model is calculated at the ensemble integration layer based on a plurality of NN outputs, the ensemble output indicating a classification of the sample to a class selected from at least two classes.
19. The non-transitory computer readable medium of claim 18, wherein the method comprises one or more of the following: a) applying at least two different activation functions, so that at least two different NN threads each apply a different activation function; b) applying said dropout on said set of input features received in said set input layer such that at least two different subsets include different numbers of features.
20. The non-transitory computer readable medium of claim 18, wherein the method comprises applying a unique activation function at each NN thread.