Training classification models using different training sources and applying their inference engine

Through iterative programs and STIR loops, the imbalance problem of deep neural network classification engines in uneven distribution training data is solved, and the stability and efficiency of training complex classification engines on small training data sources are achieved.

CN113743437BActive Publication Date: 2025-05-16MACRONIX INTERNATIONAL CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010782604.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-28
Filing Date
2020-08-06
Publication Date
2025-05-16
Estimated Expiration
2040-08-06

AI Technical Summary

Technical Problem

When training the classification engine of deep neural networks, we face the problem of learning algorithm imbalance caused by uneven distribution of data categories, which in turn affects the object recognition performance of certain categories.

Method used

Using an iterative program, the model is continuously optimized through STIR loops (sampling, training, inference, inspection results) by selecting small samples from the training data source for model training, and using the model to infer and check results on larger samples until the specified criteria are met.

Benefits of technology

Effectively use smaller training data sources to train complex classification engines, saving computing resources, and avoiding instability due to small training data sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113743437B_ABST
    Figure CN113743437B_ABST
Patent Text Reader

Abstract

The present invention provides a method for generating a classification model using a training data set. An iterative procedure for training an ANN model, wherein an iteration includes selecting a small sample of training data from a training data source, training the model using the sample, inferring the model using a large sample of training data, and checking the inference results. The results are evaluated to determine whether the model meets a condition, and if the model does not meet the specified standard, the cycle of sampling, training, inference and checking the results in the iterative procedure (STIR cycle) is repeated until the standard is met. A classification engine trained by the method described herein is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to training a classification engine including a neural network, and more particularly to an application to training data having multiple object categories and uneven distribution. Background Art

[0002] The subject matter discussed in this section should not be considered as prior art simply because it is mentioned in this section. Similarly, the problems mentioned in this section or related to the subject matter and provided as technical background should not be considered as mentioned in the prior art. The subject matter in this section only represents different methods, which themselves may also correspond to implementing the claimed technology.

[0003] A deep neural network is an artificial neural network (ANN) that uses multiple nonlinear and complex transformation layers to model high-order features sequentially. Deep neural networks provide feedback through backpropagation, which carries the difference between the observed and predicted outputs, to adjust parameters. With the availability of large training data sets, the power of parallel and distributed computing, and sophisticated training algorithms, deep neural networks have developed rapidly. Deep neural networks have facilitated significant advances in many fields, such as computer vision, speech recognition, and natural language processing.

[0004] Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) can be configured as deep neural networks. Convolutional neural networks have been successfully applied to image recognition through an architecture that includes convolutional layers, nonlinear layers, and pooling layers. Recurrent neural networks are designed to use the sequential information of input data by establishing recurrent connections between blocks such as perceptrons, long short-term memory units, and gated recurrent units. In addition, many other emerging deep neural networks have been proposed for limited background information, such as deep spatiotemporal neural networks, multidimensional recurrent neural networks, and convolutional autoencoders.

[0005] The goal of training a deep neural network is to optimize the weight parameters of each layer, which gradually combines simpler features into complex features so that the most appropriate hierarchical representation can be learned from the data. A single loop of the optimization procedure is as follows. First, given a training dataset, the output in each layer is calculated sequentially through a forward pass, and the function signal is propagated forward through the network. In the final output layer, the target loss function measures the error between the inferred output and the given label. In order to minimize the training error, the error signal is propagated backward using the chain rule through a backward pass, and the gradient of all weights in the entire neural network is calculated. Finally, the weight parameters are updated using an optimization algorithm based on stochastic gradient descent. Batch gradient descent performs parameter updates for each complete dataset, while stochastic gradient descent provides a random approximation by performing updates for each small dataset. There are currently several optimization algorithms developed based on stochastic gradient descent. For example, the Adagrad and Adam training algorithms perform stochastic gradient descent while modifying the learning rate appropriately based on the update frequency and gradient moment for each parameter separately.

[0006] In machine learning, classification engines including ANNs are trained using a database of objects labeled according to multiple categories to be recognized by the classification engine. In some databases, the number of objects in each of the different categories to be recognized may vary greatly. The uneven distribution of objects among the categories may cause an imbalance in the learning algorithm, resulting in a decrease in the performance of object recognition for certain categories. One way to solve this problem is to use larger and larger training sets to level out the imbalance or include a sufficient number of objects from rare categories. This will result in these large training sets requiring a lot of computing resources to be applied to the training of the classification engine.

[0007] There is therefore a need for a technique to improve the training of classification engines using reasonably sized databases of labeled objects. Summary of the invention

[0008] The present invention relates to a computer-implemented method that improves computer-implemented techniques for training classification engines including artificial neural networks.

[0009] According to one embodiment of the present invention, a method is provided for improving the manufacture of integrated circuits by detecting and classifying defects in integrated circuit components during the manufacturing process using the techniques described herein.

[0010] The technology generally includes an iterative procedure, which uses a training data source to train an ANN model or other classification engine model, wherein the training data source may include multiple objects, such as images, audio data, text, and other types of information, which may be used alone or in a variety of combinations, and these objects may be classified using a large number of categories, resulting in uneven distribution. The iterative procedure includes selecting a small sample of training data from the training data source, using the sample to train the model, using the inference mode of the model on a larger sample of the training data, and checking the inference results. The results are evaluated to determine whether the model is satisfactory. If the model does not meet the specified criteria, the cycle of sampling, training, inference, and checking the results in the iterative procedure (STIR cycle) is repeated until the criteria are met.

[0011] The techniques described in this article enable training complex classification engines using smaller data sources for training and enable efficient use of computing resources while overcoming the instability caused by reliance on small data sources for training.

[0012] In order to better understand the above and other aspects of the present invention, embodiments are given below and described in detail with reference to the accompanying drawings: BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 A simplified diagram of the iterative training algorithm described in the present invention is shown.

[0014] Figure 2A , 2B 2C, 2D, 2E, and 2F illustrate various stages of an iterative training algorithm according to an embodiment of the present invention.

[0015] Figure 3 Draw something like Figures 2A to 2F A flowchart of one embodiment of an iterative training algorithm is shown.

[0016] Figure 4A , 4B 4C, 4D, 4E, and 4F illustrate various stages of an iterative training algorithm according to another embodiment of the present invention.

[0017] Figure 5 A flowchart similar to that shown in FIGS. 4A to 4F is shown for another embodiment of an iterative training algorithm.

[0018] Figure 6 A table showing a classification of training sources is included, which is used to explain the technique of selecting training subsets after the first training subset.

[0019] Fig. 7A and 7B Illustrated are the stages of an iterative training algorithm according to an alternative technique for selecting a training subset.

[0020] Figure 8 A simplified diagram of a classification engine trained using the method described in the present invention during integrated circuit manufacturing is shown.

[0021] Fig. 9 is an image of a graphical user interface of a computer system executed by a system configured to perform a training program according to the present invention during an initial stage.

[0022] Fig.10 is an image of a graphical user interface of a computer system executed by a system configured to perform, at a subsequent stage, a training program as described herein.

[0023] Fig.11 is a simplified block diagram of a computer system configured to perform the training program described herein.

[0024]

Explanation of symbols

[0025] 10~14,300~312,500~512:Steps

[0026] 20: Source dataset

[0027] 60: Processing Station X

[0028] 61: Inspection Tools

[0029] 62: Processing station X+1

[0030] 63: Classification Engine (ANN Model M)

[0031] 600: Diagonal area

[0032] 601, 602: Area

[0033] 1200: Calculator Systems

[0034] 1210: Storage subsystem

[0035] 1222: Memory Subsystem

[0036] 1232: Read-only memory

[0037] 1234: Random Access Memory

[0038] 1236: File storage subsystem

[0039] 1238: User interface input device

[0040] 1255: Bus subsystem

[0041] 1272: Central Processing Unit

[0042] 1274: Network interface

[0043] 1276: User interface output device

[0044] 1278: Deep learning processor (graphics processing unit, field programmable gate array, coarse-grained reconfigurable architecture)

[0045] S: source dataset

[0046] SE1: First evaluation subset

[0047] SE2: Second evaluation subset

[0048] SE(n+1), SE(n+2): Evaluation subsets

[0049] ST1: First training subset

[0050] ST2: Second training subset

[0051] ST3: The third training subset

[0052] ST+, ST-: object

[0053] ER1: First error subset

[0054] ER2: Second error subset

[0055] ER3: Third error subset

[0056] ER(n), ER(n+1): Error subset

[0057] B1~B4:Block

[0058] M1, M2, M3: Model DETAILED DESCRIPTION

[0059] Figures 1 to 11 provide a detailed description of embodiments of the present technology.

[0060] Figure 1 A processing loop (STIR loop) is shown that is able to iteratively learn the errors in each previous loop, relying on a small training source.

[0061] exist Figure 1In the process, a computer system accesses a database that stores a source data set S20 of labeled training objects that can be used to train a classification engine. The loop begins by selecting a training subset ST(i)10 from the data set S that is substantially smaller than the entire data set S(20). The computer system uses the training subset ST(i)10 to train a model M(i)11 that includes parameters and coefficients used by the classification engine. Next, the inference engine uses the model M(i)11 to classify objects in an evaluation subset SE(i)12 of the source data set S(20). The inference results are evaluated based on the model M(i)11, and an error subset ER(i)13 of objects that are misclassified is identified in the evaluation subset SE(i)12. The size and characteristics of the error subset ER(i)13 are then compared to predetermined parameters that represent the performance of the model M(i)11 on the classified categories. If the model M(i)11 meets the parameters, it is saved for use by the inference engine in the field (14). If the model M(i)11 does not meet the parameters, the loop returns to select the training subset for the next cycle. The loop is iterated until the parameters are met and the final model is developed.

[0062] Figures 2A to 2F The various stages of the training algorithm described in the present invention are shown. Figure 2A In the example, a portion of the source data set S is accessed to select a first training subset ST(1), and a portion of the source data set S (excluding the first training subset) is accessed to use as a first evaluation subset SE(1). The first training subset ST(1) may be 1% or less of the source data set S.

[0063] Use ST(1) to generate model M(1), and use model M(1) to classify the objects in the first evaluation subset SE(1).

[0064] like Figure 2B As shown, an error subset ER(1) of objects that are misclassified by the model M(1) is identified in the first evaluation subset.

[0065] like Figure 2C As shown, a portion of the error subset ER(1) is accessed to be used as a second training subset ST(2). ST(2) may include a portion or all of the error subset ER(1).

[0066] Model M(2) is generated using ST(1) and ST(2), and model M(2) is used to classify objects in the second evaluation subset SE(2) (excluding ST(1) and ST(2)).

[0067] like Figure 2DAs shown, an error subset ER(2) of objects misclassified by the model M(2) is identified in the second evaluation subset. For the training subset ST(i) with i=2, less than half of the objects in the error subset ER(1) may be included.

[0068] like Figure 2E As shown, a portion of the error subset ER(2) is accessed to be used as a third training subset ST(3). ST(3) may include a portion or all of the error subset ER(2).

[0069] Model M(3) is generated using ST(1), ST(2) and ST(3), and model M(3) is used to classify objects in the third evaluation subset SE(3) (excluding ST(1), ST(2) and ST(3)).

[0070] like Figure 2F As shown, an error subset ER(3) of objects misclassified by the model M(3) is identified in the third evaluation subset. In this embodiment, the error subset ER(3) is small and the model M(3) can be used in the field. This process can be performed iteratively until the proportion of classification errors reaches a satisfactory point, such as less than 10% or less than 1%, depending on the use of the training model.

[0071] Figure 3 is similar to Figures 2A to 2F A flow chart of a training algorithm is shown, which may be executed in a computer system.

[0072] The algorithm begins by accessing a source dataset S of training data (300). For a first loop with index (i) = 1, a training subset ST(1) is obtained from the set S (301). A model M(1) of a classification engine is trained using the training subset ST(1) (302). Next, an evaluation subset SE(1) of the source dataset S is accessed (303). The evaluation subset is classified using the model M(1), and an error subset ER(1) of objects misclassified by the model M(1) is identified in the evaluation subset (304).

[0073] After the first loop, index (i) is incremented (305) and the next loop of iteration begins. The next loop includes selecting a training subset ST(i), which includes a portion of the error subset ER(i-1) in the previous loop (306). Next, the model M(i) is trained using a combination ST(i) of (i) from 1 to (i), which includes the training subsets of the current loop and all existing loops (307). An evaluation subset SE(i) for the current loop is selected, which does not include the training subsets of the current loop and all existing loops (308). Using the model M(i), the evaluation subset SE(i) is classified and the error subset ER(i) of the model M(i) of the current loop is identified (309). The error subset ER(i) of the current loop is evaluated using the expected parameters of a successful model (310). If the evaluation satisfies the condition (311), the model M(i) of the current loop is stored (312). At step 311, if the evaluation does not satisfy the condition, the algorithm loop will return to step 305 to increment the index (i), and then repeat the loop until a final model is provided.

[0074] Figures 4A to 4F An optional procedure for training a classification engine is shown, the procedure including the step of partitioning a source data set. Figure 4A As shown, the source data set S is divided into blocks B1 to B4, which can be non-overlapping blocks B1 to B4 of objects in the source data set S. In one embodiment, the source data set S may include 10,000 defects of category C1 and 98 defects of category C2. The blocks can be selected in a manner that maintains the relative number in the category. Thus, in this embodiment, it is divided into 10 blocks, and the selected objects in each block include approximately 10% of the defects in category C1 (approximately 100) and approximately 10% of the defects in category C2 (approximately 10). If the distribution is uneven, the model performance may decrease. For example, if objects of a category appear in a short time interval, and the objects in the training data set are partitioned only by time intervals, the category may be unreasonably divided into one block or two blocks.

[0075] Figure 4B The objects of the first block B1 are accessed to provide a first training subset ST1, which contains far fewer objects than all the objects in the first block. The model M1 is trained using the first training subset ST1. ST1 may include only a small portion of the training data set S, for example less than 1%.

[0076] like Figure 4CAs shown, the first model M1 is applied to a first evaluation subset SE1, which includes part or all of the objects in the second block B2, and a first error subset ER1 that identifies objects misclassified by the first model M1 in the first evaluation subset SE1.

[0077] Figure 4D A second training subset ST2 is shown selected from the objects in the first error subset ER1 and used in combination with the first training subset (indicated by the + on the arrow) to develop a model M2. ST2 may include some or all of the objects in ER1.

[0078] like Figure 4E As shown, the model M2 is applied to a second evaluation subset SE2, which includes some or all of the objects in the third block B3. A second error subset ER2 of objects misclassified by the second model M2 is identified in the second evaluation subset SE2. And, a third training subset ST3 is identified, which includes some or all of the objects in the second error subset ER2. Using the first, second and third training subsets, a third model M3 is trained.

[0079] like Figure 4F As shown, the third model M3 is applied to the third evaluation subset SE3, which includes some or all of the objects in the fourth block B4. A third error subset ER3 is identified, which includes objects that are misclassified by the model M3 in the third evaluation subset SE3. In this embodiment, the third error subset ER3 is very small, representing that the model M3 can be used as the final model.

[0080] Figure 5 is similar to Figures 4A to 4F A flow chart of a training algorithm is shown, which may be executed in a computer system.

[0081] The algorithm begins by accessing a source dataset S of training data and dividing it into non-overlapping blocks B(i) (500). For the first loop with index (i) = 1, a training subset ST(1) is obtained, which includes a small number of objects from block B(i) (501). The model M(1) of the classification engine is trained using the training subset ST(1) (502). Next, an evaluation subset SE(1) is obtained, which includes some or all of the objects in block B2 from the source dataset S (503). The evaluation subset SE(1) is classified using the model M(1), and an error subset ER(1) of objects misclassified by the model M(1) is identified in the evaluation subset SE(1) (504).

[0082] After the first loop, index (i) is incremented (505) and the next loop of iteration begins. The next loop includes selecting a training subset ST(i) that includes part of the error subset ER(i-1) in the previous loop (506). Next, the model M(i) is trained using the combination ST(i) of (i) from 1 to (i), which includes the training subsets of the current loop and all existing loops (507). An evaluation subset SE(i) for the current loop is selected, which includes part or all of the objects from the next block B(i+1) and does not include the training subsets of the current loop and all existing loops (508). Using the model M(i), the evaluation subset SE(i) is classified and the error subset ER(i) of the model M(i) of the current loop is identified (509). The error subset ER(i) of the current loop is evaluated using the expected parameters of a successful model (510). If the evaluation satisfies the condition (511), the model M(i) of the current loop is stored (512). At step 511, if the evaluation does not satisfy the condition, the algorithm loop will return to step 505 to increment the index (i), and then repeat the loop until a final model is provided.

[0083] In both of the above alternatives, the initial training subset ST1 is selected so that it is much smaller than the training data set S. For example, less than 1% of the training data set S can be used as the initial training subset ST1 to train the first model M1. This has the effect of improving the efficiency of the training procedure. The selection of the initial training subset can apply a random selection technique so that the distribution of objects of the categories in the training subset is similar to the distribution in the source training data set S. However, in some cases, the training data set S includes objects of multiple categories that are unevenly distributed. For example, with categories C1 to C9, the data set S may have the following object distribution:

[0084] C1: 125600

[0085] C2:5680

[0086] C3:4008

[0087] C4:254

[0088] C5:56

[0089] C6:32

[0090] C7:14

[0091] C8:7

[0092] C9:2

[0093] Under this distribution, a training algorithm using a training subset similar to this distribution can only produce models that perform well for the first three classes C1, C2, and C3.

[0094] To improve performance, the initial training subset can be selected according to a distribution balancing procedure. For example, the method for selecting the initial training subset can be to set parameters for the categories to be trained. In one method, the parameters can include a maximum number of objects in each category. In this method, the training subset can be limited to, for example, a maximum of 50 objects per category, so that, in the above example, categories C1 to C5 are limited to 50 objects, while categories C6 to C9 remain unchanged.

[0095] Another parameter can be the minimum number of objects in each category, used in conjunction with the maximum number discussed above. For example, if the minimum value is set to 5, categories C6 to C8 are included, while category C9 is ignored. Thus, the training subset includes 50 objects each in categories C1 to C5, 32 objects in category C6, 14 objects in category C7, and 7 objects in category C8, for a total of 303 objects. It is worth noting that because the initial training subset ST1 is small, it is expected that the size of the error subset is relatively large. For example, the accuracy of model M1 can be about 60%. The training procedure requires additional training data to resolve 40% of the misclassifications.

[0096] A program can be set to select the second training subset ST2 so that 50% of ST1>ST2>3% of ST1. Similarly, for the third training subset, 50% of (ST1+ST2)>ST3>3% of (ST1+ST2), and so on using a range of relative sizes between 3% and 50% until the final training subset STN.

[0097] A range of relative sizes between 5% and 20% is suggested.

[0098] The object in extra training subset ST2 to STN can also be selected in many ways. In one embodiment, use random selection to select extra training subset from the corresponding error subset. In this random method, the size and distribution of the training subset are the same as the object size and object distribution of the error subset selected by it. The risk of this method is that the signal of some categories may be very small in the training subset.

[0099] In another approach, objects are selected based on a class-aware procedure. For example, the training subset may include the maximum number of objects in each class. For example, the maximum for a particular cycle may be 20 objects per class, with classes with fewer than 20 objects being included.

[0100] In another approach, some objects can be selected randomly, while some objects can be selected through a category-aware procedure.

[0101] Figure 6A graphical user interface that can be used in a computer system configured to perform training in each loop to select the next training subset as described in the present invention is shown. In the interface, a table is displayed, which has multiple columns corresponding to the counts of the reference standard (ground truth) classifications of object categories C1 to C14, the column "Total" is the sum of the corresponding rows, the column "Precision" is the measured precision of the inference engine, which represents the percentage of correct classifications in all categories, and the column "Source" represents the number of objects in the training data source used to form the model in the inference engine. This table has multiple rows, which correspond to the number of objects in each object category that the inference engine classified using the model, the row "Total" represents the total number of objects in the corresponding column, and the row "Recall" represents the percentage of correct classifications in each category.

[0102] Thus, the row labeled C1 correctly classified 946 objects, and incorrectly classified 8 objects in C2 as C1, 3 objects in C4 as C1, 4 objects in C5 as C1, etc. The column labeled C1 shows that 946 C1 objects were correctly classified, 5 C1 objects were classified as C2, 4 C1 objects were classified as C3, 10 C1 objects were classified as C4, and so on.

[0103] The diagonal region 600 includes the correct classifications. Regions 601 and 602 include the erroneous subsets.

[0104] The example here shows the results after several cycles, using a training source with 1229 classified objects, which comes from the combination of the training subsets from multiple cycles as described above, to produce a model that correctly classified 1436 of the 1570 objects in the evaluation subset. Therefore, the error subset at this stage is relatively small (134 errors). In earlier cycles, the error subset may be larger.

[0105] To select the next training subset, the program can select a random combination of objects from regions 601 and 602. Thus, to add approximately 5% of the objects (62 objects) to the training subset of 1229 objects, approximately half of the error subset (134 / 2 = 77 objects) can be identified as the next training subset and combined with the training subset from the existing cycle. Of course, the number of objects can be selected more precisely if desired.

[0106] In order to select the next training subset using a class-aware procedure, the present invention provides two optional methods.

[0107] First, if the goal of the model is to provide higher accuracy on all categories, the number of misclassified objects in each row (e.g. 10) can be included in the training subset, where all objects with less than 10 misclassified objects should be included in the training subset. Here, C10 has 2 misclassified objects in this row to be included in the training subset for the current cycle.

[0108] Second, if the purpose of the model is to provide a higher recall on a particular class, the number of misclassified objects in each column (e.g. 10) can be included in the training subset, where all objects with less than 10 misclassified objects are included in the training subset. Here, C10 has 10 misclassified objects in the column to be included in the training subset of the current cycle.

[0109] In one approach, a random selection method is applied in the first and other early cycles, a hybrid method is applied in the middle cycles, and a class-aware method is applied in the last cycle when the error subset is small.

[0110] In some embodiments, the following may be applied: Fig. 7A and 7B Please refer to Figure 6 , in the example shown in the table, there are few objects in categories C11 to C14. Model stability may be poor, especially for classifications in these categories. Therefore, when selecting a training subset for a particular cycle, correctly classified objects (true positives) can be added to the training subset. For example, to achieve a target number of 10 objects, this method can add three true positive objects of category C12 to the training subset, two true positive objects of category C13 to the training subset, and one true positive object of category C14 to the training subset.

[0111] Fig. 7A Shown is the selection of a training subset ST(n+1), which is selected from an error subset ER(n) generated using a model M(n) in an evaluation subset SE(n+1), using part of the objects (ST-) from the error subset ER(n) and part of the correctly classified objects (ST+) from the evaluation subset SE(n+1). Figure 7B The objects used in the next cycle are shown. The training subset ST(n+2) is selected from an error subset ER(n+1) generated by using the model M(n+1) in an evaluation subset SE(n+2). Part of the objects (ST-) from the error subset ER(n+1) and part of the correctly classified objects (ST+) from the evaluation subset SE(n+2) are used.

[0112] Defect images of circuit components on a body captured in a manufacturing line can be classified into many categories. For a particular manufacturing process, the number of these defects varies greatly, so the distribution of training data is uneven and the amount of data involved is large. An embodiment of the technology described herein can be used to train an ANN to identify and classify these defects, thereby improving the manufacturing process.

[0113] There are many types of defects, and defects with similar shapes may originate from different sources. For example, if the appearance of a defective image is caused by a problem in the previous or next layer, part of the pattern will not appear in the category. For example, there may be problems such as embedded defects or hole-like cracks in the layer after the current layer. However, a pattern that disappears in an image of a different category may be a problem in the current layer. Therefore, it is desirable to build a neural network model that can classify all types of defects.

[0114] We need to monitor defects in online processes to evaluate the stability and quality of online products, or the life of manufacturing tools.

[0115] Figure 8 A simplified diagram of a manufacturing production line is shown, which includes a processing station 60, an image sensor 61, and a processing station 62. In this production line, an integrated circuit wafer is input to a processing station X, and is processed such as deposition or etching, and output to an image sensor 61. From the image sensor, the wafer is input to a processing station X+1, where a process such as deposition, etching, or packaging is performed. The wafer is then output to the next stage. The image from the image sensor is provided to a classification engine, which includes an ANN trained according to the techniques described herein, which can identify and classify defects in the wafer. The classification engine can also receive images from other stages in the manufacturing process. Information about defects in the wafer sensed by the inspection tool 61 with an image sensor can be used to improve the manufacturing process, for example by adjusting the procedures performed at the processing station X or at other processing stations.

[0116] Fig. 9 and 10 A graphical user interface executed by a computer system configured to execute the training program of the present invention is shown. Fig. 9In FIG. 1 , five blocks are displayed on the graphical user interface. The “Original Database” block includes a list of categories and the number of objects in each category. The interface driver can automatically add data to this block by analyzing the training data set S. The second block (first data extraction) includes fields associated with each category, which can be added to the data of this block by direct user input or in response to the parameters for the corresponding error subset to set the number of objects to be used for each category in the first training subset ST1. The third block (second data extraction) includes fields associated with each category, which can be added to the data of this block by direct user input or in response to the parameters for the corresponding error subset to set the number of objects to be used for each category in the second training subset ST2. The fourth block (third data extraction) includes fields associated with each category, which can be added to the data of this block by direct user input or in response to the parameters for the corresponding error subset to set the number of objects to be used for each category in the third training subset ST3. The fifth block (Please select a calculation model) includes a drop-down menu for selecting which ANN structure model is to be trained. An example called “CNN” is shown here. The graphical user interface also includes an "execute" button component which, when selected, executes the program using the parameters provided through the interface.

[0117] Fig.10 Draw Fig. 9 A graphical user interface is shown in which the contents of the second to fourth blocks are filled in. In this example, the contents of the second to fourth blocks correspond to the training data used in subsequent cycles, including a combination of training subsets ST1, ST2 and ST3.

[0118] Described herein are a plurality of flow charts executed by a computer configured to execute a training program. These processes can be implemented using processor programming, which is programmed by a computer program stored in a memory accessible by a computer system, and can be executed by a processor, a dedicated logic hardware including a field programmable integrated circuit, and a combination of dedicated logic hardware and a computer program. For all flow charts herein, it is understood that the steps can be combined, executed in parallel, or executed in a different order without affecting the functions implemented. In some cases, as the reader understands, the rearrangement of the steps can also achieve the same result if only certain other changes are made. In other cases, as the reader understands, the rearrangement of the steps can achieve the same result only if certain conditions are met. In addition, it should be understood that the flow charts herein only show steps related to understanding the present invention, and it should be understood that multiple additional steps can be performed before, after, and between the steps shown herein to complete other functions.

[0119] As used herein, a subset of a set excludes the degenerate case of an empty subset and subsets that include all members of the set.

[0120] Fig.11 is a simplified block diagram of a computer system 1200, one or more of which in a network can be programmed to implement the techniques disclosed herein. Computer system 1200 includes one or more central processing units (CPUs) 1272 that communicate with a plurality of peripheral devices via a bus subsystem 1255. These peripheral devices may include a storage subsystem 1210, such as memory devices and a file storage subsystem 1236, a user interface input device 1238, a user interface output device 1276, and a network interface subsystem 1274. The input and output devices allow a user to interact with computer system 1200. Network interface subsystem 1274 provides an interface to connect to an external network, including interfaces to corresponding interface devices in other computer systems.

[0121] User interface input devices 1238 may include a keyboard; a pointing device, such as a mouse, trackball, touch pad, or graphics tablet; a scanner; a touch panel included in a display; audio input devices, such as a voice recognition system and a microphone; and other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways of entering information into the computer system 1200.

[0122] The user interface output device 1276 may include a display subsystem, a printer, a fax machine, or a non-visual display such as a sound output device. The display subsystem may include an LED display, a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or other device for producing a visible image. The display subsystem may also provide a non-visual display, such as a sound output device. In general, the use of the term "output device" is intended to include all possible device types and ways of outputting information from the computer system 1200 to a user or another machine or computer system.

[0123] The storage subsystem 1210 stores programming and data structures that provide some or all of the functions of the modules and methods described herein to train ANN models. These models are typically used in ANNs executed by the deep learning processor 1278.

[0124] In one implementation, the neural network is implemented using a deep learning processor 1278, which can be a configurable and reconfigurable processor, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA) and a graphics processing unit (GPU) and other configured devices. The deep learning processor 1278 can be hosted by a deep learning cloud platform, such as Google Cloud Platform TM 、Xilinx TM and Cirrascale TM Examples of deep learning processors 1278 include Google's Tensor Processing Unit (TPU) TM , rack-mount solutions such as the GX4 Rackmount Series TM 、GX149 Rackmount Series TM 、NVIDIA DGX-1 TM , Microsoft's Stratix V FPGA TM Graphcore’s Intelligent Processor Unit (IPU) TM , with a Snapdragon processor TM Qualcomm's Zeroth Platform TM , NVIDIA's Volta TM NVIDIA's DRIVE PX TM , NVIDIA's JETSON TX1 / TX2 MODULE TM , Intel's Nirvana TM , Movidius VPU TM , Fujitsu DPI TM , ARM's DynamicIQ TM , IBM TrueNorth TM wait.

[0125] The memory subsystem 1222 used in the storage subsystem 1210 may include multiple memories, including a primary random access memory (RAM) 1234 for storing instructions and data during program execution, and a read-only memory (ROM) 1232 for storing fixed instructions. The file storage subsystem 1236 may be used for program and data files (including Figure 1 , 35) provides permanent storage and may include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM, an optical drive, or a removable media cassette. Modules implementing certain functions may be stored in the file storage subsystem 1236 of the storage subsystem 1210 or in other machines accessible to the processor.

[0126] The bus subsystem 1255 provides a mechanism for communication among the various components and subsystems of the computer system 1200. Although the bus subsystem 1255 is shown as only a single bus, the bus subsystem may use multiple busses.

[0127] Computer system 1200 itself may be of various types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe computer, a server farm, a widely distributed group of loosely coupled network computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, Fig.11 The description of the computer system 1200 shown in FIG. 1 is provided as a specific example only and is intended to illustrate a preferred embodiment of the present invention. Fig.11 Other configurations of computer system 1200 may have more or fewer components than the one shown.

[0128] Embodiments of the techniques described herein include computer programs stored on a memory of a non-transitory computer-readable medium that can be accessed and read by a computer, including, for example, Figure 1 , 3 and 5 describe the programs and data files.

[0129] In summary, although the present invention has been disclosed by way of embodiments, it is not intended to limit the present invention. A person skilled in the art of the present invention may make various modifications and improvements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be based on the scope defined by the claims.

Claims

1. A method for generating a classification model using a training data set S of a plurality of objects to classify the objects into a plurality of categories, the objects in the training data set S comprising images of defects on an integrated circuit component in an integrated circuit manufacturing process, the defects comprising defects of a plurality of categories, the method comprising one or more programmed computers: For an index i=1, obtain a first training subset ST(i), which includes some of the objects in the training data set; Using the first training subset ST(i) to train a first model M(i); Using the first model M(i) to classify a first evaluation subset SE(i) in the training data set that does not include the first training subset ST(i), and identifying an error subset ER(i) of the objects that are misclassified in the first evaluation subset SE(i); (a) incrementing the index i and obtaining another training subset ST(i) comprising the objects that are part of the error subset ER(i-1); (b) training a model M(i) using the combination of the training subsets ST(i), where i is 1 to i; (c) using the model M(i) to classify an evaluation subset SE(i) in the training data set S that does not include the training subset ST(i), where i is 1 to i, and identifying an error subset ER(i) of the objects that are misclassified in the evaluation subset SE(i); as well as (d) evaluating the error subset ER(i) to estimate a performance of the model M(i), and if the performance satisfies a condition, storing the model M(i), and if the performance does not satisfy the condition, repeating steps (a) to (d), The step of obtaining the other training subset ST(i) with i>1, wherein the other training subset ST(i) includes some of the objects in the error subset ER(i-1), comprises: For a given one of these categories, when the number of these objects misclassified in the error subset is less than a minimum number M, the objects in one or more of the error subsets ER(i) i=i-2 to 1 are added to increase the number of these objects of the given category in the training subset.

2. The method of claim 1, wherein the evaluating comprises determining a number of the objects that are misclassified and comparing the number to a threshold value.

3. The method of claim 1, wherein the evaluating comprises determining a number of the objects misclassified in the error subset ER(i) and comparing the number with the number of the objects misclassified in a previous error subset ER(i-1). 4 . The method of claim 1 , wherein the first training subset ST(i) with i=1 comprises 10% or less of the objects in the training dataset S. 5 . 5 . The method of claim 1 , wherein the first training subset ST(i) with i=1 comprises 1% or less of the objects in the training dataset S. 6 . The method of claim 1 , wherein the training subset ST(i) with i=2 includes less than half of the objects in the error subset ER( 1 ).

7. The method of claim 1 , comprising dividing the training data set S into a plurality of blocks of training data, wherein the first training subset ST(1) is obtained from a first block among the blocks, and the first evaluation subset includes part or all of a second block among the blocks and does not include the first block. The method of claim 7 , wherein the first block and the second block are of the same size.

9. The method as claimed in claim 1 comprises dividing the training data set S into a plurality of blocks of training data, wherein the blocks have the same size, and wherein the training subset ST(i) where i is a given value and the evaluation subset SE(i) where i is the given value are obtained from different blocks among the blocks.

10. The method of claim 9, comprising determining a distribution of the objects of the categories in the training data set, and partitioning the training data set so that some or all of the blocks have the determined distribution.

11. The method of claim 1, comprising: accessing a database comprising the objects classified according to the categories; as well as The database is filtered according to these categories to generate the training data set S.

12. The method of claim 11, wherein the filtering comprises setting a maximum limit on the number of objects of a given category to obtain the objects to be included in the training data set S.

13. The method of claim 12, wherein the filtering comprises setting a minimum limit on the number of objects of a given category to obtain the objects to be included in the training data set S.

14. The method of claim 1, wherein the training subset ST(i) with i=1 has a number N1 of these objects, and the training subset ST(i) with i=2 has a number N2 of these objects, and the number N2 is between 50% and 3% of the number N1.

15. The method of claim 14, wherein the number N2 is between 20% and 5% of the number N1.

16. The method of claim 1, wherein the training subset ST(i) for combinations where i is 1 to A-1 has a number NA of these objects, and the training subset ST(i) for combinations where i is A has a number NB of these objects, and the number NB is between 50% and 3% of the number NA. The method of claim 16 , wherein the number NB is between 20% and 5% of the number NA.

18. The method of claim 1, wherein the step of obtaining the other training subset ST(i) with i>1, the other training subset ST(i) including some of the objects in the error subset ER(i-1) comprises: A target number of these objects are obtained in the error subset ER(i-1) without taking into account the categories contained in the training subset.

19. The method of claim 1, wherein the step of obtaining the other training subset ST(i) with i>1, the other training subset ST(i) including some of the objects in the error subset ER(i-1) comprises: The objects are obtained so that no more than a maximum number M of objects misclassified for each of the categories are included in the training subset.

20. The method of claim 1, wherein the step of obtaining the other training subset ST(i) with i>1, the other training subset ST(i) including some of the objects in the error subset ER(i-1) comprises: The objects are obtained such that at least a minimum number M of objects misclassified for each of the categories are contained in the training subset.

21. The method of claim 1, wherein the step of obtaining the other training subset ST(i) with i>1, the other training subset ST(i) including some of the objects in the error subset ER(i-1) comprises: Obtain a fraction of the target number of these objects in the error subset ER(i-1) without considering these categories; as well as A difference of the target number is obtained so that no more than a maximum number M of the objects misclassified per category is included in the difference of the target number, and the portion of the target number and the difference are included in the training subset.

22. The method of claim 1, comprising applying the stored model M(i) to an inference engine to detect and classify defects in an integrated circuit manufacturing process.

23. The method of claim 1, comprising executing a user interface to provide an interactive tool to display information about the categories of the objects in the training dataset S, setting parameters to configure the training dataset S, and setting parameters to obtain the training subset ST(i) from the error subset ER(i).

24. The method of claim 23, wherein the user interface provides interactive tools to display information about the categories of the objects in the training subset ST(i) and information about the objects in the error subset ER(i).

25. A computer system comprising: One or more processors access a memory storing a classification engine trained by the method of claim 1.

26. A computer program product comprising: A non-transitory computer readable memory storing a computer program, the computer program including a logic to execute a program, the program including: For an index i=1, a first training subset ST(i) is obtained, the first training subset includes a portion of the objects in the training data set, the objects in the training data set include images of defects on integrated circuit components in an integrated circuit manufacturing process, the defects include defects of multiple categories; Using the first training subset ST(i) to train a first model M(i); Using the first model M(i) to classify a first evaluation subset SE(i) in the training data set that does not include the first training subset ST(i), and identifying an error subset ER(i) of the objects that are misclassified in the first evaluation subset SE(i); (a) incrementing the index i and obtaining another training subset ST(i) comprising the objects that are part of the error subset ER(i-1); (b) training a model M(i) using the combination of the training subsets ST(i), where i is 1 to i; (c) using the model M(i) to classify an evaluation subset SE(i) in the training data set S that does not include the training subset ST(i), where i is 1 to i, and identifying an error subset ER(i) of those objects that are misclassified in the evaluation subset SE(i); and (d) evaluating the error subset ER(i) to estimate a performance of the model M(i), and if the performance satisfies a condition, storing the model M(i), and if the performance does not satisfy the condition, repeating steps (a) to (d), The step of obtaining the other training subset ST(i) with i>1, wherein the other training subset ST(i) includes some of the objects in the error subset ER(i-1), comprises: For a given one of these categories, when the number of these objects misclassified in the error subset is less than a minimum number M, the objects in one or more of the error subsets ER(i) i=i-2 to 1 are added to increase the number of these objects of the given category in the training subset.

Citation Information

Patent Citations

  • Classifier training method and device

    CN105320957A

  • Real estate evaluating platform methods, apparatuses, and media

    US20150242747A1

  • Consistent filtering of machine learning data

    US20150379425A1