Computer implemented method, computer system and computer programming product
By dividing the label training data set into two subsets and using the optimization model to filter the subset during the loop, the problem of poor model performance when training neural networks using the training set of wrong label elements in the prior art is solved, and a higher accuracy model training is achieved.
Patent Information
- Application Number
- CN202010782038.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-09
- Filing Date
- 2020-08-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-08-06
AI Technical Summary
The prior art is difficult to effectively use training sets containing error tag elements to train neural networks, resulting in increased training errors and poor model performance.
By dividing the tag training dataset into two subsets and gradually cleaning the data using an optimized model filtered subset during the loop, an optimized model filtered subset is generated to improve data cleanliness and model accuracy.
The training set with error label elements is implemented to train neural networks, improving the accuracy and performance of the model, and the generated cleaned training data set can be used to train target neural networks with higher accuracy.
Smart Images

Figure CN113780547B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a computer implemented method, a computer system and a computer programming product for training a neural network and cleaning data for training a neural network using cleaned data. Background Art
[0002] The subject matter discussed in this section should not be considered prior art simply because it is mentioned in this section. Similarly, the problems mentioned in this section or related to the subject matter provided as background to the invention should not be considered to be recognized in the prior art. The subject matter in this section merely represents different methods, which themselves may also correspond to the claimed technical implementations.
[0003] Neural networks, including deep neural networks, are a type of artificial neural network (ANN) that uses multiple nonlinear and complex transforming layers to continuously model high-level features. Neural networks provide feedback through backpropagation, which carries the difference between the observed output and the predicted output to adjust parameters. Neural networks have developed with the availability of large training datasets, the ability of parallel and distributed computing, and sophisticated training algorithms. Neural networks have facilitated significant advances in many fields, such as computer vision, speech recognition, and natural language processing.
[0004] Convolutional neural networks (CNN) and recurrent neural networks (RNN) can be configured as deep neural networks. In particular, convolutional neural networks have achieved success in image recognition. Convolutional neural networks include a structure of convolution layers, nonlinear layers, and pooling layers. Recurrent neural networks are designed to use the sequence information of input data to establish recurrent connections between blocks such as perceptrons, long short-term memory units, and gated recurrent units. In addition, many other emerging deep neural networks have been proposed in limited environments, such as deep spatio-temporal neural networks, multi-dimensional recurrent neural networks, and convolutional auto-encoders.
[0005] The goal of training a deep neural network is to optimize the weight parameters of each layer, which progressively combines simpler features into complex features so that the most appropriate hierarchical representation can be learned from the data. A single loop of the optimization process is arranged as follows. First, given a training dataset, the forward pass sequentially computes the output of each layer and propagates the functional signal forward through the network. In the final output layer, the objective loss function measures the error between the inferred output and a specific label. To minimize the training error, the backward pass backpropagates the error signal using the chain rule and computes the gradient of all weights in the entire neural network. Finally, an optimization algorithm based on stochastic gradient descent is used to update the weight parameters. While batch gradient descent performs parameter updates for each complete dataset, stochastic gradient descent provides a stochastic approximation by performing updates for each small set of data examples. Several optimization algorithms are derived from stochastic gradient descent. For example, the Adagrad and Adam training algorithms perform stochastic gradient descent while adaptively modifying the learning rate based on the update frequency and moments of the gradients of each parameter, respectively.
[0006] In machine learning, a classification engine including an ANN is trained using a training set that includes a database of data examples labeled according to features to be recognized by the classification engine. Often, some of the data examples used as elements of the training set are incorrectly labeled. In some training sets, a large number of elements are incorrectly labeled. Incorrect labeling can interfere with the learning algorithm used to generate the model, resulting in poor performance.
[0007] The present invention seeks to provide a technique to improve the training of ANNs using training sets with mislabeled elements. Summary of the invention
[0008] The present invention describes a computer-implemented method for cleaning a training data set for a neural network, as well as a computer system and a computer programming product. The computer system and the computer programming product include computer instructions for performing the method. The present invention provides a neural network disposed in an inference engine, the neural network being trained using the techniques described herein.
[0009] A technique for cleaning a training data set for a neural network using dirty training data begins by accessing a dirty labeled training data set. The labeled training data set is divided into a first subset A and a second subset B. A process includes looping between the first subset A and the second subset B, the process including generating a refined model-filtered subset of the first subset A and the second subset B to provide a cleaned data set. Each refined model-filtered subset may improve cleanliness and increase the number of elements.
[0010] In general, the process described herein includes accessing a labeled training data set (S) that includes relatively dirty labeled data elements. The labeled training data set is divided into a first subset A and a second subset B. The process includes: in loop A, using the first subset A to train a model MODEL_A of a neural network; and using the model MODEL_A to filter the second subset B of the labeled training data set. A first model-filtered subset B1F of subset B is provided, which has a plurality of elements depending on the accuracy of MODEL_A. Then, the next loop (i.e., loop AB) includes: using the first model-filtered subset B1F to train the model MODEL_B1F; and using the model MODEL_B1F to filter the first subset A of the labeled training data set. Model MODEL_B1F may have better accuracy than model MODEL_A, and a first optimized model-filtered subset A1F of the first subset A is generated, which has a plurality of elements depending on the accuracy of MODEL_B1F. The execution of another loop (i.e., loop ABA) may include: training model MODEL_A1F using the first optimized model filtering subset A1F; and filtering a second subset B of the label training data set using model MODEL_A1F. Model MODEL_A1F may have better accuracy than model MODEL_A. And generating a second optimized model filtering subset B2F of the second subset B, which has a plurality of elements depending on the accuracy of MODEL_A1F, and the second optimized model filtering subset B2F may have a greater number of elements than the first model filtering subset B1F.
[0011] In the embodiments described herein, looping may continue until an optimized model filtered subset satisfies an iteration criterion based on, for example, data quality or maximum cycle number.
[0012] The combination of the optimized model filter subsets from the first subset A and the second subset B can be combined to provide a cleaned training data set. The cleaned training data set can be used to train an output model of the target neural network, which has a higher accuracy than when trained using the original training data set. The target neural network with the output model can be configured in an inference engine.
[0013] As used herein, a "subset" of a set excludes the null subset and the degenerate case of a subset that includes all members of the set.
[0014] In order to better understand the above and other aspects of the present invention, embodiments are given below and described in detail with reference to the accompanying drawings: BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a simplified diagram of a production line configured with an artificial neural network for defect classification.
[0016] Figure 2A A flow chart of a method for filtering a training data set and training an ANN model using the filtered training data set is shown.
[0017] Figure 2B A flowchart of another method for filtering a training data set and training an ANN model using the filtered training data set is shown.
[0018] Figure 3 Draw for use with a pattern similar to Figure 2A and Figure 2B The method shown is a technique for filtering intermediate data subsets.
[0019] Figure 4 It is a graph that plots the cleanliness of the data versus the number of elements in the training data. It is used to illustrate the correlation between cleanliness, number of elements, and the resulting training model performance.
[0020] Figure 5 is similar to Figure 4 A graph with contour lines showing the accuracy of the trained model across the graph.
[0021] Figure 6 and Figure 7 is similar to Figure 4 The respective drawings have Figure 5 A plot of the first subset A and the second subset B of the contour line label dataset.
[0022] Figure 8 Plotting the case with 80% clean data Figure 6 The first subset A of .
[0023] Fig. 9 Draw as Figure 2A and Figure 2B The first model filtered subset B1F generated is described.
[0024] Fig.10 Draw as Figure 2A and Figure 2B The generated first optimized model filter subset A1F is described.
[0025] Fig.11 Draw as Figure 2A and Figure 2B The generated second optimized model filter subset B2F is described.
[0026] Fig.12 Draw as Figure 2Aand Figure 2B The resulting model filtered subset A2F is described.
[0027] Fig.13 A combination of the modeled filtered subset A2F and the second optimized modeled filtered subset B2F for training the output model is shown.
[0028] Fig.14 is a simplified diagram of a computer system as described herein.
[0029] Fig.15 Embodiments of an inference engine as described herein are shown deployed in a camera, a smartphone, and a car.
[0030]
Explanation of symbols
[0031] 60, 62, X, X+1: Processing station
[0032] M: Artificial neural network model
[0033] 61: Image Sensor
[0034] 63: Inference Engine
[0035] 100, 101, 102, 102, 104, 105, 106, 107, 108, 109, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 170, 171: Step 200: Optimal Point
[0036] 202, 212: Label points
[0037] 210, 220, 222, 230, 233: points
[0038] 600: Training server
[0039] 601: Camera
[0040] 602: Smartphone
[0041] 603: Automobile
[0042] A: The first subset
[0043] B: Second subset
[0044] AB, ABA, ABAB, ABABA: loop
[0045] A1F: First Optimized Model Filter Subset
[0046] A2F: Model filtered subset
[0047] B1F: The first model-filtered subset
[0048] B2F: Second Optimized Model Filtering Subset
[0049] M: ANN model DETAILED DESCRIPTION
[0050] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0051] The embodiments of the present invention are described with reference to Figure 1-Figure 15 Provide detailed explanation.
[0052] Figure 1 60, an image sensor 61, and a processing station 62. In this production line, an integrated circuit wafer is input to processing station X (i.e., processing station 60), and is subjected to processes such as deposition or etching, and is output to image sensor 61. The chip is input from the image sensor to processing station X+1 (i.e., processing station 62), and is subjected to processes such as deposition, etching, or packaging at processing station X+1. Then, the chip is output to the next stage. The image from the image sensor is provided to an inference engine 63, which includes an artificial neural network (ANN) model M trained according to the techniques described herein, which identifies and classifies defects in the chip. The inference engine may also receive images from other stages of the manufacturing process. This information about chip defects sensed at image sensor 61 can be used to improve the manufacturing process (e.g., by adjusting the process performed at processing station X or at other stations).
[0053] As described above, a method for training a neural network to classify defects in a production line or for other classification functions may include a computer-implemented process for cleaning a training data set by removing mislabeled elements.
[0054] Defective images on integrated circuit components acquired in a production line can be classified into many categories and used as elements of a training dataset. These defects vary greatly in the number of specific manufacturing processes, so the training data will have an uneven distribution and contain a large data size. In addition, the labeling process of such images can be done by humans who make a lot of mistakes. For example, in order to build a new neural network model to classify defect categories or types, first, we need to provide a labeled image database to train. The image database includes defect information. One of them may have 50,000 defective images in the database, and each image is labeled by human classification. Therefore, one image in the set will be classified as category 9, while another image in the set will be classified as category 15... and so on. However, human errors and ambiguous cases lead to labeling errors. For example, an image that should be classified as defect category 7 is incorrectly classified as category 3. A dataset with incorrectly classified elements can be called a dirty data set or a noisy data set.
[0055] Embodiments of the techniques described herein can be used to clean dirty data sets and use the cleaned data sets to train ANNs to identify and classify defects, thereby improving manufacturing processes. Such trained ANNs can be used to monitor defects in an in-line process (e.g., to assess the stability and quality of in-line products or the life of manufacturing tools).
[0056] Figure 2A A flow chart of a computer-implemented process for training a neural network ANN starting with "dirty" training data is depicted. The flow chart begins with providing a labeled training data set S (step 100), which may be stored in a database that is accessible by a processor or processors executing the process. An exemplary labeled training data set may include thousands or tens of thousands (or more) of images labeled as described above, or any other type of training data selected depending on the task function of the neural network to be implemented.
[0057] The computer-implemented process accesses a database to retrieve a first subset A and a second subset B of a training data set S (step 101). In one approach, the first subset A and the second subset B are selected so that the distribution of dirty data elements in the subsets is approximately equal to the distribution in the entire data set S. In addition, the first subset A and the second subset B may be selected so that the number of data elements in each subset is approximately the same. Since it is desirable to maximize the number of clean data elements used in a training algorithm, the first subset A and the second subset B may be selected by equally dividing the training data set S. The elements of the first subset A and the second subset B are randomly selected so that the distribution of dirty elements in the two subsets is at least statistically equal.
[0058] Next, in the flowchart (loop A), one of the two subsets (e.g., the first subset A) is used to train the neural network to generate a model MODEL_A (step 102). The second subset B is filtered using the model MODEL_A to generate a first model-filtered subset B1F of the second subset B, and the first model-filtered subset B1F is stored in a memory (step 103) (filtering the second subset B). Figure 3 An example of the technique of using model-filtered subsets is shown. The first model-filtered subset B1F contains the elements of the second subset B whose labels match the inference results from the neural network executing the model MODEL_A. Due to this filtering, the first model-filtered subset B1F should have fewer overall elements and a lower percentage of mislabeled elements than the second subset B.
[0059] Next (loop AB), the neural network is trained using the first modeled filtered subset B1F to produce a refined model MODEL_B1F (step 104). As used herein, the term "refined" is used to refer to a model produced using a modeled filtered subset (or an optimized model filtered subset as in the following example), and does not represent any relative quality measure of the model. Then, using Figure 3The described technique (filtering the first subset A) filters the first subset A using the optimization model MODEL_B1F to generate a first optimized model-filtered subset A1F of the first subset A, and stores the first optimized model-filtered subset A1F in a memory (step 105). The first optimized model-filtered subset A1F includes elements of the first subset A whose labels match the inference results from the neural network executing the optimization model MODEL_B1F. Due to this filtering, the first optimized model-filtered subset A1F may have fewer overall elements and a lower percentage of mislabeled elements compared to the first subset A.
[0060] In the next iteration (loop ABA), the neural network is trained using the first optimized model filter subset A1F to generate an optimized model MODEL_A1F, and the optimized model MODEL_A1F is stored in the memory (step 106). Then, the following method is used: Figure 3 The technique (filtering the second subset B) uses the optimized subset MODEL_A1F to filter the second subset B to generate a second optimized model-filtered subset B2F of the second subset B, and stores the second optimized model-filtered subset B2F in a memory (step 107). Compared to the first model-filtered subset B1F of the second subset B, the second optimized model-filtered subset B2F may have a greater number of elements and a lower percentage of mislabeled elements.
[0061] In this example, no additional filtering cycle is required to provide a cleaned training data set for generating the final output model. For example, the cleaned training data set at this stage may include a combination of the second optimized model filtered subset B2F of the second subset B and the first optimized model filtered subset A1F of the first subset A.
[0062] If no additional filtering cycles are performed, a computer implemented algorithm may train a neural network using a combination of the optimized model filter subsets (e.g., the union of the first optimized model filter subset A1F and the second optimized model filter subset B2F) to produce an output model of the neural network (step 108). The neural network trained at this stage using the cleaned data set may be the same as the neural network used in steps 102, 104, and 106 (or the neural network trained at this stage using the cleaned data set may be a different neural network) to produce the optimized model filter subset. The output model may then be stored in an inference engine for use in the field or in a memory (e.g., a database) (step 109) for later use.
[0063] In some embodiments, Figure 2A In steps 102, 104 and 106, only a subset or a part of a filtered subset may be used as training data to reduce the required processing resources.
[0064] Figure 2B A flowchart is shown of a computer-implemented process for training a neural network ANN starting with "dirty" training data, which iteratively extends the process to other cycles AB, ABA, ABAB, ABABA. The flowchart begins with providing a labeled training data set S (step 150), which may be stored in a database that is accessible by a processor or processors executing the process. An exemplary labeled training data set may include thousands or tens of thousands (or more) of images labeled as described above, or any other type of training data selected depending on the task function of the neural network to be implemented.
[0065] The computer-implemented process accesses a database to obtain a first subset A and a second subset B of a training data set S (step 151). In one approach, the first subset A and the second subset B are selected so that the distribution of dirty data elements in the subsets is approximately equal to the distribution in the entire data set S. In addition, the first subset A and the second subset B can be selected so that the number of data elements in each subset is the same or approximately the same. Since it is desirable to maximize the number of clean data elements used in the training algorithm, the first subset A and the second subset B can be selected by equally dividing the training data set S. The elements of the first subset A and the second subset B are randomly selected so that the distribution of dirty elements in the two subsets at least statistically tends to remain relatively equal. Other techniques for selecting elements of the first subset A and the second subset B can be used to take into account the number of elements in each category and other data content-aware selection techniques.
[0066] Next, in the flowchart, one of the two subsets (e.g., the first subset A) is used to train the neural network to generate a model MODEL_A(n-1) (where n=1) and set the index of the tracking loop (step 152). The second subset B is filtered using the model MODEL_A(n-1) to generate a first model-filtered subset BmF (where m=1) of the second subset B, and the first model-filtered subset BmF is stored in a memory (step 153). Figure 3 An example of the technique of using model-filtered subsets is shown. The first model-filtered subset BmF contains the elements of the second subset B whose labels match the inference results from the neural network executing the model MODEL_A(n-1). Due to this filtering, the first model-filtered subset BmF should have fewer overall elements and a lower percentage of mislabeled elements than the second subset B.
[0067] Next, the neural network is trained using the first model-filtered subset BmF to generate an optimized model MODEL_BmF (step 154). Figure 3 The described technique uses the optimization model MODEL_BmF to filter the first subset A to generate an optimized model-filtered subset AnF of the first subset A, and stores the optimized model-filtered subset AnF in a memory (step 155). The optimized model-filtered subset AnF includes elements of the first subset A whose labels match the inference results from the neural network executing the optimization model MODEL_BmF. Due to this filtering, the optimized model-filtered subset AmF may have fewer total elements and a lower percentage of mislabeled elements compared to the first subset A.
[0068] At this stage, the process determines whether an iteration criterion is satisfied. For example, the iteration criterion may be a maximum number of cycles (as represented by whether index n or index m exceeds a threshold). Alternatively, the iteration criterion may be whether the size (i.e., the number of elements) of the optimized modeled filtered subsets AnF and BmF (i.e., the first modeled filtered subset) converges to the size of the filtered subsets A(n-1)F and B(m-1)F, respectively (step 156). For example, convergence may be indicated if the difference in sizes is less than a threshold, where the threshold may be selected based on the specific application and the training data set used. For example, the threshold may be on the order of 0.1% to 5%.
[0069] As reference Figure 2A As explained, the loops may have a fixed number without requiring an iteration criterion, thereby providing at least one optimized model filter subset, and preferably providing at least one optimized model filter subset for each of the first subset A and the second subset B.
[0070] exist Figure 2B In the case of, if the size of the optimized model filtering subsets AnF and BmF (i.e., the first model filtering subset) has not converged or has not met other iteration criteria, the neural network is trained using the optimized model filtering subset AnF to generate the optimized model MODEL_AnF, and the optimized model MODEL_AnF is stored in the memory (step 157). This process is performed to increase the indexes n and m (step 158), and returns to step 153, where the just generated optimized model MODEL_A(n-1)F is used to filter the second subset B.
[0071] This process continues until the iteration criteria of step 156 are met. If the criteria are met at step 156, the optimized model filter subsets of the first subset A and the second subset B are selected. For example, the optimized model filter subset with the largest number of elements may be selected. The selected model filter subsets of the first subset A and the second subset B are combined to provide a cleaned data set, and a target neural network is trained using the combination of the selected model filter subsets of the first subset A and the second subset B to produce an output model (step 159). The target neural network trained at this stage using the cleaned data set may be the same as the neural network used in steps 152, 154, and 157 (or the target neural network trained at this stage using the cleaned data set may be a different neural network) to produce the optimized model filter subsets.
[0072] The output model may then be stored in an inference engine for use in the field or in storage (e.g., a database) (step 160) for later use.
[0073] In some embodiments, Figure 2B In steps 152, 154, 157, only a subset or a part of a filtered subset may be used as training data to reduce the required processing resources.
[0074] Generally speaking, Figure 2B The process shown includes the following example process, which includes:
[0075] Step S1: using a pre-provided optimized model of one of the first subset and the second subset to filter the subset to train a real-time (instant) optimization model of a neural network;
[0076] Step S2: filtering the other of the first subset and the second subset using the real-time optimization model to provide a real-time optimization model-filtered subset of the other of the first subset and the second subset; and
[0077] Step S3: Determine whether the iteration criteria are met. If the iteration criteria are not met, then execute steps S1 to S3; if the iteration criteria are met, then use the selected model filtering optimization subset of the first subset A and the selected model filtering subset of the second subset B to generate a training model of the neural network.
[0078] Figure 3 Draw a technique for filtering subsets of a training set using a neural network model, such as Figure 2A and Figure 2B Steps 103, 105, 107, 153 and 155 are performed.
[0079] Assuming MODEL_X is provided, the process uses MODEL_X (trained using a subset of the training data set) on subset Y and executes the neural network (step 170). MODEL_X can be MODEL_A, MODEL_B1F, MODEL_A1F, or in general MODEL_A(n)F or MODEL_B(m)F. Subset Y is a subset that is not used to train MODEL_X (another subset).
[0080] Then, elements of subset Y having labels matching the classified data output by the neural network are selected as members of the model-filtered subset of subset Y (step 171).
[0081] This technology can refer to Figure 4-Figure 13 The figure is further described. Figure 4-Figure 13 Plot the characteristics of the training dataset and subsets.
[0082] Figure 4 A graph representing a training dataset S (e.g., a benchmark file CIFAR 10 with 20% noise), which shows the cleanliness of the data on the y-axis and the number of elements of the training dataset S on the x-axis. For example, Figure 4 Can represent a data set of 50,000 elements, Figure 4 The cleanliness of the data can range from 0 to 100%. Any particular data point X, Y is an indication of the number of elements in the set and the cleanliness of the set data. In general, training data sets with more elements (i.e., extending outward along the x-axis) produce more accurate models. Similarly, training data sets with higher cleanliness (i.e., extending upward along the y-axis) can produce more accurate models. The training data set will have an optimal point 200, where the maximum number of data elements with maximum data cleanliness includes 100% of the training data set S. Ideally, if the training data set can be represented by the optimal point 200, the quality of the neural network trained using the training data set will be the best based on the training set.
[0083] Figure 5 is a heuristic contour line Figure 4A reproduction of a graph where the contour lines correspond to the accuracy of the trained neural network based on the training data that falls along the contour lines. Of course, different models will have different contour lines. Thus, the contour lines for the resulting model with an accuracy of 25% intersect at the top of the graph closer to the start of the x-axis, and intersect on the right side of the graph at a relatively low cleanliness of the data. Labeled point 201 represents the location in the graph where the model will have an accuracy of approximately 68%. Labeled point 202 represents the location in the graph where the model will have an accuracy of less than approximately 68%. Labeled point 212 represents the location in the graph where the model will have an accuracy ranging between 68% and 77%. In certain applications, it may be desirable to use a training set to train a model that has an accuracy above the 85% contour line in the upper right corner of the graph.
[0084] Figure 6 The effect of dividing the training dataset into a first subset of approximately 50% of the elements is shown. This corresponds to Figure 2A and Figure 2B The first subset A is referenced in the process shown. This means that using only half of the dataset does not achieve a model accuracy greater than 85%.
[0085] Figure 7 The second subset B is shown, which is generated by splitting the training set in half. Ideally, the first subset A will have nearly identical cleanliness characteristics as the second subset B, so at least conceptually the same contours can be applied. Again, the second subset B cannot be used alone in this conceptual example to achieve a high model accuracy, such as in the 85% range.
[0086] Figure 8 The effect of a cleanliness value of about 80% on the data is plotted. As shown, for the first subset A of 25,000 elements in this example, 20,000 elements can be selected to obtain 100% cleanliness. Ideally, the algorithm used to filter the first subset A can identify these 20,000 correctly labeled elements. In this example, as shown at point 210, the accuracy of the model using the training data set with about 80% cleanliness and 25,000 elements is about 68%.
[0087] Fig. 9 Draw as Figure 2B In this case, the first subset A is used to train the model MODEL_A(0), and MODEL_A(0) is used to filter the second subset B to generate the model-filtered subset B(1)F. Figure 8As described above, model MODEL_A(0) will have an accuracy of about 68%. Therefore, the second subset B with an accuracy of about 68% will identify nearly 99% of the clean data by filtering the second subset B using MODEL_A(0). This clean data is the first model-filtered subset B1F (the result from loop AB). Fig. 9 As shown by the contour lines in , this first model filtered subset B1F containing 68% of the data with nearly 99% accuracy can be expected to produce model MODEL_B1F, which is represented at point 220 with approximately 77% accuracy. Point 222 represents the required accuracy to produce a model with 77% accuracy using the full second subset B.
[0088] Fig.10 Draw as Figure 2B The first optimized model filter subset A1F is generated using the model MODEL_B1F for n=1 in the process of . Since MODEL_B1F has an accuracy of about 77%, the first optimized model filter subset A1F will have about 77% of the elements of the first subset A. As shown by the contour lines, this first optimized model filter subset A1F can be expected to produce a model MODEL_A1F with an accuracy of nearly 79%.
[0089] Fig.11 Draw as Figure 2B The second optimized model filter subset B2F generated by using the model MODEL_A1F for m=2 in the process of . Since MODEL_A1F has an accuracy of nearly 79%, the second optimized model filter subset B2F has nearly 79% of the elements of the second subset B. As shown by the contour lines, this second optimized model filter subset B2F can be expected to produce an improved model MODEL_B2F with an accuracy of nearly 79%.
[0090] Fig.12 Draw as Figure 2B The modeled filtered subset A2F generated by using the model MODEL_B2F for n=2 in the process of . Since MODEL_B2F has an accuracy of about 79%, the modeled filtered subset A2F will have about 79% of the elements of the second subset B. As shown by the contour lines, this modeled filtered subset A2F can be expected to produce a model with an accuracy of about 79% (close to 80%).
[0091] This loop can continue as described above. However, for this training data set, it can be seen that the number of elements in the model-filtered subset is converging to a maximum value of 80%. Therefore, the loop can be stopped and a final training can be selected.
[0092] Fig.13The combination of the largest modeled filtered subset A2F of the first subset A and the largest second optimized modeled filtered subset B2F of the second subset B is shown, which includes nearly 80% of the elements of the first subset A and the second subset B and has a cleanliness of nearly 99%. As a result, the output model trained using this combination can be expected to have an accuracy of about 85% (point 233), which is much higher than the estimated accuracy of between 68% and 77% (point 230) of the model trained using the uncleaned training set A (i.e., the first subset).
[0093] Fig.14 is a simplified block diagram of a computer system 1200, one or more of which may be programmed to implement the techniques disclosed herein. The computer system 1200 includes one or more central processing units (CPUs) 1272 that communicate with a plurality of peripheral devices via a bus subsystem 1255. For example, these peripheral devices may include a storage subsystem 1210 including a storage device and a file storage subsystem 1236, a user interface input device 1238, a user interface output device 1276, and a network interface subsystem 1274. The input and output devices allow a user to interact with the computer system 1200. The network interface subsystem 1274 provides an interface to an external network, including interfaces to corresponding interface devices in other computer systems.
[0094] The user interface input devices 1238 may include: a keyboard; a pointing device (e.g., a mouse, a trackball, a touchpad, or a graphics tablet); a scanner; a touch screen that is incorporated into a display; an audio input device (e.g., a voice recognition system and a microphone); and other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computer system 1200.
[0095] The user interface output device 1276 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include an LED display, a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD), a projection device, or some other machinery for generating visual images. The display subsystem may also provide a non-visual display, such as an audio output device. In general, the use of the term "output device" is intended to include all possible device types and methods for outputting information from the computer system 1200 to a user, another machine, or a computer system.
[0096] The storage subsystem 1210 stores programming and data structures that provide the functionality of some or all of the modules and methods described herein to train ANN models. These models are typically applied to ANNs executed by the deep learning processor 1278.
[0097] In one embodiment, a neural network is implemented using a deep learning processor 1278, which can be a processor that can be configured and reconfigured, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA) and a graphics processing unit (GPU), as well as other configured devices. The deep learning processor 1278 can be hosted by a deep learning cloud platform (e.g., Google Cloud Platform, Xilinx Cloud Platform, and Cirrascale Cloud Platform). Exemplary deep learning processors 1278 include Google's Tensor Processing Unit (TPU), rack-mount solutions such as the GX4 rack-mount server series, the GX149 rack-mount server series, NVIDIA DGX-1, Microsoft's Stratix V FPGA, Graphcore's Intelligent Processor Unit (IPU) TM , Qualcomm's Zeroth platform with Snapdragon processor, NVIDIA's Volta, NVIDIA's DRIVE PX, NVIDIA's JETSON TX1 / TX2 modules, Intel's Nirvana, Movidius vision processing unit (VPU), Fujitsu DPI, ARM's DynamicIQ, IBM TrueNorth, etc.
[0098] The storage subsystem 1222 for the storage subsystem 1210 may include a plurality of memories including a main random access memory (RAM) 1234 for storing instructions and data during programming execution and a read only memory (ROM) 1232 for storing fixed instructions. The instructions include a process for cleaning a training data set and for using, for example, Figure 2A , Figure 2B , Figure 3and Figure 4-13 The process of training a neural network using a cleaned dataset is shown in the figure.
[0099] File storage subsystem 1236 can store programming and data files (including Figure 2A , Figure 2B , Figure 3 The program and data files described above provide persistent storage and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functions of some embodiments may be stored in the storage subsystem 1210 or in other machines accessed by the processor via the file storage subsystem 1236.
[0100] The bus subsystem 1255 provides a mechanism for the various elements and subsystems of the computer system 1200 to communicate with each other as intended. Although the bus subsystem 1255 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple busses.
[0101] The computer system 1200 itself can be of various types, including a personal computer, a portable computer, a workstation, a terminal, a network computer, a television, a mainframe, a server farm, a widely distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, for the purpose of describing the preferred embodiment of the present invention, Fig.14 The description of the computer system 1200 shown in FIG. 1 is intended only as a specific example. Fig.14 The computer system is depicted; many other configurations of computer system 1200 are possible with more or fewer components.
[0102] Embodiments of the technology described herein include a computer program stored on a non-transitory computer readable medium, wherein the non-transitory computer readable medium is configured as a memory that can be accessed and read by a computer, the computer program including Figure 2A , Figure 2B and Figure 3 The programming and data files described.
[0103] Other embodiments of the methods described in this section may include a non-transitory computer readable storage medium storing instructions executed by a processor to perform any of the methods described above. Another embodiment of the methods described in this section may include a system comprising a memory and one or more processors operable to execute instructions stored in the memory to perform any of the methods described above.
[0104] According to many embodiments, any data structures and codes described or referenced above are stored on a computer-readable storage medium, which can be any device or medium capable of storing code and / or data for use by a computer system. However, this includes, but is not limited to, volatile memory, non-volatile memory, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), magnetic and optical storage devices (such as hard disks, magnetic tapes, CDs), DVDs (digital versatile discs or digital laser discs), or other computer-readable media known or later developed that can store data. Compared to Fig.14 The computer system is depicted; many other configurations of computer system 1200 are possible with more or fewer components.
[0105] The thin platform inference engine may include a processor (e.g., a CPU 1272 such as a microcomputer), optionally coupled to a deep learning processor 1278 storing parameters of a trained output model, and input and output ports for receiving input and sending output generated by executing the model. For example, the processor may include a LINUX kernel and ANN programming implemented using executable instructions stored in non-transitory memory accessed by the processor and the deep learning processor, and use model parameters during inference operations.
[0106] As described herein, a device used or included by an inference engine includes: logic for performing ANN operations on input data and a trained model, wherein the model includes a set of model parameters; and a memory that stores the trained model operably coupled to the logic, having a set of trained parameters having values calculated using a training algorithm to compensate for a dirty training set as described herein.
[0107] Fig.15 The present technology is shown to be applied to an inference engine that is suitable for use in an "edge device" such as an Internet of Things (IoT) model. For example, the inference engine can be configured to implement Fig.14 The training server 600 generates trained sets of memory models of ANN for the camera 601, the smart phone 602 and the car 603. Figure 2A and Figure 2B As described above, the trained model can be applied in semiconductor manufacturing.
[0108] Included herein are a number of flow charts illustrating logic for cleaning trained data sets and for trained neural networks. The logic may be implemented using a processor programmed using computer programming stored in memory accessed by the processor and executed by the processor, dedicated logic hardware including field programmable integrated circuits, and a combination of dedicated logic hardware and computer programming. For all of the flow charts herein, it should be understood that many steps may be combined, performed in parallel, or performed in a different order without affecting the functionality achieved. In some cases, as the reader will understand, the rearrangement of steps may also achieve the same result only if certain other changes are made. In other cases, as the reader will understand, the rearrangement of steps may achieve the same result only if certain other changes are also made. In addition, it should be understood that the flow charts herein only show steps relevant to the understanding of the present invention; and it should be understood that many additional steps for accomplishing other functions may be performed before, after, and between those steps shown.
[0109] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A computer-implemented method for cleaning training data of a neural network, wherein: include: Accessing a labeled training data set; wherein the elements in the labeled training data set are defect images on integrated circuit components obtained in a production line; Using a first subset of the labeled training data set to train a first model of the neural network; Using the first model to filter a second subset of the labeled training data set to provide a first model-filtered subset of the second subset; Training a first optimized model of the neural network using the first model-filtered subset of the second subset; Filtering the first subset using the first optimized model to provide a first optimized model-filtered subset of the first subset; Using the first optimized model of the first subset to filter a subset to train a second optimized model of the neural network; and Using the second optimized model to filter the second subset of the labeled training data set to provide a second optimized model-filtered subset of the second subset; combining the first optimized model-filtered subset with the second optimized model-filtered subset of the second subset to provide a filtered training set; Using the filtered training set to train an output model of a target neural network; and storing the output model in a memory; The output model is used to identify and classify defects in the chip.
2. The computer-implemented method of claim 1, wherein: Compared with the first modeled filtered subset, the second optimized modeled filtered subset has a larger number of elements.
3. The computer-implemented method of claim 1 , wherein: The first subset does not overlap with the second subset.
4. The computer-implemented method of claim 1 , wherein: The step of filtering the first subset using the first optimization model comprises: executing the neural network using the first optimization model on the first subset to generate classified data classified as a data element of the first subset; The data elements of the first subset having labels matching the classified data are selected to provide the first optimized model-filtered subset of the first subset.
5. The computer-implemented method of claim 1 , wherein: The output model is loaded into the target neural network in an inference engine.
6. The computer-implemented method of claim 1 , wherein: include: Step S1: iteratively using a pre-provided optimized model filtering subset of one of the first subset and the second subset to train a real-time optimization model of the neural network; Step S2 iteratively filters the other of the first subset and the second subset using the real-time optimization model to provide a real-time optimization model filtered subset of the other of the first subset and the second subset; as well as Step S3 iteratively determines whether an iteration criterion is met. If the iteration criterion is not met, steps S1 to S3 are executed. If the iteration criterion is met, a combination of the selected model filtering optimization subset of the first subset and the selected model filtering subset of the second subset is used to generate a training model of the neural network.
7. The computer-implemented method of claim 6, wherein: An output model is loaded into a target neural network in an inference engine.
8. A computer system, wherein: A computer system configured to clean training data of a neural network comprises: One or more processors and memory storing computer programming instructions configured to execute a process comprising: Accessing a labeled training data set; wherein the elements in the labeled training data set are defect images on integrated circuit components obtained in a production line; Using a first subset of the labeled training data set to train a first model of the neural network; Using the first model to filter a second subset of the labeled training data set to provide a first model-filtered subset of the second subset; Training a first optimized model of the neural network using the first model-filtered subset of the second subset; Filtering the first subset using the first optimized model to provide a first optimized model-filtered subset of the first subset; Using the first optimized model of the first subset to filter a subset to train a second optimized model of the neural network; and Using the second optimized model to filter the second subset of the labeled training data set to provide a second optimized model-filtered subset of the second subset; combining the first optimized model-filtered subset with the second optimized model-filtered subset of the second subset to provide a filtered training set; Using the filtered training set to train an output model of a target neural network; and storing the output model in the memory; The output model is used to identify and classify defects in the chip.
9. The computer system according to claim 8, wherein: Compared with the first modeled filtered subset, the second optimized modeled filtered subset has a larger number of elements.
10. The computer system according to claim 8, wherein: Filtering the first subset using the first optimization model includes: executing the neural network using the first optimization model on the first subset to generate classified data classified as a data element of the first subset; The data elements of the first subset having labels matching the classified data are selected to provide the first optimized model-filtered subset of the first subset.
11. A computer programming product, wherein: A computer program product configured to support training data cleaning of a neural network includes a non-transitory computer readable memory storing computer program instructions, the computer program instructions configured to perform a process, the process comprising: Accessing a labeled training data set; wherein the elements in the labeled training data set are defect images on integrated circuit components obtained in a production line; Using a first subset of the labeled training data set to train a first model of the neural network; Using the first model to filter a second subset of the labeled training data set to provide a first model-filtered subset of the second subset; Training a first optimized model of the neural network using the first model-filtered subset of the second subset; Filtering the first subset using the first optimized model to provide a first optimized model-filtered subset of the first subset; Using the first optimized model of the first subset to filter a subset to train a second optimized model of the neural network; and Using the second optimized model to filter the second subset of the labeled training data set to provide a second optimized model-filtered subset of the second subset; combining the first optimized model-filtered subset with the second optimized model-filtered subset of the second subset to provide a filtered training set; Using the filtered training set to train an output model of a target neural network; and storing the output model in a memory; The output model is used to identify and classify defects in the chip.
12. The computer program product according to claim 11, wherein: Compared with the first modeled filtered subset, the second optimized modeled filtered subset has a larger number of elements.
13. The computer program product according to claim 11, wherein: The first subset does not overlap with the second subset.
14. The computer program product of claim 11, wherein: Filtering the first subset using the first optimization model includes: executing the neural network using the first optimization model on the first subset to generate classified data classified as a data element of the first subset; The data elements of the first subset having labels matching the classified data are selected to provide the first optimized model-filtered subset of the first subset.
Citation Information
Patent Citations
Neural network training apparatus and method, and speech recognition apparatus and method
CN106683663A
Training method and device of classification model, mobile terminal, and readable storage medium
CN108875821A