Image classification model construction method and device
By constructing a tree-like image classification model, using defect characteristics and shunt rules to classify wafer surface defects, the problem of low defect detection accuracy in the existing technology is solved, and faster and more accurate defect classification is achieved.
Patent Information
- Application Number
- CN202410835228.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-08-05
Smart Images

Figure CN120431359A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of integrated circuit manufacturing, and particularly to a method and device for constructing an image classification model. Background Art
[0002] Real-time monitoring of the statistical information of wafer surface defects can achieve better yield analysis and improve the utilization rate of wafer defect detection equipment. Therefore, it is necessary to determine the types of defects after wafer surface defect detection, and an image classification algorithm model can be built to quickly and timely classify the types of defects.
[0003] Currently, defect information can be obtained by background segmentation of the wafer surface defect image, and then the defect type can be determined by template matching. However, due to multiple factors such as different parameter configurations of the lighting module, different imaging methods, and different current process layers to be detected on the wafer during the defect detection process, the image grayscale, contrast, pattern background, etc. of the wafer surface defect image vary greatly. Therefore, it is difficult to accurately classify defects by the aforementioned method. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a method and device for constructing an image classification model, to construct an image classification model for high-precision defect classification and improve the defect classification accuracy. The specific solutions are as follows:
[0005] In the first aspect, this application provides a method for constructing an image classification model, including:
[0006] Obtain a training data set, where the training data set includes multiple training images, the multiple training images have training labels, and the training images are defect images;
[0007] Construct each classification node of the image classification model based on the training data set, and the classification node has a splitting rule; during the process of constructing each classification node, if the defect image corresponding to the target node in each classification node corresponds to multiple training labels, then construct multiple child nodes of the target node; the target node is used to classify the defect image corresponding to the target node according to the splitting rule of the target node, and obtain multiple groups of defect images corresponding one-to-one to the multiple child nodes.
[0008] Optionally, the image classification model is used to classify the image to be classified to obtain a predicted classification result, and the training label is a reclassification label. The obtaining of the training data set includes:
[0009] Obtain an initial training set, where the training set includes multiple training images, and the training images have initial labels;
[0010] Clustering the initial training set to obtain reclassification labels for the multiple training images. There is a mapping relationship between the reclassification labels and the initial labels corresponding to the same training image. The image classification model is used to determine the initial classification result of the image to be classified, so as to determine the predicted classification result of the image to be classified according to the initial classification result and the mapping relationship;
[0011] During the process of clustering the initial training set, the multiple training images are successively used as target images, and multiple first feature distances between the target image and multiple similar images with the same initial label as the target image are respectively determined. The dissimilar images with different initial labels from the target image are successively used as current images, and the second feature distance between the target image and the current image is determined. If the target distance among the multiple first feature distances is greater than the second feature distance, a reclassification label is determined for the similar images corresponding to the target distance.
[0012] Optionally, the method further includes:
[0013] Obtaining a test data set, where the test data set includes multiple test images, the multiple test images have test labels, and the test images are defective images;
[0014] Determining the test classification results of the multiple test images through the image classification model;
[0015] Determining the test result according to the test classification result and the test label;
[0016] If the test result is that the test of the to-be-optimized label in the training label fails, an optimization operation is performed to reconstruct each classification node of the image classification model;
[0017] The optimization operation includes: adding the test images with the test label being the to-be-optimized label to the initial training set as training images, and / or setting a limiting condition for the diversion rule for a preset node in the classification node.
[0018] Optionally, the defect features include at least one of position feature, shape feature, gray feature, alignment-related feature, and texture feature; the first feature distance is calculated based on the defect features of the target image and the defect features of the similar images, and the second feature distance is calculated based on the defect features of the target image and the defect features of the current image.
[0019] Optionally, the diversion rule includes defect attribute and classification threshold.
[0020] Optionally, the diversion rule corresponding to the target node includes a target defect attribute and a target classification threshold, and the method further includes:
[0021] Successively use the candidate defect attributes as the current defect attribute, and determine the amplitude and total bandwidth of the current defect attribute. The amplitude of the current defect attribute is the sum of the amplitudes of the current defect attribute in the training images corresponding to multiple training labels. The amplitude of the current defect attribute in the training image corresponding to the current label among the multiple training labels is the difference between the maximum value and the minimum value of the current defect attribute in the training image corresponding to the current label. The total bandwidth of the current defect attribute is the difference between the maximum value and the minimum value of the current defect attribute in the multiple training images.
[0022] Determine the current fitness of the current defect attribute to the training dataset according to the ratio of the amplitude and total bandwidth of the current defect attribute.
[0023] Use the candidate defect attribute whose current fitness meets the first condition as the target defect attribute.
[0024] Optionally, the method further includes:
[0025] Arrange the eigenvalues of the target defect attribute in the training image in ascending order to obtain an eigenvalue sequence.
[0026] Use the average value of two adjacent eigenvalues in the eigenvalue sequence as the candidate threshold.
[0027] Successively use the candidate threshold as the current threshold, and classify the defect images corresponding to the target node based on the current threshold to obtain multiple groups of initial defect images corresponding one-to-one to the multiple child nodes of the target node.
[0028] Determine the current classification purity of the current threshold according to the number of training images corresponding to each training label in the multiple groups of initial defect images.
[0029] Use the candidate threshold whose current classification purity meets the second condition as the target classification threshold.
[0030] Optionally, the method further includes:
[0031] Obtain the modification information for the node to be modified in each classification node.
[0032] Modify at least one of the defect attribute and classification threshold of the node to be modified according to the modification information.
[0033] Optionally, the method further includes:
[0034] If the defect attribute of the target node is the same as the defect attribute of the target child node among the multiple child nodes, increase the classification threshold of the target child node to the classification threshold of the target node, and replace the target child node with the multiple child nodes of the target child node.
[0035] In a second aspect, an embodiment of the present application further provides an image classification model construction device, including:
[0036] A training data acquisition unit, configured to acquire a training data set, the training data set includes multiple training images, the multiple training images have training labels, and the training images are defect images;
[0037] A model construction unit, configured to construct each classification node of the image classification model based on the training data set, the classification node has a shunting rule; during the process of constructing each classification node, if the defect image corresponding to the target node in each classification node corresponds to multiple training labels, construct multiple child nodes of the target node; the target node is configured to classify the defect image corresponding to the target node according to the shunting rule of the target node, and obtain multiple groups of defect images corresponding one-to-one to the multiple child nodes.
[0038] Optionally, the image classification model is used to classify the image to be classified to obtain a predicted classification result, the training label is a reclassification label, and the training data acquisition unit includes:
[0039] An initial training set acquisition unit, configured to acquire an initial training set, the training set includes multiple training images, and the training images have initial labels;
[0040] A reclassification unit, configured to perform clustering on the initial training set to obtain the reclassification labels of the multiple training images, and there is a mapping relationship between the reclassification label and the initial label corresponding to the same training image, and the image classification model is used to determine the initial classification result of the image to be classified, so as to determine the predicted classification result of the image to be classified according to the initial classification result and the mapping relationship;
[0041] During the process of clustering the initial training set, the multiple training images are successively used as target images, and multiple first feature distances between the target image and multiple similar images having the same initial label as the target image are respectively determined. The dissimilar images having different initial labels from the target image are successively used as current images, and the second feature distance between the target image and the current image is determined. If the target distance among the multiple first feature distances is greater than the second feature distance, a reclassification label is determined for the similar images corresponding to the target distance.
[0042] Optionally, the device further includes:
[0043] A test data set acquisition unit for acquiring a test data set, where the test data set includes a plurality of test images, the plurality of test images have test labels, and the test images are defective images;
[0044] A test classification unit for determining test classification results of the plurality of test images through the image classification model;
[0045] A test result determination unit for determining a test result according to the test classification result and the test label;
[0046] An optimization unit for performing an optimization operation to reconstruct each classification node of the image classification model if the test result fails for a label to be optimized in the training labels;
[0047] The optimization operation includes: adding the test image with the test label being the label to be optimized to the initial training set as a training image, and / or setting a limiting condition for the shunt rule for a preset node in the classification node.
[0048] Optionally, the defect features include at least one of position feature, shape feature, gray level feature, alignment-related feature, and texture feature; the first feature distance is calculated based on the defect features of the target image and the defect features of the same-class images, and the second feature distance is calculated based on the defect features of the target image and the defect features of the current image.
[0049] Optionally, the shunt rule includes defect attributes and classification thresholds.
[0050] Optionally, the shunt rule corresponding to the target node includes target defect attributes and target classification thresholds, and the device further includes:
[0051] A data determination unit for successively taking candidate defect attributes as current defect attributes, and determining the amplitude and total bandwidth of the current defect attribute. The amplitude of the current defect attribute is the sum of the amplitudes of the current defect attribute in the training images corresponding to multiple training labels. The amplitude of the current defect attribute in the training image corresponding to the current label among the multiple training labels is the difference between the maximum value and the minimum value of the current defect attribute in the training image corresponding to the current label. The total bandwidth of the current defect attribute is the difference between the maximum value and the minimum value of the current defect attribute in the multiple training images;
[0052] A fitness calculation unit for determining the current fitness of the current defect attribute to the training data set according to the ratio of the amplitude and the total bandwidth of the current defect attribute;
[0053] A defect attribute determination unit, configured to use a candidate defect attribute whose current fitness meets a first condition as the target defect attribute.
[0054] Optionally, the device further includes:
[0055] A sorting unit, configured to sort the feature values of the target defect attribute in the training image in ascending order to obtain a feature value sequence;
[0056] A candidate threshold determination unit, configured to use the average value of two adjacent feature values in the feature value sequence as a candidate threshold;
[0057] A classification unit, configured to use the candidate threshold as the current threshold one by one, and classify the defect images corresponding to the target node based on the current threshold to obtain multiple groups of initial defect images corresponding one by one to multiple child nodes of the target node;
[0058] A classification purity calculation unit, configured to determine the current classification purity of the current threshold according to the number of training images corresponding to each training label in the multiple groups of initial defect images;
[0059] A classification threshold determination unit, configured to use a candidate threshold whose current classification purity meets a second condition as the target classification threshold.
[0060] Optionally, the device further includes:
[0061] A modification information determination unit, configured to obtain modification information for a node to be modified in each classification node;
[0062] A modification unit, configured to modify at least one of the defect attribute and the classification threshold of the node to be modified according to the modification information.
[0063] Optionally, the device further includes:
[0064] A model simplification unit, configured to, if the defect attribute of the target node is the same as the defect attribute of a target child node among the multiple child nodes, increase the classification threshold of the target child node to the classification threshold of the target node, and replace the target child node with multiple child nodes of the target child node.
[0065] In a third aspect, an embodiment of the present application discloses a computer device, which includes a processor and a memory:
[0066] The memory is configured to store program code and transmit the program code to the processor;
[0067] The processor is configured to execute the image classification model construction method as described in the first aspect according to the instructions in the program code.
[0068] In a fourth aspect, an embodiment of the present application discloses a computer-readable storage medium for storing a computer program, which is used to execute the image classification model construction method as described in the first aspect when executed by a processor.
[0069] The embodiment of the present application provides an image classification model construction method and apparatus. The training data set may include multiple training images with training labels. Based on the training data set, each classification node of the image classification model can be constructed, and each classification node has a splitting rule. During the process of constructing each classification node, if the defective images corresponding to the target node in each classification node correspond to multiple training labels, multiple child nodes of the target node are constructed. That is to say, the image classification model is a tree structure, and the number of classification nodes is determined according to the training process, having a lower number of parameters compared to the deep learning algorithm model, and thus having a faster classification speed. The target node is used to classify the defective images corresponding to the target node according to the splitting rule of the target node, obtaining multiple groups of defective images corresponding one-to-one to multiple child nodes. In this way, the defective images can be classified sequentially through each classification node, thereby accurately and quickly determining the defect category of the defective images and improving the classification accuracy and speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0071] Figure 1 FIG. shows a flowchart of an image classification model construction method provided by an embodiment of the present application;
[0072] Figure 2 FIG. shows a structural diagram of an image classification model provided by an embodiment of the present application;
[0073] Figure 3 FIG. shows a structural diagram of another image classification model provided by an embodiment of the present application;
[0074] Figure 4 FIG. is a structural diagram of a defect classification apparatus provided by an embodiment of the present application;
[0075] Figure 5 FIG. is a structural block diagram of an image classification model construction apparatus provided by an embodiment of the present application;
[0076] Figure 6Structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0077] To make the above objects, features, and advantages of the present application more obvious and understandable, the following will describe the detailed implementation manners of the present application in conjunction with the accompanying drawings.
[0078] In the following description, many specific details are set forth to facilitate a thorough understanding of the present application. However, the present application may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.
[0079] For ease of understanding, the following will describe in detail an image classification model construction method and device provided by an embodiment of the present application in conjunction with the accompanying drawings.
[0080] Reference Figure 1 As shown, it is a flowchart of an image classification model construction method provided by an embodiment of the present application, and the method may include the following steps.
[0081] S101, Obtain a training data set, where the training data set includes multiple training images, and the multiple training images have training labels.
[0082] In an embodiment of the present application, a training data set can be obtained for constructing an image classification model. The training data set can include multiple training images, and the multiple training images can have their respective training labels. The training images are defective images, and the training labels are used to indicate the defect categories of the training images.
[0083] The training images are test images of wafers, which show the defect information on the wafer surface and are obtained by completing the defect detection task through a wafer surface defect detection system. Based on the training images, defect features can be determined. The defect features include at least one of, but are not limited to, position features, shape features, gray-scale features, alignment-related features, and texture features. The defect features can be calculated by an image attribute calculation program in a computer. The position feature represents the position information of the defect in each coordinate system of the wafer. The shape feature represents the morphological features of each defect area, such as aspect ratio and roundness. The gray-scale feature represents the probability statistical information in the defect gray-scale image or the pattern background information of its defect area. The alignment-related feature represents the image alignment similarity between the detection image and the comparison image. The texture feature represents the noise level of the defect gray-scale image or the pattern background information of its defect area.
[0084] The training labels can be manually labeled labels or re-classification labels. Among them, the re-classification labels are determined based on the defect features of the defective images, can better match the image classification model, and are beneficial to improving the accuracy of the image classification model.
[0085] The reclassification labels are obtained through reclassification operations. The reclassification operations can specifically be as follows: obtain an initial training set, where the initial training set includes multiple training images, and the training images have initial labels. Then, perform clustering on the initial training set to obtain the reclassification labels of the multiple training images. There is a mapping relationship between the reclassification labels and the initial labels corresponding to the same training image. The initial labels can be manually labeled labels. In this way, during the application process of the model, the image classification model can determine the initial classification result of the image to be classified. The initial classification result matches the reclassification label. According to the initial classification result and the mapping relationship, the predicted classification result of the image to be classified can be determined, and the predicted classification result matches the initial label. In this way, during the construction and application process of the image classification model, reclassification labels are used. The reclassified labels better match the image classification model, making the image classification model more accurate. The manual labeling process and the result output process match the initial labels. The initial labels are obtained through manual labeling and are adapted to the classification requirements, enabling image classification to be adapted to different classification requirements. Compared with deep learning models, it can use a model of a general size to adapt to all classification tasks.
[0086] The clustering performed on the initial training set can be nearest neighbor clustering. During the process of clustering the initial training set, multiple training images can be successively used as target images, and multiple first feature distances between the target image and multiple similar images with the same initial label as the target image can be determined respectively; after determining the target image, the dissimilar images with different initial labels from the target image can be successively used as the current image, and the second feature distance between the target image and the current image can be determined; if the target distance among the multiple first feature distances is greater than the second feature distance, it indicates that although the target image and the similar image corresponding to the target distance are classified by the user with the same label, from the perspective of image features, the similar image corresponding to the target distance of the target image is not the same classification. Then, a reclassification label is determined for the similar image corresponding to the target distance, and during the clustering process, this reclassification label will be used as a new initial label to participate in the reclassification of other training images. In this way, by taking all training images as target images and traversing all training images as the current image, the reclassification of the initial training set is achieved, enabling the training images in the initial training set to have reclassification labels related to image features.
[0087] The first feature distance is calculated based on the defect features of the target image and the defect features of the similar images, and the second feature distance is calculated based on the defect features of the target image and the defect features of the current image. The first feature distance can be a cosine distance or an L2 distance, and the second feature distance can be a cosine distance or an L2 distance.
[0088] In addition to the training data set, the embodiments of the present application may also have a test data set. The test data set includes multiple test images, and the multiple test images have test labels. The test images are defective images. The test labels may be manually labeled labels. If the training labels are manually labeled labels, the manually labeled training sample set may be divided into a training data set and a test data set according to a ratio; if the training labels are reclassified labels and the initial labels of the training images in the initial training set are manually labeled labels, the manually labeled training sample set may be divided into an initial training set and a test data set according to a ratio.
[0089] S102. Construct each classification node of the image classification model based on the training data set, and the classification node has a shunt rule.
[0090] In the embodiments of the present application, each classification node of the image classification model can be constructed based on the training data set. The classification node has a shunt rule. The shunt rule includes a defect attribute and a classification threshold. The image classification model can be a tree structure. The number of parameters of the image classification model is lower than that of a large-volume deep learning algorithm model, indicating that it has less computational complexity and faster classification speed.
[0091] The defect attribute is used to match the defect features of the defective image. It includes the attributes corresponding to the defect features. The defect features are the feature values corresponding to the defect attributes. The defect attribute can be, for example, the aspect ratio attribute, roundness attribute, etc. corresponding to the shape feature. The classification threshold is used to define the classification condition based on the defect attribute. Each classification condition corresponds to a child node of the classification node. For example, if the classification threshold is 5, the classification conditions can be the first classification condition that the feature value corresponding to the defect attribute is greater than or equal to 5, and the second classification condition that the feature value corresponding to the defect attribute is less than 5. At this time, the classification node has 2 child nodes corresponding to 2 classification conditions respectively. Or if the classification thresholds are 5 and 10, the classification conditions can be the first classification condition that the feature value corresponding to the defect attribute is less than 5, the second classification condition that the feature value corresponding to the defect attribute is greater than or equal to 5 and less than 10, and the third classification condition that the feature value corresponding to the defect attribute is greater than or equal to 10. At this time, the classification node has 3 child nodes corresponding to 3 classification conditions respectively.
[0092] Taking the target node in each classification node as an example, the classification rule of the target node may include a target defect attribute and a target classification threshold. Then, the target node can classify the defect image corresponding to the target node according to the shunt rule of the target node, and obtain multiple groups of defect images corresponding one by one to multiple child nodes of the target node. Specifically, the target node can classify the defect image whose target defect attribute satisfies the first classification condition based on the target classification threshold into the first child node among the multiple child nodes, and classify the defect image whose target defect attribute satisfies the second classification condition based on the target classification threshold into the second child node among the multiple child nodes, realizing the classification based on the target defect attribute and the target classification threshold.
[0093] In the process of constructing each classification node, if the defect image corresponding to the target node corresponds to multiple training labels, multiple child nodes of the target node are constructed. On the contrary, if the defect image corresponding to the target node corresponds to the same training label, the target node is used as the end node, and this training label is used as the classification result of the target node. That is to say, for the composition of the image classification model, the classification nodes can be constructed one by one from top to bottom, and the shunt rules of the classification nodes can be determined.
[0094] After constructing the target node, the target defect attribute of the target node can be determined. The target defect attribute can be determined according to the adaptability of all defect attributes to the training dataset to ensure better inter-class separation of the target node.
[0095] Specifically, the candidate defect attribute can be successively used as the current defect attribute, the amplitude and total bandwidth of the current defect attribute are determined, and according to the ratio of the amplitude and total bandwidth of the current defect attribute, the current adaptability of the current defect attribute to the training dataset is determined; the candidate defect attribute whose current adaptability satisfies the first condition is used as the target defect attribute. The current adaptability of the current defect attribute to the training dataset = amplitude of the current defect attribute / total bandwidth. The candidate defect attribute whose current adaptability satisfies the first condition can be, for example, the candidate defect attribute with the highest current adaptability among all current adaptabilities.
[0096] Among them, the amplitude of the current defect attribute is the sum of multiple amplitudes of the current defect attribute in the training images corresponding to multiple training labels respectively. That is, the amplitude of the current defect attribute in the training images corresponding to the same training label is the amplitude of one type of defect attribute, and the multiple amplitudes corresponding to multiple training labels respectively are the amplitudes of multiple types of defect attributes. The amplitude of the current defect attribute is the sum of the amplitudes of multiple types of defect attributes. Among them, the amplitude of the current defect attribute in the training image corresponding to the current label among the multiple training labels is the difference between the maximum value and the minimum value of the current defect attribute in the training image corresponding to the current label. The total bandwidth of the current defect attribute is the difference between the maximum value and the minimum value of the current defect attribute in the multiple training images.
[0097] After determining the target defect attribute of the target node, the target classification threshold of the target node can be determined. The target classification threshold can be determined by the classification purity of the candidate thresholds to ensure a better inter-class separation degree of the target node.
[0098] Specifically, the feature values of the target defect attribute in the training image can be sorted in ascending order to obtain a feature value sequence; the average value of two adjacent feature values in the feature value sequence is used as a candidate threshold; the candidate threshold is successively used as the current threshold, and the defect images corresponding to the target node are classified based on the current threshold to obtain multiple groups of initial defect images corresponding to the multiple child nodes of the target node one by one; according to the number of training images corresponding to each training label in the multiple groups of initial defect images, the current classification purity of the current threshold is determined; the candidate threshold whose current classification purity meets the second condition is used as the target classification threshold. The candidate threshold whose current classification purity meets the second condition can be the candidate threshold with the highest current classification purity among all current classification purities.
[0099] Taking the number of multiple child nodes as 2 as an example, the calculation process of the current classification purity purity can be expressed as:
[0100] purity = Corr([b1, b2,..., b i , [c1, c2,..., c i )
[0101] where b i is the number of training images corresponding to the i-th training label in the first group of defect images in the multiple groups of initial defect images, c i is the number of training images corresponding to the i-th training label in the second group of defect images in the multiple groups of initial defect images, and Corr is the calculation formula of matrix correlation. The sum of all b i and all c i is the total number N of the defect images corresponding to the target node.
[0102] After determining the target classification threshold of the target node, the target node is constructed. After that, the construction of the child nodes of the target node can be carried out. The number of constructed child nodes is equal to the number of groups of defective images obtained after classifying the target node, so that the multiple groups of defective images correspond to multiple child nodes of the target node one by one. The construction process of the child nodes of the target node can refer to the construction process of the target node. If the defective images corresponding to the child nodes of the target node correspond to multiple training labels, the child nodes of the child nodes of the target node can be constructed. On the contrary, if the defective images corresponding to the child nodes of the target node correspond to the same training label, the child nodes of the target node can be used as the end nodes, and the training label can be used as the classification result of the child nodes. According to the aforementioned construction principle of the classification node, the classification nodes can be constructed one by one from top to bottom, and the classification rules of the classification nodes can be determined until all the end nodes are constructed. The defective images corresponding to all the end nodes only contain the defective images corresponding to one training label, and then the image classification model is constructed.
[0103] Reference Figure 2 As shown, it is a schematic structural diagram of an image classification model provided by an embodiment of the present application. The image classification model may include a total of 9 classification nodes, namely Node 1 - Node 9. The defective attribute of Node 1 is the defective shape, which is used to distinguish between square and circular. The defective images with the square defective state are classified into the first group, corresponding to Node 2. The defective images with the circular defective shape are classified into the second group, corresponding to Node 3. Node 2 and Node 3 are the child nodes of Node 1.
[0104] The defective attribute of Node 2 is the defective size, and the classification threshold is 5, which is used to distinguish the size of the defective. Among the defective images in the first group, the defective images with a defective size less than 5nm are classified into the third group, corresponding to Node 4. Among the defective images in the first group, the defective images with a defective size greater than or equal to 5nm are classified into the fourth group, corresponding to Node 5. Node 4 and Node 5 are the child nodes of Node 2.
[0105] The defective attribute of Node 3 is the defective size, and the classification threshold is 5, which is used to distinguish the size of the defective. Among the defective images in the second group, the defective images with a defective size less than 5nm are classified into the fifth group, corresponding to Node 6. Among the defective images in the second group, the defective images with a defective size greater than or equal to 5nm are classified into the sixth group, corresponding to Node 7. Node 6 and Node 7 are the child nodes of Node 3.
[0106] The defective attribute of Node 7 is the defective size, and the classification threshold is 10, which is used to distinguish the size of the defective. Among the defective images in the sixth group, the defective images with a defective size less than 10nm are classified into the seventh group, corresponding to Node 8. Among the defective images in the sixth group, the defective images with a defective size greater than or equal to 10nm are classified into the eighth group, corresponding to Node 9. Node 8 and Node 9 are the child nodes of Node 7.
[0107] In the embodiments of the present application, if the training label is a manually labeled label, the image classification model classifies the image to be classified and obtains a classification result that is one of the training labels, realizing the classification of the image to be classified; if the training label is a reclassification label, the image classification model classifies the image to be classified and obtains a classification result that is one of the reclassification labels, and maps the reclassification label to one of the initial labels according to the mapping relationship, realizing the classification of the image to be classified.
[0108] Since the reclassification label is determined based on the defect features of the defect image, the image classification model is constructed directionally based on the mechanism of the difference interval between the defect feature values of the defect images, so it has a high classification accuracy. The classification accuracy of the image classification model formed by this method has been verified in the actual application scenario, and the classification accuracy can reach more than 98%. The image classification model can be applied to the field of integrated circuit manufacturing yield monitoring, which has been less studied in China, and realizes a method for intelligently generating a wafer surface defect image classification model according to the application scenario. This method can be applied to all wafer surface defect detection devices based on the principle of photochemical detection, such as a deep ultraviolet light source defect detection system based on bright field.
[0109] In the embodiments of the present application, after constructing the classification nodes of the image classification model, the image classification model can be simplified. Specifically, if the defect attribute of the target node is the same as the defect attribute of the target child node among the multiple child nodes of the target node, the classification threshold of the target child node can be increased to the classification threshold of the target node, and the target child node is replaced by the multiple child nodes of the target child node, so that the target child node is deleted, the overall number of classification nodes is reduced, the number of child nodes of the target node is increased, the node layer is decreased, and the image classification model is simplified.
[0110] Reference Figure 2 As shown, the defect attributes of node 3 and node 7 are both defect shapes, which are used to distinguish the defect size. Node 3 is used as the target node, and node 7 is used as the target child node. Then the classification threshold of node 7 can be increased to the classification threshold of node ③, obtaining a new classification threshold of 10. Node 7 is replaced by the child nodes of node 7: node 8 and node 9, realizing the simplification of the image classification model. Reference Figure 3 As shown, it is a schematic structural diagram of another image classification model provided by the embodiments of the present application. The defect attribute of node 3 is the defect shape, the classification thresholds are 5 and 10, and the child nodes are node 6, node & and node 9. In this way, node 7 is deleted, the overall number of classification nodes is reduced, the number of child nodes of node 3 is increased, the node layer is decreased, and the image classification model is simplified.
[0111] In the embodiments of this application, the image classification model can also be verified. Specifically, a test data set is obtained, the test classification results of multiple test images are determined through the image classification model, and the test results are determined according to the test classification results and test labels. The test results can include the classification accuracy rates corresponding to each training label. If the classification accuracy rates corresponding to each training label are all higher than the preset value, it indicates that the image classification model passes the test. If the test result shows that the test for the to-be-optimized label in the training labels fails, an optimization operation can be performed to reconstruct each classification node of the image classification model, that is, S102 is executed again.
[0112] The optimization operation can include adding the test images with the to-be-optimized label as test labels to the initial training set (when the training label is a manually labeled label, adding the test images to the training data set), that is, increasing the number of training images in the training data set to reconstruct the classification nodes. The optimization operation can also include setting limiting conditions for the shunt rules for preset nodes in the classification nodes, such as configuring a list of candidates for defect attributes for the preset nodes, thereby affecting the determination result of the defect attributes of the preset nodes, and then changing the structure of the image classification model to reconstruct the classification nodes. The above two optimization operations can be executed alternatively or both can be executed.
[0113] In the embodiments of this application, the image classification model includes multiple classification nodes. The classification nodes have shunt rules. The shunt rules include defect attributes and classification thresholds. The shunt rules determine the classification function of the classification nodes, and the shunt rules of each classification node are independent of each other with low coupling and intuitive effects. Therefore, the shunt rules can be adjusted, and by adjusting the shunt rules, the function of the classification nodes can be adjusted, and then the function of the image classification model can be adjusted. Specifically, the modification information for the to-be-modified nodes in each classification node can be obtained, and at least one of the defect attributes and classification thresholds of the to-be-modified nodes can be modified according to the modification information. The modification information can be generated through the user's setting instruction, which is used to indicate the modification object and modification content of the to-be-modified node, facilitating the user to fine-tune the intermediate parameters of the image classification model and enhancing the operator's freedom to modify the classification requirements in real time. Therefore, the model construction method in the embodiments of this application can implement a high-precision defect classification model while reducing the number of model parameters and can meet the diverse classification requirements of different users.
[0114] In the embodiments of this application, defect classification can be achieved through a defect classification device. Refer to Figure 4As shown in the figure, it is a schematic structural diagram of a defect classification device provided by an embodiment of the present application. The defect classification device includes a data acquisition module 11, a data processing module 12, a defect classification module 13, and a custom module 14. The data acquisition module 11 is used to acquire an initial training set and a validation data set. The data processing module 12 is used to cluster the initial training set to obtain a training data set. The defect classification module 13 includes an image classification model. The custom module 14 is used to adjust the shunt rule and set a limit condition for the shunt rule of a preset node in the classification node.
[0115] An embodiment of the present application provides a method for constructing an image classification model. The training data set may include multiple training images, and the training images have training labels. Based on the training data set, each classification node of the image classification model can be constructed. The classification node has a shunt rule, and the shunt rule includes a defect attribute and a classification threshold. The obtained image classification model can classify the to-be-classified image to obtain a predicted classification result. During the process of constructing each classification node, if the defect images corresponding to the target node in each classification node correspond to multiple training labels, then multiple child nodes of the target node are constructed. That is to say, the image classification model is a tree structure, and the number of classification nodes is determined according to the training process. Compared with the deep learning algorithm model, it has a lower number of parameters and thus has a faster classification speed. The target node is used to classify the defect images corresponding to the target node according to the shunt rule of the target node, and obtain multiple groups of defect images corresponding one-to-one to the multiple child nodes. In this way, the defect images can be classified sequentially by each classification node, so as to accurately and quickly determine the defect category of the defect images and improve the classification accuracy and speed.
[0116] Based on the above method for constructing an image classification model, an embodiment of the present application further provides an image classification model construction device. Refer to Figure 5 As shown in the figure, it is a structural block diagram of an image classification model construction device provided by an embodiment of the present application. The device may include:
[0117] A training data acquisition unit 110, configured to acquire a training data set, where the training data set includes multiple training images, the multiple training images have training labels, and the training images are defect images;
[0118] A model construction unit 120, configured to construct each classification node of the image classification model based on the training data set, where the classification node has a shunt rule; during the process of constructing each classification node, if the defect images corresponding to the target node in each classification node correspond to multiple training labels, then multiple child nodes of the target node are constructed; the target node is used to classify the defect images corresponding to the target node according to the shunt rule of the target node, and obtain multiple groups of defect images corresponding one-to-one to the multiple child nodes.
[0119] Optionally, the image classification model is used to classify the image to be classified to obtain a predicted classification result, and the training label is a reclassification label. The training data acquisition unit includes:
[0120] An initial training set acquisition unit for acquiring an initial training set, where the training set includes multiple training images, and the training images have initial labels;
[0121] A reclassification unit for clustering the initial training set to obtain reclassification labels of the multiple training images. There is a mapping relationship between the reclassification label and the initial label corresponding to the same training image. The image classification model is used to determine an initial classification result of the image to be classified, so as to determine the predicted classification result of the image to be classified according to the initial classification result and the mapping relationship;
[0122] During the process of clustering the initial training set, the multiple training images are successively used as target images, and multiple first feature distances between the target image and multiple homogeneous images having the same initial label as the target image are respectively determined. The heterogeneous images having different initial labels from the target image are successively used as current images, and a second feature distance between the target image and the current image is determined. If the target distance among the multiple first feature distances is greater than the second feature distance, a reclassification label is determined for the homogeneous images corresponding to the target distance.
[0123] Optionally, the apparatus further includes:
[0124] A test data set acquisition unit for acquiring a test data set, where the test data set includes multiple test images, the multiple test images have test labels, and the test images are defective images;
[0125] A test classification unit for determining a test classification result of the multiple test images through the image classification model;
[0126] A test result determination unit for determining a test result according to the test classification result and the test label;
[0127] An optimization unit for, if the test result fails the test for the to-be-optimized label in the training label, performing an optimization operation to reconstruct each classification node of the image classification model;
[0128] The optimization operation includes: adding the test images with the test label being the to-be-optimized label to the initial training set as training images, and / or setting a limiting condition for the diversion rule for a preset node in the classification node.
[0129] Optionally, the defect attribute includes an attribute corresponding to a defect feature, and the defect feature includes at least one of a position feature, a shape feature, a grayscale feature, an alignment-related feature, and a texture feature; the first feature distance is calculated based on the defect features of the target image and the defect features of the same-class image, and the second feature distance is calculated based on the defect features of the target image and the defect features of the current image.
[0130] Optionally, the shunting rule includes a defect attribute and a classification threshold.
[0131] Optionally, the shunting rule corresponding to the target node includes a target defect attribute and a target classification threshold, and the apparatus further includes:
[0132] A data determination unit, configured to sequentially use a candidate defect attribute as the current defect attribute, and determine the amplitude and total bandwidth of the current defect attribute. The amplitude of the current defect attribute is the sum of the amplitudes of the current defect attribute in the training images corresponding to multiple training labels. The amplitude of the current defect attribute in the training image corresponding to the current label among the multiple training labels is the difference between the maximum value and the minimum value of the current defect attribute in the training image corresponding to the current label. The total bandwidth of the current defect attribute is the difference between the maximum value and the minimum value of the current defect attribute in the multiple training images.
[0133] A fitness calculation unit, configured to determine the current fitness of the current defect attribute to the training data set according to the ratio of the amplitude and the total bandwidth of the current defect attribute.
[0134] A defect attribute determination unit, configured to use the candidate defect attribute whose current fitness meets the first condition as the target defect attribute.
[0135] Optionally, the apparatus further includes:
[0136] A sorting unit, configured to sort the feature values of the target defect attribute in the training image in ascending order to obtain a feature value sequence.
[0137] A candidate threshold determination unit, configured to use the average value of two adjacent feature values in the feature value sequence as the candidate threshold.
[0138] A classification unit, configured to sequentially use the candidate threshold as the current threshold, classify the defect images corresponding to the target node based on the current threshold, and obtain multiple groups of initial defect images corresponding one-to-one to the multiple sub-nodes of the target node.
[0139] A classification purity calculation unit, configured to determine the current classification purity of the current threshold according to the number of training images corresponding to each training label in the multiple groups of initial defect images.
[0140] A classification threshold determination unit, configured to use a candidate threshold whose current classification purity meets a second condition as the target classification threshold.
[0141] Optionally, the apparatus further includes:
[0142] A modification information determination unit, configured to obtain modification information of a node to be modified in each classification node;
[0143] A modification unit, configured to modify at least one of the defect attribute and the classification threshold of the node to be modified according to the modification information.
[0144] Optionally, the apparatus further includes:
[0145] A model simplification unit, configured to, if the defect attribute of the target node is the same as the defect attribute of a target child node among the multiple child nodes, increase the classification threshold of the target child node to the classification threshold of the target node, and replace the target child node with multiple child nodes of the target child node.
[0146] In another aspect, an embodiment of the present application provides a computer device. Refer to Figure 6 , which shows a structural diagram of a computer device provided by an embodiment of the present application. As Figure 6 shown, the device includes a processor 310 and a memory 320:
[0147] The memory 310 is configured to store program code and transmit the program code to the processor; [[ID=2,7]]
[0148] The processor 320 is configured to execute the image classification model construction method provided in the above embodiment according to the instructions in the program code.
[0149] This computer device may include a terminal device or a server, and the foregoing image classification model construction apparatus may be configured in this computer device.
[0150] In another aspect, an embodiment of the present application further provides a storage medium, which is used to store a computer program, and the computer program is used to execute the image classification model construction method provided in the above embodiment.
[0151] In addition, an embodiment of the present application further provides a computer program product including instructions, which, when running on a computer, causes the computer to execute the image classification model construction method provided in the above embodiment.
[0152] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by program instructions and hardware. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments. The foregoing storage medium can be at least one of the following media: read-only memory (ROM), RAM, magnetic disk, optical disk, or other media that can store program code.
[0153] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the partial description of the method embodiments.
[0154] The above is only the preferred embodiment of the present application. Although the present application has been disclosed above with preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make many possible changes and modifications to the technical solution of the present application, or modify it into an equivalent embodiment with equivalent changes, without departing from the scope of the technical solution of the present application. Therefore, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application still fall within the scope of the protection of the technical solution of the present application.
Claims
1. A method for constructing an image classification model, characterized in that: include: Acquire a training data set, where the training data set includes a plurality of training images, the plurality of training images have training labels, and the training images are defect images; Constructing classification nodes of an image classification model based on the training data set, wherein the classification nodes have a diversion rule; in the process of constructing the classification nodes, if a defect image corresponding to a target node in the classification nodes corresponds to multiple training labels, constructing multiple child nodes of the target node; The target node is used to classify the defect images corresponding to the target node according to the diversion rule of the target node, so as to obtain multiple groups of defect images corresponding to the multiple child nodes one by one.
2. The method according to claim 1, characterized in that The image classification model is used to classify the image to be classified to obtain a predicted classification result, the training label is a reclassification label, and the obtaining of the training data set includes: Obtain an initial training set, where the training set includes a plurality of training images, each of which has an initial label; Clustering the initial training set to obtain reclassification labels for the plurality of training images, wherein the reclassification labels and the initial labels corresponding to the same training image have a mapping relationship, and the image classification model is used to determine an initial classification result of the image to be classified, so as to determine a predicted classification result of the image to be classified based on the initial classification result and the mapping relationship; In the process of clustering the initial training set, the multiple training images are successively used as target images, and multiple first feature distances between the target image and multiple similar images with the same initial label as the target image are respectively determined. Different types of images with different initial labels from the target image are successively used as current images, and the second feature distance between the target image and the current image is determined. If the target distance among the multiple first feature distances is greater than the second feature distance, a reclassification label is determined for the similar image corresponding to the target distance.
3. The method according to claim 2, characterized in that The method further comprises: Acquire a test data set, where the test data set includes a plurality of test images, the plurality of test images have test labels, and the test images are defect images; Determining test classification results of the plurality of test images using the image classification model; determining a test result based on the test classification result and the test label; If the test result is that the to-be-optimized label in the training label fails the test, performing an optimization operation to reconstruct each classification node of the image classification model; The optimization operation includes: adding the test image whose test label is the label to be optimized as a training image to the initial training set, and / or setting restriction conditions of the diversion rule for the preset nodes in the classification node.
4. The method according to claim 2, characterized in that The defect features include at least one of position features, shape features, grayscale features, alignment-related features and texture features; the first feature distance is calculated based on the defect features of the target image and the defect features of the similar image, and the second feature distance is calculated based on the defect features of the target image and the defect features of the current image.
5. The method according to any one of claims 1 to 4, characterized in that The diversion rules include defect attributes and classification thresholds.
6. The method according to claim 5, characterized in that The diversion rule corresponding to the target node includes a target defect attribute and a target classification threshold, and the method further includes: Taking candidate defect attributes one by one as current defect attributes, determining the amplitude and total bandwidth of the current defect attribute, where the amplitude of the current defect attribute is the sum of multiple amplitudes of the current defect attribute in training images corresponding to multiple training labels, the amplitude of the current defect attribute in the training image corresponding to the current label among the multiple training labels is the difference between the maximum and minimum values of the current defect attribute in the training image corresponding to the current label, and the total bandwidth of the current defect attribute is the difference between the maximum and minimum values of the current defect attribute in the multiple training images; determining a current degree of adaptation of the current defect attribute to the training data set according to a ratio of the amplitude of the current defect attribute to the total bandwidth; The candidate defect attribute whose current fitness meets the first condition is used as the target defect attribute.
7. The method according to claim 6, characterized in that The method further comprises: Arranging the eigenvalues of the target defect attributes in the training image in positive order to obtain a eigenvalue sequence; The average value of two adjacent eigenvalues in the eigenvalue sequence is used as a threshold to be selected; The candidate thresholds are successively used as current thresholds, and based on the current thresholds, defect images corresponding to the target node are determined and classified to obtain multiple groups of initial defect images corresponding to multiple child nodes of the target node; determining a current classification purity of the current threshold value according to the number of training images corresponding to each training label in the plurality of groups of initial defect images; The candidate threshold whose current classification purity satisfies the second condition is used as the target classification threshold.
8. The method according to claim 5, characterized in that The method further comprises: Acquire modification information of the nodes to be modified in each classification node; At least one of the defect attribute and the classification threshold of the node to be modified is modified according to the modification information.
9. The method according to claim 5, characterized in that The method further comprises: If the defect attribute of the target node is the same as the defect attribute of a target child node among the multiple child nodes, the classification threshold of the target child node is increased to the classification threshold of the target node, and the target child node is replaced with multiple child nodes of the target child node.
10. An image classification model construction device, characterized in that: include: A training data acquisition unit is used to acquire a training data set, wherein the training data set includes a plurality of training images, the plurality of training images have training labels, and the training images are defect images; A model construction unit is configured to construct classification nodes of an image classification model based on the training data set, wherein the classification nodes have a diversion rule; in the process of constructing the classification nodes, if a defect image corresponding to a target node in the classification nodes corresponds to multiple training labels, construct multiple child nodes of the target node; The target node is used to classify the defect images corresponding to the target node according to the diversion rule of the target node, so as to obtain multiple groups of defect images corresponding to the multiple child nodes one by one.
11. A computer device, characterized in that: The computer device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the image classification model construction method described in any one of claims 1-9 according to the instructions in the program code.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, which, when executed by a processor, is used to execute the image classification model construction method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Training method and pedestrian re-identification method of multi-task classification network
CA3166088A1
Fruit classification system based on computer vision
CN110728664A
Image classification method and device, server and computer readable storage medium
CN112036514A