Image classification model construction method and device
Patent Information
- Application Number
- CN202410835228.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-06-25
AI Technical Summary
然而,由于缺陷检测过程中照明模块参数配置不同、成像方式不同,以及晶圆的当前待检测工艺层不同等多重因素,导致晶圆表面缺陷图像的图像灰度、对比度、图案背景等信息变化很大,因此难以通过前述方式进行准确的缺陷分类
[0069]本申请实施例提供了一种图像分类模型构建方法及装置,训练数据集可以包括多个训练图像,训练图像具有训练标签,基于训练数据集可以构建图像分类模型的各个分类节点,分类节点具有分流规则。在构建各个分类节点的过程中,若各个分类节点中的目标节点对应的缺陷图像对应多个训练标签,则构建目标节点的多个子节点,也就是说,图像分类模型为树状结构,分类节点的数量根据训练过程确定,相比于深度学习算法模型具有较低的参数量,因此具有更快的分类速度。目标节点用于根据目标节点的分流规则,对目标节点对应的缺陷图像进行分类,得到与多个子节点一一对应的多组缺陷图像,这样可以通过各个分类节点依次对缺陷图像进行分类,从而准确快速确定缺陷图像的缺陷类别,提高分类精度和速度。
Smart Images

Figure CN120431359B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of integrated circuit manufacturing, and in particular to a method and apparatus for constructing an image classification model. Background Technology
[0002] Real-time monitoring of statistical information on wafer surface defects can enable better yield analysis and improve the utilization rate of wafer defect detection equipment. Therefore, it is necessary to determine the type of defect after wafer surface defect detection. By building an image classification algorithm model, defect type classification can be performed quickly and timely.
[0003] Currently, defect information can be obtained from wafer surface defect images through background segmentation, and then the defect type can be determined through template matching. However, due to multiple factors such as different illumination module parameter configurations, different imaging methods, and different current process layers to be inspected on the wafer during defect detection, the image grayscale, contrast, pattern background, and other information of wafer surface defect images vary greatly, making it difficult to accurately classify defects using the aforementioned methods. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide an image classification model construction method and apparatus to realize the construction of a high-precision defect classification image classification model and improve the defect classification accuracy. The specific solution is as follows:
[0005] Firstly, this application provides a method for constructing an image classification model, including:
[0006] Obtain a training dataset, which includes multiple training images with training labels, and the training images are defect images;
[0007] Based on the training dataset, each classification node of the image classification model is constructed, and the classification node has a diversion rule. During the construction of each classification node, if the defect image corresponding to the target node in each classification node corresponds to multiple training labels, then multiple child nodes of the target node are constructed. The target node is used to classify the defect image corresponding to the target node according to the diversion rule of the target node, so as to obtain multiple sets of defect images that correspond one-to-one with the multiple child nodes.
[0008] Optionally, the image classification model is used to classify the image to be classified to obtain a predicted classification result, the training label is a reclassification label, and obtaining the training dataset includes:
[0009] Obtain an initial training set, which includes multiple training images, each with an initial label;
[0010] Clustering the initial training set yields reclassification labels for the multiple training images. The reclassification labels and initial labels for the same training image have a mapping relationship. The image classification model is used to determine the initial classification result of the image to be classified, so as to determine the predicted classification result of the image to be classified based on the initial classification result and the mapping relationship.
[0011] During the clustering process of the initial training set, the multiple training images are successively used as target images, and multiple first feature distances are determined between the target image and multiple similar images with the same initial label as the target image. Different images with different initial labels as the target image are successively used as current images, and a second feature distance is determined between the target image and the current image. If the target distance among the multiple first feature distances is greater than the second feature distance, a reclassification label is determined for the similar image corresponding to the target distance.
[0012] Optionally, the method further includes:
[0013] Obtain a test dataset, which includes multiple test images, each with a test label, and the test images are defect images.
[0014] The test classification results of the multiple test images are determined using the image classification model.
[0015] The test results are determined based on the test classification results and the test labels;
[0016] If the test result is that the label to be optimized in the training label fails the test, then an optimization operation is performed to reconstruct each classification node of the image classification model;
[0017] The optimization operation includes: adding the test image with the test label to be optimized as a training image to the initial training set, and / or setting the limiting conditions of the diversion rule for the preset nodes in the classification nodes.
[0018] Optionally, the defect features include at least one of location features, shape features, grayscale features, alignment-related features, and texture features; the first feature distance is calculated based on the defect features of the target image and the defect features of similar images, and the second feature distance is calculated based on the defect features of the target image and the defect features of the current image.
[0019] Optionally, the triage rules include defect attributes and classification thresholds.
[0020] Optionally, the traffic splitting rule corresponding to the target node includes target defect attributes and target classification threshold, and the method further includes:
[0021] Candidate defect attributes are successively used as current defect attributes. The amplitude and total bandwidth of the current defect attribute are determined. The amplitude of the current defect attribute is the sum of the amplitudes of the current defect attribute in the training images corresponding to the multiple training labels. The amplitude of the current defect attribute in the training image corresponding to the current label is the difference between the maximum and minimum values of the current defect attribute in the training image corresponding to the current label. The total bandwidth of the current defect attribute is the difference between the maximum and minimum values of the current defect attribute in the multiple training images.
[0022] The current fit of the current defect attribute to the training dataset is determined based on the ratio of the magnitude of the current defect attribute to the total bandwidth.
[0023] The candidate defect attribute whose current fit satisfies the first condition is taken as the target defect attribute.
[0024] Optionally, the method further includes:
[0025] The feature values of the target defect attribute in the training image are sorted in ascending order to obtain a feature value sequence;
[0026] The average of two adjacent feature values in the feature value sequence is used as the candidate threshold.
[0027] The candidate threshold is used as the current threshold in turn. Based on the current threshold, the defect image corresponding to the target node is classified to obtain multiple sets of initial defect images that correspond one-to-one with multiple child nodes of the target node.
[0028] The current classification purity of the current threshold is determined based on the number of training images corresponding to each training label in the multiple sets of initial defect images.
[0029] The candidate threshold whose current classification purity satisfies the second condition is used as the target classification threshold.
[0030] Optionally, the method further includes:
[0031] Obtain modification information for the nodes to be modified in each of the classification nodes;
[0032] Modify at least one of the defect attributes and classification thresholds of the node to be modified based on the modification information.
[0033] Optionally, the method further includes:
[0034] If the defect attribute of the target node is the same as the defect attribute of the target child node among the plurality of child nodes, then the classification threshold of the target child node is increased to the classification threshold of the target node, and the target child node is replaced with the plurality of child nodes of the target child node.
[0035] Secondly, embodiments of this application also provide an image classification model construction apparatus, comprising:
[0036] A training data acquisition unit is used to acquire a training dataset, which includes multiple training images, each training image having a training label, and the training images being defect images.
[0037] The model building unit is used to build various classification nodes of the image classification model based on the training dataset. The classification nodes have a splitting rule. During the construction of the various classification nodes, if the defect image corresponding to the target node in each classification node corresponds to multiple training labels, then multiple child nodes of the target node are constructed. The target node is used to classify the defect image corresponding to the target node according to the splitting rule of the target node, so as to obtain multiple sets of defect images that correspond one-to-one with the multiple child nodes.
[0038] Optionally, the image classification model is used to classify the image to be classified to obtain a predicted classification result, the training label is a reclassification label, and the training data acquisition unit includes:
[0039] An initial training set acquisition unit is used to acquire an initial training set, which includes multiple training images, each training image having an initial label.
[0040] The reclassification unit is used to cluster the initial training set to obtain reclassification labels for the multiple training images. The reclassification labels and initial labels of the same training image have a mapping relationship. The image classification model is used to determine the initial classification result of the image to be classified, so as to determine the predicted classification result of the image to be classified based on the initial classification result and the mapping relationship.
[0041] During the clustering process of the initial training set, the multiple training images are successively used as target images, and multiple first feature distances are determined between the target image and multiple similar images with the same initial label as the target image. Different images with different initial labels as the target image are successively used as current images, and a second feature distance is determined between the target image and the current image. If the target distance among the multiple first feature distances is greater than the second feature distance, a reclassification label is determined for the similar image corresponding to the target distance.
[0042] Optionally, the device further includes:
[0043] A test dataset acquisition unit is used to acquire a test dataset, which includes multiple test images, each test image having a test label, and the test images being defect images.
[0044] A test classification unit is used to determine the test classification result of the plurality of test images through the image classification model;
[0045] A test result determination unit is used to determine the test result based on the test classification result and the test label;
[0046] An optimization unit is configured to perform an optimization operation to reconstruct each classification node of the image classification model if the test result is that the label to be optimized in the training labels fails the test.
[0047] The optimization operation includes: adding the test image with the test label to be optimized as a training image to the initial training set, and / or setting the limiting conditions of the diversion rule for the preset nodes in the classification nodes.
[0048] Optionally, the defect features include at least one of location features, shape features, grayscale features, alignment-related features, and texture features; the first feature distance is calculated based on the defect features of the target image and the defect features of similar images, and the second feature distance is calculated based on the defect features of the target image and the defect features of the current image.
[0049] Optionally, the triage rules include defect attributes and classification thresholds.
[0050] Optionally, the triage rule corresponding to the target node includes target defect attributes and target classification threshold, and the device further includes:
[0051] The data determination unit is used to successively use candidate defect attributes as current defect attributes, determine the amplitude and total bandwidth of the current defect attribute, wherein the amplitude of the current defect attribute is the sum of multiple amplitudes of the current defect attribute in training images corresponding to multiple training labels respectively, the amplitude of the current defect attribute in the training image corresponding to the current label in the multiple training labels is the difference between the maximum and minimum values of the current defect attribute in the training image corresponding to the current label, and the total bandwidth of the current defect attribute is the difference between the maximum and minimum values of the current defect attribute in the multiple training images;
[0052] The fitness calculation unit is used to determine the current fitness of the current defect attribute to the training dataset based on the ratio of the magnitude of the current defect attribute to the total bandwidth.
[0053] The defect attribute determination unit is used to take the candidate defect attribute whose current fit satisfies the first condition as the target defect attribute.
[0054] Optionally, the device further includes:
[0055] A sorting unit is used to sort the feature values of the target defect attribute in the training image in ascending order to obtain a feature value sequence;
[0056] The candidate threshold determination unit is used to take the average of two adjacent feature values in the feature value sequence as the candidate threshold.
[0057] The classification unit is used to successively use the candidate threshold as the current threshold, and classify the defect image corresponding to the target node based on the current threshold to obtain multiple sets of initial defect images that correspond one-to-one with multiple child nodes of the target node.
[0058] The classification purity calculation unit is used to determine the current classification purity of the current threshold based on the number of training images corresponding to each training label in the multiple sets of initial defect images.
[0059] The classification threshold determination unit is used to select the candidate threshold whose current classification purity satisfies the second condition as the target classification threshold.
[0060] Optionally, the device further includes:
[0061] The modification information determination unit is used to obtain modification information for the nodes to be modified in each of the classification nodes;
[0062] The modification unit is used to modify at least one of the defect attributes and classification thresholds of the node to be modified according to the modification information.
[0063] Optionally, the device further includes:
[0064] The model simplification unit is used to, if the defect attribute of the target node is the same as the defect attribute of the target child node among the plurality of child nodes, increase the classification threshold of the target child node to the classification threshold of the target node, and replace the target child node with the plurality of child nodes of the target child node.
[0065] Thirdly, embodiments of this application disclose a computer device, which includes a processor and a memory:
[0066] The memory is used to store program code and transmit the program code to the processor;
[0067] The processor is used to execute the image classification model construction method as described in the first aspect according to the instructions in the program code.
[0068] Fourthly, embodiments of this application disclose a computer-readable storage medium for storing a computer program, which, when executed by a processor, performs the image classification model construction method as described in the first aspect.
[0069] This application provides a method and apparatus for constructing an image classification model. The training dataset may include multiple training images with training labels. Based on the training dataset, various classification nodes of the image classification model can be constructed, and each classification node has a diversion rule. During the construction of each classification node, if the defect image corresponding to the target node in each classification node corresponds to multiple training labels, then multiple child nodes of the target node are constructed. That is, the image classification model has a tree structure, and the number of classification nodes is determined according to the training process. Compared with deep learning algorithm models, it has a lower number of parameters, thus achieving a faster classification speed. The target node is used to classify the defect image corresponding to the target node according to the diversion rule of the target node, obtaining multiple sets of defect images that correspond one-to-one with multiple child nodes. In this way, defect images can be classified sequentially through each classification node, thereby accurately and quickly determining the defect category of the defect image, improving classification accuracy and speed. Attached Figure Description
[0070] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0071] Figure 1 A flowchart illustrating an image classification model construction method provided in an embodiment of this application is shown.
[0072] Figure 2 This illustration shows a structural diagram of an image classification model provided in an embodiment of this application;
[0073] Figure 3 This paper shows a schematic diagram of the structure of another image classification model provided in an embodiment of this application;
[0074] Figure 4 This is a schematic diagram of the structure of a defect classification device provided in an embodiment of this application;
[0075] Figure 5 A structural block diagram of an image classification model construction device provided in an embodiment of this application;
[0076] Figure 6This is a structural diagram of a computer device provided in an embodiment of this application. Detailed Implementation
[0077] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the specific embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0078] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.
[0079] For ease of understanding, the following detailed description, in conjunction with the accompanying drawings, provides an image classification model construction method and apparatus provided in this application.
[0080] refer to Figure 1 The diagram shown is a flowchart illustrating an image classification model construction method provided in an embodiment of this application. The method may include the following steps.
[0081] S101, Obtain a training dataset, which includes multiple training images and has training labels.
[0082] In this embodiment of the application, a training dataset can be obtained for the construction of an image classification model. The training dataset may include multiple training images, each of which may have its own training label. The training images are defect images, and the training labels are used to indicate the defect category of the training images.
[0083] The training images are test images of the wafer, displaying defect information on the wafer surface. These images are obtained by performing defect detection tasks using a wafer surface defect detection system. Based on the training images, defect features can be determined. These features include, but are not limited to, at least one of the following: location features, shape features, grayscale features, alignment-related features, and texture features. These features can be calculated using image attribute calculation programs in a computer. Location features represent the positional information of the defect in various coordinate systems of the wafer. Shape features represent the morphological characteristics of each defect region, such as aspect ratio and roundness. Grayscale features represent the probability statistics of the defect grayscale image or the background information of its defect region pattern. Alignment-related features represent the image alignment similarity between the detection image and the comparison image. Texture features represent the noise level of the defect grayscale image or the background information of its defect region pattern.
[0084] Training labels can be manually labeled or reclassification labels. Reclassification labels are determined based on the defect features of defective images, which can better match the image classification model and improve the accuracy of the image classification model.
[0085] Reclassification labels are obtained through a reclassification operation. Specifically, this operation involves obtaining an initial training set, which includes multiple training images with initial labels. Then, the initial training set is clustered to obtain reclassification labels for these training images. A mapping relationship exists between the reclassification labels and the initial labels for the same training image. The initial labels can be manually labeled. During model application, the image classification model can determine the initial classification result of the image to be classified. This initial classification result matches the reclassification label. Based on the initial classification result and the mapping relationship, the predicted classification result of the image to be classified can be determined, and this predicted classification result matches the initial label. Thus, the use of reclassification labels in the construction and application of the image classification model allows for better matching of the labels with the model, resulting in higher accuracy. The manual labeling process and the output process match the initial labels, which are obtained manually and adapted to classification requirements. This enables image classification to adapt to different classification needs, allowing it to utilize a universally sized model to adapt to all classification tasks compared to deep learning models.
[0086] Clustering of the initial training set can be done using nearest neighbor clustering. During this clustering process, multiple training images can be successively used as target images, and multiple first feature distances are determined between each target image and multiple similar images with the same initial label. After determining the target images, dissimilar images with different initial labels are successively used as current images, and second feature distances are determined between the target image and the current images. If the target distance among the multiple first feature distances is greater than the second feature distance, it indicates that although the target image and the similar images corresponding to the target distance have been classified by the user into the same label, from the perspective of image features, the similar images corresponding to the target distance are not in the same category. Therefore, a reclassification label is determined for the similar images corresponding to the target distance, and this reclassification label will be used as a new initial label in the reclassification of other training images during the clustering process. In this way, by using all training images as target images and traversing all training images as current images, the reclassification of the initial training set is achieved, giving the training images in the initial training set reclassification labels related to image features.
[0087] The first feature distance is calculated based on the defect features of the target image and the defect features of similar images, and the second feature distance is calculated based on the defect features of the target image and the defect features of the current image. The first feature distance can be a cosine distance or an L2 distance, and the second feature distance can be a cosine distance or an L2 distance.
[0088] In addition to the training dataset, this embodiment may also include a test dataset, which includes multiple test images with test labels. The test images are defective images. The test labels can be manually labeled. If the training labels are manually labeled, the manually labeled training sample set can be proportionally divided into a training dataset and a test dataset. If the training labels are reclassification labels, and the initial labels of the training images in the initial training set are manually labeled, the manually labeled training sample set can be proportionally divided into an initial training set and a test dataset.
[0089] S102, construct each classification node of the image classification model based on the training dataset, and the classification node has a splitting rule.
[0090] In this embodiment, classification nodes of an image classification model can be constructed based on a training dataset. Each classification node has a splitting rule, which includes defect attributes and a classification threshold. The image classification model can be a tree structure. The number of parameters in the image classification model is lower than that of large-volume deep learning algorithm models, meaning it requires less computation and has a faster classification speed.
[0091] Defect attributes are used to match defect features in a defect image. These include attributes corresponding to the defect features, where the defect features are the feature values corresponding to the defect attributes. For example, defect attributes can be aspect ratio, roundness, etc., corresponding to shape features. Classification thresholds are used to define classification conditions based on defect attributes. Each classification condition corresponds to a child node of the classification node. For example, if the classification threshold is 5, the classification conditions can be the first classification condition where the feature value of the defect attribute is greater than or equal to 5, and the second classification condition where the feature value of the defect attribute is less than 5. In this case, the classification node has two child nodes corresponding to the two classification conditions. Alternatively, if the classification thresholds are 5 and 10, the classification conditions can be the first classification condition where the feature value of the defect attribute is less than 5, the second classification condition where the feature value of the defect attribute is greater than or equal to 5 and less than 10, and the third classification condition where the feature value of the defect attribute is greater than or equal to 10. In this case, the classification node has three child nodes corresponding to the three classification conditions.
[0092] Taking the target node in each classification node as an example, the classification rules of the target node can include target defect attributes and target classification threshold. Then, the target node can classify the defect image corresponding to the target node according to the target node's diversion rules, and obtain multiple sets of defect images that correspond one-to-one with the multiple child nodes of the target node. Specifically, the target node can classify the defect image whose target defect attributes meet the first classification condition based on the target classification threshold into the first child node among the multiple child nodes, and classify the defect image whose target defect attributes meet the second classification condition based on the target classification threshold into the second child node among the multiple child nodes, thus realizing classification based on target defect attributes and target classification threshold.
[0093] In constructing each classification node, if the defect image corresponding to the target node corresponds to multiple training labels, then multiple child nodes of the target node are constructed. Conversely, if the defect image corresponding to the target node corresponds to the same training label, then the target node is used as the terminal node, and the training label is used as the classification result of the target node. In other words, for the structure of the image classification model, classification nodes can be constructed one by one from top to bottom, and the classification node splitting rules can be determined.
[0094] After constructing the target node, the target defect attributes of the target node can be determined. The target defect attributes can be determined based on the fit of all defect attributes to the training dataset to ensure that the target node has a better inter-class separation.
[0095] Specifically, candidate defect attributes can be successively used as current defect attributes. The amplitude and total bandwidth of the current defect attribute are determined. Based on the ratio of the amplitude of the current defect attribute to the total bandwidth, the current fit of the current defect attribute to the training dataset is determined. Candidate defect attributes whose current fit satisfies a first condition are used as the target defect attributes. The current fit of the current defect attribute to the training dataset = amplitude of the current defect attribute / total bandwidth. For example, a candidate defect attribute whose current fit satisfies the first condition can be a candidate defect attribute with the highest current fit among all current fit attributes.
[0096] The amplitude of the current defect attribute is the sum of multiple amplitudes of the current defect attribute in the training images corresponding to multiple training labels. Specifically, the amplitude of the current defect attribute in the training image corresponding to the same training label represents one type of defect attribute amplitude, while the amplitudes corresponding to multiple training labels represent multiple types of defect attribute amplitudes, and the amplitude of the current defect attribute is the sum of the amplitudes of these multiple types of defect attributes. Furthermore, the amplitude of the current defect attribute in the training image corresponding to the current label is the difference between the maximum and minimum values of the current defect attribute in that training image. The total bandwidth of the current defect attribute is the difference between the maximum and minimum values of the current defect attribute in the multiple training images.
[0097] After determining the target defect attributes of the target node, the target classification threshold of the target node can be determined. The target classification threshold can be determined by the classification purity of the candidate threshold to ensure a better inter-class separation of the target node.
[0098] Specifically, the feature values of the target defect attribute in the training images can be sorted in ascending order to obtain a feature value sequence; the average of two adjacent feature values in the feature value sequence is used as a candidate threshold; the candidate threshold is successively used as the current threshold, and the defect image corresponding to the target node is classified based on the current threshold to obtain multiple sets of initial defect images that correspond one-to-one with multiple child nodes of the target node; the current classification purity of the current threshold is determined according to the number of training images corresponding to each training label in the multiple sets of initial defect images; the candidate threshold whose current classification purity satisfies the second condition is used as the target classification threshold. The candidate threshold whose current classification purity satisfies the second condition can be the candidate threshold with the highest current classification purity among all current classification purities.
[0099] Taking a case where there are two child nodes, the calculation process for the current classification purity can be represented as follows:
[0100] purity=Corr([b1,b2,...,b i ],[c1,c2,...,c i ])
[0101] Among them, b i c represents the number of training images in the first group of defect images that correspond to the i-th training label. i Let be the number of training images in the second set of defect images from multiple initial defect images that correspond to the i-th training label, and Corr be the formula for calculating matrix correlation. All b i and all c i The sum is the total number N of defect images corresponding to the target node.
[0102] After determining the target classification threshold for the target node, the target node construction is complete. Next, child nodes of the target node can be constructed. The number of child nodes constructed equals the number of defect images obtained after classification by the target node, ensuring a one-to-one correspondence between the multiple defect images and the multiple child nodes of the target node. The construction process of child nodes can refer to the construction process of the target node. If the defect image corresponding to a child node of the target node corresponds to multiple training labels, then a child node of the child node of the target node can be constructed. Conversely, if the defect image corresponding to a child node of the target node corresponds to the same training label, then the child node of the target node can be used as the terminal node, and the training label can be used as the classification result of that child node. Based on the aforementioned principles for constructing classification nodes, classification nodes can be constructed one by one from top to bottom, and the classification rules for classification nodes can be determined, until all terminal nodes are constructed. If all terminal nodes correspond to defect images containing only one type of training label, then the image classification model is complete.
[0103] refer to Figure 2 The diagram shown is a structural schematic of an image classification model provided in an embodiment of this application. The image classification model may include 9 classification nodes, from node 1 to node 9. The defect attribute of node 1 is the defect shape, which is used to distinguish between square and circular defects. Defect images with square defects are classified into the first group, corresponding to node 2. Defect images with circular defects are classified into the second group, corresponding to node 3. Nodes 2 and 3 are child nodes of node 1.
[0104] The defect attribute of node 2 is defect size, and the classification threshold is 5. It is used to distinguish the size of defects. In the first group of defect images, defect images with a defect size of less than 5nm are classified into the third group, corresponding to node 4. In the first group of defect images, defect images with a defect size of greater than or equal to 5nm are classified into the fourth group, corresponding to node 5. Nodes 4 and 5 are child nodes of node 2.
[0105] The defect attribute of node 3 is defect size, and the classification threshold is 5. It is used to distinguish the size of defects. In the second group of defect images, defect images with a defect size of less than 5nm are classified into the fifth group, corresponding to node 6. Defect images with a defect size of greater than or equal to 5nm in the second group of defect images are classified into the sixth group, corresponding to node 7. Nodes 6 and 7 are child nodes of node 3.
[0106] The defect attribute of node 7 is defect size, and the classification threshold is 10. It is used to distinguish the size of defects. In the sixth group of defect images, defect images with a defect size of less than 10nm are classified into the seventh group, corresponding to node 8. In the sixth group of defect images, defect images with a defect size of greater than or equal to 10nm are classified into the eighth group, corresponding to node 9. Nodes 8 and 9 are child nodes of node 7.
[0107] In this embodiment, if the training label is a manually labeled label, the image classification model classifies the image to be classified and obtains a classification result of one of the training labels, thus realizing the classification of the image to be classified; if the training label is a reclassification label, the image classification model classifies the image to be classified and obtains a classification result of one of the reclassification labels, and maps the reclassification label to one of the initial labels according to the mapping relationship, thus realizing the classification of the image to be classified.
[0108] Since the reclassification labels are determined based on the defect features of the defect images, the image classification model is constructed based on the mechanism of the difference range between defect feature values between defect images, thus achieving high classification accuracy. The classification accuracy of the image classification model formed by this method has been verified in practical application scenarios, reaching over 98%. This image classification model can be applied to the field of integrated circuit manufacturing yield monitoring, which has received relatively little research in China. It realizes a method for intelligently generating wafer surface defect image classification models according to application scenarios. This method is applicable to all wafer surface defect detection devices based on photochemical detection principles, such as bright-field deep ultraviolet light source defect detection systems.
[0109] In this embodiment of the application, after constructing the classification nodes of the image classification model, the image classification model can be simplified. Specifically, if the defect attribute of the target node is the same as the defect attribute of the target child node among the multiple child nodes of the target node, the classification threshold of the target child node can be increased to the classification threshold of the target node, the target child node can be replaced with multiple child nodes of the target child node, the target child node can be deleted, the overall number of classification nodes is reduced, the number of child nodes of the target node is increased, the number of node layers is reduced, and the image classification model is simplified.
[0110] refer to Figure 2 As shown, the defect attributes of nodes 3 and 7 are both defect shape, used to distinguish defect size. With node 3 as the target node and node 7 as the target child node, the classification threshold of node 7 can be added to the classification threshold of node 3, resulting in a new classification threshold of 10. Replacing node 7 with its child nodes 8 and 9 simplifies the image classification model. (Reference) Figure 3 The diagram shown is a structural schematic of another image classification model provided in this application embodiment. The defect attribute of node 3 is defect shape, the classification threshold is 5 and 10, and the child nodes are node 6, node 8 and node 9. Thus, node 7 is deleted, the overall number of classification nodes is reduced, the number of child nodes of node 3 is increased, the number of node layers is reduced, and the image classification model is simplified.
[0111] In this embodiment, the image classification model can also be validated. Specifically, a test dataset is obtained, and the test classification results of multiple test images are determined using the image classification model. The test result is determined based on the test classification results and test labels. The test result may include the classification accuracy corresponding to each training label. If the classification accuracy corresponding to each training label is higher than a preset value, it indicates that the image classification model has passed the test. If the test result is that the label to be optimized in the training labels fails the test, an optimization operation can be performed to reconstruct the various classification nodes of the image classification model, i.e., S102 is re-executed.
[0112] Optimization operations can include adding test images with the desired label to the initial training set (if the training label is manually labeled, add the test images to the training dataset), i.e., increasing the number of training images in the training dataset to rebuild the classification nodes. Optimization operations can also include setting restrictions on the splitting rules for preset nodes in the classification system, such as configuring a candidate list of defect attributes for preset nodes, thereby affecting the determination of defect attributes for preset nodes and thus changing the structure of the image classification model to rebuild the classification nodes. Either of these two optimization operations can be performed, or both can be performed.
[0113] In this embodiment, the image classification model includes multiple classification nodes. Each classification node has a diversion rule, which includes defect attributes and classification thresholds. The diversion rule determines the classification function of the classification node, and the diversion rules of each classification node are independent of each other, with low coupling and intuitive effect. Therefore, the diversion rules can be adjusted, and by adjusting the diversion rules, the function of the classification node can be adjusted, thereby adjusting the function of the image classification model. Specifically, modification information for the node to be modified in each classification node can be obtained, and at least one of the defect attributes and classification thresholds of the node to be modified can be modified according to the modification information. The modification information can be generated by user setting instructions to indicate the modification object and modification content of the node to be modified, which facilitates the user to fine-tune the intermediate parameters of the image classification model and improves the operator's freedom to modify the classification requirements in real time. Therefore, the model construction method in this embodiment can achieve a high-precision defect classification model while reducing the number of model parameters, and can meet the diverse classification needs of different users.
[0114] In this embodiment of the application, defect classification can be achieved through a defect classification device, as shown in the reference. Figure 4The diagram shown is a structural schematic of a defect classification device provided in an embodiment of this application. The defect classification device includes a data acquisition module 11, a data processing module 12, a defect classification module 13, and a custom module 14. The data acquisition module 11 is used to acquire an initial training set and a validation dataset. The data processing module 12 is used to cluster the initial training set to obtain a training dataset. The defect classification module 13 includes an image classification model. The custom module 14 is used to adjust the diversion rules and set the diversion rule constraints for preset nodes in the classification nodes.
[0115] This application provides a method for constructing an image classification model. The training dataset can include multiple training images with training labels. Based on the training dataset, various classification nodes of the image classification model can be constructed. Each classification node has a splitting rule, including defect attributes and classification thresholds. The resulting image classification model can classify the images to be classified and obtain predicted classification results. During the construction of each classification node, if the defect image corresponding to the target node in each classification node corresponds to multiple training labels, then multiple child nodes of the target node are constructed. That is, the image classification model has a tree structure, and the number of classification nodes is determined according to the training process. Compared with deep learning algorithm models, it has a lower number of parameters, thus achieving faster classification speed. The target node is used to classify the defect image corresponding to the target node according to the splitting rule of the target node, obtaining multiple sets of defect images that correspond one-to-one with multiple child nodes. This allows for the sequential classification of defect images through each classification node, thereby accurately and quickly determining the defect category of the defect image, improving classification accuracy and speed.
[0116] Based on the above image classification model construction method, this application embodiment also provides an image classification model construction device, referencing... Figure 5 The diagram shown is a structural block diagram of an image classification model construction device provided in an embodiment of this application. The device may include:
[0117] The training data acquisition unit 110 is used to acquire a training dataset, which includes multiple training images, each training image having a training label, and the training images being defect images.
[0118] The model building unit 120 is used to build various classification nodes of the image classification model based on the training dataset. The classification nodes have diversion rules. In the process of building the various classification nodes, if the defect image corresponding to the target node in each classification node corresponds to multiple training labels, then multiple child nodes of the target node are built. The target node is used to classify the defect image corresponding to the target node according to the diversion rules of the target node, so as to obtain multiple sets of defect images that correspond one-to-one with the multiple child nodes.
[0119] Optionally, the image classification model is used to classify the image to be classified to obtain a predicted classification result, the training label is a reclassification label, and the training data acquisition unit includes:
[0120] An initial training set acquisition unit is used to acquire an initial training set, which includes multiple training images, each training image having an initial label.
[0121] The reclassification unit is used to cluster the initial training set to obtain reclassification labels for the multiple training images. The reclassification labels and initial labels of the same training image have a mapping relationship. The image classification model is used to determine the initial classification result of the image to be classified, so as to determine the predicted classification result of the image to be classified based on the initial classification result and the mapping relationship.
[0122] During the clustering process of the initial training set, the multiple training images are successively used as target images, and multiple first feature distances are determined between the target image and multiple similar images with the same initial label as the target image. Different images with different initial labels as the target image are successively used as current images, and a second feature distance is determined between the target image and the current image. If the target distance among the multiple first feature distances is greater than the second feature distance, a reclassification label is determined for the similar image corresponding to the target distance.
[0123] Optionally, the device further includes:
[0124] A test dataset acquisition unit is used to acquire a test dataset, which includes multiple test images, each test image having a test label, and the test images being defect images.
[0125] A test classification unit is used to determine the test classification result of the plurality of test images through the image classification model;
[0126] A test result determination unit is used to determine the test result based on the test classification result and the test label;
[0127] An optimization unit is configured to perform an optimization operation to reconstruct each classification node of the image classification model if the test result is that the label to be optimized in the training labels fails the test.
[0128] The optimization operation includes: adding the test image with the test label to be optimized as a training image to the initial training set, and / or setting the limiting conditions of the diversion rule for the preset nodes in the classification nodes.
[0129] Optionally, the defect attributes include attributes corresponding to defect features, and the defect features include at least one of position features, shape features, grayscale features, alignment-related features, and texture features; the first feature distance is calculated based on the defect features of the target image and the defect features of similar images, and the second feature distance is calculated based on the defect features of the target image and the defect features of the current image.
[0130] Optionally, the triage rules include defect attributes and classification thresholds.
[0131] Optionally, the triage rule corresponding to the target node includes target defect attributes and target classification threshold, and the device further includes:
[0132] The data determination unit is used to successively use candidate defect attributes as current defect attributes, determine the amplitude and total bandwidth of the current defect attribute, wherein the amplitude of the current defect attribute is the sum of multiple amplitudes of the current defect attribute in training images corresponding to multiple training labels respectively, the amplitude of the current defect attribute in the training image corresponding to the current label in the multiple training labels is the difference between the maximum and minimum values of the current defect attribute in the training image corresponding to the current label, and the total bandwidth of the current defect attribute is the difference between the maximum and minimum values of the current defect attribute in the multiple training images;
[0133] The fitness calculation unit is used to determine the current fitness of the current defect attribute to the training dataset based on the ratio of the magnitude of the current defect attribute to the total bandwidth.
[0134] The defect attribute determination unit is used to take the candidate defect attribute whose current fit satisfies the first condition as the target defect attribute.
[0135] Optionally, the device further includes:
[0136] A sorting unit is used to sort the feature values of the target defect attribute in the training image in ascending order to obtain a feature value sequence;
[0137] The candidate threshold determination unit is used to take the average of two adjacent feature values in the feature value sequence as the candidate threshold.
[0138] The classification unit is used to successively use the candidate threshold as the current threshold, and classify the defect image corresponding to the target node based on the current threshold to obtain multiple sets of initial defect images that correspond one-to-one with multiple child nodes of the target node.
[0139] The classification purity calculation unit is used to determine the current classification purity of the current threshold based on the number of training images corresponding to each training label in the multiple sets of initial defect images.
[0140] The classification threshold determination unit is used to select the candidate threshold whose current classification purity satisfies the second condition as the target classification threshold.
[0141] Optionally, the device further includes:
[0142] The modification information determination unit is used to obtain modification information for the nodes to be modified in each of the classification nodes;
[0143] The modification unit is used to modify at least one of the defect attributes and classification thresholds of the node to be modified according to the modification information.
[0144] Optionally, the device further includes:
[0145] The model simplification unit is used to, if the defect attribute of the target node is the same as the defect attribute of the target child node among the plurality of child nodes, increase the classification threshold of the target child node to the classification threshold of the target node, and replace the target child node with the plurality of child nodes of the target child node.
[0146] On another front, embodiments of this application provide a computer device, see [link to relevant documentation]. Figure 6 The figure illustrates a structural diagram of a computer device provided in an embodiment of this application, such as... Figure 6 As shown, the device includes a processor 310 and a memory 320:
[0147] The memory 310 is used to store program code and transmit the program code to the processor;
[0148] The processor 320 is used to execute the image classification model construction method provided in the above embodiments according to the instructions in the program code.
[0149] The computer device may include a terminal device or a server, and the aforementioned image classification model building apparatus may be configured in the computer device.
[0150] In another aspect, embodiments of this application also provide a storage medium for storing a computer program for executing the image classification model construction method provided in the above embodiments.
[0151] In addition, this application also provides a computer program product including instructions, which, when run on a computer, causes the computer to execute the image classification model construction method provided in the above embodiments.
[0152] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by program instructions in hardware. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium can be at least one of the following media: read-only memory (ROM), RAM, magnetic disk, or optical disk, etc., and other media capable of storing program code.
[0153] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0154] The above description is merely a preferred embodiment of this application. Although this application has disclosed preferred embodiments above, it is not intended to limit this application. Any person skilled in the art can make many possible variations and modifications to the technical solutions of this application using the methods and techniques disclosed above, or modify them into equivalent embodiments with equivalent changes, without departing from the scope of the technical solutions of this application. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of this application without departing from the content of the technical solutions of this application shall still fall within the protection scope of the technical solutions of this application.
Claims
1. A method for constructing an image classification model, characterized in that, include: Obtain a training dataset, which includes multiple training images with training labels, and the training images are defect images; Based on the training dataset, each classification node of the image classification model is constructed, and the classification node has a diversion rule; in the process of constructing each classification node, if the defect image corresponding to the target node in each classification node corresponds to multiple training labels, then multiple child nodes of the target node are constructed. The target node is used to classify the defect images corresponding to the target node according to the target node's diversion rules, so as to obtain multiple sets of defect images that correspond one-to-one with the multiple child nodes; The image classification model is used to classify the image to be classified and obtain a predicted classification result. The training label is a reclassification label. Obtaining the training dataset includes: Obtain an initial training set, which includes multiple training images, each with an initial label; Clustering the initial training set yields reclassification labels for the multiple training images. The reclassification labels and initial labels for the same training image have a mapping relationship. The image classification model is used to determine the initial classification result of the image to be classified, so as to determine the predicted classification result of the image to be classified based on the initial classification result and the mapping relationship. During the clustering process of the initial training set, the multiple training images are successively used as target images, and multiple first feature distances are determined between the target image and multiple similar images with the same initial label as the target image. Different images with different initial labels as the target image are successively used as current images, and a second feature distance is determined between the target image and the current image. If the target distance among the multiple first feature distances is greater than the second feature distance, a reclassification label is determined for the similar image corresponding to the target distance. The traffic splitting rule corresponding to the target node includes the target defect attribute and the target classification threshold, and the method further includes: Candidate defect attributes are successively used as current defect attributes. The amplitude and total bandwidth of the current defect attribute are determined. The amplitude of the current defect attribute is the sum of the amplitudes of the current defect attribute in the training images corresponding to the multiple training labels. The amplitude of the current defect attribute in the training image corresponding to the current label is the difference between the maximum and minimum values of the current defect attribute in the training image corresponding to the current label. The total bandwidth of the current defect attribute is the difference between the maximum and minimum values of the current defect attribute in the multiple training images. The current fit of the current defect attribute to the training dataset is determined based on the ratio of the magnitude of the current defect attribute to the total bandwidth. The candidate defect attribute whose current fit satisfies the first condition is taken as the target defect attribute; The feature values of the target defect attribute in the training image are sorted in ascending order to obtain a feature value sequence; The average of two adjacent feature values in the feature value sequence is used as the candidate threshold. The candidate threshold is used as the current threshold in turn. Based on the current threshold, the defect image corresponding to the target node is classified to obtain multiple sets of initial defect images that correspond one-to-one with multiple child nodes of the target node. The current classification purity of the current threshold is determined based on the number of training images corresponding to each training label in the multiple sets of initial defect images. The candidate threshold whose current classification purity satisfies the second condition is used as the target classification threshold.
2. The method according to claim 1, characterized in that, The method further includes: Obtain a test dataset, which includes multiple test images, each with a test label, and the test images are defect images. The test classification results of the multiple test images are determined using the image classification model. The test results are determined based on the test classification results and the test labels; If the test result is that the label to be optimized in the training label fails the test, then an optimization operation is performed to reconstruct each classification node of the image classification model; The optimization operation includes: adding the test image with the test label to be optimized as a training image to the initial training set, and / or setting the limiting conditions of the diversion rule for the preset nodes in the classification nodes.
3. The method according to claim 1, characterized in that, The defect features include at least one of location features, shape features, grayscale features, alignment-related features, and texture features; the first feature distance is calculated based on the defect features of the target image and the defect features of similar images, and the second feature distance is calculated based on the defect features of the target image and the defect features of the current image.
4. The method according to claim 1, characterized in that, The method further includes: Obtain modification information for the nodes to be modified in each of the classification nodes; Modify at least one of the defect attributes and classification thresholds of the node to be modified based on the modification information.
5. The method according to claim 1, characterized in that, The method further includes: If the defect attribute of the target node is the same as the defect attribute of the target child node among the plurality of child nodes, then the classification threshold of the target child node is increased to the classification threshold of the target node, and the target child node is replaced with the plurality of child nodes of the target child node.
6. An image classification model construction device, characterized in that, include: A training data acquisition unit is used to acquire a training dataset, which includes multiple training images, each training image having a training label, and the training images being defect images. The model building unit is used to build various classification nodes of the image classification model based on the training dataset. The classification nodes have a splitting rule. In the process of building the various classification nodes, if the defect image corresponding to the target node in each classification node corresponds to multiple training labels, then multiple child nodes of the target node are built. The target node is used to classify the defect images corresponding to the target node according to the target node's diversion rules, so as to obtain multiple sets of defect images that correspond one-to-one with the multiple child nodes; The image classification model is used to classify the image to be classified and obtain a predicted classification result. The training label is a reclassification label. The training data acquisition unit includes: An initial training set acquisition unit is used to acquire an initial training set, which includes multiple training images, each training image having an initial label. The reclassification unit is used to cluster the initial training set to obtain reclassification labels for the multiple training images. The reclassification labels and initial labels of the same training image have a mapping relationship. The image classification model is used to determine the initial classification result of the image to be classified, so as to determine the predicted classification result of the image to be classified based on the initial classification result and the mapping relationship. During the clustering process of the initial training set, the multiple training images are successively used as target images, and multiple first feature distances are determined between the target image and multiple similar images with the same initial label as the target image. Different images with different initial labels as the target image are successively used as current images, and a second feature distance is determined between the target image and the current image. If the target distance among the multiple first feature distances is greater than the second feature distance, a reclassification label is determined for the similar image corresponding to the target distance. The traffic splitting rules corresponding to the target node include target defect attributes and target classification thresholds, and the device further includes: The data determination unit is used to successively use candidate defect attributes as current defect attributes, determine the amplitude and total bandwidth of the current defect attribute, wherein the amplitude of the current defect attribute is the sum of multiple amplitudes of the current defect attribute in training images corresponding to multiple training labels respectively, the amplitude of the current defect attribute in the training image corresponding to the current label in the multiple training labels is the difference between the maximum and minimum values of the current defect attribute in the training image corresponding to the current label, and the total bandwidth of the current defect attribute is the difference between the maximum and minimum values of the current defect attribute in the multiple training images; The fitness calculation unit is used to determine the current fitness of the current defect attribute to the training dataset based on the ratio of the magnitude of the current defect attribute to the total bandwidth. A defect attribute determination unit is used to take the candidate defect attribute whose current fit satisfies the first condition as the target defect attribute; A sorting unit is used to sort the feature values of the target defect attribute in the training image in ascending order to obtain a feature value sequence; The candidate threshold determination unit is used to take the average of two adjacent feature values in the feature value sequence as the candidate threshold. The classification unit is used to successively use the candidate threshold as the current threshold, and classify the defect image corresponding to the target node based on the current threshold to obtain multiple sets of initial defect images that correspond one-to-one with multiple child nodes of the target node. The classification purity calculation unit is used to determine the current classification purity of the current threshold based on the number of training images corresponding to each training label in the multiple sets of initial defect images. The classification threshold determination unit is used to select the candidate threshold whose current classification purity satisfies the second condition as the target classification threshold.
7. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the image classification model construction method according to any one of claims 1-5 according to the instructions in the program code.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which, when executed by a processor, is used to perform the image classification model construction method according to any one of claims 1-5.
Citation Information
Patent Citations
Training method and pedestrian re-identification method of multi-task classification network
CA3166088A1
Fruit classification system based on computer vision
CN110728664A