Classification method and device of high-resolution chip images based on multi-scale fusion
By segmenting high-resolution chip images into small images of different scales and fusing the classification results, the problem of low prediction credibility when classification of high-resolution images in the prior art is solved, and higher classification accuracy is achieved.
Patent Information
- Application Number
- CN202111284365.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-01
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-11-01
AI Technical Summary
When classifying high-resolution chip images, the prior art has problems such as high image resolution and small defect area leading to low reliability in model prediction.
By segmenting high-resolution images into small maps of different scales, the classification model is trained on small maps of different scales, the classification results of each group of small maps are obtained, and the classification results of small maps are fused to obtain the classification results of high-resolution chip images.
It significantly improves the accuracy of classifying high-resolution chip images, making the results after fusion more accurate than the prediction results of a single-scale model, and is suitable for high-precision chip image classification scenarios.
Smart Images

Figure CN114140671B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image detection technology, and in particular to a classification method and device for high-resolution chip images based on multi-scale fusion. Background Art
[0002] As electronic chips are used more and more frequently in various industries, people pay more and more attention to chip quality, so various types of tests are needed on the produced chips.
[0003] In the related art, when detecting whether a chip contains surface defects, a deep learning-based image classification method is used to classify the acquired chip surface image, and whether the chip surface has defects is determined based on the result of the image classification task. Among them, due to the powerful feature extraction capability of deep learning, the chip image classification task has been continuously developed from traditional methods to methods based on deep learning. For example, AlexNet proposed by Alex Krizhevsky et al. in 2012 deepened the network depth compared to LeNet, adopted the ReLU activation function to solve the gradient vanishing problem of the sigmoid function when the network layer is deep, used the dropout method to prevent the model from overfitting, and used a variety of data enhancement methods to improve the generalization ability of the model; GoogLeNet proposed by Christian Szegedy et al. in 2014 adopted the Inception module, used three convolution kernels of different sizes to perform convolution and maximum pooling operations on the input image, spliced the outputs of the four operations along the channel dimension, and fused information of different scales to obtain a better representation of the image; VGGNe proposed by Karen Simonyan et al. in 2014 replaced the receptive field of large convolution kernels with multiple small convolution kernels and reduced the number of parameters, and improved performance by stacking 3*3 small convolution kernels and 2*2 pooling kernels to deepen the network depth; ResNet proposed by Kaiming He et al. in 2015 used the residual structure to solve the degradation problem of deep models, and created a new model record with a 152-layer network architecture.
[0004] However, the applicant has found that when classifying high-resolution chip images using the image classification method in the related art, there are the following two main disadvantages:
[0005] First, because the chip is small in size and the defect area on the chip is even smaller, high-resolution camera equipment is needed for shooting, so the image resolution is usually high, such as 5472*3648, which is about 100:1 compared to the area ratio of the images in the datasets of existing models (such as VOC and COCO). If the image size is reduced to the level of the VOC and COCO datasets, the defective part of the chip surface will almost disappear. If the image is not reduced, the current convolutional neural network classification effect is poor and the amount of calculation is large.
[0006] Second, since the proportion of the defective part to the entire image is very small, after data analysis, the smallest defect area is only 900 pixels (length and width are only 30 pixels), accounting for 0.0045% of the entire image. The current convolutional neural network obtains the final features through multiple layers of downsampling layers, so that the defect information cannot be correctly expressed in the final features, resulting in low credibility of the model prediction.
[0007] Therefore, due to the characteristics of high resolution and small defect area of chip images, the classification model in the related art cannot be directly applied to the classification of high-resolution chip images. Summary of the invention
[0008] The present application aims to solve one of the technical problems in the related art at least to some extent.
[0009] To this end, the first purpose of the present application is to propose a classification method for high-resolution chip images based on multi-scale fusion. The method divides the high-resolution image into small images of different scales, and then trains classification models for the small images of different scales respectively to obtain the classification results of each group of small images. Finally, the classification results of the small images are fused to obtain the classification results of the high-resolution chip image. The characteristic of small-resolution images having higher classification accuracy than large-resolution images is utilized, so that the fused results are more accurate than the prediction results of a single-scale model, thereby significantly improving the accuracy of classifying high-resolution chip images, and is suitable for high-precision chip image classification scenarios.
[0010] The second objective of the present application is to propose a classification device for high-resolution chip images based on multi-scale fusion.
[0011] A third object of the present application is to provide a non-transitory computer-readable storage medium.
[0012] To achieve the above-mentioned purpose, the first embodiment of the present application is to propose a classification method for high-resolution chip images based on multi-scale fusion, the method comprising the following steps:
[0013] Dividing the high-resolution chip image to be classified multiple times according to different proportions to obtain multiple groups of images of different sizes, wherein the size of each image in any group of images is the same;
[0014] A classification model is trained for each group of images, and the classification model outputs a prediction score of each image in the corresponding image group being a positive example;
[0015] Determine the positive example credibility threshold and the negative example credibility threshold of each classification model according to the credibility of the positive examples and the negative examples corresponding to different prediction thresholds, and determine the prediction result corresponding to each group of images according to the prediction score of each image in each group of images and the positive example credibility threshold and the negative example credibility threshold of the classification model corresponding to each group of images;
[0016] Combine any two of the multiple groups of images in sequence, fuse the prediction results corresponding to each two groups of images, and determine two groups of target images according to the accuracy of each fused prediction result, and use the fused prediction results corresponding to the two groups of target images as the final classification results of the high-resolution chip image to be classified.
[0017] Optionally, in one embodiment of the present application, the high-resolution chip image to be classified is divided multiple times according to different ratios, including: dividing the high-resolution chip image to be classified into multiple groups of images of different sizes according to a preset ratio whose side length is the side length of the high-resolution chip image to be classified; taking half of two adjacent images in each group of images to form overlapping images, so that each defect is covered by at least one divided image.
[0018] Optionally, in one embodiment of the present application, the prediction result corresponding to each group of images is determined according to the prediction score of each image in each group of images, and the positive example credibility threshold and the negative example credibility threshold of the classification model corresponding to each group of images, including: comparing the prediction score of each image in each group of images with the positive example credibility threshold and the negative example credibility threshold of the corresponding classification model; if the prediction score is greater than the positive example credibility threshold, determining that the prediction result corresponding to the current group of images is defective; if the prediction score is less than the negative example credibility threshold, ignoring the current image; if the prediction score is less than or equal to the positive example credibility threshold and greater than or equal to the negative example credibility threshold, determining that the current image is an unknown image; after traversing each image in the current group of images, if there is no image with a prediction score greater than the positive example credibility threshold, calculating a first average value of the prediction scores of each of the unknown images, and comparing the first average value with a preset classification threshold to determine the prediction result of the high-resolution chip image to be classified.
[0019] Optionally, in one embodiment of the present application, the prediction results corresponding to each two groups of images are fused, including: updating the prediction score of each unknown image in the image group with a larger resolution according to the corresponding prediction result of the image group with a smaller resolution in each two groups of images; calculating a second average value of the updated prediction score of each unknown image, and comparing the second average value with the preset classification threshold to determine the fusion prediction result of the current two groups of images.
[0020] Optionally, in one embodiment of the present application, the prediction score of each unknown image in the image group with a larger resolution is updated according to the corresponding prediction result of the image group with a smaller resolution in the two groups of images, including: in the image group with a smaller resolution, obtaining multiple first images corresponding to any unknown image in the image group with a larger resolution; comparing the prediction score of each of the first images with the corresponding positive example credibility threshold, if there is any first image whose prediction score is greater than the positive example credibility threshold, determining that the prediction result of any unknown image in the image group with a larger resolution is defective, if there is no first image whose prediction score is greater than the positive example credibility threshold, calculating a third average value of the multiple first images, and merging the third average value with the prediction score of any unknown image to update the prediction score of any unknown image.
[0021] Optionally, in one embodiment of the present application, the third average value is fused with the predicted score of any unknown image, including: calculating the average value, maximum value or minimum value of the third average value and the predicted score of any unknown image.
[0022] To achieve the above purpose, the second embodiment of the present application further proposes a classification device for high-resolution chip images based on multi-scale fusion, comprising the following modules:
[0023] A division module, used for dividing the high-resolution chip image to be classified multiple times according to different proportions to obtain multiple groups of images of different sizes, wherein the size of each image in any group of images is the same;
[0024] A training module is used to train a classification model for each group of images, and output a prediction score of each image in the corresponding image group being a positive example through the classification model;
[0025] A first prediction module, used to determine a positive example credibility threshold and a negative example credibility threshold of each classification model according to the credibility of positive examples and negative examples corresponding to different prediction thresholds, and determine a prediction result corresponding to each group of images according to the prediction score of each image in each group of images and the positive example credibility threshold and the negative example credibility threshold of the classification model corresponding to each group of images;
[0026] The second prediction module is used to combine any two groups of the multiple groups of images in sequence, fuse the prediction results corresponding to each two groups of images, and determine two groups of target images according to the accuracy of each fused prediction result, and use the fused prediction results corresponding to the two groups of target images as the final classification results of the high-resolution chip image to be classified.
[0027] Optionally, in one embodiment of the present application, the division module is specifically used to: divide the high-resolution chip image to be classified into multiple groups of images of different sizes according to a preset ratio of the side length to the side length of the high-resolution chip image to be classified; take half of two adjacent images in each group of images to form overlapping images, so that each defect is covered by at least one divided image.
[0028] Optionally, in one embodiment of the present application, the first prediction module is specifically used to: compare the prediction score of each image in each group of images with the positive example credibility threshold and the negative example credibility threshold of the corresponding classification model; if the prediction score is greater than the positive example credibility threshold, determine that the prediction result corresponding to the current group of images is defective; if the prediction score is less than the negative example credibility threshold, ignore the current image; if the prediction score is less than or equal to the positive example credibility threshold and greater than or equal to the negative example credibility threshold, determine that the current image is an unknown image; after traversing each image in the current group of images, if there is no image with a prediction score greater than the positive example credibility threshold, calculate a first average value of the prediction scores of each of the unknown images, and compare the first average value with a preset classification threshold to determine the prediction result of the high-resolution chip image to be classified.
[0029] The technical solution provided by the embodiments of the present application brings at least the following beneficial effects: the solution divides the high-resolution image into small images of different scales, and then trains the classification models for the small images of different scales respectively, obtains the classification results of each group of small images, and finally fuses the classification results of the small images to obtain the classification results of the high-resolution chip image. By utilizing the higher classification accuracy of small-resolution images compared to large-resolution images, the fused results are more accurate than the prediction results of a single-scale model, thereby significantly improving the accuracy of classifying high-resolution chip images, and is suitable for high-precision chip image classification scenarios.
[0030] In order to implement the above-mentioned embodiments, the third aspect embodiment of the present application also proposes a non-temporary computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the classification method of high-resolution chip images based on multi-scale fusion in the above-mentioned embodiment is implemented.
[0031] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0033] Figure 1 A flowchart of a classification method for high-resolution chip images based on multi-scale fusion proposed in an embodiment of the present application;
[0034] Figure 2 A schematic diagram of a specific image division proposed in an embodiment of the present application;
[0035] Figure 3 A schematic diagram of a specific process of a classification method for high-resolution chip images based on multi-scale fusion proposed in an embodiment of the present application;
[0036] Figure 4 A schematic diagram of the structure of a high-resolution chip image classification device based on multi-scale fusion proposed in an embodiment of the present application. DETAILED DESCRIPTION
[0037] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0038] A classification method and device for high-resolution chip images based on multi-scale fusion proposed in an embodiment of the present invention will be described below with reference to the accompanying drawings.
[0039] Figure 1 A flowchart of the classification of high-resolution chip images based on multi-scale fusion proposed in the embodiment of the present application is shown in FIG. Figure 1 As shown, the method comprises the following steps:
[0040] Step 101 : Divide the high-resolution chip image to be classified multiple times according to different proportions to obtain multiple groups of images of different sizes, wherein the size of each image in any group of images is the same.
[0041] Specifically, the high-resolution chip image to be classified is divided multiple times according to different ratios, that is, the side length of the divided image is a preset ratio of the side length of the original high-resolution chip image, such as 1 / 2 or 1 / 4, etc. The high-resolution chip image is first divided according to one ratio to obtain a group of small images, and then the original high-resolution chip image is divided according to other ratios in turn to obtain multiple groups of small images of different scales. Among them, it is necessary to ensure that the size of each image in each group of divided images is the same, and each defect in the high-resolution chip image has at least one divided small image that can completely cover the defect. The specific division method can be set according to actual needs such as classification accuracy, and there is no restriction here.
[0042] In one embodiment of the present application, the high-resolution chip image to be classified is divided multiple times according to different ratios, including: firstly dividing the high-resolution chip image to be classified into multiple groups of images of different sizes according to a preset ratio whose side length is the side length of the high-resolution chip image to be classified, and then taking halves of two adjacent images in each group of images to form overlapping images, so that each defect is covered by at least one divided image. The two adjacent images refer to images that are adjacent to each other left and right or to each other up and down in the position of the original image.
[0043] For example, if the original image size of the high-resolution chip image to be classified is 5472*3648 pixels, the large image is divided into three scales of small images with side lengths equal to 1 / 2, 1 / 4, and 1 / 8 of the length and width of the original image, respectively recorded as three groups of small images with ratios (ratio) = 2, 4, and 8. The lengths and widths of the three groups of small images are: 2736*1824, 1368*912, and 684*456. Then, according to Figure 2 The image division processing method with a ratio of 2 shown in the figure takes half of the small images adjacent to the left and right or the upper and lower parts of the original image as overlapping small images, and obtains multiple overlapping small images, so that each defect can be fully covered by at least one small image. Therefore, after preprocessing the high-resolution chip image according to the above method, the original image can be generated respectively (2×ratio-1) 2 The number of small images corresponding to the three scales is 9, 49, and 225 respectively.
[0044] Step 102: train a classification model for each group of images, and output a prediction score of each image in the corresponding image group being a positive example through the classification model.
[0045] Among them, the trained classification model can be various convolutional neural network models in related technologies that can perform image classification tasks. For example, ResNet18 can be selected as a classification model.
[0046] Among them, the positive example refers to the classification containing defects in the image, and the corresponding negative example refers to the classification without defects in the image. The prediction score of the positive example is the probability that the currently detected image belongs to the defective category output by the trained classification model.
[0047] In the embodiment of the present application, since only whether the chip image is defective is classified, that is, the problem to be solved is a binary classification problem, it is only necessary to change the output dimension of the network fully connected layer of the classification model in the related technology to 2, which can meet the classification needs and reduce the complexity of training. Then, for the multiple groups of images of different sizes obtained in step 101, the corresponding classification model is trained for each image group. For example, continuing to refer to the above example, if three groups of small images with ratio = 2, 4, and 8 are divided, a classification model is trained for the three image groups in turn, wherein the method of training the classification model can refer to the training method of the neural network in the related technology, for example, including training through the training set in the COCO data set, which will not be repeated here. Furthermore, each image in the corresponding image group is classified by the trained classification model, and the output of the classification model is the probability that the image belongs to the defective category, that is, the prediction score of the positive example.
[0048] It should be noted that after the division, there is a correspondence between the scales of the multiple groups of images of different sizes, and the correspondence between the small images of different scales can be determined. Specifically, because there is a multiple relationship between the three selected segmentation methods, the correspondence between the original image and the small image of ratio = 2 is equivalent to the correspondence between the small images of ratio = 2 and ratio = 4, and is also equivalent to the correspondence between the small images of ratio = 4 and ratio = 8. Similarly, the correspondence between the original image and ratio = 4 is equivalent to the correspondence between ratio = 2 and ratio = 8.
[0049] Step 103, according to the credibility of positive examples and negative examples corresponding to different prediction thresholds, determine the positive example credibility threshold and the negative example credibility threshold of each classification model, and determine the prediction result corresponding to each group of images according to the prediction score of each image in each group of images, and the positive example credibility threshold and the negative example credibility threshold of the classification model corresponding to each group of images.
[0050] Specifically, for the classification model, given a prediction threshold t, the credibility of the positive and negative examples output by the model can be calculated by the formulas TP / (TP+FP) and TN / (TN+FN), respectively, where the letters T and F in TP, TN, FN and FP represent correct and incorrect, respectively, and P and N represent the prediction results of the positive and negative examples, respectively. Then TP can represent a correct positive example, that is, the prediction is a positive example and the prediction is correct. In an embodiment of the present application, after obtaining the result of the model output in step 102, the experimental results of the threshold-credibility can be analyzed to determine the positive example credibility threshold and the negative example credibility threshold of the model. For example, the output results can be verified by the verification set and the test set in the classification model data set to determine the credibility of different thresholds for positive and negative examples.
[0051] For example, taking the model with ratio = 2 as an example: when the threshold t positive =0.8, all predicted true results are positive examples, that is, if the model predicts a sample score higher than 0.8, then the sample is likely to be defective, and this application can assume that the prediction result of the model is believed; when the threshold t negative =0.4, the predicted true result has a 98.5% probability of being a negative example. Similarly, we can also trust the prediction results of the model for negative examples when the threshold is less than 0.4. Then the positive example credibility threshold and negative example credibility threshold of the model can be set to t positive =0.8, t negative =0.4. Therefore, by analyzing the threshold-credibility experimental results, we can obtain the positive example credibility thresholds t of the three models: positive and negative example confidence threshold t negative .
[0052] Furthermore, the prediction results corresponding to each group of images are determined based on the prediction scores of each image output by the classification model, as well as the positive example credibility threshold and negative example credibility threshold of the classification model corresponding to each group of images. Since a single model corresponds to a corresponding group of images, the prediction results corresponding to a group of images are the prediction results output by the corresponding single model.
[0053] In one embodiment of the present application, obtaining the prediction result of a single classification model includes: comparing the prediction score of each image in each group of images with the positive example credibility threshold and the negative example credibility threshold of the corresponding classification model; if the prediction score is greater than the positive example credibility threshold, determining that the prediction result corresponding to the current group of images is defective; if the prediction score is less than the negative example credibility threshold, ignoring the current image; if the prediction score is less than or equal to the positive example credibility threshold and greater than or equal to the negative example credibility threshold, determining that the current image is an unknown image; and then, after traversing each image in the current group of images, if there is no image with a prediction score greater than the positive example credibility threshold, calculating a first average value of the prediction scores of each unknown image, and comparing the first average value with a preset classification threshold to determine the prediction result of the high-resolution chip image to be classified.
[0054] Specifically, for a large image, we traverse all the small images after it is divided into one of the proportions. If there is a small image with a predicted score higher than t positive , we can directly judge that the high-resolution chip image is defective. If the score of the small image is lower than t negative , then ignore this small picture, for the score in t negative and t positive The small images between them are selected and their average scores are calculated as the final score of the high-resolution chip image output by the model. Finally, an appropriate classification threshold is selected to compare with the final score to determine the classification result of the large image.
[0055] Therefore, in this embodiment, a single model prediction result is output, and the prediction result can also include the prediction score of each small image.
[0056] Step 104, combine any two groups of the multiple groups of images in sequence, fuse the prediction results corresponding to each two groups of images, and determine two groups of target images according to the accuracy of each fused prediction result, and use the fused prediction results corresponding to the two groups of target images as the final classification results of the high-resolution chip image to be classified.
[0057] In an embodiment of the present application, in order to improve the accuracy of image classification, after obtaining the prediction results of a single model, the prediction results corresponding to each two groups of images are fused. The specific process includes updating the prediction score of each unknown image in the image group with a larger resolution according to the corresponding prediction result of the image group with a smaller resolution in each two groups of images, and then calculating the second average value of the updated prediction score of each unknown image, and comparing the second average value with a preset classification threshold to determine the fusion prediction result of the current two groups of images.
[0058] As one possible implementation, the prediction score of each unknown image in the image group with a larger resolution is updated according to the corresponding prediction result of the image group with a smaller resolution in the two groups of images, including obtaining multiple first images corresponding to any unknown image in the image group with a larger resolution in the image group with a smaller resolution, and comparing the prediction score of each first image with the corresponding positive example credibility threshold, wherein the prediction score of the first image and its corresponding credibility threshold can be obtained when the prediction result of the group of images with a smaller resolution is determined in step 103 in the same manner. Further, through comparison, if it is determined that there is any first image whose prediction score is greater than its corresponding positive example credibility threshold, then it is determined that the prediction result of any unknown image in the image group with a larger resolution is defective, and if there is no first image whose prediction score is greater than the positive example credibility threshold, that is, if there is no first image whose prediction score is greater than the positive example threshold, then a third average value is calculated for the multiple first images corresponding to the unknown image, and the third average value is merged with the prediction score of any unknown image to update the prediction score of any unknown image.
[0059] For example, continuing to refer to the above image division example, the three scale models with ratio = 2, 4, and 8 are combined in pairs, and there are three combinations of 2-4, 2-8, and 4-8. The image result with a smaller resolution (that is, the image size is smaller and the ratio value is larger) is used to assist in determining the uncertain result in the larger resolution image, that is, the prediction score of the unknown image.
[0060] In any combination, if the prediction score of an unknown image in the larger resolution image group output by the classification model is t1, all the small images in the smaller resolution image group corresponding to it are first calculated according to the correspondence between small images of different scales, that is, in the image group with smaller resolution, multiple first images corresponding to any unknown image in the image group with larger resolution are obtained. Then, the average score of the corresponding multiple first images is calculated as t2 (i.e., the third average value) in the manner of calculating the first average value in step 103, and then the corresponding fusion method is selected to fuse the two scores. In this example, three methods can be selected: calculating the average, maximum or minimum value of the third average value and the prediction score of any unknown image, that is, the final score after the prediction score of any unknown image is updated is t=(t1+t2) / 2 or t=max(t1,t2) or t=min(t1,t2). After the prediction scores of all unknown images in the larger image group are updated in the above manner, the average score of the updated prediction scores of all unknown images is calculated (i.e., the second average value), and the second average value is compared with the preset classification threshold to determine the fusion prediction result of the current two groups of images. For example, if the second average value is less than the classification threshold, the classification result of the high-resolution chip image is determined to be a negative example, that is, it does not contain defects. The specific setting method of the classification threshold can be calibrated through experiments according to actual needs and is not limited here.
[0061] Furthermore, through a large number of experiments, the accuracy of the output prediction results under different combinations and the fusion of the third average value and the prediction score of the unknown image is verified, and the data shown in Table 1 below is obtained. Therefore, through a large number of experiments conducted by this application, it can be concluded that the classification effect of model fusion selecting 4-8 is the best, and the fusion method with the minimum value is the best. Therefore, according to the accuracy of each fusion prediction result, the two groups of target images, i.e., the image groups with ratios of 4 and 8, are determined.
[0062] Table 1
[0063]
[0064] Furthermore, the fusion prediction results corresponding to the image groups with ratios of 4 and 8 are the final classification results of the high-resolution chip images to be classified.
[0065] Therefore, the multi-scale fusion high-resolution chip image classification method proposed in this application fuses the outputs of small images of different scales as the final classification result of the large image, and utilizes the higher classification accuracy of small-resolution images compared to large-resolution images, so that the fused results are more accurate than the prediction results of a single-scale model.
[0066] To summarize, the classification method for high-resolution chip images based on multi-scale fusion in the embodiment of the present application divides the high-resolution image into small images of different scales, and then trains the classification models for the small images of different scales respectively to obtain the classification results of each group of small images, and finally fuses the classification results of the small images to obtain the classification results of the high-resolution chip image. The characteristic of small-resolution images having higher classification accuracy than large-resolution images is utilized, so that the fused results are more accurate than the prediction results of a single-scale model, thereby significantly improving the accuracy of classifying high-resolution chip images, and is suitable for high-precision chip image classification scenarios.
[0067] In order to more clearly illustrate the classification method of high-resolution chip images based on multi-scale fusion in the embodiment of the present application, a specific embodiment is described in detail below.
[0068] First, if Figure 3 As shown in the figure, taking the model combination of ratio = 4 and ratio = 8 as an example, the specific steps include: for a large image of a high-resolution chip to be classified, traverse all its small images with ratio = 4: if there is a small image with a score higher than t positive , we can directly judge that the large image is defective; if the score of the small image is lower than t negative , then ignore this small picture; for ratio=4, the score is in t negative and t positive For the small graph between , let the score be t1, calculate which small graphs with ratio = 8 it corresponds to, and then calculate the average score t2 of these small graphs with ratio = 8 according to the single model method (i.e., the calculation method in the embodiment of obtaining the prediction result of a single classification model in step 103), then the final score of the uncertain small graph with ratio = 4 is t = min(t1, t2). After obtaining the scores of all uncertain small graphs with ratio = 4, take their average as the score of the large graph, and finally select a suitable threshold to determine the result of the large graph.
[0069] In order to implement the above embodiment, the present application also proposes a classification device for high-resolution chip images based on multi-scale fusion. Figure 4 A schematic diagram of the structure of a classification device for high-resolution chip images based on multi-scale fusion proposed in an embodiment of the present application is shown in FIG. Figure 4 As shown, the classification device for high-resolution chip images based on multi-scale fusion includes a division module 100 , a training module 200 , a first prediction module 300 and a second prediction module 400 .
[0070] The division module 100 is used to divide the high-resolution chip image to be classified multiple times according to different proportions to obtain multiple groups of images of different sizes, wherein the size of each image in any group of images is the same.
[0071] The training module 200 is used to train a classification model for each group of images, and output a prediction score of each image in the corresponding image group being a positive example through the classification model.
[0072] The first prediction module 300 is used to determine the positive example credibility threshold and the negative example credibility threshold of each classification model according to the credibility of the positive examples and negative examples corresponding to different prediction thresholds, and determine the prediction result corresponding to each group of images according to the prediction score of each image in each group of images, and the positive example credibility threshold and the negative example credibility threshold of the classification model corresponding to each group of images.
[0073] The second prediction module 400 is used to combine any two groups of the multiple groups of images in sequence, fuse the prediction results corresponding to each two groups of images, and determine two groups of target images according to the accuracy of each fused prediction result, and use the fused prediction results corresponding to the two groups of target images as the final classification results of the high-resolution chip image to be classified.
[0074] Optionally, in one embodiment of the present application, the segmentation module 100 is specifically used to: segment the high-resolution chip image to be classified into multiple groups of images of different sizes according to a preset ratio of the side length to the side length of the high-resolution chip image to be classified; take half of two adjacent images in each group of images to form overlapping images, so that each defect is covered by at least one divided image
[0075] Optionally, in one embodiment of the present application, the first prediction module 300 is further used to: compare the prediction score of each image in each group of images with the positive example credibility threshold and the negative example credibility threshold of the corresponding classification model; if the prediction score is greater than the positive example credibility threshold, determine that the prediction result corresponding to the current group of images is defective; if the prediction score is less than the negative example credibility threshold, ignore the current image; if the prediction score is less than or equal to the positive example credibility threshold and greater than or equal to the negative example credibility threshold, determine that the current image is an unknown image; after traversing each image in the current group of images, if there is no image with a prediction score greater than the positive example credibility threshold, calculate a first average value of the prediction scores of each unknown image, and compare the first average value with a preset classification threshold to determine the prediction result of the high-resolution chip image to be classified.
[0076] Optionally, in one embodiment of the present application, the second prediction module 400 is specifically used to: update the prediction score of each unknown image in the image group with a larger resolution according to the corresponding prediction result of the image group with a smaller resolution in each two groups of images; calculate the second average value of the updated prediction score of each unknown image, and compare the second average value with a preset classification threshold to determine the fusion prediction result of the current two groups of images.
[0077] Optionally, in one embodiment of the present application, the second prediction module 400 is also used to: in an image group with a smaller resolution, obtain multiple first images corresponding to any unknown image in the image group with a larger resolution; compare the prediction score of each first image with the corresponding positive example credibility threshold; if there is any first image whose prediction score is greater than the positive example credibility threshold, determine that the prediction result of any unknown image in the image group with a larger resolution is defective; if there is no first image whose prediction score is greater than the positive example credibility threshold, calculate a third average value of the multiple first images, and merge the third average value with the prediction score of any unknown image to update the prediction score of any unknown image.
[0078] Optionally, in one embodiment of the present application, the second prediction module 400 is further used to: calculate the average value, maximum value or minimum value of the third average value and the prediction score of any unknown image.
[0079] It should be noted that the above explanation of the embodiment of the classification method of high-resolution chip images based on multi-scale fusion is also applicable to the device of this embodiment, and will not be repeated here.
[0080] To summarize, the classification device for high-resolution chip images based on multi-scale fusion in the embodiment of the present application divides the high-resolution image into small images of different scales, and then trains the classification models for the small images of different scales respectively to obtain the classification results of each group of small images, and finally fuses the classification results of the small images to obtain the classification results of the high-resolution chip image. By utilizing the characteristic that the classification accuracy of small-resolution images is higher than that of large-resolution images, the fused results are more accurate than the prediction results of a single-scale model, thereby significantly improving the accuracy of classifying high-resolution chip images, and is suitable for high-precision chip image classification scenarios.
[0081] In order to implement the above embodiments, the present application also proposes a non-temporary computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the classification of high-resolution chip images based on multi-scale fusion as described in any of the above embodiments is implemented.
[0082] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0083] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0084] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0085] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0086] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0087] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0088] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0089] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A classification method for high-resolution chip images based on multi-scale fusion, characterized in that: The following steps are involved: Dividing the high-resolution chip image to be classified multiple times according to different proportions to obtain multiple groups of images of different sizes, wherein the size of each image in any group of images is the same; A classification model is trained for each group of images, and the classification model outputs a prediction score of each image in the corresponding image group being a positive example; According to the credibility of positive and negative examples corresponding to different prediction thresholds, the positive credibility threshold and negative credibility threshold of each classification model are determined, and according to the prediction score of each image in each group of images, and the positive credibility threshold and negative credibility threshold of the classification model corresponding to each group of images, the prediction result corresponding to each group of images is determined; wherein, the determination of the prediction result corresponding to each group of images includes: comparing the prediction score of each image in each group of images with the positive credibility threshold and negative credibility threshold of the corresponding classification model; if the prediction score is greater than the positive credibility threshold, determining that the prediction result corresponding to the current group of images is defective; if the prediction score is less than the negative credibility threshold, ignoring the current image; if the prediction score is less than or equal to the positive credibility threshold and greater than or equal to the negative credibility threshold, determining that the current image is an unknown image; after traversing each image in the current group of images, if there is no image with a prediction score greater than the positive credibility threshold, calculating a first average value of the prediction scores of each unknown image, and comparing the first average value with a preset classification threshold to determine the prediction result of the high-resolution chip image to be classified; Combine any two of the multiple groups of images in sequence, fuse the prediction results corresponding to each two groups of images, and determine two groups of target images according to the accuracy of each fused prediction result, and use the fused prediction results corresponding to the two groups of target images as the final classification results of the high-resolution chip images to be classified; wherein, the fusion of the prediction results corresponding to each two groups of images includes: updating the prediction score of each unknown image in the image group with a larger resolution according to the corresponding prediction result of the image group with a smaller resolution in each two groups of images; calculating the second average value of the updated prediction score of each unknown image, and comparing the second average value with the preset classification threshold to determine the fusion prediction result of the current two groups of images.
2. The classification method according to claim 1, characterized in that: The high-resolution chip image to be classified is divided multiple times according to different proportions, including: Dividing the high-resolution chip image to be classified into a plurality of groups of images of different sizes according to a preset ratio whose side length is the side length of the high-resolution chip image to be classified; In each group of images, half of two adjacent images are taken to form overlapping images, so that each defect is covered by at least one divided image.
3. The classification method according to claim 1, characterized in that: The method of updating the prediction score of each unknown image in the image group with a larger resolution according to the corresponding prediction result of the image group with a smaller resolution in the two groups of images comprises: In the image group with a smaller resolution, a plurality of first images corresponding to any unknown image in the image group with a larger resolution are obtained; The prediction score of each of the first images is compared with the corresponding positive example credibility threshold. If there is any first image whose prediction score is greater than the positive example credibility threshold, the prediction result of any unknown image in the image group with a larger resolution is determined to be defective. If there is no first image whose prediction score is greater than the positive example credibility threshold, a third average value of the multiple first images is calculated, and the third average value is merged with the prediction score of any unknown image to update the prediction score of any unknown image.
4. The classification method according to claim 3, characterized in that: The fusing the third average value with the predicted score of any unknown image comprises: Calculate the average, maximum or minimum value of the third average value and the prediction score of any unknown image.
5. A classification device for high-resolution chip images based on multi-scale fusion, characterized in that: include: A division module, used for dividing the high-resolution chip image to be classified multiple times according to different proportions to obtain multiple groups of images of different sizes, wherein the size of each image in any group of images is the same; A training module is used to train a classification model for each group of images, and output a prediction score of each image in the corresponding image group being a positive example through the classification model; A first prediction module is used to determine the positive example credibility threshold and the negative example credibility threshold of each classification model according to the credibility of the positive and negative examples corresponding to different prediction thresholds, and determine the prediction result corresponding to each group of images according to the prediction score of each image in each group of images and the positive example credibility threshold and the negative example credibility threshold of the classification model corresponding to each group of images; wherein the first prediction module is specifically used to: compare the prediction score of each image in each group of images with the positive example credibility threshold and the negative example credibility threshold of the corresponding classification model; if the prediction score is greater than the positive example credibility threshold, determine that the prediction result corresponding to the current group of images is defective; if the prediction score is less than the negative example credibility threshold, ignore the current image; if the prediction score is less than or equal to the positive example credibility threshold and greater than or equal to the negative example credibility threshold, determine that the current image is an unknown image; after traversing each image in the current group of images, if there is no image with a prediction score greater than the positive example credibility threshold, calculate the first average value of the prediction score of each unknown image, and compare the first average value with the preset classification threshold to determine the prediction result of the high-resolution chip image to be classified; The second prediction module is used to combine any two groups of the multiple groups of images in sequence, fuse the prediction results corresponding to each two groups of images, and determine two groups of target images according to the accuracy of each fused prediction result, and use the fused prediction results corresponding to the two groups of target images as the final classification results of the high-resolution chip images to be classified; wherein the second prediction module is specifically used to: update the prediction score of each unknown image in the image group with a larger resolution according to the corresponding prediction results of the image group with a smaller resolution in each two groups of images; calculate the second average value of the updated prediction score of each unknown image, and compare the second average value with the preset classification threshold to determine the fused prediction result of the current two groups of images.
6. The classification device according to claim 5, characterized in that: The division module is specifically used for: Dividing the high-resolution chip image to be classified into a plurality of groups of images of different sizes according to a preset ratio whose side length is the side length of the high-resolution chip image to be classified; In each group of images, half of two adjacent images are taken to form overlapping images, so that each defect is covered by at least one divided image.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the classification method of high-resolution chip images based on multi-scale fusion as described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Multi-scale fusion food image classification model training and image classification method
CN111222546A
Method for automatically detecting small targets in high-resolution image based on computer vision and deep learning
CN111582093A