A method for industrial image defect detection based on attraction-repulsion contrast learning
By adopting a method based on attractive and exclusion contrast learning in industrial defect detection, synthesized anomaly images are generated and a teacher-student network model is constructed, and the problems of uncertain identification ability and insufficient discrimination in the existing technology are solved, and more accurate anomaly detection and positioning effects are achieved.
Patent Information
- Application Number
- CN202410928145.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-11
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-07-11
AI Technical Summary
The existing unsupervised anomaly detection methods have problems with uncertain identification capabilities and insufficient discrimination in industrial defect detection, especially in the absence of sufficient anomaly samples.
Using a method based on attraction and exclusion contrast learning, a synthetic anomaly image is generated by data augmentation of normal image data, and a teacher-student network model is constructed for training, and the model's discrimination ability is enhanced by using attraction and exclusion processing.
It improves the robustness and generalization performance of the model, can more accurately identify and locate abnormal areas, and enhances the effect of industrial defect detection.
Smart Images

Figure CN118941504B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial defect detection, and in particular to an industrial image defect detection method based on attraction-repulsion contrast learning. Background Art
[0002] Image anomaly detection plays an important role in product quality inspection, medical image analysis, intelligent security systems and other fields. The key to this task lies in two points: whether the anomaly can be determined and whether the anomaly can be accurately located. In recent years, with the development of deep learning, computer vision and other fields, more and more models have been put into practical application. However, training models requires a large amount of sample data, including not only normal samples, but also abnormal samples and corresponding labels. Since abnormal samples are often difficult to obtain, in other words, the number and types of abnormal samples are far from enough for model training, and abnormalities are usually unknown, the cost of obtaining abnormal samples is not only high, but also it is difficult to create unknown anomalies. For supervised learning methods, it is impractical to obtain enough training data with human annotations in order to accurately locate anomalies, especially weak anomalies. Because obtaining the true label annotation of abnormal data is a very time-consuming and laborious task, and requires the participation of people with expert knowledge in this field to accurately annotate, this work is indeed a big challenge. Therefore, the method of unsupervised anomaly detection and positioning using only normal samples has attracted widespread attention from researchers.
[0003] Most unsupervised anomaly detection methods focus on normal samples and only use normal samples for training. Due to the high generalization of convolutional neural networks, the model's ability to recognize abnormal samples is uncertain. Under this paradigm, the model only learns normal features and lacks sufficient understanding of abnormal features, resulting in insufficient discriminability. If synthetic abnormal samples can be introduced to participate in model training, this problem can be effectively alleviated.
[0004] Existing unsupervised anomaly detection methods often use models pre-trained on the ImageNet dataset for multi-level feature extraction. Unsupervised anomaly detection methods based on knowledge distillation contain two networks, one is the teacher network and the other is the student network. The teacher network uses the model pre-trained on ImageNet for multi-level feature extraction, while the student network is generally a network with the same structure as the teacher network but not pre-trained. Students try to learn from the teacher as much as possible and gradually acquire detection capabilities through learning. However, this often overlooks an important issue, that is, in the process of learning from the teacher, students make the normal sample features generated in the two networks as close as possible, and the abnormal sample features generated will also be as close as possible, so that the students' identification ability is limited. Therefore, the research on industrial defect detection is still in its infancy, and the corresponding basic theory and method framework are still lacking. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide an industrial image defect detection method based on attraction and repulsion contrast learning to achieve industrial defect detection in view of the deficiencies of the above-mentioned prior art.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: an industrial image defect detection method based on attraction and repulsion contrast learning, comprising the following steps:
[0007] Step 1: Obtain normal image data from existing industrial image datasets;
[0008] The existing industrial image dataset contains a total of M image datasets of different categories, of which m are object categories and n are texture categories, m+n=M; each category dataset is divided into a training set, a test set and a label dataset, wherein the training set contains only normal images, and the test set contains both normal images and real abnormal images; the normal image data obtained from the existing image dataset refers to the normal image data I obtained from the training set;
[0009] Step 2: Obtain a real abnormal image dataset TA and a label dataset TAGT corresponding to the real abnormal image data;
[0010] Step 2.1, obtaining a real abnormal image dataset TA from an existing image dataset;
[0011] First, find the test set of each category of data set, then remove the normal image part from the test set, and finally the set of remaining real abnormal images is the real abnormal image data set TA;
[0012] Step 2.2, obtain the label dataset TAGT corresponding to the real abnormal image data from the existing image dataset;
[0013] The specific method of obtaining the label data set TAGT corresponding to the real abnormal image data from the existing image data set is: first find the label data set in each category data set, and then collect all the defect class data sets in the label data set to form the label data set TAGT;
[0014] Step 3: Based on the real abnormal image dataset TA and its corresponding label dataset TAGT, the acquired normal image data I is enhanced to obtain the synthetic abnormal image I. a ;
[0015] Step 3.1, performing background detection on the acquired normal image data I, and generating an initial mask at the same time;
[0016] The background detection of the acquired normal image data I refers to the background detection of the image data of m object categories, while the background detection of the image data of n texture categories does not need to be performed;
[0017] The specific method for performing background detection on the acquired normal image data I is as follows: firstly, the image color is adjusted to a certain extent by adaptive histogram equalization and global histogram equalization, then Gaussian filtering is performed to remove some noise and then thresholding is performed, and opening, closing, corrosion and dilation operations are performed to obtain a black and white image, and the object part is set to black;
[0018] The generating an initial mask refers to generating a completely black image with the same size as the image I as the initial mask corresponding to the image I;
[0019] Step 3.2, randomly select a point (x, y) in the selected area of the normal image I;
[0020] For the object class, the selected area refers to the black area in the black and white image obtained in step 3.1 in image I; for the texture class, it refers to any position in image I;
[0021] Step 3.3, randomly select a point from the abnormal value in the real abnormal image, multiply it by the transparency factor α and assign it to the point (x, y) in image I;
[0022] Randomly select a real abnormal image from the real abnormal image dataset TA obtained in step 2.1. According to the corresponding label data in the label dataset TAGT, randomly select a point from the abnormal value position in the real abnormal image and multiply it by the transparency factor α and assign it to the point (x, y) in image I. At the same time, assign 255 to the (x, y) point in the initial mask corresponding to image I in step 3.1.
[0023] Step 3.4: Perform restricted random sampling on image I and continue to assign values until the number of abnormal points is met;
[0024] The restricted random sampling refers to randomly selecting a new point (x', y') from the four positions above, below, left and right of the point (x, y) in the image I, and then assigning a value to the point according to step 3.3, and assigning a value of 255 to the point (x', y') in the initial mask, and so on, until the abnormal point number requirement is met. All the selected points can constitute a defective block. At this time, the points assigned a value of 255 in the initial mask are the abnormal pixel positions corresponding to the defective block in the image I, and the remaining positions are normal pixel positions;
[0025] Step 3.5, obtaining a synthesized abnormal image;
[0026] Repeat steps 3.2 to 3.4 multiple times to obtain multiple defect blocks and multiple initial masks; synthesize the multiple defect blocks, and the final image obtained is the synthesized abnormal image I a By performing an OR operation on a plurality of initial masks, a final mask image corresponding to the synthesized abnormal image can be obtained;
[0027] Step 4: Build an industrial image defect detection model and use the synthetic abnormal image obtained in step 3 to train the model;
[0028] Step 4.1: The synthesized abnormal image I a Send it to the teacher network to obtain the feature f after multi-level feature fusion t , and normalize it into features
[0029] Step 4.1.1: The synthesized abnormal image I a Send to the teacher network to extract multi-level feature parameters;
[0030] The teacher network refers to the main network composed of a Resnet18 network pre-trained on ImageNet, in which the parameters remain unchanged during training and testing;
[0031] The multi-level feature parameters extracted include the first, second and third layer features output by the Resnet18 network;
[0032] Step 4.1.2, fuse the multi-level feature parameters extracted by the teacher network;
[0033] The specific method of fusing the extracted multi-level feature parameters is as follows: convolution and pooling operations are performed on the features of the first and second layers of the Resnet18 network output respectively to align them with the features of the third layer; the third layer features are then subjected to the block space attention mechanism to extract the information that needs to be focused on, and finally the element-by-element addition operation of the three-level features is performed to obtain the feature f after the fusion of the multi-level features. t ;
[0034] The specific method of extracting the information that needs to be focused on from the third-layer features through the block convolution attention mechanism is as follows: performing r*r block adaptive maximum pooling and block adaptive average pooling operations on the third-layer features, and then performing channel average pooling and channel maximum pooling respectively, and then adding the results of the two poolings, performing a convolution operation with a kernel size of 7*7 on the result of the addition, and performing interpolation and activation operations to obtain the importance weight coefficient, and then adding the weight coefficient to the third-layer features to obtain weighted third-layer feature parameters, and finally sending the weighted third-layer feature parameters to the channel attention module to further enhance the feature representation of different channels;
[0035] Step 4.1.3: The feature f is obtained by fusing the multi-level features extracted by the teacher network t Normalize to obtain normalized feature parameters
[0036] The feature f is obtained by fusing the multi-level features extracted by the teacher network t The specific method for normalization is:
[0037]
[0038] Among them, f t (i,j) represents feature f t The feature at (i,j) in the graph, w is the feature f t Width;
[0039] Step 4.2: The synthesized abnormal image I a Send it to the student network to obtain the feature f after multi-level feature fusion s ';
[0040] The student network refers to the main network composed of a Resnet18 network with adjustable parameters, which uses the same method as the teacher network to obtain the feature f after multi-level feature fusion s ';
[0041] Step 4.3: The feature f is obtained by fusing the multi-level features extracted by the student network in step 4.2. s 'Perform de-differentiation processing to obtain feature f s , and normalize it into features
[0042] The feature f is obtained by fusing the multi-level features extracted by the student network in step 4.2. s 'Perform de-differentiation processing to obtain feature f s The specific method is: first in the feature f s 'Adding Gaussian noise will diversify the anomalies, and then send it to the UNet network for de-anomaly processing to obtain the repaired feature f s ;
[0043] The feature f s The specific method for normalization is:
[0044]
[0045] Among them, f s (i,j) represents feature f s The feature at (i,j) in the graph, w′ is the feature f s Width;
[0046] Step 4.4: Normalize the features and After bilinear interpolation is performed to adjust to the input image size, attraction and repulsion processing is performed to make the difference between normal pixels smaller and the difference between abnormal pixels larger;
[0047] The attraction and repulsion processing refers to minimizing the loss function of normal pixels and abnormal pixels so that the distance between normal pixels is reduced and the distance between abnormal pixels is increased;
[0048] Step 4.4.1. Calculate the characteristic parameters after bilinear interpolation and Feature difference map between ;
[0049] The characteristic parameters after calculating bilinear interpolation are and The feature difference map between refers to the feature parameters calculated after normalization and bilinear interpolation and The second norm between :
[0050]
[0051] Among them, fd is the characteristic parameter after bilinear interpolation and Feature difference map between ;
[0052] Step 4.4.2, obtain the corresponding normal pixel points and abnormal pixel points in the feature difference map according to the final mask image, and calculate the loss function of the normal pixel points and the abnormal pixel points;
[0053] The normal pixel points are normal pixel points in the entire synthetic abnormal image, and the abnormal pixel points are abnormal pixel points in the entire synthetic abnormal image;
[0054] The calculation formula of the loss function of normal pixels and abnormal pixels is as follows:
[0055]
[0056] Among them, Loss n , Loss a are the losses of normal pixels and abnormal pixels respectively, N n is the number of normal pixels, N a is the number of abnormal pixels, μ,d std are the mean and variance of the feature difference map, τ1, τ2 are adjustable parameters, and fd i and fd j is the component in the feature difference map;
[0057] Step 4.4.3, calculate the total loss function, and minimize the total loss function through the Adam optimizer to train the industrial image defect detection model constructed from step 4.1 to step 4.4;
[0058] The total loss function calculation formula is as follows:
[0059] Loss=Loss n +Loss a
[0060] Among them, Loss is the total loss;
[0061] Step 5: During the test phase, keep the parameters of the industrial image defect detection model unchanged, and locate the defect position by calculating the difference in output features between the teacher network and the student network.
[0062] The beneficial effects of the above technical scheme are as follows: the present invention provides an industrial image defect detection method based on attraction and repulsion contrast learning, (1) a new scheme for generating anomalies is proposed, which can better and more accurately segment the object parts, generate synthetic abnormal images with various shapes and sizes that are closer to real anomalies, and can also produce some unknown anomalies, so that the network can see more abnormal samples and improve the robustness of the model; (2) a teacher-student network model with symmetric structure and asymmetric function is proposed, and the pre-trained network parameters are fine-tuned outside the network to improve the generalization performance of the model. Gaussian noise and UNet network are also added to the student network at the feature level, so that the student network can use more contextual information to repair the damaged area and restore the damaged area to normal pixel values; (3) an attraction and repulsion contrast learning method is proposed, which can reduce the feature distance of normal pixels while expanding the feature distance of abnormal pixels, so that the model has the ability to better identify and locate abnormal areas. (4) According to the theory of local invariance of images, a block convolution attention mechanism is proposed, which improves the traditional attention mechanism and further improves the model's ability to capture long-distance dependencies and semantic relationships. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 A flowchart of an industrial image defect detection method based on attraction-repulsion contrast learning provided by an embodiment of the present invention;
[0064] Figure 2 This is a diagram showing the effect of defect detection on data in 15 types of data sets in MVTec AD provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0066] This embodiment takes the MVTec AD dataset as an example and adopts the industrial image defect detection method based on attraction-repulsion contrast learning of the present invention to realize defect detection.
[0067] In this embodiment, an industrial image defect detection method based on attraction and repulsion contrast learning is provided. Figure 1 As shown, the following steps are included:
[0068] Step 1: Obtain normal image data from existing industrial image datasets;
[0069] The existing industrial image dataset contains a total of M image datasets of different categories, of which m are object categories and n are texture categories, m+n=M; each category dataset is divided into a training set, a test set and a label dataset, wherein the training set contains only normal images, and the test set contains both normal images and real abnormal images; the normal image data obtained from the existing image dataset refers to the normal image data I obtained from the training set;
[0070] The MVTec AD dataset is a dataset proposed at the 2019 CVPR conference. It is specifically used to evaluate and test the performance of unsupervised image abnormal pattern detection and segmentation. It contains 15 different categories of image datasets, including 10 object categories (Objects) and 5 texture categories (Textures), with a total of 5354 high-resolution images; each category of dataset is divided into training set, test set and label dataset, where the training set only contains normal images, and the test set contains both normal images and real abnormal images; normal image data obtained from the MVTec AD dataset refers to normal image data I obtained from the training set;
[0071] The 10 object categories of the MVTec AD dataset are: Bottle, Cable, Capsule, Hazelnut, Screw, Toothbrush, Pill, Zipper, Transistor, Metal Nut; the 5 texture categories are: Carpet, Grid, Leather, Tile, and Wood.
[0072] Step 2: Obtain a real abnormal image dataset TA and a label dataset TAGT corresponding to the real abnormal image data;
[0073] Step 2.1, obtaining a real abnormal image dataset TA from an existing image dataset;
[0074] First, find the test set of each category of data set, then remove the normal image part from the test set, and finally the set of remaining real abnormal images is the real abnormal image data set TA;
[0075] Step 2.2, obtain the label dataset TAGT corresponding to the real abnormal image data from the existing image dataset;
[0076] The specific method of obtaining the label data set TAGT corresponding to the real abnormal image data from the existing image data set is: first find the label data set in each category data set, and then collect all the defect class data sets in the label data set to form the label data set TAGT;
[0077] Step 3: Based on the real abnormal image dataset TA and its corresponding label dataset TAGT, the acquired normal image data I is enhanced to obtain the synthetic abnormal image I. a ;
[0078] Step 3.1, performing background detection on the acquired normal image data I, and generating an initial mask at the same time;
[0079] The background detection of the acquired normal image data I refers to the background detection of the image data of m object categories, while the background detection of the image data of n texture categories does not need to be performed;
[0080] The specific method for performing background detection on the acquired normal image data I is as follows: firstly, the image color is adjusted to a certain extent by adaptive histogram equalization and global histogram equalization, then Gaussian filtering is performed to remove some noise and then thresholding is performed, and opening, closing, corrosion and dilation operations are performed to obtain a black and white image, and the object part is set to black;
[0081] The generating an initial mask refers to generating a completely black image with the same size as the image I as the initial mask corresponding to the image I;
[0082] Step 3.2, randomly select a point (x, y) in the selected area of the normal image I;
[0083] For the object class, the selected area refers to the black area in the black and white image obtained in step 3.1 in image I; for the texture class, it refers to any position in image I;
[0084] Step 3.3, randomly select a point from the abnormal value in the real abnormal image, multiply it by the transparency factor α and assign it to the point (x, y) in image I;
[0085] Randomly select a real abnormal image from the real abnormal image dataset TA obtained in step 2.1. According to the corresponding label data in the label dataset TAGT, randomly select a point from the abnormal value position in the real abnormal image and multiply it by the transparency factor α and assign it to the point (x, y) in image I. At the same time, assign 255 to the (x, y) point in the initial mask corresponding to image I in step 3.1.
[0086] In this embodiment, the position of the abnormal value in the real abnormal image is the corresponding position in the abnormal image corresponding to the position with a pixel value of 255 in the corresponding label; the value of the transparency factor α is 0.7;
[0087] Step 3.4: Perform restricted random sampling on image I and continue to assign values until the number of abnormal points is met;
[0088] The restricted random sampling refers to randomly selecting a new point (x', y') from the four positions above, below, left and right of the point (x, y) in the image I, and then assigning a value to the point according to step 3.3, and assigning a value of 255 to the point (x', y') in the initial mask, and so on, until the abnormal point number requirement is met. All the selected points can constitute a defective block. At this time, the points assigned a value of 255 in the initial mask are the abnormal pixel positions corresponding to the defective block in the image I, and the remaining positions are normal pixel positions;
[0089] In this embodiment, the number of abnormal points is required to be 255;
[0090] Step 3.5, obtaining a synthesized abnormal image;
[0091] Repeat steps 3.2 to 3.4 multiple times to obtain multiple defect blocks and multiple initial masks; synthesize the multiple defect blocks, and the final image obtained is the synthesized abnormal image I a By performing an OR operation on a plurality of initial masks, a final mask image corresponding to the synthesized abnormal image can be obtained;
[0092] Step 4: Build an industrial image defect detection model and use the synthetic abnormal image obtained in step 3 to train the model;
[0093] Step 4.1: The synthesized abnormal image I a Send it to the teacher network to obtain the feature f after multi-level feature fusion t , and normalize it into features
[0094] Step 4.1.1: The synthesized abnormal image I a Send to the teacher network to extract multi-level feature parameters;
[0095] The teacher network refers to the main network composed of a Resnet18 network pre-trained on ImageNet, in which the parameters remain unchanged during training and testing;
[0096] The multi-level feature parameters extracted include the first, second and third layer features output by the Resnet18 network;
[0097] Step 4.1.2, fuse the multi-level feature parameters extracted by the teacher network;
[0098] The specific method of fusing the extracted multi-level feature parameters is as follows: convolution and pooling operations are performed on the features of the first and second layers of the Resnet18 network output respectively to align them with the features of the third layer; the third layer features are then subjected to the block space attention mechanism to extract the information that needs to be focused on, and finally the element-by-element addition operation of the three-level features is performed to obtain the feature f after the fusion of the multi-level features. t ;
[0099] In this embodiment, the number of channels of the first, second and third layer feature parameters output by the Resnet18 network are 64, 128 and 256 respectively, and the feature map sizes are 64*64, 32*32 and 16*16 respectively;
[0100] The specific method of extracting the information that needs to be focused on from the third-layer features through the block convolution attention mechanism is as follows: performing r*r block adaptive maximum pooling and block adaptive average pooling operations on the third-layer features, and then performing channel average pooling and channel maximum pooling respectively, and then adding the results of the two poolings, performing a convolution operation with a kernel size of 7*7 on the result of the addition, and performing interpolation and activation operations to obtain the importance weight coefficient, and then adding the weight coefficient to the third-layer features to obtain weighted third-layer feature parameters, and finally sending the weighted third-layer feature parameters to the channel attention module to further enhance the feature representation of different channels;
[0101] Step 4.1.3: The feature f is obtained by fusing the multi-level features extracted by the teacher network t Normalize to obtain normalized feature parameters
[0102] The feature f is obtained by fusing the multi-level features extracted by the teacher network t The specific method for normalization is:
[0103]
[0104] Among them, f t (i,j) represents feature f t The feature at (i,j) in the graph, w is the feature f t width; in this embodiment, w is 16;
[0105] Step 4.2: The synthesized abnormal image I a Send it to the student network to obtain the feature f after multi-level feature fusion s ';
[0106] The student network refers to the main network composed of a Resnet18 network with adjustable parameters, which uses the same method as the teacher network to obtain the feature f after multi-level feature fusion s ';
[0107] Step 4.3: The feature f is obtained by fusing the multi-level features extracted by the student network in step 4.2. s 'Perform de-differentiation processing to obtain feature f s , and normalize it into features
[0108] The feature f is obtained by fusing the multi-level features extracted by the student network in step 4.2. s 'Perform de-differentiation processing to obtain feature f s The specific method is: first in the feature f s 'Adding Gaussian noise will diversify the anomalies, and then send it to the UNet network for de-anomaly processing to obtain the repaired feature f s ;
[0109] The feature f s The specific method for normalization is:
[0110]
[0111] Among them, f s (i,j) represents feature f s The feature at (i,j) in the graph, w′ is the feature f s width; in this embodiment, w′ is 16;
[0112] Step 4.4: Normalize the features and After bilinear interpolation is performed to adjust to the input image size, attraction and repulsion processing is performed to make the difference between normal pixels smaller and the difference between abnormal pixels larger;
[0113] The attraction and repulsion processing refers to minimizing the loss function of normal pixels and abnormal pixels so that the distance between normal pixels is reduced and the distance between abnormal pixels is increased;
[0114] Step 4.4.1. Calculate the characteristic parameters after bilinear interpolation and Feature difference map between ;
[0115] The characteristic parameters after calculating bilinear interpolation are and The feature difference map between refers to the feature parameters calculated after normalization and bilinear interpolation and The second norm between :
[0116]
[0117] Among them, fd is the characteristic parameter after bilinear interpolation and Feature difference map between ;
[0118] Step 4.4.2, obtain the corresponding normal pixel points and abnormal pixel points in the feature difference map according to the final mask image, and calculate the loss function of the normal pixel points and the abnormal pixel points;
[0119] The normal pixel points are normal pixel points in the entire synthetic abnormal image, and the abnormal pixel points are abnormal pixel points in the entire synthetic abnormal image;
[0120] The calculation basis of the loss function of normal pixels and abnormal pixels is: the more attention is paid to normal pixels that exceed the mean of the feature difference map and the abnormal pixels that are lower than the mean of the feature difference map, the greater the corresponding weight; and the other weights are set to 0 and no attention is given, because these pixels themselves meet the network's requirements for errors and will not affect the model's discrimination ability. This processing also reduces the computational burden of the model.
[0121] The calculation formula of the loss function of normal pixels and abnormal pixels is as follows:
[0122]
[0123] Among them, Loss n , Loss a are the losses of normal pixels and abnormal pixels respectively, N n is the number of normal pixels, N a is the number of abnormal pixels, μ,d std are the mean and variance of the feature difference map, τ1, τ2 are adjustable parameters, and fd i and fd j is the component in the feature difference graph; in this embodiment, τ1=1, τ2=2;
[0124] Step 4.4.3, calculate the total loss function, and minimize the total loss function through the Adam optimizer to train the industrial image defect detection model constructed from step 4.1 to step 4.4;
[0125] The total loss function calculation formula is as follows:
[0126] Loss=Loss n +Loss a
[0127] Among them, Loss is the total loss;
[0128] Step 5: During the test phase, keep the parameters of the industrial image defect detection model unchanged, and locate the defect position by calculating the difference in output features between the teacher network and the student network.
[0129] In this embodiment, the method of the present invention is also compared with the existing methods on the MVTec AD dataset. Tables 1 and 2 show the image-level and pixel-level AUC-ROC scores of different methods, and Table 3 shows the pixel-level PRO scores of different methods. It can be seen from Table 1 that compared with other methods in the table, the method of the present invention has the highest scores in 9 categories and the highest average value. The detection effect in the texture category is more obvious, of which three categories reach 100%, and the remaining two categories are also higher than 99%. From Tables 2 and 3, it can be seen that the method of the present invention is leading in pixel-level anomaly detection and positioning effects in most categories, and the average value is also leading. Compared with ST-m, in 15 categories, the method of the present invention improves the image-level average AUCROC by 4.72% (Table 1), the pixel-level average AUCROC by 1.16% (Table 2), and the PRO metric by 2.78% (Table 3), which proves that the method of the present invention is superior to most existing methods. The defect detection effect of the method of the present invention is as follows: Figure 2 As shown, the first line is the sample with abnormality, the second line is the detection effect of the method of the present invention, and the third line is the label corresponding to the abnormal image. It can be seen that no matter whether it is a large defect or a small defect, whether it is a line or a dot, whether it is a semantic information error or a logical error, the method of the present invention can effectively detect these errors, which fully demonstrates the effectiveness of model detection.
[0130] Table 1. Image-level AUC-ROC anomaly detection performance comparison on the MVTec AD dataset
[0131]
[0132]
[0133] Table 2. Comparison of pixel-level AUC-ROC anomaly detection performance on the MVTec AD dataset
[0134]
[0135] Table 3. Comparison of pixel-level AUC-PRO anomaly detection performance on the MVTec AD dataset
[0136]
[0137]
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. An industrial image defect detection method based on attraction-repulsion contrast learning, characterized in that: The following steps are involved: Step 1: Obtain normal image data I from an existing industrial image dataset; Step 2: Obtain a real abnormal image dataset TA and a label dataset TAGT corresponding to the real abnormal image data; Step 3: Based on the real abnormal image dataset TA and its corresponding label dataset TAGT, the acquired normal image data I is enhanced to obtain the synthetic abnormal image I. a ; Step 4: Build an industrial image defect detection model and use the synthetic abnormal image obtained in step 3 to train the model; Step 4.1: The synthesized abnormal image I a Send it to the teacher network to obtain the feature f after multi-level feature fusion t , and normalize it into features Step 4.2: The synthesized abnormal image I a Send it to the student network to obtain the feature f after multi-level feature fusion s '; Step 4.3: Remove the feature fs' obtained by merging the multi-level features extracted by the student network in step 4.2 and obtain the feature f s , and normalize it into features Step 4.4: Normalize the features and After bilinear interpolation is performed to adjust to the input image size, attraction and repulsion processing is performed to make the difference between normal pixels smaller and the difference between abnormal pixels larger; Step 4.4.
1. Calculate the characteristic parameters after bilinear interpolation and Feature difference map between ; The characteristic parameters after calculating bilinear interpolation are and The feature difference map between refers to the feature parameters calculated after normalization and bilinear interpolation and The second norm between : Among them, fd is the characteristic parameter after bilinear interpolation and Feature difference map between ; Step 4.4.2, obtain the corresponding normal pixel points and abnormal pixel points in the feature difference map according to the final mask image, and calculate the loss function of the normal pixel points and the abnormal pixel points; The normal pixel points are normal pixel points in the entire synthetic abnormal image, and the abnormal pixel points are abnormal pixel points in the entire synthetic abnormal image; The calculation formula of the loss function of normal pixels and abnormal pixels is as follows: Among them, Loss n , Loss a are the losses of normal pixels and abnormal pixels respectively, N n is the number of normal pixels, N a is the number of abnormal pixels, μ,d std are the mean and variance of the feature difference map, τ1, τ2 are adjustable parameters, and fd i and fd j is the component in the feature difference map; Step 4.4.3, calculate the total loss function, and minimize the total loss function through the Adam optimizer to train the industrial image defect detection model constructed from step 4.1 to step 4.4; The total loss function calculation formula is as follows: Loss=Loss n +Loss a Among them, Loss is the total loss; Step 5: During the test phase, keep the parameters of the industrial image defect detection model unchanged, and locate the defect position by calculating the difference in output features between the teacher network and the student network.
2. The industrial image defect detection method based on attraction-repulsion contrast learning according to claim 1 is characterized in that: The existing industrial image dataset described in step 1 contains a total of M image datasets of different categories, of which m are object categories and n are texture categories, m+n=M; each category dataset is divided into a training set, a test set and a label dataset, wherein the training set only contains normal images, and the test set contains both normal images and real abnormal images; the normal image data obtained from the existing image dataset refers to the normal image data I obtained from the training set.
3. The industrial image defect detection method based on attraction-repulsion contrast learning according to claim 2 is characterized in that: The specific method of step 2 is: Step 2.1, obtaining a real abnormal image dataset TA from an existing image dataset; First, find the test set of each category of data set, then remove the normal image part from the test set, and finally the set of remaining real abnormal images is the real abnormal image data set TA; Step 2.2, obtain the label dataset TAGT corresponding to the real abnormal image data from the existing image dataset; The specific method of obtaining the label data set TAGT corresponding to the real abnormal image data from the existing image data set is: first find the label data set in each category data set, and then collect all the defect class data sets in the label data set to form the label data set TAGT.
4. The industrial image defect detection method based on attraction-repulsion contrast learning according to claim 3 is characterized in that: The specific method of step 3 is: Step 3.1, performing background detection on the acquired normal image data I, and generating an initial mask at the same time; The background detection of the acquired normal image data I refers to the background detection of the image data of m object categories, while the background detection of the image data of n texture categories does not need to be performed; The specific method for performing background detection on the acquired normal image data I is as follows: firstly, the image color is adjusted to a certain extent by adaptive histogram equalization and global histogram equalization, then Gaussian filtering is performed to remove some noise and then thresholding is performed, and opening, closing, corrosion and dilation operations are performed to obtain a black and white image, and the object part is set to black; The generating an initial mask refers to generating a completely black image with the same size as the image I as the initial mask corresponding to the image I; Step 3.2, randomly select a point (x, y) in the selected area of the normal image I; For the object class, the selected area refers to the black area in the black and white image obtained in step 3.1 in image I; for the texture class, it refers to any position in image I; Step 3.3, randomly select a point from the abnormal value in the real abnormal image, multiply it by the transparency factor α and assign it to the point (x, y) in image I; Randomly select a real abnormal image from the real abnormal image dataset TA obtained in step 2.
1. According to the corresponding label data in the label dataset TAGT, randomly select a point from the abnormal value position in the real abnormal image and multiply it by the transparency factor α and assign it to the point (x, y) in image I. At the same time, assign 255 to the (x, y) point in the initial mask corresponding to image I in step 3.
1. Step 3.4: Perform restricted random sampling on image I and continue to assign values until the number of abnormal points is met; The restricted random sampling refers to randomly selecting a new point (x', y') from the four positions above, below, left and right of the point (x, y) in the image I, and then assigning a value to the point according to step 3.3, and assigning a value of 255 to the point (x', y') in the initial mask, and so on, until the abnormal point number requirement is met. All the selected points can constitute a defective block. At this time, the points assigned a value of 255 in the initial mask are the abnormal pixel positions corresponding to the defective block in the image I, and the remaining positions are normal pixel positions; Step 3.5, obtaining a synthesized abnormal image; Repeat steps 3.2 to 3.4 multiple times to obtain multiple defect blocks and multiple initial masks; synthesize the multiple defect blocks, and the final image obtained is the synthesized abnormal image I a By performing an OR operation on a plurality of initial masks, a final mask image corresponding to the synthesized abnormal image can be obtained.
5. The industrial image defect detection method based on attraction-repulsion contrast learning according to claim 4 is characterized in that: The specific method of step 4.1 is: Step 4.1.1: The synthesized abnormal image I a Send to the teacher network to extract multi-level feature parameters; The teacher network refers to the main network composed of a Resnet18 network pre-trained on ImageNet, in which the parameters remain unchanged during training and testing; The multi-level feature parameters extracted include the first, second and third layer features output by the Resnet18 network; Step 4.1.2, fuse the multi-level feature parameters extracted by the teacher network; The specific method of fusing the extracted multi-level feature parameters is as follows: convolution and pooling operations are performed on the features of the first and second layers of the Resnet18 network output respectively to align them with the features of the third layer; the third layer features are then subjected to the block space attention mechanism to extract the information that needs to be focused on, and finally the element-by-element addition operation of the three-level features is performed to obtain the feature f after the fusion of the multi-level features. t ; The specific method of extracting the information that needs to be focused on from the third-layer features through the block convolution attention mechanism is as follows: performing r*r block adaptive maximum pooling and block adaptive average pooling operations on the third-layer features, and then performing channel average pooling and channel maximum pooling respectively, and then adding the results of the two poolings, performing a convolution operation with a kernel size of 7*7 on the result of the addition, and performing interpolation and activation operations to obtain the importance weight coefficient, and then adding the weight coefficient to the third-layer features to obtain weighted third-layer feature parameters, and finally sending the weighted third-layer feature parameters to the channel attention module to further enhance the feature representation of different channels; Step 4.1.3: The feature f is obtained by fusing the multi-level features extracted by the teacher network t Normalize to obtain normalized feature parameters The feature f is obtained by fusing the multi-level features extracted by the teacher network t The specific method for normalization is: Among them, f t (i,j) represents feature f t The feature at (i,j) in the graph, w is the feature f t Width.
6. The industrial image defect detection method based on attraction-repulsion contrast learning according to claim 5 is characterized in that: The student network described in step 4.2 refers to the main network composed of a Resnet18 network with adjustable parameters, which uses the same method as the teacher network to obtain the feature f after multi-level feature fusion s '.
7. The industrial image defect detection method based on attraction-repulsion contrast learning according to claim 6 is characterized in that: The feature f is obtained by fusing the multi-level features extracted by the student network in step 4.
2. s 'Perform de-differentiation processing to obtain feature f s The specific method is: first in the feature f s 'Adding Gaussian noise will diversify the anomalies, and then send it to the UNet network for de-anomaly processing to obtain the repaired feature f s ; The feature f s The specific method for normalization is: Among them, f s (i,j) represents feature f s The feature at (i,j) in the graph, w′ is the feature f s Width.