An Industrial Product Defect Detection Method Based on Distillation Contrastive Learning
Through the method based on distillation comparison learning, an industrial product defect detection model is constructed, which solves the problems of insufficient data coverage and low detection accuracy in the existing technology, and achieves efficient and accurate defect detection and long-term data processing capabilities.
Patent Information
- Application Number
- CN202411940124.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-12-26
AI Technical Summary
The prior art is difficult to collect comprehensive labeled data sets covering all possible defect forms, cannot accurately detect abnormal data, and has problems of forgetting and drifting over long-term unupdated data.
The industrial product defect detection method based on distillation comparison learning is adopted, and the product images are collected for preprocessing, and normal images and pseudo-abnormal images are obtained. These images are trained using the teacher network to obtain the normal teacher output features and abnormal teacher output features. The student network is optimized through distillation comparison learning, and the product defect detection model is finally constructed.
It realizes the ability to cover a wide range of defect forms, high detection accuracy, and effectively process long-term unupdated data, improving the efficiency and accuracy of product defect detection.
Smart Images

Figure CN119540223B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of product defect detection, and particularly relates to an industrial product defect detection method based on distilled contrastive learning. Background Art
[0002] Detecting surface defects of industrial products is an important field in the development of industrial intelligence. Mainly in the processes of industrial production and quality control, etc., surface defects of products are detected, including scratches, damages, and color contaminations, etc., to meet the requirements of ensuring product quality and production safety.
[0003] The probability of generating abnormal products is low and the abnormal types are diverse. In traditional manual inspection, due to problems such as subjectivity and fatigue of manual inspection, the accuracy and efficiency of inspection cannot be guaranteed. While using a defect detection model can effectively solve these problems, improve the accuracy and efficiency of defect detection, and at the same time reduce costs and manual burdens. Therefore, defect detection of industrial products is very important in digital manufacturing.
[0004] Currently, there are two methods to solve defect detection. The first is the supervised method, that is, using defect images with labels (including categories, rectangular boxes or pixel-by-pixel, etc.) to input into the network for training; the second is the unsupervised method, which is to learn normal defect-free samples, learn the features of normal regions, and the network detects abnormal regions. However, since defects in industrial manufacturing are caused by uncontrollable factors in the production process, the forms of defects are diverse. Therefore, it is difficult to collect a comprehensive labeled dataset covering all possible defect forms.
[0005] Currently, the existing technologies usually use unsupervised methods in defect detection, including reconstruction-based methods and feature embedding similarity. For reconstruction-based methods, only normal images are used to train a neural network for image reconstruction, and abnormal images are detected as defect images because they cannot be well reconstructed, and the abnormal score is represented by the reconstruction error. The most common reconstruction-based methods are autoencoders and generative adversarial networks. However, this method requires a large amount of normal data as training samples. If the distribution of abnormal data is similar to that of normal data, the model may not be able to accurately detect abnormalities. Although reconstruction-based methods are very intuitive and interpretable, they can sometimes also produce good reconstruction results for abnormal images.
[0006] For feature embedding similarity, the deep neural network extracts and learns meaningful vectors for the entire image, and the anomaly score is represented by the distance between the embedding vector of the test image and the reference vector representing normality in the training dataset. Feature embedding similarity methods include memory banks, knowledge distillation, etc. The memory bank is a method that has been used more frequently in recent years. It adopts the idea of establishing a feature library, extracts the features of normal images and stores them in the library, and compares the features of the test samples with those in the feature library to perform anomaly detection. However, the anomaly detection method based on the memory bank needs to continuously update the normal data features in the memory bank, which may lead to problems of forgetting and drift for data that has not been updated for a long time. At the same time, the size of the memory bank has a decisive impact on the detection ability of the model; Knowledge distillation uses the teacher network and the student network to train the contrast model with as consistent output results as possible on the dataset. Since the model is trained uniformly on normal images, when the test sample is abnormal data, theoretically, there will be a large difference in the output of the student model and the teacher model, and thus the data is determined to be abnormal data. Summary of the Invention
[0007] Aiming at the above deficiencies in the prior art, an industrial product defect detection method based on distillation contrast learning provided by the present invention solves the problems that it is difficult to collect a comprehensive labeled dataset covering all possible defect forms in the prior art, unable to accurately detect abnormal data, and there are problems of forgetting and drift for data that has not been updated for a long time.
[0008] To achieve the above objectives, the technical solution adopted by the present invention is: an industrial product defect detection method based on distillation contrast learning, including the following steps:
[0009] S1. Collect product images, preprocess the product images to obtain normal images and pseudo-abnormal images;
[0010] S2. Obtain the true labels of the normal images and pseudo-abnormal images, and use the normal images and pseudo-abnormal images to train the teacher network to obtain the normal teacher output features and the output features of the abnormal teacher respectively;
[0011] S3. Use the normal teacher output features and the output features of the abnormal teacher as soft targets, use the true labels as hard targets, and optimize the hard targets and soft targets by using the student network according to distillation contrast learning to obtain the total distillation loss and obtain the optimal student network;
[0012] S4. Combine the normal teacher network, the abnormal teacher network and the optimal student network to construct a product defect detection model;
[0013] S5. Input the image of the product to be tested into the product defect detection model to obtain the embedded features, calculate the cosine similarity using the embedded features, compare the cosine similarity, obtain the predicted value, and analyze the predicted value and the true label to complete the industrial product defect detection.
[0014] The beneficial effects of the present invention are as follows: By using the knowledge distillation structure and data augmentation structure based on contrastive learning, the present invention constructs a product defect detection model, achieving a wide coverage of defect forms, high detection accuracy, and the ability to effectively process data that has not been updated for a long time; by establishing an end-to-end product defect detection model, the product defect detection is completed using only one model, improving the efficiency of product defect detection; and by analyzing the predicted value using the area under the receiver operating characteristic curve, the accuracy of product defect detection is improved.
[0015] Further, the S1 includes the following steps:
[0016] S101. Collect product images, randomly adjust the brightness, contrast, saturation, and hue of the product images to obtain normal images;
[0017] S102. According to the normal images, randomly select a matrix area as a patch, and perform transformation operations on the patch to obtain the transformed patch;
[0018] S103. Embed the transformed patch into a random position of the normal image to obtain a pseudo-abnormal image.
[0019] The beneficial effects of the above further solution are as follows: By randomly adjusting the basic attributes of the product images, the present invention increases the diversity of image data; and by randomly selecting and performing transformation operations to obtain pseudo-abnormal samples, the diversity of pseudo-abnormal samples is increased, making the pseudo-abnormal samples closer to the characteristics of real abnormal samples; providing effective abnormal samples for subsequent anomaly detection.
[0020] Still further, the S2 includes the following steps:
[0021] S201. Obtain the true labels of the normal images and pseudo-abnormal images, initialize the normal teacher network and the abnormal teacher network, train the normal teacher network using the normal images, and optimize the parameters of the normal teacher network according to the training results;
[0022] S202. Train the abnormal teacher network using the pseudo-abnormal images, and optimize the parameters of the abnormal teacher network according to the training results;
[0023] S203. In response to the completion of training, freeze the parameters of the normal teacher network and the abnormal teacher network, and obtain the output features of the normal teacher and the output features of the abnormal teacher.
[0024] The beneficial effects of the above further solution are as follows: During the training process of the product defect detection model of the present invention, by using two encoders with the same network structure and the same initial parameters as the teacher networks, one is trained with normal images and the other is trained with pseudo-abnormal images, respectively obtaining the teacher network for normal image distribution and the teacher network for pseudo-abnormal image classification. Since the corresponding networks have not learned the features of other distributions, the saliency of the output features is enhanced, and the accuracy and detection efficiency of anomaly detection are improved.
[0025] Furthermore, S3 includes the following steps:
[0026] S301: Take the output features of the normal teacher and the output features of the abnormal teacher as soft targets, and take the true label as the hard target;
[0027] S302: Initialize the student network. According to distillation contrast learning, use the student model to optimize the hard target and the soft target, and minimize the distillation loss function;
[0028] S303: Use the feature maps of different network layers of the normal teacher network and the abnormal teacher network to distill the student network model, and obtain the total distillation loss according to the minimized distillation loss function;
[0029] S304: Use the optimizer to perform backpropagation and optimization on the total distillation loss, adjust the parameters of the student network, and obtain the optimal student network.
[0030] Furthermore, the expression for minimizing the distillation loss function is as follows:
[0031]
[0032] where L KD represents the distillation loss, both α1 and α2 represent normalization weights, α1 + α2 ∈ {0, 1}, N represents the total number of samples, j and k both represent sample numbers, T represents the temperature, represents the prediction result of the normal teacher for sample j at temperature T, represents the prediction result of the abnormal teacher for sample j at temperature T, represents the prediction result of the student network for sample j at temperature T, v j represents the output feature obtained by the corresponding teacher network according to sample j, v k represents the output feature obtained by the corresponding teacher network according to sample k, z j represents the output feature obtained by the student network according to sample j, z k represents the output feature obtained by the student network according to sample k, c j represents the true label value.
[0033] The beneficial effects of the above further solution are as follows: Through the feature information transmitted from the teacher network to the student network in the present invention, the ability of the student network to learn and understand data distribution is improved. At the same time, by combining contrastive learning and distillation learning, the student network can learn richer and more meaningful feature representations; compared with the traditional knowledge distillation method, the present invention improves the effectiveness and discriminative ability of defect detection; the present invention uses an optimizer to perform backpropagation and optimization on the total distillation loss, effectively adjusting the parameters of the student network, enabling the student network to better fit the knowledge of the teacher network and improving the anomaly discrimination ability of the student network.
[0034] Further, S5 includes the following steps:
[0035] S501: Input the image of the product to be tested into the product defect detection model, and use the normal teacher network, abnormal teacher network, and optimal student network to process the image of the product to be tested, respectively obtaining the normal embedding features, abnormal embedding features, and student embedding features;
[0036] S502: According to the normal embedding features, abnormal embedding features, and student embedding features, through the last layer network of the product defect detection model, obtain the output features of the normal teacher network, the output features of the abnormal teacher network, and the output features of the student network;
[0037] S503: According to the output features of the normal teacher network and the output features of the student network, calculate the cosine similarity to obtain the first cosine similarity. According to the output features of the abnormal teacher network and the output features of the student network, calculate the cosine similarity to obtain the second cosine similarity;
[0038] S504: Compare the magnitudes of the first cosine similarity and the second cosine similarity to obtain a predicted value;
[0039] S505: According to the predicted value and the true label, perform analysis by calculating the area under the receiver operating characteristic curve to complete the industrial product defect detection.
[0040] The beneficial effects of the above further solution are as follows: The present invention predicts the image of the product to be tested through the constructed product defect model, calculates the output distance between the student model and the teacher model using the features of different network layers to obtain the cosine similarity, and uses the deep network layer to effectively suppress the noise and interference received by the image; and by performing anomaly detection in the semantic feature space, the significance and discriminability of the defect analysis results are improved, not only the pixel-level differences, realizing independence from pixel-level distance errors; the present invention performs analysis by calculating the area under the receiver operating characteristic curve, provides an overall performance evaluation of the product defect model at all possible thresholds, comprehensively reflects the effect of the classifier, makes the predicted value more consistent with the true label, and realizes the rapid and accurate detection of industrial product defects. Description of the Drawings
[0041] Figure 1 This is a flowchart of the method of the present invention.
[0042] Figure 2 This is a schematic diagram of generating a pseudo-abnormal image for this embodiment.
[0043] Figure 3 This is a test process diagram of using distillation contrast learning and a product defect detection model for this embodiment. Detailed Embodiments
[0044] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0045] Before describing this embodiment, the following terms are first explained:
[0046] AUROC: Area Under the Receiver Operating Characteristic Curve;
[0047] ROC: Receiver Operating Characteristic;
[0048] AUC: Area Under the Curve;
[0049] Accuracy: Accuracy Rate;
[0050] Precision: Precision Rate;
[0051] Recall: Recall Rate;
[0052] F1-Score: Harmonic Mean.
[0053] Embodiment
[0054] As Figure 1 shown, the present invention provides an industrial product defect detection method based on distillation contrast learning, and its implementation method is as follows:
[0055] S1. Collect product images, preprocess the product images to obtain normal images and pseudo-abnormal images, and the specific steps are as follows:
[0056] S101. Collect product images, randomly adjust the brightness, contrast, saturation, and hue of the product images to obtain normal images;
[0057] S102. Randomly select a matrix region from the normal image as a patch, and perform transformation operations on the patch to obtain a transformed patch.
[0058] S103. Embed the transformed patch into a random position of the normal image to obtain a pseudo-abnormal image.
[0059] In this embodiment, the present invention provides a new detection model. The knowledge distillation structure and data augmentation structure used in the product defect detection model are based on the idea of contrast learning, which is different from the traditional knowledge distillation model. First, this model uses the features of normal images and pseudo-defect images, corresponding to two different teacher models in knowledge distillation, and simultaneously distills the student model. When detecting a test image, using the idea of contrast learning, by comparing the distances between the output features of the student model and the two teacher models, it is determined whether the test image is normal or abnormal. Second, the present invention does not rely on pixel-level distance errors, but performs anomaly detection in the semantic feature space. By operating on a higher-level representation, a more meaningful and discriminative analysis of the defect can be carried out, rather than just pixel-level differences. Compared with the traditional knowledge distillation method, these improvements enhance the effectiveness and discriminative ability of defect detection.
[0060] In this embodiment, the present invention is divided into two parts: the construction and training of the product defect detection model, and the testing and defect detection of the defect detection model. The construction and training of the product defect detection model include collecting product images and preprocessing, training the teacher model, and training the student model. The testing and defect detection of the defect detection model include testing the product defect detection model on the product images to be tested, calculating the cosine similarity, and obtaining a prediction value. Analyzing the prediction value to complete the industrial product defect detection.
[0061] Collect product images, randomly adjust the brightness, contrast, saturation, and hue of the product images to obtain normal images. Then randomly select a small matrix region from the normal image as a patch. The size and aspect ratio of this patch can be adjusted according to needs to adapt to different abnormal sample shapes and sizes. By randomly selecting the patch region, it is possible to simulate the situation where abnormalities occur in different parts of the real world. And perform transformation operations on the patch. The transformation operations can be: rotation or modifying pixel values. The rotation operation can make the patch have different angles when pasted back to the original image, increasing the diversity of the samples. Modifying pixel values can change the appearance of the patch to make it closer to the characteristics of real abnormal samples. As Figure 2 shown, paste the transformed patch back to a random position of the normal image to generate a pseudo-abnormal sample.
[0062] S2. Obtain the true labels of the normal images and pseudo-abnormal images, and use the normal images and pseudo-abnormal images to train the teacher network to obtain the output features of the normal teacher and the output features of the abnormal teacher respectively. The specific steps are as follows:
[0063] S201. Obtain the true labels of the normal images and pseudo-abnormal images, initialize the normal teacher network and the abnormal teacher network, use the normal images to train the normal teacher network, and optimize the parameters of the normal teacher network according to the training results;
[0064] S202. Use the pseudo-abnormal images to train the abnormal teacher network, and optimize the parameters of the abnormal teacher network according to the training results;
[0065] S203. In response to the completion of training, freeze the parameters of the normal teacher network and the abnormal teacher network, and obtain the output features of the normal teacher and the output features of the abnormal teacher.
[0066] In this embodiment, two encoders with the same network structure and the same initial parameters are used as the teacher network. One teacher network is trained with normal images and is the normal teacher network, and the other teacher network is trained with pseudo-abnormal images and is the abnormal teacher network; according to the obtained normal images and pseudo-abnormal images, convert the normal images and pseudo-abnormal images into tensors and perform normalization processing to obtain the true labels of the normal images and pseudo-abnormal images. Since the normal images and pseudo-abnormal images are known, the true label of the normal image is 0, and the true label of the pseudo-abnormal image is 1;
[0067] Initialize the normal teacher network and the abnormal teacher network, use binary cross-entropy as the loss function for supervised training, use the normal images to train the normal teacher network, and optimize the parameters of the normal teacher network according to the training results, use the pseudo-abnormal images to train the abnormal teacher network, and optimize the parameters of the abnormal teacher network according to the training results;
[0068] The formula for the normal teacher loss function is:
[0069]
[0070] where L NT represents the normal teacher loss, N represents the total number of samples, i represents the sample number, and p i represents the probability that sample i is predicted as the positive class;
[0071] The formula for the abnormal teacher loss function is:
[0072]
[0073] where L AT represents the abnormal teacher loss, N represents the total number of samples, i represents the sample number, and pi Represents the probability that sample i is predicted as the positive class; in response to the completion of training, freeze the parameters of the normal teacher network and the abnormal teacher network to obtain the trained normal teacher network and abnormal teacher network; and use the trained normal teacher network and abnormal teacher network to process the normal images and pseudo-abnormal images to obtain the output features of the normal teacher and the output features of the abnormal teacher.
[0074] S3. Use the output features of the normal teacher and the output features of the abnormal teacher as soft targets, use the true labels as hard targets, and optimize the hard targets and soft targets using the student network according to distillation contrast learning to obtain the total distillation loss and the optimal student network. The specific steps are as follows:
[0075] S301. Use the output features of the normal teacher and the output features of the abnormal teacher as soft targets, and use the true labels as hard targets;
[0076] S302. Initialize the student network, optimize the hard targets and soft targets using the student model according to distillation contrast learning, and minimize the distillation loss function;
[0077] S303. Distill the student network model using the feature maps of different network layers of the normal teacher network and the abnormal teacher network, and obtain the total distillation loss according to the minimized distillation loss function;
[0078] S304. Use the optimizer to perform backpropagation and optimization on the total distillation loss, adjust the parameters of the student network, and obtain the optimal student network.
[0079] In this embodiment, the student network adopts the same initial parameters and network structure as the teacher network, uses the output features of the normal teacher and the output features of the abnormal teacher as soft targets, and uses the true labels as hard targets; initialize the student network, and optimize the hard targets and soft targets simultaneously using the student network according to distillation contrast learning to minimize the distillation loss function. The expression for minimizing the distillation loss function is as follows:
[0080]
[0081] where L KD represents the distillation loss, both α1 and α2 represent normalization weights, α1 + α2 ∈ {0, 1}, and L NT-soft represents the normal teacher loss in the soft target, L AT-soft represents the abnormal teacher loss in the soft target, L hard represents the hard target loss, N represents the total number of samples, j and k both represent sample numbers, represents the prediction result of the normal teacher for sample j at temperature T, represents the prediction result of the abnormal teacher for sample j at temperature T, vj denotes the output feature obtained by the corresponding teacher network according to sample j, v k denotes the output feature obtained by the corresponding teacher network according to sample k, z j denotes the output feature obtained by the student network according to sample j, z k denotes the output feature obtained by the student network according to sample k, denotes the prediction result of the student network at temperature T, T ∈ (0, ∞). When T > 1, the difference in predicted values between the target class and the non-target class decreases. On the contrary, when T < 1, the difference in predicted values between the target class and the non-target class will be further enlarged, c j denotes the true label value, c j ∈ {0, 1};
[0082] The student network model is distilled using the feature maps of different network layers of two teacher networks. The outputs of different network layers can capture features at different abstraction levels. Shallow network layers may pay more attention to low-level features such as edges and textures, while deep network layers may pay more attention to high-level semantic features such as object shapes and structures. By using the outputs of multiple different network layers, richer and more diverse feature representations can be obtained, which helps to improve the performance of the student model. Using the outputs of multiple different network layers for distillation can enhance the robustness of the model. Features at different levels may have different sensitivities to different types of anomalies. By considering features at multiple levels, the student model can more comprehensively learn the representations of abnormal features, thereby improving the detection ability for various anomalies. In real-world scenarios, images may be affected by noise and interference. Shallow network layers may be more sensitive to these noises and interferences, while deep network layers may have stronger suppression capabilities. By using the outputs of multiple different network layers, noise and interference can be effectively suppressed, improving the robustness and accuracy of the model.
[0083] As Figure 3 shown in (a) below, during the distillation process, a normal image is input into the student network for forward propagation to obtain the embedding features and prediction results of the normal distillation student network in the last four layers. Similarly, the same normal image is input into the normal teacher network for forward propagation to obtain the corresponding output of the normal distillation teacher network. The embedding features of the student network for the abnormal image in the last four layers are used as the output of the abnormal distillation student network, and the outputs and embedding features of the abnormal teacher network at the same layers are used as the output of the abnormal distillation teacher network. The distillation loss function is used to calculate the losses of normal distillation and abnormal distillation respectively. The abnormal distillation loss and the normal distillation loss are added together to obtain the total distillation loss; and the optimizer is used to perform backpropagation and optimization on the total distillation loss to adjust the parameters of the student network.
[0084] S4. Combine the normal teacher network, the abnormal teacher network, and the optimal student network to construct a product defect detection model.
[0085] In this embodiment, a product defect detection model based on distillation contrast learning is constructed by combining the normal teacher network, the abnormal teacher network, and the optimal student network. The model is constructed by the network obtained from distillation contrast learning. Therefore, the loss function of the product defect detection model is also the same distillation loss function.
[0086] S5. Input the image of the product to be tested into the product defect detection model to obtain the embedded features, calculate the cosine similarity using the embedded features, compare the cosine similarities, obtain the prediction value, and analyze the prediction value and the true label to complete the industrial product defect detection. The specific steps are as follows:
[0087] S501. Input the image of the product to be tested into the product defect detection model, and use the normal teacher network, the abnormal teacher network, and the optimal student network to process the image of the product to be tested, respectively obtaining the normal embedded features, the abnormal embedded features, and the student embedded features.
[0088] S502. According to the normal embedded features, the abnormal embedded features, and the student embedded features, through the last layer network of the product defect detection model, obtain the output features of the normal teacher network, the output features of the abnormal teacher network, and the output features of the student network.
[0089] S503. Calculate the cosine similarity according to the output features of the normal teacher network and the output features of the student network to obtain the first cosine similarity, and calculate the cosine similarity according to the output features of the abnormal teacher network and the output features of the student network to obtain the second cosine similarity.
[0090] S504. Compare the magnitudes of the first cosine similarity and the second cosine similarity to obtain the prediction value.
[0091] S505. According to the prediction value and the true label, analyze by calculating the area under the receiver operating characteristic curve to complete the industrial product defect detection.
[0092] In this embodiment, as Figure 3As shown in (b), the product image currently being tested is input into the product defect detection model. The normal teacher network, the abnormal teacher network, and the optimal student network are respectively used to process the product image to be tested, and the normal embedding feature, the abnormal embedding feature, and the student embedding feature are respectively obtained. The output distances between the student model and the teacher models are calculated using the output features of different network layers, that is, the output features of the last layer of the network are used to obtain the output features of the normal teacher network, the output features of the abnormal teacher network, and the output features of the student network. These output features are all included in the semantic feature space and are high-dimensional deep features. The cosine similarities between the student model and the outputs of different teacher models are calculated to obtain the normal cosine similarity and the abnormal cosine similarity, realizing the calculation of the cosine similarity without relying on the pixel-level distance error. The expression for calculating the cosine similarity is as follows:
[0093]
[0094] Among them, cos(q,k) represents the cosine similarity between the student network and the teacher network, q represents the output feature of the teacher network, k represents the output feature of the student network, n represents the total number of product images currently being tested, q i represents the i-th output feature of the corresponding teacher network, k i represents the i-th output feature of the student network. When the output features of the corresponding teacher network and the student network are more similar, the cosine similarity value is closer to 1, and vice versa, it is closer to -1;
[0095] According to the calculated normal cosine similarity and abnormal cosine similarity, when the normal cosine similarity is greater than the abnormal cosine similarity, a prediction value is obtained, and it can be known that the product image currently being tested is a normal image; when the abnormal cosine similarity is greater than the normal cosine similarity, a prediction value is obtained, and it can be known that the product image currently being tested is an abnormal image;
[0096] Combined with the true label, the obtained prediction value is directly used to calculate the AUROC, and the accuracy, precision, recall, and harmonic mean need to be calculated; the calculation expressions are as follows:
[0097]
[0098] Among them, TP represents the number of positive classes predicted as positive classes, TN represents the number of negative classes predicted as negative classes, FP represents the number of negative classes predicted as positive classes, FN represents the number of positive classes predicted as negative classes, Accuracy represents the accuracy rate, Precision represents the precision rate, Recall represents the recall rate, and F1-Score represents the harmonic mean; ROC is the Receiver Operating Characteristic; the abscissa of the ROC is the false positive rate (also called the false positive class rate, False Positive Rate), and the ordinate is the true positive rate (true positive class rate, True Positive Rate); among them, TPR represents the ratio of samples that are actually positive and are correctly judged as positive, and its calculation formula is the same as that of the recall rate formula; FPR represents the ratio of samples that are actually negative and are wrongly judged as positive, and its calculation formula is Given a binary classification model and its threshold, a coordinate point (X = FPR, Y = TPR) can be calculated from the true values and predicted values of all samples; the AUROC curve is formed by plotting the relationship between the true positive rate (TPR) and the false positive rate (FPR). The closer the AUROC curve is to the upper left corner, the better the prediction performance of the present invention; when AUROC is used for classification tasks, the area enclosed by the ROC curve and the coordinate axes is between 0.1 and 1. As a numerical value, AUC can intuitively evaluate the quality of the classifier, and the larger the value, the better; through the overall performance evaluation of the present invention under all possible thresholds, industrial product defect detection is completed.
Claims
1. A method for industrial product defect detection based on distillation contrastive learning, characterized in that: The following steps are involved: S1, collecting product images, preprocessing the product images, and obtaining normal images and pseudo-abnormal images; S2. Obtain the true labels of normal images and pseudo-abnormal images, use the normal images and pseudo-abnormal images to train the teacher network, and obtain the output features of the normal teacher and the output features of the abnormal teacher respectively; S3. Take the output features of the normal teacher and the output features of the abnormal teacher as soft targets, and the true labels as hard targets. According to the distillation contrast learning, the student network is used to optimize the hard targets and soft targets, and the total distillation loss is obtained to obtain the optimal student network, which is specifically: S301, taking the output features of the normal teacher and the output features of the abnormal teacher as soft targets, and taking the true label as hard target; S302, initialize the student network, optimize the hard targets and soft targets using the student model according to distillation contrast learning, and minimize the distillation loss function; S303, distilling the student network model using feature graphs of different network layers of the normal teacher network and the abnormal teacher network, and obtaining the total distillation loss according to the minimized distillation loss function; S304, using the optimizer to back-propagate and optimize the total distillation loss, adjusting the student network parameters, and obtaining the optimal student network; S4. Combining the normal teacher network, the abnormal teacher network and the optimal student network, a product defect detection model is constructed; S5. Input the image of the product to be tested into the product defect detection model to obtain embedded features, calculate the cosine similarity using the embedded features, and compare the cosine similarity to obtain the predicted value, analyze the predicted value and the true label, and complete the industrial product defect detection, specifically: S501, input the image of the product to be tested into the product defect detection model, use the normal teacher network, the abnormal teacher network and the optimal student network to process the image of the product to be tested, and obtain normal embedding features, abnormal embedding features and student embedding features respectively; S502, according to the normal embedding features, the abnormal embedding features and the student embedding features, through the last layer network of the product defect detection model, obtain the output features of the normal teacher network, the output features of the abnormal teacher network and the output features of the student network; S503, calculating the cosine similarity according to the output features of the normal teacher network and the output features of the student network to obtain a first cosine similarity, and calculating the cosine similarity according to the output features of the abnormal teacher network and the output features of the student network to obtain a second cosine similarity; S504, comparing the first cosine similarity and the second cosine similarity to obtain a predicted value; S505. According to the predicted value and the true label, the area under the receiver operating characteristic curve is calculated for analysis to complete the industrial product defect detection.
2. The industrial product defect detection method based on distillation contrast learning according to claim 1, characterized in that: The S1 comprises the following steps: S101, collecting product images, and randomly adjusting the brightness, contrast, saturation, and hue of the product images to obtain normal images; S102, randomly selecting a matrix area as a patch according to the normal image, and performing a transformation operation on the patch to obtain a transformed patch; S103 , embedding the transformed patch into a random position of the normal image to obtain a pseudo abnormal image.
3. The industrial product defect detection method based on distillation contrast learning according to claim 1, characterized in that: The S2 comprises the following steps: S201, obtaining the true labels of normal images and pseudo-abnormal images, initializing the normal teacher network and the abnormal teacher network, training the normal teacher network using normal images, and optimizing the parameters of the normal teacher network according to the training results; S202, training the abnormal teacher network using pseudo abnormal images, and optimizing the parameters of the abnormal teacher network according to the training results; S203. In response to the training being completed, parameters of the normal teacher network and the abnormal teacher network are frozen, and output features of the normal teacher and output features of the abnormal teacher are obtained.
4. The industrial product defect detection method based on distillation contrast learning according to claim 1, characterized in that: The expression for minimizing the distillation loss function is as follows: Among them, L KD represents the distillation loss, α1 and α2 represent the normalized weights, α1+α2∈{0,1}, N represents the total number of samples, j and k represent the sample numbers, T represents the temperature, represents the prediction result of the normal teacher based on sample j at temperature T, represents the prediction result of the abnormal teacher based on sample j at temperature T, represents the prediction result of the student network based on sample j at temperature T, v j represents the output feature obtained by the corresponding teacher network based on sample j, v k represents the output feature obtained by the corresponding teacher network based on sample k, z j represents the output features obtained by the student network based on sample j, z k represents the output features obtained by the student network based on sample k, c j represents the true label value, c j ∈{0,1}.
Citation Information
Patent Citations
Glass defect detection algorithm based on multi-mode teacher and student framework
CN118468230A
Image defect detection method based on double-branch inverse distillation and multi-input image
CN118521570A