Semi-Supervised Industrial Defect Detection Method and System Based on Feature Comparison
Through the semi-supervised learning method of feature comparison, the teacher network is used to generate pseudo-labels and optimize the student network, solving the problem of insufficient labeling data in industrial defect detection and achieving efficient and accurate defect detection.
Patent Information
- Application Number
- CN202210094096.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-01-26
AI Technical Summary
The prior art requires a large amount of labeled data in industrial defect detection, which is costly and deep learning methods rely on specific image features and cannot cope with complex backgrounds, resulting in low detection efficiency and error-prone.
A semi-supervised learning method based on feature comparison is adopted to generate pseudo-labels through the teacher network and filter reliable pixels. Cross-entropy and feature comparison are used to optimize the student network, reduce the need for labeling data, and improve detection accuracy.
Achieve high-precision defect detection under a small number of supervised samples, reduce production costs, and improve the robustness and accuracy of the detection model.
Smart Images

Figure CN114494780B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and deep learning, and in particular, to a semi-supervised industrial defect detection method and system based on feature comparison. Background Art
[0002] In industrial production, almost all products need to be quality inspected, and most of the quality inspection processes are completed by quality inspectors using visual inspection to detect product defects (hereinafter referred to as visual inspection). Due to the diversity of products and defects, the workload and difficulty of quality inspectors are greatly increased, resulting in a decrease in the efficiency of manual visual inspection and easy omission and misjudgment due to the fatigue and mistakes of quality inspectors, increasing the time cost of the production line and possibly affecting the quality of products on the market. Therefore, the adoption of automated detection technology has great value.
[0003] Early automated detection methods tended to extract specific manual image features according to the type of defects, and used digital image processing methods such as threshold segmentation, elliptical Gabor filters, RGB histograms, etc. to select specific image features. The recognition rate of digital image processing methods is very sensitive to various factors such as lighting and contrast, and it is too dependent on the extracted specific image features, unable to handle complex backgrounds and the recognition tasks of multiple defects, and does not have universality.
[0004] In recent years, with the development of deep learning research in the field of machine learning, introducing deep learning methods into the detection of various product defect images can greatly improve the recognition accuracy, reduce the omission rate, and improve the robustness. However, existing deep learning methods require a large amount of training data, which is difficult to obtain in the industrial field, and the cost of annotating these massive data is extremely high, including the cost of manual annotation, the cost of training graphics card resources, and the cost of extremely long training time.
[0005] In order to achieve the purpose of using a small number of samples and a large number of unlabeled samples, semi-supervised learning methods have received extensive attention. However, most semi-supervised learning methods are applied to image classification, and there are few semi-supervised learning methods combined with image segmentation and applied to industrial defect detection. Patent CN113436169A proposes a method and system for detecting surface cracks of industrial equipment based on semi-supervised semantic segmentation of GAN. However, it is significantly different from the core of the present invention - additional supervision based on feature comparison. The method of generating annotations based on GAN used in this patent has disadvantages such as difficult convergence and instability during training, while the supervision method based on feature comparison in the present invention is a method that can ensure stable convergence both in theory and experiments.
[0006] Patent document CN112801962A (application number: CN202110066948.7) discloses a semi-supervised industrial product defect detection method and system based on positive sample learning. An image restoration network and a defect segmentation prediction network are trained through positive sample images. An image containing a defect to be detected is input into the image restoration network to obtain a restored genuine product image. After calculating the absolute value of the difference between the two, the three images are spliced to obtain a detection tensor. The segmentation prediction network generates a binary segmentation mask image based on the detection tensor to obtain the defect area. However, this invention is not based on feature comparison, cannot avoid semantic confusion, and has limited accuracy in defect detection. Summary of the Invention
[0007] Aiming at the defects in the prior art, the purpose of the present invention is to provide a semi-supervised industrial defect detection method and system based on feature comparison.
[0008] A semi-supervised industrial defect detection method based on feature comparison provided by the present invention includes:
[0009] Step S1: Collect pictures of products to be detected and randomly label the pictures with tags;
[0010] Step S2: Classify the products to be detected into labeled inputs and unlabeled inputs;
[0011] Step S3: For labeled inputs, use the pictures and the corresponding tags to train the student network;
[0012] For unlabeled inputs, input them into the teacher network to generate corresponding pseudo-labels and representations;
[0013] Step S4: Screen the pseudo-labels to distinguish reliable pixels and unreliable pixels;
[0014] Step S5: Feed the reliable pixels into the student network for supervision; for the unreliable pixels, based on their feature encoding information, perform feature optimization on the student network based on contrast learning.
[0015] Preferably, in the said step S3:
[0016] The steps of training the student network with the pictures and the corresponding tags for labeled inputs are as follows:
[0017] Step S3.1: Perform data augmentation on the input pictures;
[0018] Step S3.2: Use the input pictures after data augmentation to train a teacher network, and then the teacher network does not perform gradient update;
[0019] Step S3.3: Use the same input pictures in step S3.2 to train a student network;
[0020] Step S3.4: Train the teacher network and the student network based on cross-entropy.
[0021] Preferably, the cross-entropy is used to measure the difference in semantic information between the input image and the labels of the true annotations, and its calculation process is as follows:
[0022]
[0023]
[0024] where y i is the difference in semantic information in the i-th feature vector; x i is the i-th value of the feature vector, x j is the j-th value of the feature vector. First, normalize it through the softmax function to convert the values of each dimension of the feature vector into probability form, and then calculate its cross-entropy; H y′ (y) is the cross-entropy; y' i is the ideal result, the correct label vector.
[0025] Preferably, in the step S4:
[0026] Use the entropy value to perform pixel-level screening on the pseudo-labels. The specific steps are as follows:
[0027] Step S4.1: For the predicted probability distribution of each pixel, the entropy value is calculated according to the following formula:
[0028] Entropy(p) = -p i log(p i )
[0029] where Entropy(p) is the entropy value, p is the pixel to be calculated, and p i is the probability that the pixel p belongs to the category i;
[0030] Step S4.2: Calculate the entropy values of all pixels. The pixels with entropy values ranked in the latter 50% are regarded as reliable pixels, and the pixels with entropy values ranked in the former 50% are regarded as unreliable pixels;
[0031] Step S4.3: Regard the reliable pixels and the pseudo-labels of the reliable pixels as labeled inputs and use them as the supervision information of the student network. The loss function is cross-entropy.
[0032] Preferably, in the step S5:
[0033] The steps for optimizing the features of the student network based on contrast learning according to the feature coding information for low-confidence labels are as follows:
[0034] Step S5.1: Sort the predicted probability distributions of the non-pixels according to the probability values;
[0035] Step S5.2: For each category of reliable pixels, if this category does not appear in the top three categories of unreliable pixel probabilities, perform loss calculation optimization for feature comparison, and its calculation process is as follows:
[0036]
[0037]
[0038] Among them, C is the probability category, M represents the position information on the picture, z i represents the representation output by the teacher network at the corresponding position, τ is a preset temperature coefficient, N is a preset number of negative samples, is the representation of the corresponding positive sample, is the representation of the corresponding negative sample, and <,> represents the operation of taking the inner product between vectors.
[0039] A semi-supervised industrial defect detection system based on feature comparison provided by the present invention includes:
[0040] Module M1: Collect pictures of products to be tested and randomly label the pictures;
[0041] Module M2: Classify the products to be tested into labeled inputs and unlabeled inputs;
[0042] Module M3: For labeled inputs, use the pictures and corresponding labels to train the student network;
[0043] For unlabeled inputs, input them into the teacher network to generate corresponding pseudo-labels and representations;
[0044] Module M4: Screen the pseudo-labels to distinguish reliable pixels and unreliable pixels;
[0045] Module M5: Feed the reliable pixels into the student network for supervision; for the unreliable pixels, based on their feature coding information, perform feature optimization on the student network based on contrast learning.
[0046] Preferably, in the module M3:
[0047] The steps of training the student network with the pictures and corresponding labels for the labeled inputs are as follows:
[0048] Module M3.1: Perform data augmentation on the input pictures;
[0049] Module M3.2: Use the input pictures after data augmentation to train a teacher network, and then the teacher network does not perform gradient update;
[0050] Module M3.3: Train a student network with the same input images as in Module M3.2;
[0051] Module M3.4: Train the teacher network and the student network based on cross - entropy.
[0052] Preferably, the cross - entropy is used to measure the difference in semantic information between the input image and the labels of the true annotations, and its calculation process is as follows:
[0053]
[0054]
[0055] where y i is the difference in semantic information in the i - th feature vector; x i is the i - th value of the feature vector, x j is the j - th value of the feature vector. First, it is normalized through the softmax function to convert the values of each dimension of the feature vector into probability form, and then its cross - entropy is obtained; H y′ (y) is the cross - entropy; y' i is the ideal result, the correct label vector.
[0056] Preferably, in Module M4:
[0057] Use the entropy value to perform pixel - level screening on the pseudo - labels, and the specific steps are as follows:
[0058] Module M4.1: For the predicted probability distribution of each pixel, the entropy value is calculated according to the following formula:
[0059] Entropy(p)=-p i log(p i )
[0060] where Entropy(p) is the entropy value, p is the pixel to be calculated, and p i is the probability that the pixel p belongs to class i;
[0061] Module M4.2: Calculate the entropy values of all pixels. The pixels with entropy values ranked in the lower 50% are regarded as reliable pixels, and the pixels with entropy values ranked in the upper 50% are regarded as unreliable pixels;
[0062] Module M4.3: Regard the reliable pixels and the pseudo - labels of the reliable pixels as labeled inputs, which are used as the supervision information for the student network, and the loss function is cross - entropy.
[0063] Preferably, in Module M5:
[0064] The steps for optimizing the features of the student network based on contrast learning for low - confidence labels according to the feature encoding information are as follows:
[0065] Module M5.1: Sort the predicted probability distribution of non-pixelizable pixels according to probability values;
[0066] Module M5.2: For reliable pixels of each category, when this category does not appear in the top three categories of unreliable pixel probabilities, optimize the loss calculation of feature comparison. The calculation process is as follows:
[0067]
[0068]
[0069] Among them, C is the probability category, M represents the position information on the picture, z i represents the representation output by the teacher network at the corresponding position, τ is a preset temperature coefficient, N is the preset number of negative samples, is the representation of the corresponding positive sample, is the representation of the corresponding negative sample, and <,> represents the operation of taking the inner product between vectors.
[0070] Compared with the prior art, the present invention has the following beneficial effects:
[0071] 1. The present invention has the ability to train the model in the case of semi-supervised input, so that the training process of the industrial defect detection model does not require a large amount of labeled data, greatly saving the production cost;
[0072] 2. The present invention is based on feature comparison, avoiding semantic confusion, greatly utilizing unlabeled data, and greatly improving the accuracy of defect detection;
[0073] 3. Based on the distinguishability of features of different categories in the segmentation results, the present invention provides additional constraints for training in the case of a small number of supervised samples, so that the defect detection model trained in the semi-supervised case still has strong detection ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects and advantages of the present invention will become more apparent:
[0075] Figure 1 is a flowchart of the brief process of obtaining the model after defect detection by the training network model of the present invention;
[0076] Figure 2 is a structural block diagram of the detection system and the module inference system of the embodiment of the present invention;
[0077] Figure 3 is a brief illustration diagram of the algorithm of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0078] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.
[0079] Example 1:
[0080] A semi-supervised industrial defect detection method based on feature comparison provided by the present invention, as Figures 1 - 3 shown, includes:
[0081] Step S1: Collect pictures of products to be tested and randomly label the pictures with tags;
[0082] Step S2: Classify the products to be tested into labeled inputs and unlabeled inputs;
[0083] Step S3: For labeled inputs, use the pictures and corresponding tags to train the student network;
[0084] For unlabeled inputs, input them into the teacher network to generate corresponding pseudo-labels and representations;
[0085] Step S4: Screen the pseudo-labels to distinguish reliable pixels and unreliable pixels;
[0086] Step S5: Feed the reliable pixels into the student network for supervision; for the unreliable pixels, based on their feature coding information, perform feature optimization on the student network based on contrastive learning.
[0087] Specifically, in the said step S3:
[0088] The steps of training the student network with the pictures and corresponding tags for labeled inputs are as follows:
[0089] Step S3.1: Perform data augmentation on the input pictures;
[0090] Step S3.2: Use the input pictures after data augmentation to train a teacher network, and then the teacher network does not perform gradient update;
[0091] Step S3.3: Use the same input pictures in step S3.2 to train a student network;
[0092] Step S3.4: Based on cross-entropy, train the teacher network and the student network.
[0093] Specifically, the cross-entropy is used to measure the difference in semantic information between the input image and the labels of the true annotations, and its calculation process is as follows:
[0094]
[0095]
[0096] Among them, y i is the difference in semantic information in the i-th feature vector; x i is the i-th value of the feature vector, x j is the j-th value of the feature vector. First, it is normalized through the softmax function to convert the values of each dimension of the feature vector into probability form, and then its cross-entropy is obtained; H y′ (y) is the cross-entropy; y' i is the ideal result, the correct label vector.
[0097] Specifically, in the step S4:
[0098] The entropy value is used to perform pixel-level screening on the pseudo-labels, and the specific steps are as follows:
[0099] Step S4.1: For the predicted probability distribution of each pixel, the entropy value is calculated according to the following formula:
[0100] Entropy(p) = -p i log(p i )
[0101] where Entropy(p) is the entropy value, p is the pixel to be calculated, and p i is the probability that the pixel p belongs to the category i;
[0102] Step S4.2: Calculate the entropy values of all pixels. The pixels with entropy values ranked in the latter 50% are regarded as reliable pixels, and the pixels with entropy values ranked in the former 50% are regarded as unreliable pixels;
[0103] Step S4.3: Regard the reliable pixels and the pseudo-labels of the reliable pixels as labeled inputs and use them as the supervision information of the student network. The loss function is cross-entropy.
[0104] Specifically, in the step S5:
[0105] The steps for optimizing the features of the student network based on contrast learning according to the feature coding information for the low-confidence labels are as follows:
[0106] Step S5.1: Sort the predicted probability distributions of the non-pixels according to the probability values;
[0107] Step S5.2: For the reliable pixels of each category, when this category does not appear in the top three categories of unreliable pixel probabilities, the loss calculation optimization of feature comparison is performed, and its calculation process is as follows:
[0108]
[0109]
[0110] Among them, C is the probability category, M represents the position information on the picture, z i represents the representation output by the teacher network at the corresponding position, τ is a preset temperature coefficient, N is a preset number of negative samples, is the representation of the corresponding positive sample, is the representation of the corresponding negative sample, and <,> represents the operation of taking the inner product between vectors.
[0111] Example 2:
[0112] Embodiment 2 is a preferred example of Embodiment 1 to illustrate the present invention more specifically.
[0113] Those skilled in the art can understand a semi-supervised industrial defect detection method based on feature comparison provided by the present invention as a specific implementation manner of a semi-supervised industrial defect detection system based on feature comparison, that is, the semi-supervised industrial defect detection system based on feature comparison can be implemented by executing the step process of the semi-supervised industrial defect detection method based on feature comparison.
[0114] A semi-supervised industrial defect detection system based on feature comparison provided by the present invention includes:
[0115] Module M1: Collect pictures of products to be tested and randomly label the pictures;
[0116] Module M2: Classify the products to be tested into labeled input and unlabeled input;
[0117] Module M3: For the labeled input, train the student network with the pictures and the corresponding labels;
[0118] For the unlabeled input, input it into the teacher network to generate corresponding pseudo-labels and representations;
[0119] Module M4: Screen the pseudo-labels to distinguish reliable pixels and unreliable pixels;
[0120] Module M5: Send the reliable pixels into the student network for supervision; for the unreliable pixels, optimize the features of the student network based on contrast learning according to their feature coding information.
[0121] Specifically, in the module M3:
[0122] The steps for training the student network with the labeled input images and the corresponding labels are as follows:
[0123] Module M3.1: Perform data augmentation on the input images;
[0124] Module M3.2: Train a teacher network with the input images after data augmentation, and then the teacher network does not perform gradient update;
[0125] Module M3.3: Train a student network with the same input images as in Module M3.2;
[0126] Module M3.4: Train the teacher network and the student network based on cross - entropy.
[0127] Specifically, the cross - entropy is used to measure the difference in semantic information between the input image and the true labeled tags, and its calculation process is as follows:
[0128]
[0129]
[0130] where y i is the difference in semantic information in the i - th feature vector; x i is the i - th value of the feature vector, x j is the j - th value of the feature vector. First, it is normalized through the softmax function to convert the values of each dimension of the feature vector into probability form, and then its cross - entropy is obtained; H y′ (y) is the cross - entropy; y' i is the ideal result, the correct label vector.
[0131] Specifically, in the module M4:
[0132] Use the entropy value to perform pixel - level screening on the pseudo - labels, and the specific steps are as follows:
[0133] Module M4.1: For the predicted probability distribution of each pixel, the entropy value is calculated according to the following formula:
[0134] Entropy(p) = - p i log(p i )
[0135] where Entropy(p) is the entropy value, p is the pixel to be calculated, and p i is the probability that the pixel p belongs to the category i;
[0136] Module M4.2: Calculate the entropy value of all pixels. Pixels with entropy values ranked in the lower 50% are regarded as reliable pixels, and pixels with entropy values ranked in the upper 50% are regarded as unreliable pixels;
[0137] Module M4.3: Consider the reliable pixels and the pseudo - labels of reliable pixels as the labeled input, which serves as the supervision information for the student network, and the loss function is cross - entropy.
[0138] Specifically, in the said Module M5:
[0139] The steps for optimizing the features of the student network based on contrastive learning for low - confidence labels according to the feature encoding information are as follows:
[0140] Module M5.1: Sort the predicted probability distributions of the unreliable pixels according to the probability values;
[0141] Module M5.2: For the reliable pixels of each category, when this category does not appear in the top three categories of the probability ranking of the unreliable pixels, calculate and optimize the loss of feature contrast, and its calculation process is as follows:
[0142]
[0143]
[0144] Among them, C is the probability category, M represents the position information on the picture, z i represents the representation output by the teacher network at the corresponding position, τ is a preset temperature coefficient, N is the preset number of negative samples, is the representation of the corresponding positive sample, is the representation of the corresponding negative sample, and <,> represents the operation of taking the inner product between vectors.
[0145] Example 3:
[0146] Embodiment 3 is a preferred example of Embodiment 1 to illustrate the present invention more specifically.
[0147] 1. A semi - supervised industrial defect detection system and method based on feature contrast, including:
[0148] Step A: Collect pictures of the product to be measured and perform pixel - level annotation on part of the pictures to form an input sample with partial supervision information.
[0149] Step B: According to whether the product to be measured has a picture, divide the input sample into "labeled input" and "unlabeled input".
[0150] Step C: For the "labeled input", directly use these pictures and the corresponding labels to train the student network.
[0151] Step D: For the "unlabeled input", input it into the teacher network to generate corresponding pseudo-labels and representations.
[0152] Step E: Using the entropy value, perform pixel-level screening on the pseudo-labels to distinguish high-confidence and low-confidence pseudo-labels.
[0153] Step F: For the high-confidence labels, directly input them into the student network for supervision.
[0154] Step G: For the low-confidence labels, optimize the features of the student network based on contrastive learning according to their feature encoding information.
[0155] 2. The semi-supervised industrial defect detection system and method based on feature contrast according to claim 1, wherein the step C includes:
[0156] Step S1: Perform data augmentation on the input image, such as flipping, cropping, changing the size, and Gaussian blurring.
[0157] Step S2: Use the input image after data augmentation to preliminarily train a teacher network, and then this network does not perform gradient update.
[0158] Step S3: Train a student network with the same input image as the previous step.
[0159] Step S4: The definition of the loss function during training is based on cross-entropy to train the teacher network and the student network.
[0160] 3. The semi-supervised industrial defect detection system and method based on feature contrast according to claim 1, wherein the cross-entropy is used to measure the difference in semantic information between the input image and the labels of the true annotations, and its calculation process is as follows:
[0161]
[0162]
[0163] where y i is the difference in semantic information in the i-th feature vector;
[0164] x i is the i-th value of the feature vector. First, it is normalized through the softmax function to convert the values of each dimension of the feature vector into probability form, and then its cross-entropy is obtained.
[0165] H y′ (y) is the cross-entropy;
[0166] y i ' is the ideal result, that is, the correct label vector.
[0167] 4. The semi-supervised industrial defect detection system and method based on feature comparison according to claim 1, wherein the step E includes:
[0168] Step S1: For the predicted probability distribution of each pixel, its entropy value is calculated according to the following formula:
[0169] Entropy(p) = -p i log(p i )
[0170] where Entropy(p) is the entropy value, p is the pixel to be calculated, and p i is the probability that the category of pixel p is i
[0171] Step S2: Calculate the entropy values of all pixels. The pixels with smaller entropy values, which are in the latter 50%, are regarded as reliable pixels, and those in the former 50% are regarded as unreliable pixels.
[0172] Step S3: Take the reliable pixels and their pseudo-labels as the labeled input and use them as the supervision information of the student network. The loss function is also the cross-entropy.
[0173] 5. The semi-supervised industrial defect detection system and method based on feature comparison according to claim 1, wherein the step G includes:
[0174] Step S1: Sort the predicted probability distributions of the unreliable pixels according to the probability values.
[0175] Step S2: For the reliable pixels of each category, when this category does not appear in the top three categories of the probability sorting of the unreliable pixels, calculate and optimize the loss of feature comparison. The calculation process is as follows:
[0176]
[0177]
[0178] where C is the probability category, M represents the position information on the picture, z i represents the representation output by the teacher network at the corresponding position, τ is a preset temperature coefficient, N is a preset number of negative samples, is the representation of the corresponding positive sample, is the representation of the corresponding negative sample, <,> represents the operation of taking the inner product between vectors, log(x) and e x are common mathematical functions.
[0179] Example 4:
[0180] Example 4 is a preferred example of Example 1 to illustrate the present invention more specifically.
[0181] The present invention provides a semi-supervised industrial defect detection system and method based on feature comparison, which relates to the hardware system and software algorithm for defect detection. The hardware system includes a detection table, an imaging device, and a model inference system. The software algorithm adopts a deep learning segmentation scheme, uses a feature encoding-decoding network, and improves the segmentation accuracy through the discriminability of feature encoding in different categories.
[0182] In particular, aiming at the problem of too high cost of defect sample annotation in industry, the present invention uses a small number of labeled samples and a large number of unlabeled samples to improve the accuracy of defect detection, provides a semi-supervised industrial defect detection system and method based on feature comparison, and improves the performance of the detection method, and performs the following step-by-step operations:
[0183] Step A: Collect pictures of products to be detected, and perform pixel-level annotation on some pictures to form input samples with partial supervision information.
[0184] Step B: According to whether the product to be detected has a picture, divide the input samples into "labeled input" and "unlabeled input".
[0185] Step C: For the "labeled input", directly use these pictures and corresponding labels to train the student network.
[0186] Step D: For the "unlabeled input", input it into the teacher network to generate corresponding pseudo-labels and representations.
[0187] Step E: Use entropy value to perform pixel-level screening on the pseudo-labels to distinguish high-confidence and low-confidence pseudo-labels.
[0188] Step F: For the high-confidence labels, directly send them into the student network for supervision.
[0189] Step G: For the low-confidence labels, perform feature optimization based on contrast learning on the student network according to their feature encoding information.
[0190] Among them, in order to achieve better detection effects, a variety of data augmentation techniques need to be used in combination for data preprocessing. For the labeled input, data augmentation such as rotation, flipping, Gaussian blur, and coloring is required, and specific data augmentation operations are performed according to a certain probability. In the specific processing process, some functions and methods in Opencv need to be used. For the unlabeled pictures, data augmentation should be performed on their pseudo-labels and the original Figure 1 together, and in the order of first performing weak data augmentation and then performing strong data augmentation. Strong data augmentation here includes but is not limited to random cropping, random category replacement, etc., and is performed according to a uniform probability distribution.
[0191] In specific implementation, the feature extraction neural network used in this invention includes but is not limited to typical neural networks such as Resnet, VGG, MobileNet, and ICNet. The semi-supervised industrial defect detection system and method based on feature comparison uses a residual network for feature extraction. The ResNet network structure diagram consists of four residual blocks. Each residual block contains two convolutional layers, both using 3*3 convolutional kernels. Through forward identity mapping, ResNet solves the problem of gradient disappearance that many neural networks encounter when developing deeper, providing a technical basis for implementing deeper network structures. In this invention, the network structure of ResNet is introduced taking ResNet18 as an example. In actual implementation, deeper residual neural networks such as ResNet34, ResNet50, ResNet101, and ResNet152 have been successfully used. Feature extraction is carried out through the following steps:
[0192] Step S1. Divide the existing dataset. It is divided into a training set, a validation set, and a test set in a ratio of about 50%-25%-25%. The training set is used to train the network, the validation set is used to adjust and select network parameters, and the test set is used to determine the performance of the model.
[0193] Step S2. Both labeled data and unlabeled data can be input into the ResNet network to extract features, and the output features are in a vectorized representation form.
[0194] Step S3. After obtaining the vectorized features, input them into the segmentation label generation network DeepLab v3+. This is a complex structure that uses the features of the shallowest and deepest layers of ResNet as inputs, and uses the ASPP module, and finally upsamples to the original image size.
[0195] Step S4. According to the segmentation information in the previous step, we splice the feature vectors generated by DeepLab v3+. The spliced feature vectors enter a three-layer fully connected layer. The number of neurons in the three-layer fully connected layer is 256, 128, and labelnumber in sequence, that is, the number of output channels of the last fully connected layer is the number of labels to be classified. This step reflects the targeting ability of this method for multi-classification problems.
[0196] Step S5. The definition of the loss function in training is based on cross-entropy, which is used to measure the difference in semantic information between two input images, namely the feature vectors of the image to be measured (various defect images during training) and the template image. Its calculation process is as follows:
[0197]
[0198]
[0199] Let \(x_i\) be the \(i\)-th value of the feature vector. First, it is normalized through the softmax function to convert the values of each dimension of the feature vector into a probability form, and then its cross-entropy is obtained. The smaller the cross-entropy, the more accurate the prediction result, and this loss is used to guide the training process of the network.
[0200] As Figure 1 shown, the training process of the semi-supervised industrial defect detection system and method based on feature comparison implemented by the present invention specifically includes the following steps:
[0201] (1) Collect real defect image samples, including a special fixture for fixing the object to be detected, a light source and a camera fixed on the fixture
[0202] (2) Perform partial manual annotation on the real image samples to be detected and the defect image samples generated by using data augmentation technology.
[0203] (3) Train the student-teacher network. Extract image features through ResNet, splice the feature vectors, input them into the fully connected layer, and complete the training of the entire network under the guidance of the Loss function based on cross-entropy.
[0204] The system of the present invention is implemented through three parts: a detection table, image acquisition, and algorithm inference. As Figure 2 shown, among them, the detection table includes part of the production line, which is a platform for manually picking up products for shooting or can be embedded in the automated production line for installing the image acquisition module, and is provided with a camera mounting bracket and necessary positioning and fastening devices. Image acquisition includes a camera, a light source, and related accessories, which are installed on the bracket of the platform. Algorithm inference includes a host computer and corresponding neural network models and algorithms. The detection table is provided with a robotic arm device with multiple degrees of freedom of movement for obtaining pictures of workpieces to be detected in all directions, and its control system is also adapted to detect at any angle.
[0205] An example of detecting keyboard defects by the semi-supervised industrial defect detection system and method based on feature comparison implemented by the present invention specifically includes the following steps:
[0206] (1) Keyboard picture acquisition: The keyboard is fixed at the same position by a fixture, and the position of the keyboard in the collected pictures has extremely high consistency, so that each keycap to be detected can be accurately intercepted, and the situation of the keyboard can be accurately photographed.
[0207] (2) Perform pixel-level manual annotation on all samples of the keyboard images and the keyboard defect image samples generated by using data augmentation technology, and count the defect categories as: stains, double images, blind keys, and reverse keys.
[0208] (3) By using the structure of the teacher-student network, extracting image features with the encoder, concatenating the feature vectors, and then inputting them into the fully connected layer, the training of the entire network is completed under the guidance of the Loss function based on cross-entropy and the InfoNce Loss.
[0209] In practical applications, since the types of defects are not sufficient, data augmentation can be performed in the following form. Through image processing, we can artificially generate more defect images of the same type by adding these defects to normal keycap images, thereby increasing the training data. And their colors are all random, so the generated processing defect images have sufficient diversity, and the generation methods of other defects are similar.
[0210] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same program. Therefore, the systems, devices, and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the methods or the structures within the hardware component.
[0211] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be arbitrarily combined with each other.
Claims
1. A semi-supervised industrial defect detection method based on feature comparison, characterized in that Including: Step S1: Collect pictures of products to be tested and randomly label the pictures with tags; Step S2: Classify the products to be tested into labeled input and unlabeled input; Step S3: For the labeled input, use the pictures and the corresponding tags to train the student network; For the unlabeled input, input it into the teacher network to generate corresponding pseudo-labels and representations; Step S4: Screen the pseudo-labels to distinguish reliable pixels and unreliable pixels; Step S5: Feed the reliable pixels into the student network for supervision; for the unreliable pixels, according to their feature encoding information, perform feature optimization on the student network based on contrastive learning; In the said Step S5: The step of performing feature optimization on the student network based on contrastive learning according to the feature encoding information for the low-confidence labels is as follows: Step S5.1: Sort the predicted probability distributions of the unreliable pixels according to the probability values; Step S5.2: For the reliable pixels of each category, when this category does not appear in the top three categories of the probability ranking of the unreliable pixels, calculate and optimize the loss of feature contrast.
2. The semi-supervised industrial defect detection method based on feature comparison according to claim 1, wherein In the said Step S3: The steps of training the student network with the pictures and the corresponding tags for the labeled input are as follows: Step S3.1: Perform data augmentation on the input pictures; Step S3.2: Use the input pictures after data augmentation to train a teacher network, and then the teacher network does not perform gradient update; Step S3.3: Use the same input pictures in Step S3.2 to train a student network; Step S3.4: Based on cross-entropy, train the teacher network and the student network.
3. The semi-supervised industrial defect detection method based on feature contrast according to claim 2, characterized in that: The said cross-entropy is used to measure the difference in semantic information between the input image and the true labeled tags, and its calculation process is as follows: where y i is the difference in semantic information in the i-th eigenvector; x i is the i-th value of the eigenvector, and x j is the j-th value of the eigenvector. First, it is normalized through the softmax function to convert the values of each dimension of the eigenvector into a probability form, and then its cross-entropy is obtained; H y′ (y) is the cross-entropy; y′ i is the ideal result, the correct label vector.
4. The semi-supervised industrial defect detection method based on feature comparison according to claim 1, wherein In the said Step S4: Use the entropy value to screen the pseudo-labels at the pixel level, and the specific steps are as follows: Step S4.1: For the predicted probability distribution of each pixel, the entropy value is calculated according to the following formula: Entropy(p)=-p i log(p i ) where Entropy(p) is the entropy value, p is the pixel to be calculated, and p i is the probability that the pixel p belongs to the category i; Step S4.2: Calculate the entropy values of all pixels. The pixels with entropy values ranked in the latter 50% are regarded as reliable pixels, and the pixels with entropy values ranked in the former 50% are regarded as unreliable pixels; Step S4.3: Regard the reliable pixels and their pseudo-labels as labeled input and use them as the supervision information of the student network, and the loss function is cross-entropy.
5. The semi-supervised industrial defect detection method based on feature contrast according to claim 1, characterized in that In the said Step S5.2: Calculate and optimize the loss of feature contrast, and its calculation process is as follows: Among them, C is the probability category, M represents the position information on the image, and z i represents the representation output by the teacher network at the corresponding position, τ is a preset temperature coefficient, and N is the preset number of negative samples. is the representation of the corresponding positive sample, is the representation of the corresponding negative sample, and <,> represents the operation of taking the inner product between vectors.
6. A semi-supervised industrial defect detection system based on feature comparison, characterized in that Including: Module M1: Collect pictures of products to be tested and randomly label the pictures with tags; Module M2: Classify the products to be tested into labeled input and unlabeled input; Module M3: For the labeled input, use the pictures and the corresponding tags to train the student network; For the unlabeled input, input it into the teacher network to generate corresponding pseudo-labels and representations; Module M4: Screen the pseudo-labels to distinguish reliable pixels and unreliable pixels; Module M5: For reliable pixels, send them into the student network for supervision; for unreliable pixels, according to their feature encoding information, optimize the features of the student network based on contrastive learning; In the said module M5: The steps of optimizing the features of the student network based on contrastive learning according to the feature encoding information for low-confidence labels are as follows: Module M5.1: Sort the predicted probability distributions of unreliable pixels according to the probability values; Module M5.2: For reliable pixels of each category, when this category does not appear among the top three categories with the highest probabilities of unreliable pixels, calculate and optimize the loss of feature contrast.
7. The semi-supervised industrial defect detection system based on feature comparison according to claim 6, wherein In the said module M3: The steps of training the student network with labeled input images and corresponding labels are as follows: Module M3.1: Augment the input images; Module M3.2: Train a teacher network with the augmented input images, and then the teacher network does not update the gradients; Module M3.3: Train a student network with the same input images as in Module M3.2; Module M3.4: Train the teacher network and the student network based on cross-entropy.
8. The semi-supervised industrial defect detection system based on feature contrast according to claim 7, characterized in that: The said cross-entropy is used to measure the difference in semantic information between the input image and the true labeled label, and its calculation process is as follows: Among them, y i is the difference in semantic information in the i-th eigenvector; x i is the i-th value of the eigenvector, and x j is the j-th value of the eigenvector. First, it is normalized through the softmax function to convert the values of each dimension of the eigenvector into a probability form, and then its cross-entropy is obtained; H y′ (y) is the cross-entropy; y′ i is the ideal result, the correct label vector.
9. The semi-supervised industrial defect detection system based on feature comparison according to claim 6, wherein In the said module M4: Use the entropy value to perform pixel-level screening on the pseudo-labels, and the specific steps are as follows: Module M4.1: For the predicted probability distribution of each pixel, the entropy value is calculated according to the following formula: Entropy(p)=-p i log(p i ) where Entropy(p) is the entropy value, p is the pixel to be calculated, and p i is the probability that the pixel p belongs to class i; Module M4.2: Calculate the entropy values of all pixels. The pixels with entropy values ranked in the last 50% are regarded as reliable pixels, and the pixels with entropy values ranked in the top 50% are regarded as unreliable pixels; Module M4.3: Regard the reliable pixels and the pseudo-labels of reliable pixels as labeled inputs and use them as the supervision information of the student network. The loss function is cross-entropy.
10. The semi-supervised industrial defect detection system based on feature contrast according to claim 6, characterized in that In the said module M5.2: Calculate and optimize the loss of feature contrast, and its calculation process is as follows: Among them, C is the probability category, M represents the position information on the image, and z i represents the representation output by the teacher network at the corresponding position, τ is a preset temperature coefficient, and N is the preset number of negative samples. is the representation of the corresponding positive sample, is the representation of the corresponding negative sample, and <,> represents the operation of taking the inner product between vectors.
Citation Information
Patent Citations
Semi-supervised industrial product flaw detection method and system based on positive sample learning
CN112801962A
Industrial equipment surface crack detection method and system based on semi-supervised semantic segmentation
CN113436169A