A semi-supervised deep learning defect detection method
By adopting a semi-supervised deep learning method in industrial manufacturing, the architecture of student convolutional neural network and teacher convolutional neural network is utilized, combined with the SNAM attention module, the problem of inefficient automatic defect detection in the existing technology is solved, high-precision defect detection is achieved, and fewer labeled samples are used.
Patent Information
- Application Number
- CN202210446071.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-04-26
AI Technical Summary
The prior art is inefficient in automatic defect detection in industrial manufacturing, and the reliability and accuracy of deep unsupervised learning methods are insufficient, which cannot effectively alleviate the lack of large number of labeled samples.
The semi-supervised deep learning defect detection method is adopted to initialize the student convolutional neural network and the teacher convolutional neural network, and use the architecture of Fixmatch and average teacher model, combined with the SNAM attention module, to train to achieve high-precision defect detection.
It achieves the accuracy similar or even better than supervised learning with fewer labeled samples, and improves the automatic detection efficiency of surface defects of industrial products.
Smart Images

Figure CN114998202B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of automatic defect detection, and relates to a semi-supervised deep learning defect detection method. Background Art
[0002] Industrial manufacturing requires high-precision automatic defect detection technology (ASI) to detect different types of defects on the surface of products, such as scratches, holes, pits, and protrusions on the steel surface. These defects will affect the performance and aesthetics of the products, causing considerable economic losses.
[0003] Manual detection methods are inefficient and costly in terms of manpower. The latest progress in deep learning has produced new methods that can automatically learn high-level features from training samples while classifying defects without the need to manually design a feature set.
[0004] Deep learning methods can be divided into supervised learning, unsupervised learning, and semi-supervised learning. Supervised learning methods are based on CNN and can achieve high-precision defects given a large amount of training data. However, supervised learning heavily relies on manpower to collect and label training samples. Through unsupervised or semi-supervised learning methods, the lack of a large number of labeled samples can be alleviated. Popular ASI deep unsupervised learning methods are deep autoencoders and generative adversarial networks (GANs). The disadvantage of unsupervised learning is that it is usually not as reliable or accurate as supervised learning. Semi-supervised learning combines the advantages of supervised learning and unsupervised learning, and can obtain accuracy similar to or even better than supervised learning, but uses fewer labeled samples. Summary of the Invention
[0005] To solve the above problems, the present invention provides a semi-supervised deep learning defect detection method, including the following steps:
[0006] S10, classifying the training sample data according to whether there is a label;
[0007] S20, initializing the weight parameter m of the student convolutional neural network Fs(m);
[0008] S30, initializing the teacher convolutional neural network parameter Ft(m) = Copy(Fs(m)), and both the teacher convolutional neural network and the parameter are copied from the student convolutional neural network;
[0009] S40, after obtaining the student convolutional neural network Fs(m) and the teacher convolutional neural network Ft(m) and initializing the weight parameters, training the student convolutional neural network and the teacher convolutional neural network;
[0010] S50, obtaining the trained student convolutional neural network and teacher convolutional neural network, and now the student convolutional neural network can be used for defect detection, the data to be detected is input into the student convolutional neural network, and the student convolutional neural network predicts whether it has defects or what type of defects it belongs to.
[0011] Preferably, S10 specifically includes dividing the training sample data into labeled data samples X = {(xb, yb): b(1,...., B1)} and unlabeled data samples U = {ub: b(1,...., B2)}, where xb is the image data of the labeled data sample, yb is its label data, ub is the image data of its unlabeled sample data, and setting the training batch Bi, Bi represents the number of the i-th batch, including B1 = 32 or B2 = 128.
[0012] Preferably, the S20 middle school student convolutional neural network Fs(m) adopts the resnet34 network architecture, and is combined with the SNAM attention module to enhance the network's ability to extract image features, that is, a SNAM attention module is added to the end of the resnet residual module, specifically, the image feature input passes through two convolutional layers Conv1 and Conv2 and then enters the SNAM attention module, and then the feature output is added to the input feature.
[0013] Preferably, the SNAM attention module uses a batch normalized scaling factor γ on the NAM attention mechanism to indicate the importance of the weight. The scaling factor γ measures the variance. The larger the variance, the richer the information contained in the channel, and the more important the channel information is. The specific formula is:
[0014]
[0015] Among them, B in is the image feature input, μB and σ 2 B are the mean and variance of mini-batch B, respectively; γ and β distributions represent trainable scale factors and shifts.
[0016] Preferably, the output feature Mc of the SNAM attention module is obtained by the following formula:
[0017] M C =sigmoid(Wγ(BN(F1))) (2)
[0018] Among them, F1 is the input feature, BN is the batch normalization layer, W is the weight vector, sigmoid is the sigmoid activation function, γ is the scale factor of each channel, and each channel is multiplied by a weight after batch normalization BN calculation. Among them, T is a hyperparameter that determines the degree of sharpening, C is the number of channels, and γi and γ j Represents the scaling factors of the i-th channel and the j-th channel respectively. After the input feature F1 passes through the BN layer, each channel is multiplied by a weight w i It is then input into the sigmoid activation function. At this point, the SNAM attention module is calculated.
[0019] Preferably, the network weight of the teacher convolutional neural network training at time t in S30 is:
[0020] θ t ^=αθ t-1 ^+(1-α)θ t (3)
[0021] Among them, α is the coefficient, θ t-1 ^ is the weight of the teacher convolutional neural network at time t-1, θ t is the weight of the student convolutional neural network at time t, θ t ^ is the weight of the teacher convolutional neural network at time t.
[0022] Preferably, the training in S40 is performed by inputting the training data into the neural network in batches of data, that is, B1 labeled data and B2 unlabeled data are input into the neural network during one training process.
[0023] Preferably, the current description defined in S40 is a training process at time t, in which the fixmatch data enhancement method is used, that is, a weak enhancement data augmentation method is adopted for labeled data and pseudo labels for unlabeled data, and a strong enhancement data augmentation method is adopted for calculating the predicted value of unlabeled data.
[0024] Preferably, the weak enhancement data augmentation method includes performing a random horizontal flip or random cropping operation on the image with a probability of 50%.
[0025] Preferably, the strongly enhanced data augmentation method includes: providing a series of conversion functions, the conversion functions including color inversion, translation, contrast adjustment, rotation, adjusting image sharpness, blurring the image, adjusting image smoothness, overexposure or cropping, and the data randomly selects two conversion functions from them.
[0026] The beneficial effects of the present invention include at least: the present invention adopts semi-supervised learning to combine the advantages of supervised learning and unsupervised learning, and can obtain accuracy similar to or even better than supervised learning, but using fewer labeled samples.
[0027] The defect detection method of the present invention based on the semi-supervised deep learning architecture Fixmatch and the average teacher model can realize high-precision automatic detection of surface defects of industrial products with only a small amount of labeled data. Description of the Drawings
[0028] Figure 1 It is a flowchart of the steps of the semi-supervised deep learning defect detection method of the present invention;
[0029] Figure 2 It is a schematic diagram of the residual module and SNAM attention module of the semi-supervised deep learning defect detection method of the present invention;
[0030] Figure 3 It is a schematic diagram of the principle of the SNAM attention module of the semi-supervised deep learning defect detection method of the present invention;
[0031] Figure 4 It is a flowchart of the training of the convolutional network of the semi-supervised deep learning defect detection method of the present invention. Detailed Embodiments
[0032] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0033] On the contrary, the present invention covers any alternatives, modifications, equivalent methods and solutions made within the spirit and scope of the present invention as defined by the claims. Further, in order to enable the public to have a better understanding of the present invention, some specific details are described in detail in the following detailed description of the present invention. Those skilled in the art can fully understand the present invention without the description of these details.
[0034] See Figure 1 , which is a schematic diagram of the technical solution of the present invention in the embodiment of the present invention. The technical solution of the present invention is a semi-supervised deep learning defect detection method, including the following steps:
[0035] S10. Classify the training sample data according to whether there is a label;
[0036] S20. Initialize the weight parameter m of the student convolutional neural network Fs(m);
[0037] S30. Initialize the teacher convolutional neural network parameter Ft(m) = Copy(Fs(m)), and the teacher convolutional neural network and parameters are copied from the student convolutional neural network;
[0038] S40. After obtaining the student convolutional neural network Fs(m), the teacher convolutional neural network Ft(m) and the initialized weight parameters, train the student convolutional neural network and the teacher convolutional neural network;
[0039] S50, obtaining the trained student convolutional neural network and teacher convolutional neural network, and now the student convolutional neural network can be used for defect detection, the data to be detected is input into the student convolutional neural network, and the student convolutional neural network predicts whether it has defects or what type of defects it belongs to.
[0040] S10 specifically includes dividing all labeled data into a test data set T = {(xb, yb): b(1, ...., B0)} and a labeled training data set X = {(xb, yb): b(1, ...., B1)} in a ratio of 3:7, and all unlabeled data into an unlabeled training data set U = {ub: b(1, ...., B2)}, where the data of the test data set T and the labeled training data set X both contain image data xb and label data yb, and the unlabeled training data set U only contains image data ub without label data. Set the training batch Bi, where Bi represents the number of the i-th batch, including B1 = 32 or B2 = 128, and B0 is the number of all samples in the test data set.
[0041] The S20 middle school student convolutional neural network Fs(m) adopts the resnet34 network architecture and combines the SNAM attention module to improve the network's ability to extract image features. That is, the SNAM attention module is added at the end of the resnet residual module. Specifically, the image feature input enters the SNAM attention module after passing through two convolutional layers Conv1 and Conv2, and then the feature output is added to the input feature. Figure 2 .
[0042] The SNAM attention module uses the batch normalized scaling factor γ on the NAM attention mechanism to indicate the importance of the weight. The scaling factor γ measures the variance. The larger the variance, the richer the information contained in the channel, and the more important the channel information is. The specific formula is:
[0043]
[0044] Among them, B in is the image feature input, μB and σ 2 B are the mean and variance of mini-batch B, respectively; γ and β distributions represent trainable scale factors and shifts.
[0045] See also Figure 3 , the output feature Mc of the SNAM attention module is obtained as follows:
[0046] M C =sigmoid(Wγ(BN(F1))) (2)
[0047] Among them, F1 is the input feature, BN is the batch normalization layer, W is the weight vector, sigmoid is the sigmoid activation function, γ is the scale factor of each channel, and each channel is multiplied by a weight after batch normalization BN calculation. Among them, T is a hyperparameter that determines the degree of sharpening, C is the number of channels, and γ i and γ j Represents the scaling factors of the i-th channel and the j-th channel respectively. After the input feature F1 passes through the BN layer, each channel is multiplied by a weight w i It is then input into the sigmoid activation function. At this point, the SNAM attention module is calculated.
[0048] The SNAM attention module of the present invention improves the weight calculation formula of the existing NAM. SNAM draws on the sharpening function and proposes a new weight calculation formula (where T is the temperature T, C is the number of channels, i and j represent the i-th channel and j-th channel respectively). This formula can sharpen the data of a vector, that is, larger data becomes larger, smaller data becomes smaller, and the gap between data can be magnified. In the existing NAM method, each channel information is multiplied by a weight, and the weight calculation formula is The weight w i is a number less than 1, so that each channel information is multiplied by the weight to suppress the insignificant information. In SNAM, the weight w i Some are greater than 1, some are less than 1, for the proportional factor γ i The larger the channel, the i is greater than 1. Similarly, for channels with smaller proportional factors, its w i is less than 0. SNAM can highlight salient information and suppress insignificant information.
[0049] The teacher convolutional neural network in S30 is based on the average teacher model. The teacher model and the student model use the same network architecture, where the teacher's network weight is the exponential moving average EMA of the student convolutional neural network weight. Specifically, the network weight of the teacher convolutional neural network at time t is:
[0050] θ t ^=αθ t-1 ^+(1-α)θ t (3)
[0051] Among them, α is the coefficient, θ t-1 ^ is the weight of the teacher convolutional neural network at time t-1, θ t is the weight of the student convolutional neural network at time t, θ t^ is the weight of the teacher convolutional neural network at time t.
[0052] In S40, the training takes the batch data as the unit and inputs the training data into the neural network for training, that is, in one training process, B1 labeled data and B2 unlabeled data are input into the neural network.
[0053] In S40, it is defined that the current description is the training process at time t, where the data augmentation method of fixmatch is used, that is, the weak augmentation data augmentation method is adopted for the labeled data and the pseudo-labels of the unlabeled data are calculated, and the strong augmentation data augmentation method is adopted for the predicted values of the unlabeled data.
[0054] The weak augmentation data augmentation method includes performing random horizontal flipping or random cropping operations on the pictures with a probability of 50%. Both of these operations can increase the data quantity and alleviate the overfitting problem in the neural network training process, but will not cause serious distortion of the pictures.
[0055] The strong augmentation data augmentation method includes: giving a series of transformation functions, and the transformation functions include color inversion, translation, contrast adjustment, rotation, adjustment of image sharpness, blurring of the image, adjustment of image smoothness, overexposure or cropping, and two transformation functions are randomly selected from them. The strong augmentation operation can effectively expand the picture data and alleviate the overfitting problem in the neural network training process. In addition, the picture data often has a large difference from the original picture after these transformations, so it is called strong augmentation.
[0056] S40 specifically includes:
[0057] S41, X_a = a(xb)
[0058] First, perform weak augmentation data augmentation on the picture data xb of the labeled sample X. Define a() as the weak augmentation data augmentation, and the detailed description of the weak augmentation data augmentation refers to the foregoing description. X_a is the picture data of the labeled data sample after weak augmentation data augmentation.
[0059] S42, Predict_X_a = Fs_t-1(X_a)
[0060] The labeled data X_a after weak augmentation data augmentation is input into the student convolutional neural network to calculate the predicted value Predict_X_a. See Figure 4 , the student convolutional neural network is Fs_t-1(m), where t - 1 indicates that the weight parameters of the current network are learned at time t - 1.
[0061] S43, Loss_X = H(yb, Predict_X_a)
[0062] Calculate the loss value of the labeled data, that is, use the cross-entropy function H to calculate the cross-entropy between the predicted value of the labeled data and the label yb of the labeled data, and obtain the loss value of the labeled data. The specific calculation formula of the cross-entropy H is p(x) and q(x) represent the true probability distribution and the predicted probability distribution respectively.
[0063] S44, U_A = A(ub)
[0064] Perform strong data augmentation on the image data ub of the unlabeled samples. For the details of strong data augmentation, refer to the previous description. Define A() as strong data augmentation, and U_A as the image data of the unlabeled samples after strong data augmentation.
[0065] S45, Predict_U_A = Fs_t-1(U_A)
[0066] The unlabeled data after strong data augmentation is input into the student convolutional neural network, and its predicted value is calculated through the student convolutional neural network. Predict_U_A is its predicted value.
[0067] S46, U_a = a(ub)
[0068] Perform weak data augmentation on the image data ub of the unlabeled samples. For the details of weak data augmentation, refer to the previous description. U_a is the image data of the unlabeled samples after weak data augmentation.
[0069] S47, Predict_U_a = Ft_t-1(U_a)
[0070] Input the unlabeled data after weak data augmentation into the teacher convolutional neural network, and calculate its predicted value through the teacher convolutional neural network.
[0071] S48, If Predict_U_a > T:
[0072] Pseudo_label = Predict_U_a
[0073] After S47, the predicted value Predict_U_a of the teacher convolutional neural network for the weakly augmented unlabeled data is obtained. If its predicted value is higher than the threshold T, then retain the predicted value as the pseudo-label of the unlabeled data; if its maximum value does not exceed the threshold T, then discard this sample in this training and do not allow this sample to participate in this training.
[0074] We set the threshold T to 0.95. If the maximum value exceeds the threshold T, then we consider that the prediction result of this sample is consistent with its true label. We will retain this sample to participate in the calculation of the unlabeled data loss term, and use the prediction of the teacher convolutional neural network as the pseudo-label of this sample.
[0075] S49, Loss_U = H(Pseudo_label, Predict_U_A)
[0076] The predicted value Predict_U_A of the unlabeled data obtained through strong data augmentation is obtained by S45, and the pseudo-label Pseudo_label of the unlabeled data obtained by S48 is obtained. The cross-entropy of the predicted value and the pseudo-label of the unlabeled data is calculated using the cross-entropy loss function as the loss value Loss_U of the unlabeled data.
[0077] S410, Loss_X + Loss_U
[0078] Add the labeled data loss value and the unlabeled data loss value as the total loss value.
[0079] S411, Update the weight parameter m of the student convolutional neural network Fs_t-1(m) using Stochastic Gradient Descent (SGD).
[0080] Stochastic Gradient Descent (SGD) is to apply the gradient descent algorithm on a batch of samples to calculate the mean of their gradients to obtain the gradient, and then update the neural network weight parameters.
[0081] S412, After updating the weights of the student convolutional neural network in S411, use the updated weight parameter of the student convolutional neural network Fs_t(m) to update the teacher convolutional neural network Ft_t-1(m). The update method of the teacher convolutional neural network refers to S30.
[0082] S413, The training of the weight parameters of the student convolutional neural network and the teacher convolutional neural network at time t is completed through S411 and S412 respectively. Directly input all the image data of the test dataset T into the updated student convolutional neural network for prediction, and compare the prediction result with its label to calculate the accuracy. The test formula for accuracy is as follows:
[0083]
[0084] ACC represents the accuracy, T represents the number of samples where the predicted value of the student convolutional neural network is consistent with the image label, and B0 represents the number of all test datasets.
[0085] Repeat S40 to train the student convolutional neural network and the teacher network and calculate the accuracy of the test data set. When the accuracy of the test data set can meet the requirement, such as 100%, the training of the student convolutional neural network and the teacher convolutional neural network is completed.
[0086] S50, after S40, we obtain the trained student convolutional neural network and teacher convolutional neural network. At this point, the student convolutional neural network can be used for defect detection. The data to be detected is input into the student convolutional neural network, and the student convolutional neural network can predict whether it has defects or what type of defects it belongs to.
[0087] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A semi-supervised deep learning defect detection method, characterized in that, The following steps are involved: S10, classifying the training sample data according to whether they have labels or not; S20, initialize the weight parameter m of the student convolutional neural network Fs(m); S30, initialize the teacher convolutional neural network parameters Ft(m)=Copy(Fs(m)), the teacher convolutional neural network and parameters are copied from the student convolutional neural network; S40, after obtaining the student convolutional neural network Fs(m) and the teacher convolutional neural network Ft(m) and the initialization weight parameters, the student convolutional neural network and the teacher convolutional neural network are trained; S50, obtaining a trained student convolutional neural network and a teacher convolutional neural network, so that the student convolutional neural network can be used to perform defect detection, inputting the image data to be detected into the student convolutional neural network, and the student convolutional neural network predicts whether it has defects or which type of defects it belongs to; The S10 specifically includes dividing the training sample data into labeled data samples X={(xb,yb):b (1,....,B1)} and unlabeled data samples U={ub:b (1,....,B2)}, where xb is the image data of the labeled data sample, yb is its label data, and ub is the image data of its unlabeled sample data, and setting a training batch Bi, where Bi represents the number of the i-th batch, including B1=32 or B2=128; The S20 middle school student convolutional neural network Fs(m) adopts the resnet34 network architecture and combines the SNAM attention module to improve the network's ability to extract image features, that is, the SNAM attention module is added at the end of the residual module of resnet. Specifically, the image feature input enters the SNAM attention module after passing through two convolutional layers Conv1 and Conv2, and then the feature output is added to the input feature; The SNAM attention module uses the scale factor of batch normalization in the NAM attention mechanism to represent the importance of weights, and the scale factor measures the variance. The larger the variance, the richer the information contained in the channel, and the more important the channel information is. The specific formula is as follows: (1) Among them, is the input of image features, and are the mean and variance of the mini-batch B respectively; γ and β The distribution represents the trainable scale factor and displacement; The output features of the SNAM attention module are obtained by the following formula: (2) in, F 1 are input features, BN It is the batch normalization layer, i.e., the BN layer. W is the weight vector, sigmoid is the sigmoid activation function, γ is the scaling factor for each channel. Each channel is calculated by batch normalization BN and then multiplied by a weight, whose weight is , where T is a hyperparameter that determines the degree of sharpening, and C is the number of channels. γ i and γ j Represent the scale factors of the i-th channel and the j-th channel respectively, and the input features F 1 After the BN layer, each channel is multiplied by a weight w i Then it is input into the sigmoid activation function. At this point, the SNAM attention module is calculated. The training in S40 is to input the training data into the neural network in batches of data, that is, B1 labeled data and B2 unlabeled data are input into the neural network in one training process; The definition in S40 is currently describing the training process at time t, in which the fixmatch data enhancement method is used, that is, a weak enhancement data augmentation method is adopted to calculate the predicted value of the labeled data and the pseudo label of the unlabeled data, and a strong enhancement data augmentation method is adopted to calculate the predicted value of the unlabeled data.
2. The semi-supervised deep learning defect detection method according to claim 1, characterized in that, The network weight of the teacher convolutional neural network training at time t in S30 is: (3) Among them, is a coefficient, is the weight of the teacher convolutional neural network at time t-1, is the weight of the student convolutional neural network at time t, is the weight of the teacher convolutional neural network at time t.
3. The semi-supervised deep learning defect detection method according to claim 1, characterized in that, The weak enhancement data augmentation method includes performing a random horizontal flip or a random cropping operation on the image with a probability of 50%.
4. The semi-supervised deep learning defect detection method according to claim 3, characterized in that, The strong enhancement data augmentation method includes: providing a series of conversion functions, the conversion functions including color inversion, translation, contrast adjustment, rotation, adjusting image sharpness, blurring the image, adjusting image smoothness, overexposure or cropping, and randomly selecting two conversion functions from them.