Industrial defect detection method based on distillation network
By adopting an attention-based distillation network method in industrial defect detection, using channel and spatial attention to enhance information transmission, and optimizing model parameters through unsupervised training and hard feature loss, the detection accuracy problems caused by the lack of defect samples in the prior art are solved, and a more efficient defect detection effect is achieved.
Patent Information
- Application Number
- CN202510148633.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
The existing industrial defect detection methods lack defect samples and cannot conduct sufficient supervised training, which leads to the inability of student networks to effectively learn the ability of teachers networks during distillation, and thus cannot accurately judge defect samples.
The attention-based distillation network method is adopted to build industrial defect detection models of teacher network, CBAM module, student network and autoencoder module, and use channel attention and spatial attention to enhance the information transmission from teacher network to student network, and optimize model parameters through unsupervised training and hard feature loss.
It improves the accuracy and ability of students' networks in defect detection, enhances attention to important parts of teachers' network characteristics, and improves the accuracy of defect detection without increasing excessive reasoning overhead.
Smart Images

Figure CN120070990A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial defect detection, and specifically to a defect detection method based on an attention-based distillation network. Background Art
[0002] In recent years, the field of industrial defect detection has attracted extensive attention and gradually become a research hotspot. The main problem in the field of industrial defect detection is the lack of defect samples, making it impossible to conduct sufficient supervised training. Therefore, unsupervised methods that only use normal samples for training have become the main research direction.
[0003] One method is to generate some defect samples, which can be achieved by image splicing or using a generative model to generate some defective images. However, the forged defect samples can never simulate the true defect sample distribution.
[0004] Another method is to use distillation. The teacher network uses a pre-trained model and freezes its parameters. The student network is trained through distillation and only uses normal samples for training. In this way, the student network only learns the distribution of normal samples. During testing, when encountering a defect sample, there will be a large difference between the student network and the teacher network, and thus the defect sample can be identified. In most distillation-based methods, the student network often fails to learn the capabilities of the teacher network well, resulting in an inability to accurately identify defect samples. Summary of the Invention
[0005] The purpose of the present invention is to propose an industrial defect detection method based on a distillation network for the deficiencies existing in the current field of industrial defect detection.
[0006] The specific technical solution adopted by the present invention is as follows:
[0007] An industrial defect detection method based on a distillation network, which includes:
[0008] S1. Obtain an industrial defect detection data set composed of normal sample images without industrial defects;
[0009] S2. Construct an industrial defect detection model composed of a teacher network, a CBAM (Convolutional Block Attention Module) module, a student network, and an AutoEncoder module. The teacher network is a pre-trained model that can be used for industrial defect detection and its model parameters are fixed. The student network has the same network structure as the teacher network;
[0010] S3. After initializing the network parameters of the student network and the autoencoder module as learnable parameters, perform unsupervised training on the industrial defect detection model by distillation until the iteration termination condition is reached. During the unsupervised training process, input the normal sample images into the teacher network, the student network, and the autoencoder module. First, the teacher network extracts feature maps at different levels from the normal sample images. Then, each level of feature map extracted by the teacher network is input into the CBAM module, processed through channel attention and spatial attention respectively, and then injected into the corresponding network layer of the student network to be fused with the feature maps at different levels extracted by the student network itself from the normal sample images. The fusion result serves as the feature map output by the corresponding network layer of the student network. Then, the autoencoder module extracts feature maps from the normal sample images. Finally, calculate the hard feature loss pairwise between the last layer of feature maps extracted by the teacher network, the last layer of feature maps output by the student network, and the feature maps extracted by the autoencoder module, and sum them up as the total loss. Based on the total loss value, update the learnable parameters in the industrial defect detection model in the reverse direction.
[0011] S4. Input the industrial defect image to be detected into the industrial defect detection model that has completed unsupervised training, obtain the last layer of feature maps extracted by the teacher network and the last layer of feature maps output by the student network, and detect whether there are industrial defects in the industrial defect image by comparing the feature distances at different positions of the images.
[0012] Preferably, both the teacher network and the student network adopt a PDN (patch description network) network composed of four convolutional layers.
[0013] Preferably, the input of the PDN network is an image patch of 33×33. The four convolutional layers are 4×4 convolutional layer, 4×4 convolutional layer, 3×3 convolutional layer, and 4×4 convolutional layer in sequence, and there are 2×2 average pooling operations after the first two 4×4 convolutional layers.
[0014] Preferably, the teacher network is obtained by distilling the PDN network based on WideResNet-101.
[0015] Preferably, in the student network, the input of any i-th network layer is the feature map output by the (i - 1)-th layer in the student network and the result of the feature map extracted by the i-th network layer in the teacher network after passing through the CBAM module. The feature map output by any i-th network layer is where Layer i-1() represents the (i - 1)-th network layer in the student network, and CA() and SA() represent channel attention and spatial attention in the CBAM module respectively.
[0016] Preferably, the iteration termination condition is that the current iteration number reaches a preset maximum iteration number.
[0017] Preferably, the industrial defect detection dataset adopts the Mvtec or VisA dataset.
[0018] Preferably, in S4, it is detected whether there are industrial defects in the industrial defect image. If the feature distance of a certain image patch in the two feature maps exceeds the threshold, it is considered that there is an industrial defect at this image patch.
[0019] Compared with the prior art, the present invention has the following beneficial effects:
[0020] In view of the deficiencies in the existing industrial defect detection field, the present invention proposes a defect detection method based on an attention-based distillation network. On the basis of the traditional teacher-student distillation model, channel attention and spatial attention are added to strengthen the information transmission from the teacher network to the student network, so that the student network can pay more attention to the important parts of the features of the teacher network during the distillation process. In addition to attention, the entire defect detection model architecture includes three main modules: the teacher network, the student network, and the AutoEncoder module. The teacher network and the student network are used to solve visual defect problems, while the teacher network and the AutoEncoder are used to solve logical defect problems. The method of the present invention can be added as a plug-in to the industrial defect detection algorithm based on the teacher-student network to improve the accuracy of defect detection without generating excessive inference overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic diagram of the steps of an industrial defect detection method based on a distillation network;
[0022] Figure 2 It is a schematic diagram of the framework for unsupervised training of an industrial defect detection model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific implementation manners of the present invention in conjunction with the accompanying drawings. Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in each embodiment of the present invention can be combined correspondingly without conflict.
[0024] As Figure 1 shown, in a preferred embodiment of the present invention, an industrial defect detection method based on a distillation network is provided, which includes four steps S1 to S4. The following separately introduces the specific implementation manners of each step.
[0025] S1. Obtain an industrial defect detection data set composed of normal sample images without industrial defects.
[0026] It should be noted that in the scenario of industrial defect detection, there are generally a large number of normal sample data sets, but the defect samples are lacking. Therefore, the present invention adopts an unsupervised method, and only normal samples need to be added to the training set for training. In the embodiments of the present invention, the industrial defect detection data set used adopts the Mvtec and VisA data sets. In these two data sets, there are only normal sample images without industrial defects in the training set, but there are some defect sample images in the test set. Therefore, during the training process, only the normal sample images without industrial defects are used as the training set for training, and the test set is used to test the model performance subsequently.
[0027] S2. Construct an industrial defect detection model composed of a teacher network, a CBAM (Convolutional Block Attention Module) module, a student network, and an AutoEncoder module. The teacher network is a pre-trained model that can be used for industrial defect detection and the model parameters are fixed. The student network has the same network structure as the teacher network.
[0028] In an embodiment of the present invention, referring to the practice in EfficientAD, both the teacher network and the student network adopt a PDN (patch description network) composed of four convolutional layers. The input of this PDN network is an image patch of 33×33. The four convolutional layers are a 4×4 convolutional layer, a 4×4 convolutional layer, a 3×3 convolutional layer, and a 4×4 convolutional layer in sequence. And there is a 2×2 average pooling operation after the first two 4×4 convolutional layers. The network parameters (stride, kernel size, number of kernels, padding, activation function) of the specific four convolutional layers and two pooling layers, namely Conv-1, AvgPool-1, Conv-2, AvgPool-2, Conv-3, and Conv-4, are as follows:
[0029]
[0030] In addition, the above teacher network is a pre-trained model, which is obtained by distilling the PDN network based on WideResNet-101. The dataset used is ImageNet, and the loss is MSE. When performing distillation training with the student model, the model parameters are fixed. The model structure of the above student network is the same as that of the teacher network, but its model parameters need to be initialized to facilitate re-learning during subsequent training.
[0031] S3. After initializing the network parameters of the student network and the autoencoder module as learnable parameters, the industrial defect detection model is trained unsupervised by distillation until the iteration termination condition is reached.
[0032] Figure 2 The framework principle of the above unsupervised training process is shown. During the unsupervised training process, the input normal sample images are sent into the teacher network, the student network, and the autoencoder module. First, the teacher network extracts feature maps at different levels from the normal sample images. Then, each level of feature map extracted by the teacher network is input into the CBAM module, and after being processed by channel attention and spatial attention respectively, it is injected into the corresponding network layer of the student network and fused with the feature maps at different levels extracted by the student network itself from the normal sample images. The fusion result is used as the feature map output by the corresponding network layer of the student network. Then, the autoencoder module extracts feature maps from the normal sample images. Finally, the hard feature loss is calculated pairwise for the last layer of feature map extracted by the teacher network, the last layer of feature map output by the student network, and the feature map extracted by the autoencoder module, and the sum is used as the total loss. Based on the total loss value, the learnable parameters in the industrial defect detection model are updated backward.
[0033] In an embodiment of the present invention, in order to represent the feature output process in the student network in a general way, for any i-th network layer in the student network, assume that its input is the feature map output by the (i - 1)-th layer in the student network and the feature map extracted by the i-th network layer in the teacher network After passing through the CBAM module, the feature map output by any i-th network layer is In the formula, Layer i-1 () represents the (i - 1)-th network layer in the student network, CA() and SA() respectively represent the channel attention and spatial attention in the CBAM module, and when i = 1, represents the input normal sample image
[0034] Thus, based on the above PDN network composed of four convolutional layers, the training process of this model framework can be expressed as follows:
[0035] Step 1: Set the training parameters, the learning rate R = 0.0001, and the maximum number of iterations G = 40000
[0036] Step 2: Import the pre-trained weights into the teacher network and freeze the parameters
[0037] Step 3: Initialize the student network and the AutoEncoder
[0038] Step 4: Perform iterative training by distillation until the training iteration times reach the maximum number of iterations G. The overall training iteration process is as follows:
[0039] Step 4.1: In the training stage, input the normal sample image I. The image I passes through the teacher network to extract features, and the four convolutional layers respectively extract the corresponding level feature maps, so as to obtain the feature maps of each layer where respectively represent the feature maps extracted by the four convolutional layers from front to back
[0040] Step 4.2: Input the picture I and F t into the student network at the same time. The feature of each layer of the student network is
[0041]
[0042] After passing through the entire student network, the feature of the last layer is obtained which is the final feature map F s .
[0043] It should be noted that the specific calculation processes of the channel attention and spatial attention in the above CBAM module belong to the prior art of CBAM, and their calculation formulas can be expressed as follows:
[0044] CA(F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F)))
[0045] SA(F) = σ(f 7x7 ([AvgPool(F); MaxPool(F)]))
[0046] where f 7x7 represents a 7x7 convolutional operation, [a; b] represents concatenating a and b, σ() represents the Sigmoid activation function, AvgPool and MaxPool represent average pooling and max pooling respectively, and MLP is a multi-layer perceptron.
[0047] Step 4.3: Input the image I into the AutoEncoder to obtain the feature map F a .
[0048] Step 4.4: Calculate the loss function, whose form is:
[0049] Loss = Loss st + Loss sa + Loss ta
[0050] where Loss st represents and the hard feature loss between F s , Loss sa represents the hard feature loss between F s and F a , Loss ta represents and the hard feature loss between F a , and their calculation formulas are as follows:
[0051]
[0052] In the formula, MSE d-hard represents the hard feature loss calculation function with p_hard as the threshold. The hard feature loss is an existing loss function proposed in EfficientAD, where the threshold p_hard ∈ [0, 1] can be used as a quantile to extract the image patches participating in the total loss calculation. Denote the two feature maps for which the hard feature loss needs to be calculated as A and B, then MSE p_hardThe calculation method of (A - B) is as follows: First, pair the image patches at the corresponding positions in the two feature maps. For each pair of image patches i, calculate the squared error L between the corresponding local feature maps. i , and then sort them in ascending order according to the squared error L i , and determine the squared error value L at the quantile p_hard. Then, sum up the squared errors L in the sequence that are not less than L i as the calculation result. In this way, only the parts with large feature differences in the feature map can be extracted for gradient backpropagation. i
[0053] Step 4.5: Calculate the Loss gradient of the total loss and update the model parameters through backpropagation.
[0054] Step 4.6: Continue the iteration until the number of training iterations reaches the preset maximum number of iterations G, save the model parameters, and complete the unsupervised training.
[0055] S4. Input the industrial defect image to be detected into the industrial defect detection model that has completed unsupervised training, obtain the last-layer feature map extracted by the teacher network and the last-layer feature map output by the student network, and detect whether there are industrial defects in the industrial defect image by comparing the feature distances at different positions of the image.
[0056] It should be noted that to finally detect whether there are industrial defects in the industrial defect image, it is determined by comparing the feature distances between the last-layer feature map extracted by the teacher network and the last-layer feature map F s output by the student network. If the feature distance of a certain Patch in the two feature maps of the image exceeds a certain threshold, it is considered that there is an industrial defect in this Patch. The specific threshold can be optimized according to the actual dataset and is not limited.
[0057] Next, the industrial defect detection method based on the distillation network described in S1 - S4 above will be tested on a specific dataset to demonstrate its technical effects.
[0058] Embodiment
[0059] To prove the effectiveness of the present invention, experiments are respectively carried out on two datasets, Mvtec and VisA. At the same time, to compare the effect of adding the CBAM module to the present invention, an ablation experiment without the CBAM module is set up, denoted as EfficientAD for short. The test metrics are selected as Image AUROC and Pixel AUROC. Image AUROC can reflect the ability of the model in the classification task, and Pixel AUROC can reflect the ability of the model in the segmentation task. The overall situation is shown in Tables 1 and 2 as follows:
[0060] Table 1 Experimental Results of Mvtec Dataset
[0061] Image AUROC Pixel AUROC EfficientAD 99.1 97.5 The present invention 99.3 97.8
[0062] Table 2 Experimental Results of VisA Dataset
[0063] Image AUROC Pixel AUROC EfficientAD 98.1 98.4 EfficientAD+Attention 98.4 98.6
[0064] As can be seen from Table 1 and Table 2, compared with the traditional method of EfficientAD, in the Mvtec dataset, the Image AUROC of the present invention is increased by 0.2, and the Pixel AUROC is increased by 0.3; in the VisA dataset, the Image AUROC is increased by 0.3, and the Pixel AUROC is increased by 0.2.
[0065] In addition, the results of each category in Mvtec and VisA of the present invention are shown in Table 3 and Table 4.
[0066] Table 3 Results of Each Category in Mvtec
[0067]
[0068]
[0069] Table 4 Results of Each Category in VisA
[0070]
[0071] The above-described embodiments are only some preferred implementation solutions of the present invention, but are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. An industrial defect detection method based on a distillation network, characterized in that: include: S1. Obtain an industrial defect detection data set consisting of normal sample images without industrial defects; S2. Construct an industrial defect detection model consisting of a teacher network, a CBAM (Convolutional Block Attention Module) module, a student network and an AutoEncoder module, wherein the teacher network is a pre-trained model that can be used for industrial defect detection and the model parameters are fixed, and the student network has the same network structure as the teacher network; S3. Initialize the network parameters of the student network and the autoencoder module as learnable parameters, and perform unsupervised training on the industrial defect detection model by distillation until the iteration termination condition is reached; during the unsupervised training process, the input normal sample image is sent to the teacher network, the student network and the autoencoder module, and the teacher network first extracts feature maps of different levels from the normal sample image, and then each level feature map extracted by the teacher network is input into the CBAM module, and then injected into the network layer corresponding to the student network after being processed by channel attention and spatial attention respectively, and fused with the different levels of feature maps extracted by the student network from the normal sample image, and the fusion result is used as the feature map output by the corresponding network layer of the student network; then the autoencoder module extracts the feature map from the normal sample image; finally, the hard feature loss (Hard Feature Loss) is calculated for the last layer feature map extracted by the teacher network, the last layer feature map output by the student network, and the feature map extracted by the autoencoder module, and the sum is used as the total loss, and the learnable parameters in the industrial defect detection model are reversely updated based on the total loss value; S4. Input the industrial defect image to be detected into the industrial defect detection model that has completed unsupervised training, obtain the last layer feature map extracted by the teacher network and the last layer feature map output by the student network, and detect whether there are industrial defects in the industrial defect image by comparing the feature distances at different positions of the image.
2. The industrial defect detection method based on distillation network according to claim 1, characterized in that: The teacher network and the student network both adopt a PDN (patchd esciption network) network consisting of four layers of convolution.
3. The industrial defect detection method based on distillation network according to claim 2, characterized in that: The input of the PDN network is a 33×33 image block. The four convolutional layers are 4×4 convolutional layer, 4×4 convolutional layer, 3×3 convolutional layer, and 4×4 convolutional layer, and the first two 4×4 convolutional layers are followed by 2×2 average pooling operations.
4. The industrial defect detection method based on distillation network according to claim 2, characterized in that: The teacher network is obtained by distilling the PDN network based on WideResNet-101.
5. The industrial defect detection method based on distillation network according to claim 1, characterized in that: In the student network, the input of any i-th layer is the feature map output by the i-1th layer in the student network. And the feature map extracted by the i-th network layer in the teacher network After the CBAM module, the feature map output by any i-th network layer is In the formula, Layer i-1 () represents the i-1th network layer in the student network, CA() and SA() represent the channel attention and spatial attention in the CBAM module respectively.
6. The industrial defect detection method based on distillation network according to claim 1, characterized in that: The iteration termination condition is that the current iteration number reaches a preset maximum iteration number.
7. The industrial defect detection method based on distillation network according to claim 1, characterized in that: The industrial defect detection dataset uses the Mvtec or VisA dataset.
8. The industrial defect detection method based on distillation network according to claim 1, characterized in that: In the S4, it is detected whether there is an industrial defect in the industrial defect image. If the feature distance of an image block in the image between two feature images exceeds a threshold, it is considered that there is an industrial defect in the image block.
Citation Information
Patent Citations
Knowledge distillation-based positive sample industrial defect detection method
CN112991330A
Surface defect detection method based on multi-scale attention guidance and knowledge distillation
CN113947590A
Abnormality detection method based on combination of knowledge distillation and image reconstruction
CN115861256A
Steel surface defect detection method based on semi-supervised target detection algorithm
CN117726628A
Inverse distillation defect detection method based on position-aware cyclic convolution
CN118365616A