Knowledge distillation-based unsupervised industrial anomaly detection method and system
By constructing an unsupervised industrial anomaly detection method of abnormal synthesis network, teacher feature extraction network, variable noise denoising student network and pixel-level abnormal segmentation network, the problems of performance limitations caused by the lack of abnormal samples in the prior art and difficulty in capturing the global context structure of local features are solved, and high accuracy detection and positioning of surface defects of industrial products are achieved.
Patent Information
- Application Number
- CN202411883573.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The existing industrial anomaly detection methods have performance limitations in the absence of abnormal samples, and traditional convolutional neural networks are difficult to effectively capture the global context structure of local features, resulting in limited detection effects.
Unsupervised industrial anomaly detection method based on knowledge distillation is adopted to integrate the abnormality detection model by constructing anomaly synthesis network, teacher feature extraction network, variable noise denoising student network and pixel-level anomaly segmentation network. This method uses Perlin noise to generate simulated anomaly images, combines the denoising module of the teacher feature extraction network and the student network, a self-supervised reconstruction module and a pixel-level anomaly segmentation network to realize the detection and positioning of surface defects of industrial products.
It significantly improves the accuracy and robustness of industrial image abnormality detection, can more accurately identify and locate product surface defects, and achieve more reliable detection and identification of industrial product surface defects.
Smart Images

Figure CN119991555A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of visual detection technology, relates to a neural network model, and in particular to model construction, and specifically is an unsupervised industrial anomaly detection method and system based on knowledge distillation. Background Art
[0002] Unsupervised anomaly detection is an important task in the field of computer vision, which involves identifying and locating abnormal samples in data. This technology has a wide range of applications in industrial inspection, medical diagnosis and other fields. However, due to the scarcity and diversity of abnormal events, it is extremely challenging to obtain enough abnormal samples for supervised training. Therefore, most current methods focus on unsupervised learning and only use normal samples for training.
[0003] Existing industrial anomaly detection methods usually use knowledge distillation technology, in which a pre-trained convolutional neural network is used as a teacher network, and another student network with the same structure but randomly initialized is optimized during the training process to generate feature representations similar to those of the teacher network. Since the student network is only trained on normal data, it will produce significantly different feature representations from the teacher network when faced with abnormal data, thereby achieving anomaly detection. However, this method has some limitations. First, due to the lack of constraints on abnormal samples, the student network may over-generalize. Second, since traditional convolutional neural networks cannot recognize the global context structure of local features, the detection effect of the model may be limited when detecting more complex global structural anomalies.
[0004] In summary, how to reduce the performance limitations caused by the lack of abnormal samples in the process of industrial defect detection, and how to effectively capture the global structure of local features and make full use of them to improve detection accuracy have become technical problems to be solved by existing methods. Summary of the invention
[0005] The purpose of the present invention is to provide an unsupervised industrial anomaly detection method and system based on knowledge distillation, which is used to solve the problem of low efficiency of surface anomaly detection of industrial products in the prior art.
[0006] In a first aspect, the present invention provides an unsupervised industrial anomaly detection method based on knowledge distillation, the method comprising: obtaining an industrial product surface image, preprocessing the industrial product surface image to obtain a training set and a test set respectively; constructing an anomaly synthesis network, a teacher feature extraction network, a variable denoising student network and a pixel-level anomaly segmentation network, and integrating the anomaly synthesis network, the teacher feature extraction network, the variable denoising student network and the pixel-level anomaly segmentation network in sequence to form an anomaly detection model; training the anomaly detection model through the training set; verifying the trained anomaly detection model using the test set, and outputting an anomaly detection target model after verification that it meets the requirements; wherein the anomaly synthesis network is used to synthesize forged abnormal images, the teacher feature extraction network is used to extract multi-scale features of the image and perform feature guidance, the variable denoising student network is used to denoise the abnormal area of the image, and the pixel-level anomaly segmentation network is used to classify and locate anomalies in the image.
[0007] In an implementation of the first aspect, the preprocessing of the industrial product surface images to obtain a training set and a test set respectively includes: classifying the industrial product surface images according to image categories to obtain images of multiple categories; performing data enhancement processing on the industrial product surface images according to different categories to obtain enhanced images; selecting images without abnormalities in the enhanced images as training sets, and selecting the enhanced images with abnormalities and the corresponding abnormal mask images as test sets.
[0008] In an implementation of the first aspect, the anomaly synthesis network generates a random two-dimensional Perlin noise and binarizes it to create an anomaly mask, and mixes the non-abnormal image with the external data source image at random opacity to generate a simulated abnormal image.
[0009] In an implementation of the first aspect, the teacher feature extraction network includes a ResNet18 convolutional neural network pre-trained on a large natural image dataset, wherein the ResNet18 convolutional neural network has fixed weights and the final convolution module is removed.
[0010] In an implementation of the first aspect, the variable denoising student network includes a denoising module and a self-supervised reconstruction module. The denoising module is used to denoise the input image in the feature dimension to obtain a preliminary denoised feature map, and the self-supervised reconstruction module is used to further reconstruct the preliminary denoised feature map and generate discriminant features for features with abnormal contextual structure information.
[0011] In an implementation of the first aspect, the denoising module includes an encoder and a decoder, the encoder adopts a randomly initialized ResNet18 convolutional neural network, the decoder implements an inverse structure through bilinear upsampling, and the encoder and the decoder both include four residual units; the self-supervised reconstruction module includes a variable mask convolution and a channel attention module, the channel attention module uses a mask operation at the center of the convolution, and uses context information to predict or reconstruct the masked information, thereby forcing the model to learn the global context information of normal local features.
[0012] In an implementation manner of the first aspect, the pixel-level anomaly segmentation network includes two residual blocks and a dilated spatial pyramid pooling module, and the dilated spatial pyramid pooling module is used to adaptively focus on important feature channels and generate an anomaly segmentation map for anomaly detection.
[0013] In an implementation of the first aspect, the training of the anomaly detection model using the training set includes: keeping the weights in the anomaly detection model fixed, synthesizing a forged abnormal image through the anomaly synthesis network and inputting it into the deformable denoising student network, and inputting the normal image in the training set into the teacher feature extraction network to calculate a first loss function; keeping the deformable denoising student network fixed, synthesizing a forged abnormal image through the anomaly synthesis network and inputting it into the deformable denoising student network and the teacher feature extraction network respectively, so as to train the pixel-level anomaly segmentation network by comparing the feature map differences between the deformable denoising student network and the teacher feature extraction network to identify abnormal areas; calculating the focal loss and the L1 loss respectively according to the pixel-level anomaly segmentation network, and calculating the second loss function according to the focal loss and the L1 loss.
[0014] In a second aspect, the present invention also provides an unsupervised industrial anomaly detection method based on knowledge distillation, and performs surface defect detection on industrial products using the anomaly detection target model established by the above-mentioned unsupervised industrial anomaly detection method based on knowledge distillation.
[0015] In a third aspect, the present invention also discloses an unsupervised industrial anomaly detection system based on knowledge distillation, the method comprising: a preprocessing unit, used to obtain an industrial product surface image, preprocessing the industrial product surface image to obtain a training set and a test set respectively; an integration unit, used to construct an anomaly synthesis network, a teacher feature extraction network, a variable denoising student network and a pixel-level anomaly segmentation network, and sequentially integrate the anomaly synthesis network, the teacher feature extraction network, the variable denoising student network and the pixel-level anomaly segmentation network to form an anomaly detection model; a training unit, used to train the anomaly detection model through the training set; a verification unit, used to verify the trained anomaly detection model using the test set, and output an anomaly detection target model after verification that it meets the requirements; wherein the anomaly synthesis network is used to synthesize forged abnormal images, the teacher feature extraction network is used to extract multi-scale features of the image and perform feature guidance, the variable denoising student network is used to denoise the abnormal area of the image, and the pixel-level anomaly segmentation network is used to classify and locate anomalies in the image.
[0016] As described above, the unsupervised industrial anomaly detection method and system based on knowledge distillation described in the present invention have the following beneficial effects:
[0017] Compared with the prior art, the present invention significantly improves the accuracy and robustness of industrial image anomaly detection and positioning by introducing unsupervised knowledge distillation technology. By using Perlin noise and external data sets to generate simulated abnormal images to increase the abnormal sample data required for training, the teacher feature extraction network extracts the natural feature representation of the input image, while the denoising module of the student network generates anomaly-free feature representation. The self-supervised reconstruction module can capture the global context structure of local features and generate discriminative features when structural anomalies are found. Finally, anomaly detection and positioning are performed through a pixel-level anomaly segmentation network to obtain more accurate anomaly scores, thereby significantly improving the sensitivity and accuracy of industrial image anomaly detection and achieving more reliable detection, identification and positioning of product surface defects. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Shown is a flowchart of the unsupervised industrial anomaly detection method based on knowledge distillation according to an embodiment of the present invention.
[0019] Figure 2 Shown is a model diagram of the anomaly detection model described in an embodiment of the present invention.
[0020] Figure 3 Shown is a structural block diagram of the self-supervised reconstruction module according to an embodiment of the present invention.
[0021] Figure 4Shown is a structural block diagram of the unsupervised industrial anomaly detection system based on knowledge distillation according to an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0023] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention. The illustrations only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0024] See also Figures 1 to 4 . The present invention significantly improves the accuracy and robustness of industrial image anomaly detection and localization by introducing unsupervised knowledge distillation technology. By using Perlin noise and external data sets to generate simulated abnormal images to increase the abnormal sample data required for training, the teacher feature extraction network extracts the natural feature representation of the input image, while the denoising module of the student network generates anomaly-free feature representation. The self-supervised reconstruction module can capture the global context structure of local features and generate discriminative features when structural anomalies are found. Finally, anomaly detection and localization are performed through a pixel-level anomaly segmentation network to obtain more accurate anomaly scores, thereby significantly improving the sensitivity and accuracy of industrial image anomaly detection and achieving more reliable detection, identification and localization of product surface defects.
[0025] The technical solutions in the embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0026] like Figure 1 As shown, in one embodiment, the present invention provides an unsupervised industrial anomaly detection method based on knowledge distillation, the method comprising the following steps:
[0027] S100: Acquire surface images of industrial products, and preprocess the surface images of the industrial products to obtain a training set and a test set, respectively.
[0028] In one embodiment, the preprocessing of the industrial product surface images to obtain a training set and a test set respectively includes:
[0029] Classifying the surface image of the industrial product according to image categories to obtain images of multiple categories;
[0030] Performing data enhancement processing on the surface image of the industrial product according to different categories to obtain an enhanced image;
[0031] The images without abnormalities in the enhanced images are selected as training sets, and the enhanced images with abnormalities and the corresponding abnormal mask images are selected as test sets.
[0032] In this embodiment, in order to subsequently train the model, after obtaining the original industrial product surface images, the industrial product surface images are classified according to image categories, so as to facilitate subsequent enhancement processing of the acquired industrial product surface images according to different categories to obtain enhanced images of different categories, and select images without abnormalities from the enhanced images of different categories as training sets, and select images with abnormalities as test sets, so as to facilitate subsequent training of the model.
[0033] Specifically, in order to better simulate the changes of objects in actual scenes, the images are first classified according to the industrial product categories, which can be divided into texture category images and object category images. For texture category images, discrete angles are used to rotate the images for data enhancement, such as 0 degrees, 90 degrees, 180 degrees, and 270 degrees. This processing method can increase the diversity of samples, and the rotated images have the same texture information as the original images. For object category images, too large a rotation angle may cause the structure and logical information of the object to be distorted or block key information. Therefore, a smaller random angle range is set, such as (-5 degrees, 5 degrees), and angle rotation is performed to obtain more diverse data enhancement images. After that, the pixel size of all images is adjusted to a preset size, such as 224×224, and normal images without abnormalities are selected as training sets in the processed images, and abnormal images with abnormalities are selected as test sets, so as to facilitate the subsequent training of the model through the training set and the testing of the trained model through the test set.
[0034] S200, construct an abnormal synthesis network, a teacher feature extraction network, a variable denoising student network and a pixel-level abnormal segmentation network, and sequentially integrate the abnormal synthesis network, the teacher feature extraction network, the variable denoising student network and the pixel-level abnormal segmentation network to form an abnormal detection model. The abnormal synthesis network is used to synthesize forged abnormal images, the teacher feature extraction network is used to extract multi-scale features of the image and perform feature guidance, the variable denoising student network is used to denoise the abnormal area of the image, and the pixel-level abnormal segmentation network is used to classify and locate abnormalities in the image.
[0035] By building an anomaly detection model, it is convenient to train and test the anomaly detection model through training sets and test sets.
[0036] like Figure 2 As shown, in one embodiment, the anomaly synthesis network generates a random two-dimensional Perlin noise and binarizes it to create an anomaly mask, and mixes the non-abnormal image with the external data source image at random opacity to generate a simulated abnormal image.
[0037] Specifically, we first randomly generate two-dimensional Perlin noise and then binarize it with a preset threshold to create anomaly mask M. a , then the anomaly image I is generated by using the image from the external data source A and blending the anomaly mask region with the non-anomaly image I with a randomly selected opacity factor β (between 0.15 and 1). a , the specific calculation process is as follows:
[0038] I a =(1-M a )⊙I+(1-β)(M a ⊙I)+β(M a ⊙A)
[0039] Among them, ⊙ represents the XOR operation.
[0040] Sufficient abnormal samples and reliable abnormal simulation can significantly improve the accuracy of anomaly detection of subsequent models. It should be noted that the anomaly synthesis network only plays a role in the training phase.
[0041] In one embodiment, the teacher feature extraction network includes a ResNet18 convolutional neural network pre-trained on a large natural image dataset, wherein the ResNet18 convolutional neural network has fixed weights and the final convolution module is removed.
[0042] Specifically, the teacher feature extraction network is a ResNet18 network pre-trained on ImageNet with fixed weights, and the output feature maps are selected from the output feature maps of the first three convolutional layers (i.e., conv2, conv3, and conv4) of ResNet18, denoted as T1, T2, and T3, respectively, to serve as supervision for the variable denoising student network.
[0043] In one embodiment, the variable denoising student network includes a denoising module and a self-supervised reconstruction module. The denoising module is used to denoise the input image in the feature dimension to obtain a preliminary denoised feature map. The self-supervised reconstruction module is used to further reconstruct the preliminary denoised feature map and generate discriminant features for features with abnormal contextual structure information. The denoising module includes an encoder and a decoder. The encoder adopts a randomly initialized ResNet18 convolutional neural network. The decoder realizes the inverse structure through bilinear upsampling, and the encoder and the decoder both include four residual units; the self-supervised reconstruction module includes a variable mask convolution and a channel attention module. The self-supervised reconstruction module uses a mask operation at the center of the convolution and uses context information to predict or reconstruct the masked information, thereby forcing the model to learn the global context information of normal local features.
[0044] Specifically, the denoising module mainly consists of an autoencoder with a unique encoder-decoder architecture. The trained autoencoder only learns the normal feature representation of the teacher feature extraction network, thereby effectively filtering out anomalies in the feature space. In the autoencoder, the encoder is a randomly initialized ResNet18 network with 4 residual blocks. The decoder is a reverse ResNet18 with four residual blocks and all downsampling operations are replaced with bilinear upsampling. The output of the final denoising module is the last three layers of feature maps extracted from the decoder, denoted as S1, S2, and S3 respectively.
[0045] like Figure 3 As shown, in one embodiment, the self-supervised reconstruction module consists of a deformable masked convolution and a channel attention module. The module uses a mask operation at the center of the convolution and then uses contextual information to predict or reconstruct the masked information, thereby forcing the model to learn the global context information of normal local features.
[0046] In this example, the convolution kernel size of the deformable mask convolution is 2 × 2. Let p be the masked feature position, x(p) be the mask feature, and the calculation process of the reconstructed feature Z(p) is as follows:
[0047]
[0048] Among them, w n and p n Represent the weight of the nth position and the initial specified offset, Δp n and Δm n are the learnable offset factor and scaling factor of the nth bit, respectively, and the scaling factor Δm n In the range [0,1]. Since p+p n +Δpn It is usually a fraction, so when calculating, bilinear interpolation is used first, and then x(p+p n +Δp n ), Δp n and Δm n are obtained by applying a single convolutional layer on the same input feature map x. This convolutional layer has the same spatial resolution and dilation as the current convolutional layer, and the output is 3N channels, where the first 2N channels correspond to the learned offset Δp n , the remaining N channels are further fed into the sigmoid layer to obtain the scaling factor Δm n The Sigmoid layer is a commonly used activation function layer, which maps the input value to the interval (0,1). The formula of the Sigmoid function is:
[0049]
[0050] In this formula, x is the input value and σ(x) is the output value. By applying the sigmoid function to the input value, a scaling factor Δm is obtained. n , which ranges from 0 to 1. These scaling factors can be used in subsequent calculations or as weights to adjust the contribution of input features.
[0051] The output of the deformable mask convolution is processed by a channel attention module, which calculates an attention score for each channel. Through this mechanism, global information is used to selectively emphasize or suppress the reconstructed features. Specifically, the channel attention module first performs global pooling on Z along the channel dimension to obtain a feature vector z∈R c , where R represents a real number set, c represents the number of channels of the feature map, z∈R c represents the set of all feature vectors consisting of c elements. Subsequently, the attention score vector s∈R c The calculation method is as follows:
[0052] s=σ(W2·δ(W1·z))
[0053] The symbol σ represents the sigmoid activation function, and the symbol δ represents the ReLU activation function, which is a commonly used activation function. Its formula is: ReLU(x)=max(0,x). and Represents the weight matrix of two fully connected layers, where the number of neurons in the first fully connected layer is The information is reduced by downscaling by a factor of r. Then, the attention score vector s is replicated in the spatial dimension to generate a tensor S of the same size as Z. Finally, S and Z are element-wise multiplied to obtain the final reconstructed feature map Y∈R h×w×c , where h represents the height, w represents the width, and c represents the number of channels.
[0054] In one embodiment, the pixel-level anomaly segmentation network includes two residual blocks and a dilated spatial pyramid pooling module, and the dilated spatial pyramid pooling module is used to adaptively focus on important feature channels and generate an anomaly segmentation map for anomaly detection.
[0055] Among them, channel-aware atrous spatial pyramid pooling is used to adaptively focus on important feature channels, thereby enhancing the understanding and attention to key feature channels to more effectively handle small target anomaly detection and localization tasks. Specifically, after receiving the input feature map, the channel-aware atrous spatial pyramid pooling passes through four atrous convolution modules with a convolution kernel size of 3 and a dilation rate of 1, 6, 12, and 18, respectively. Then, these feature maps are spliced, and the spliced feature maps are subjected to global average pooling and global maximum pooling operations, spliced again, and processed through a multi-layer perceptron. Finally, the processed feature map is multiplied with the feature map before the global average pooling and global maximum pooling operations, and input into the scoring head.
[0056] The scoring head consists of a convolution module with a kernel size of 3, a normalization layer, an activation function, and a convolution module with a kernel size of 1. In the pixel-level anomaly scoring network, the role of the scoring head is to generate a scoring map for anomaly detection.
[0057] S300: training the anomaly detection model using the training set.
[0058] In one embodiment, the anomaly detection model is trained using the training set, and the entire model training is divided into two stages, including:
[0059] In the first stage, only the deformable denoising student network is trained, and the forged abnormal images are synthesized by the abnormal synthesis network and then input into the deformable denoising student network, and the normal images in the training set are input into the teacher feature extraction network to minimize the output feature difference between the two, thereby training the deformable denoising student network's ability to remove abnormal features;
[0060] The total loss function L in the first stage S It is obtained by weighted addition of two losses, and its formula is as follows.
[0061] L S =L cos +λL dmcab
[0062] Among them, L cos It is used to reduce the distance between the teacher network and the student network. The formula is as follows:
[0063]
[0064] Where S k and T k They represent the output feature maps of the k-th layer’s variable denoising student network and the teacher’s feature extraction network, i and j represent the spatial coordinates on the feature map, and H k and W k Denote their height and width respectively, and D denotes the cosine distance between the two. Both S and T have three feature maps of different scales, namely S1, S2, S3 and T1, T2, T3 defined above.
[0065] L dmcab is the loss of the self-supervised reconstruction module, that is, the cosine distance between the input feature map X and the reconstructed feature map Y, and λ is a hyperparameter, λ∈(0,1), which is used to balance the influence of the two loss terms. The gradient descent algorithm is used to minimize the total loss function L S Direction, update the weight of the deformable denoising student network until the preset total number of first-stage training rounds is reached, and obtain the trained deformable denoising student network.
[0066] In the second stage, the deformable denoising student network is kept fixed, and only the pixel-level anomaly segmentation network is trained. The forged anomaly images synthesized by the anomaly synthesis network are input to the deformable denoising student network and the teacher feature extraction network respectively, and the corresponding binary anomaly mask is the true label of the segmentation network. The pixel-level anomaly segmentation network is trained by comparing the feature map differences between the deformable denoising student network and the teacher feature extraction network to identify the anomaly area;
[0067] A focal loss and an L1 loss are respectively calculated according to the pixel-level anomaly segmentation network, and a second loss function is calculated according to the focal loss and the L1 loss.
[0068] Specifically, we first calculate the feature map (S k ,T k ), k = 1, 2, 3, denoted as C k , and upsample them to one-fourth of the input size, then concatenate the upsampled feature maps in the channel dimension and input them into the pixel-level segmentation network. The output of the pixel-level segmentation network is the predicted anomaly mask, denoted as Its size is the same as the input. At the same time, the anomaly mask is downsampled to the same size, denoted as M, and the segmentation training is performed by adopting the focal loss Lfocal And L1 loss to optimize. Focus loss L focal The calculation formula is as follows.
[0069]
[0070] in λ is the focusing parameter.
[0071] The L1 loss calculation formula is as follows:
[0072]
[0073] Where M represents the real exception mask, represents the predicted anomaly mask, and i and j represent the spatial coordinates.
[0074] The total loss function of the second stage is the sum of the above two loss terms.
[0075] L seg =L facal +L1
[0076] Use the gradient descent algorithm to minimize the total loss function L seg Direction, update the weight of the pixel-level anomaly segmentation network until the preset second-stage training total rounds are reached to obtain the trained pixel-level anomaly segmentation network.
[0077] S400: Using the test set to verify the trained anomaly detection model.
[0078] In order to verify the anomaly detection model provided by this implementation, the MVTEC AD dataset is used for verification. The MVTEC AD dataset is a public dataset for industrial visual inspection tasks, specifically designed to evaluate the performance of anomaly detection algorithms. It contains 15 categories and a total of 5354 high-resolution images. Among them, 3629 images are non-anomaly images for training; 1725 images (including normal and abnormal images) are used for testing.
[0079] Using the average precision (AP) as the evaluation indicator, we draw a PR curve with the recall rate (Recall) as the horizontal axis and the precision rate (Precision) as the vertical axis. The area under the PR curve is defined as AP. The calculation formulas for precision and recall are as follows:
[0080]
[0081]
[0082] Among them, TP represents the number of defective samples that are accurately judged as defective, FP represents the number of non-defective samples that are wrongly judged as defective, and FN represents the number of defective samples that are wrongly judged as non-defective.
[0083] After verification, the anomaly detection target model is output.
[0084] After the anomaly detection model is trained, it is verified through the test set, and when the verification accuracy reaches the preset threshold, a qualified anomaly detection target model is obtained, so as to facilitate the use of the trained anomaly detection model to perform anomaly detection on industrial images.
[0085] In one embodiment, the training of the deep learning model is divided into 5000 rounds, of which the first 1000 rounds focus on training the student feature extraction network to remove anomalies. The next 4000 rounds focus on training the pixel-level anomaly segmentation network. After every 1000 rounds of training, the test set is evaluated, and the model weights are saved, and finally the model with the best performance in the test set is selected for deployment. This ensures that the entire anomaly detection system can achieve optimal performance in practical applications. After the training is completed, the best model is selected for deployment to achieve real-time detection of surface defects of industrial products.
[0086] The protection scope of the unsupervised industrial anomaly detection method based on knowledge distillation described in the embodiment of the present invention is not limited to the execution order of the steps listed in this embodiment. All solutions implemented by adding, reducing or replacing steps in the prior art based on the principles of the present invention are included in the protection scope of the present invention.
[0087] The present invention discloses an unsupervised industrial anomaly detection method based on knowledge distillation, and uses an anomaly detection target model established by the unsupervised industrial anomaly detection method based on knowledge distillation to perform surface defect detection on industrial products.
[0088] Since the detection process of the unsupervised industrial anomaly detection method based on knowledge distillation is to perform surface defect detection on industrial products based on the anomaly detection target model established by the unsupervised industrial anomaly detection method based on knowledge distillation, it will not be described in detail here.
[0089] The present invention also provides an unsupervised industrial anomaly detection system based on knowledge distillation, which can implement the unsupervised industrial anomaly detection method based on knowledge distillation described in the present invention. However, the implementation device of the unsupervised industrial anomaly detection method based on knowledge distillation described in the present invention includes but is not limited to the structure of the unsupervised industrial anomaly detection system based on knowledge distillation listed in this embodiment. All structural deformations and replacements of the prior art made according to the principles of the present invention are included in the protection scope of the present invention.
[0090] like Figure 4 As shown, in one embodiment, the present invention provides an unsupervised industrial anomaly detection system based on knowledge distillation, the system comprising:
[0091] The preprocessing unit 401 is used to obtain industrial product surface images, and preprocess the industrial product surface images to obtain a training set and a test set respectively.
[0092] The integration unit 402 is used to construct an anomaly synthesis network, a teacher feature extraction network, a variable denoising student network and a pixel-level anomaly segmentation network, and integrate the anomaly synthesis network, the teacher feature extraction network, the variable denoising student network and the pixel-level anomaly segmentation network in sequence to form an anomaly detection model.
[0093] The training unit 403 is used to train the anomaly detection model using the training set.
[0094] The verification unit 404 is used to verify the trained anomaly detection model using the test set, and output an anomaly detection target model after verification meets the requirements.
[0095] Among them, the anomaly synthesis network is used to synthesize forged abnormal images, the teacher feature extraction network is used to extract multi-scale features of the image and perform feature guidance, the variable denoising student network is used to denoise the abnormal area of the image, and the pixel-level anomaly segmentation network is used to classify and locate the anomaly of the image.
[0096] It should be noted that the structures and principles of the preprocessing unit 401, the integration unit 402, the training unit 403 and the verification unit 404 correspond one-to-one to the steps (steps S100 to S400) in the above-mentioned unsupervised industrial anomaly detection method based on knowledge distillation. The specific working principles can also be referred to the introduction of the unsupervised industrial anomaly detection method based on knowledge distillation in the above-mentioned embodiments, so they will not be repeated here.
[0097] In the several embodiments provided by the present invention, it should be understood that the disclosed system, device or method can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules / units is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules or units, which can be electrical, mechanical or other forms.
[0098] The modules / units described as separate components may or may not be physically separated, and the components displayed as modules / units may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules / units may be selected according to actual needs to achieve the purpose of the embodiments of the present invention. For example, the functional modules / units in the various embodiments of the present invention may be integrated into one processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into one module / unit.
[0099] Those of ordinary skill in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0100] The descriptions of the processes or structures corresponding to the above-mentioned figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.
[0101] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.
Claims
1. An unsupervised industrial anomaly detection method based on knowledge distillation, characterized in that: The method comprises: Acquire industrial product surface images, and preprocess the industrial product surface images to obtain a training set and a test set respectively; Constructing an anomaly synthesis network, a teacher feature extraction network, a variable denoising student network and a pixel-level anomaly segmentation network, and sequentially integrating the anomaly synthesis network, the teacher feature extraction network, the variable denoising student network and the pixel-level anomaly segmentation network to form an anomaly detection model; Training the anomaly detection model using the training set; The trained anomaly detection model is verified using the test set, and an anomaly detection target model is output after the verification meets the requirements; Among them, the anomaly synthesis network is used to synthesize forged abnormal images, the teacher feature extraction network is used to extract multi-scale features of the image and perform feature guidance, the variable denoising student network is used to denoise the abnormal area of the image, and the pixel-level anomaly segmentation network is used to classify and locate the anomaly of the image.
2. The unsupervised industrial anomaly detection method based on knowledge distillation according to claim 1 is characterized in that: The preprocessing of the industrial product surface images to obtain a training set and a test set respectively includes: Classifying the surface image of the industrial product according to image categories to obtain images of multiple categories; Performing data enhancement processing on the surface image of the industrial product according to different categories to obtain an enhanced image; The images without abnormalities in the enhanced images are selected as training sets, and the enhanced images with abnormalities and the corresponding abnormal mask images are selected as test sets.
3. The unsupervised industrial anomaly detection method based on knowledge distillation according to claim 1 is characterized in that: The anomaly synthesis network generates a random two-dimensional Perlin noise and binarizes it to create an anomaly mask, and mixes the anomaly-free image with the external data source image at random opacity to generate a simulated anomaly image.
4. The unsupervised industrial anomaly detection method based on knowledge distillation according to claim 1 is characterized in that: The teacher feature extraction network includes a ResNet18 convolutional neural network pre-trained on a large natural image dataset, wherein the ResNet18 convolutional neural network has fixed weights and the final convolution module is removed.
5. The unsupervised industrial anomaly detection method based on knowledge distillation according to claim 1 is characterized in that: The variable denoising student network includes a denoising module and a self-supervised reconstruction module. The denoising module is used to denoise the input image in the feature dimension to obtain a preliminary denoised feature map. The self-supervised reconstruction module is used to further reconstruct the preliminary denoised feature map and generate discriminant features for features with abnormal contextual structure information.
6. The unsupervised industrial anomaly detection method based on knowledge distillation according to claim 5 is characterized in that: The denoising module includes an encoder and a decoder, the encoder adopts a randomly initialized ResNet18 convolutional neural network, the decoder realizes an inverse structure through bilinear upsampling, and the encoder and the decoder both include four residual units; the self-supervised reconstruction module includes a variable mask convolution and a channel attention module, the channel attention module adopts a mask operation at the center of the convolution and uses context information to predict or reconstruct the masked information, thereby forcing the model to learn the global context information of normal local features.
7. The unsupervised industrial anomaly detection method based on knowledge distillation according to claim 1, characterized in that: The pixel-level anomaly segmentation network includes two residual blocks and a dilated spatial pyramid pooling module, wherein the dilated spatial pyramid pooling module is used to adaptively focus on important feature channels and generate anomaly segmentation maps for anomaly detection.
8. The unsupervised industrial anomaly detection method based on knowledge distillation according to claim 1, characterized in that: The training of the anomaly detection model using the training set includes: The weights in the anomaly detection model are kept fixed, a forged abnormal image is synthesized by the anomaly synthesis network and then input into the deformable denoising student network, and a normal image in the training set is input into the teacher feature extraction network to calculate a first loss function; The deformable denoising student network is kept fixed, and a forged abnormal image is synthesized by the abnormal synthesis network and then inputted into the deformable denoising student network and the teacher feature extraction network respectively, so as to train the pixel-level abnormal segmentation network by comparing the difference of the feature maps of the deformable denoising student network and the teacher feature extraction network, so as to identify abnormal areas; A focal loss and an L1 loss are respectively calculated according to the pixel-level anomaly segmentation network, and a second loss function is calculated according to the focal loss and the L1 loss.
9. An unsupervised industrial anomaly detection method based on knowledge distillation, characterized in that: Surface defect detection of industrial products is performed using an anomaly detection target model established by the unsupervised industrial anomaly detection method based on knowledge distillation as described in any one of claims 1 to 8 above.
10. An unsupervised industrial anomaly detection system based on knowledge distillation, characterized in that: The method comprises: A preprocessing unit, used to obtain an industrial product surface image, and preprocess the industrial product surface image to obtain a training set and a test set respectively; An integration unit is used to construct an anomaly synthesis network, a teacher feature extraction network, a variable denoising student network and a pixel-level anomaly segmentation network, and sequentially integrate the anomaly synthesis network, the teacher feature extraction network, the variable denoising student network and the pixel-level anomaly segmentation network to form an anomaly detection model; A training unit, used for training the anomaly detection model using the training set; A verification unit, used to verify the trained anomaly detection model using the test set, and output an anomaly detection target model after verification meets the requirements; Among them, the anomaly synthesis network is used to synthesize forged abnormal images, the teacher feature extraction network is used to extract multi-scale features of the image and perform feature guidance, the variable denoising student network is used to denoise the abnormal area of the image, and the pixel-level anomaly segmentation network is used to classify and locate the anomaly of the image.
Citation Information
Patent Citations
Gear surface defect detection method based on RDMS
CN117173098A
Industrial anomaly detection method and system based on multi-scale feature guidance and fusion
CN117710757A
Image anomaly detection method based on self-supervised learning and knowledge distillation
CN117934425A
Improved DeSTSeg-based unsupervised grain defect anomaly detection method
CN118469912A
Reverse distillation unsupervised anomaly detection method based on memory bank reinforcement
CN118823466A
Cited By
Restoration type self-supervision defect detection method and device and storage medium
CN116563250A
A method and device for defect detection based on a self-supervised restoration method and a storage medium
CN116563250B
Zero-sample industrial anomaly detection method based on knowledge distillation
CN120298397A
Zero-shot industrial anomaly detection method based on knowledge distillation
CN120298397B
Production line mask finished product abnormity detection method based on knowledge distillation
CN120598963A