An Image Defect Detection Method Based on Dual-Branch Inverse Distillation and Multi-Input Images

By using the method of double-branch inverse distillation and multi-input image in industrial defect detection, combined with the feature fusion and reconstruction technology of global branches and local branches, the problem of difficulty in detecting small objects and logical abnormalities in the prior art is solved, and a more efficient defect detection effect is achieved.

CN118521570BActive Publication Date: 2025-06-13HENAN UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410923028.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2025-06-13
Estimated Expiration
2044-07-10

AI Technical Summary

Technical Problem

In the prior art, it is difficult to effectively detect defects of small objects such as screws in industrial defect detection, and it is difficult to detect logical abnormalities, such as missing components or incorrect number of components, resulting in insufficient detection performance.

Method used

Image defect detection method based on double-branch inverse distillation and multi-input images is adopted. By constructing a teacher network and student network, combining the dual-branch bottleneck structure of global branches and local branches, feature fusion and reconstruction are carried out, and defects are detected using inverse knowledge distillation and reconstruction results.

Benefits of technology

Effective detection of structural defects and logical defects is achieved, and better detection results are obtained, especially in detecting small objects and logical abnormalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118521570B_ABST
    Figure CN118521570B_ABST
Patent Text Reader

Abstract

The present invention relates to an image defect detection method based on dual-branch inverse distillation and multi-input images in the fields of deep learning and image processing technology, comprising the following steps: S1. Construct a model, including constructing a teacher network, a student network, and a dual-branch bottleneck structure with a global branch and a local branch; S2. Model training; distill multiple feature maps e output by the teacher network with corresponding multiple feature maps d of the student network respectively, and use cosine distance as the distillation loss; S3. Model testing; obtain the two-dimensional anomaly map of the i-th layer of the model by calculating the cosine distance of each pixel, and add them up to obtain the final anomaly map M; this image defect detection method can not only detect structural defects, but also detect logical defects and obtain better detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning and image processing, and in particular to an image defect detection method based on double-branch inverse distillation and multi-input images. Background Art

[0002] Anomaly detection is a way to identify and locate anomalies with limited or no prior knowledge; anomaly detection has a wide range of application fields, including industrial defect detection, medical detection, video surveillance, etc., so it has attracted much attention; in industrial manufacturing, defect detection can automatically identify and detect defects in products, such as cracks, scratches or defects, etc. Since in most manufacturing processes, the number of normal products far exceeds the number of defective products, and at the same time, due to the rare occurrence of specific types of defect formation during the production process, it becomes more difficult to collect and label real defect samples. Therefore, it is difficult to obtain abnormal data in defect detection. This phenomenon makes it challenging to obtain a sufficient number of abnormal data to train and test the defect detection model; in view of this situation, unsupervised methods are mostly used in defect detection, that is, a set of normal samples are provided as a reference, without directly understanding the characteristics or attributes of anomalies, and the model is trained by learning the distribution of normal samples. In the inference stage, the model is directly used to evaluate the anomaly degree of new samples: if a sample significantly deviates from the distribution of normal samples, it is regarded as an abnormal product;

[0003] To address this challenge, previous studies have mostly focused on using these normal training samples to construct various self-supervised tasks, including reconstruction, pseudo-anomaly sample generation, memory bank search, knowledge distillation, etc.; however, few of the above models are designed for detecting small objects (such as screws); in most detection images, screws only account for a small part of the entire image, while most of the image is background, which brings a new challenge because small objects like screws are different from larger objects. Small objects are usually isolated and have no strong context connection with other objects, and the lack of context makes it more difficult for the model to accurately detect and locate them;

[0004] Secondly, in unsupervised anomaly detection, existing models mainly focus on detecting structural anomalies, such as dents and scratches. For example, the Chinese invention patent with the publication number CN114663392A discloses an industrial image defect detection method based on knowledge distillation, which uses the method of knowledge distillation in deep learning to detect structural defects in industry; however, when the defect involves logical anomalies, most models are difficult to achieve detection; logical anomalies mainly refer to the lack of components or incorrect number of components in industrial products. Exactly in the actual scenario, logical anomalies like component missing will more likely cause serious functional consequences, which are often more critical than small problems like scratches; therefore, there is still much room for improvement in the performance of existing methods in logical anomaly detection. Summary of the Invention

[0005] To overcome the deficiencies in the background art and solve the existing technical problems, the present invention discloses an image defect detection method based on dual-branch inverse distillation and multi-input images, which can not only detect structural defects but also logical defects and obtain better detection results.

[0006] To achieve the above invention purpose, the present invention adopts the following technical solutions:

[0007] An image defect detection method based on dual-branch inverse distillation and multi-input images, comprising the following steps:

[0008] S1. Construct a model, including constructing a teacher network, a student network, and a dual-branch bottleneck structure with a global branch and a local branch; the teacher network is used to output multiple feature maps e of different scales according to the input original image I, and input the corresponding extracted multi-scale features to the local branch, while inputting the deepest feature in the multi-scale features to the global branch; the local branch and the global branch of the dual-branch bottleneck structure are respectively used to compress the multi-scale features and the deepest features to obtain local features and global features while retaining normal features, and the local features and global features are fused by matrix multiplication to obtain dual-branch fusion features; the student network introduces a corresponding number of additional image input modules arranged in layers for the multiple feature maps e, and the topmost additional image input module fuses the dual-branch fusion features and the original image I to obtain a reconstructed feature map d, and the other layer additional image input modules respectively fuse the original image I and the feature map d obtained by the previous layer additional image input module to obtain a new reconstructed feature map d;

[0009] S2. Model training; distill the multiple feature maps e output by the teacher network and the corresponding multiple feature maps d of the student network respectively, and the distillation loss uses the cosine distance. For the i-th layer feature:

[0010]

[0011] In the formula, represents the output feature of the i-th layer of the teacher network, represents the output feature of the i-th layer of the student network, and ∙ represents the dot product; the total distillation loss is the average value of calculating the losses of all layers:

[0012]

[0013] where j represents the number of layers; the goal of model training is to minimize the distillation loss;

[0014] S3. Model testing; obtain the two-dimensional anomaly map of the i-th layer of the model by calculating the per-pixel cosine distance:

[0015]

[0016] In this formula, (h, w) represents the position of the feature vector, represents the output feature of the i-th layer of the teacher network, represents the output feature of the i-th layer of the student network, ; in order to obtain the final anomaly map M, the anomaly maps of all layers are added together: , where represents the upsampling operation on the i-th anomaly map.

[0017] Furthermore, in S1, after determining the way of initializing the parameters, the teacher network freezes the parameters and does not participate in the model update.

[0018] Furthermore, in S1, both the local branch and the global branch contain an embedding module and a feature fusion module, where the 4th residual block of ResNet is used as the embedding module.

[0019] Furthermore, in S1, before the student network fuses the original image I, the original image I is subjected to average pooling processing.

[0020] Furthermore, in S3, the anomaly map M is smoothed using Gaussian filtering.

[0021] Furthermore, in S3, the maximum anomaly value in the anomaly map M and the test sample label are used to calculate the AUC to obtain the image-level detection result, and the anomaly map and the test sample ground truth label are used to calculate the AUC and PRO respectively to obtain the pixel-level localization result and the localization accuracy.

[0022] Due to the above-mentioned technical solution, the present invention has the following beneficial effects:

[0023] The image defect detection method based on double-branch inverse distillation and multi-input images disclosed by the present invention starts from the quality detection of screws in the industrial production and manufacturing scenario, aims to maximize the detection of defective products, combines unsupervised learning to establish an industrial image defect detection model based on double-branch inverse distillation and multi-input images, takes the inverse distillation model as the benchmark model, combines the double-branch network structure, where the global branch aims to obtain global context information, while the local branch focuses on collecting local features, fuses the double-branch features to solve the problem of insufficient context information, introduces multiple additional input image modules in the reconstruction part, and detects and locates defects by comparing the reconstructed image and the input image; compared with the existing defect detection methods, the present invention adopts unsupervised learning, all the training data used are normal data without labels, and uses inverse knowledge distillation combined with the reconstruction result to detect defects, which can not only detect structural defects and logical defects, but also obtain better detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is the overall model schematic diagram of the present invention;

[0025] Figure 2 is Figure 1 the enlarged schematic diagram of the middle school student network model;

[0026] Figure 3 is Figure 1 the enlarged schematic diagram of the double-branch bottleneck structure model in the middle;

[0027] Figure 4 is Figure 1 the enlarged schematic diagram of the global branch in the middle;

[0028] Figure 5 is Figure 1 the enlarged schematic diagram of the local branch in the middle. Specific embodiments

[0029] The technical solution of the present invention will be described below in conjunction with the accompanying drawings in the embodiments of the present invention:

[0030] In the present invention, for the task scenario, a method for defect detection of industrial production and manufacturing screws is proposed; this method is based on the inverse distillation framework, combines the local branch and the global branch, introduces multiple additional input image modules in the reconstruction part, and detects and locates defects by comparing the reconstructed image and the input image.

[0031] Combined with the attached Figure 1 The image defect detection method based on double-branch inverse distillation and multi-input images includes the following steps:

[0032] Step 1, construct a model, including constructing a teacher network, a student network, and a double-branch bottleneck structure with a global branch and a local branch; the present invention mainly detects the defects of industrial production and manufacturing product images. Since the current defect sample pictures of industrial screws are scarce and the types of defects are unknown, the present invention uses unsupervised learning to train the model for identifying and locating abnormalities in the image.

[0033] The teacher network is used to output multiple feature maps \(e\) of different scales according to the input original image \(I\), and input the corresponding extracted multi-scale features to the local branch, while inputting the deepest features in the multi-scale features to the global branch; the teacher network is equivalent to the teacher encoder. To enable the teacher encoder to have good detection performance, the WideResNet-50 network is selected as the teacher network; the WideResNet-50 network is a variant of the ResNet network, introducing wider layers and a deeper network structure. The wider layers increase the capacity of the model, enabling it to capture more complex patterns and data representations, thus better capturing and learning the training data; after determining the appropriate way of initializing parameters, the parameters of the teacher encoder are frozen and do not participate in model updating.

[0034] To better extract global and local features, a dual-branch bottleneck structure is designed to hierarchically compress the embedded features; this hierarchical compression enables us to retain key normal features at multiple levels while filtering out abnormal features caused by anomalies of different sizes; both the local branch and the global branch contain an embedding module and a feature fusion module, and the 4th residual block of ResNet is used as the embedding module;

[0035] To address the detection challenges brought by large-size anomalies caused by missing or misaligned normal components, the main goal of the global branch is to retain the most representative normal semantic features while effectively filtering out abnormal semantic features. The global branch structure is as Figure 4 shown. The global branch first uses an embedding module to project the high-dimensional features into a low-dimensional space for the deep features. Its compact embedding serves as an information bottleneck, which helps prevent abnormal features from spreading to the student network, thereby enhancing the feature difference between the teacher-student network on anomalies; subsequently, global max pooling is used to compress the features spatially, effectively eliminating abnormal features, which plays a key role; to facilitate feature recovery, a deconvolution layer with a stride of 2 is used to upsample the compressed features to a feature size of 8×8 for subsequent feature recovery. However, max pooling may cause a large amount of information loss. To extract important information and retain most of the information, the features are fused with the original embedded features, and finally the spatial dimension is compressed through a 1×1 convolution operation to obtain the final global features;

[0036] The structure of the local branch is as Figure 5 shown. The local branch is used to compress multi-layer features, retain local normal features, and filter out local abnormal features; the local branch structure refers to the original inverse distillation work. The shallow features are downsampled through a 3×3 convolutional layer with a stride of 2, and after aligning and fusing multi-layer features, a 1×1 convolution is used to compress the spatial dimension, and finally the embedding module is used to compress the features to obtain local features;

[0037] Finally, matrix multiplication is adopted to fuse the local features and global features to obtain the dual-branch fusion features, which not only ensures the retention of relevant information but also effectively filters out anomalies. This comprehensive strategy not only improves the accuracy and reliability of small-size anomaly detection but also applies to large-size anomalies caused by missing or misaligned normal components.

[0038] To match the intermediate features of the teacher network, the student network uses a symmetric but opposite structure to the teacher network. This reverse design helps eliminate the response of the student network to anomalies, and the symmetry enables it to have the same feature dimension as the teacher network.

[0039] The student network introduces a corresponding number of additional image input modules arranged in layers for multiple feature maps e. For this instance, the number is three for all. The additional image input module at the highest layer fuses the dual-branch fusion features and the original image I to obtain the reconstructed feature map d. The other layer additional image input modules respectively fuse the original image I and the feature map d obtained from the previous layer additional image input module to obtain a new reconstructed feature map d. The introduction of the additional image input module can effectively avoid excessive loss of valuable information. However, since the size of the original image is (3, 256, 256), it cannot be directly aligned with the dual-branch fusion features. In this case, the initial image needs to be processed by resizing to match the required size or format for correct alignment and fusion with the dual-branch fusion features. In this module, the original input image undergoes average pooling, that is, downsampling is performed by calculating the average value of pixel values within each region, and then it is concatenated with the dual-branch fusion features. The features are processed through BatchNorm and ReLU, and finally, a 1×1 convolutional layer is used to reduce the number of channels. By introducing skip connections in the detection model using the image input module, the model downsamples the initial input image and merges it with the feature map from the final layer. The student network stacks multiple image input modules to combine the shallow appearance information with the deep detail and semantic information to obtain more comprehensive information. The structure of the entire student network is as Figure 2 shown.

[0040] Step 2, model training; during the training process, only normal images are provided. The overall process is as Figure 1 ; The three feature maps e output by the teacher network are distilled with the corresponding three feature maps d of the student network respectively. The distillation loss uses cosine distance. For the i-th layer features:

[0041]

[0042] In the formula, represents the output feature of the i-th layer of the teacher network, represents the output feature of the i-th layer of the student network, and ∙ represents the dot product. The total distillation loss is the average value of calculating the losses of all layers:

[0043]

[0044] Among them, j represents the number of layers, which is 3 in this design; by using the cosine distance, the distillation loss measures the dissimilarity between the output features of the teacher network and the output features of the student network. The goal of model training is to minimize the distillation loss, thereby encouraging the student model to imitate the behavior of the teacher model in terms of feature representation.

[0045] Step 3, model testing; the shallow layer of the neural network extracts local features to describe low-level information, such as color, edges, texture, etc., while the deep layer has a broader receptive field and can represent global semantic and structural information; since the teacher network can capture abnormal and out-of-branch features, while the student network cannot accurately reconstruct these abnormal features, therefore, the difference between the features produced by the teacher and the student can be used as an indication of the degree of sample abnormality; in other words, the low similarity of low-level and high-level features in the teacher-student model represents local abnormality and global structural abnormality respectively;

[0046] In the testing stage, the model uses the cosine distance to detect abnormalities. Since the model is only trained on normal images, the student network only learns the distribution of normal images. Therefore, due to the lack of knowledge about the abnormal parts, the student model will encounter difficulties in reconstructing the abnormal parts. By calculating the per-pixel cosine distance, a two-dimensional abnormality map of the i-th layer of the model is obtained:

[0047]

[0048] In this formula, (h, w) represents the position of the feature vector, represents the output feature of the i-th layer of the teacher network, represents the output feature of the i-th layer of the student network, ; to obtain the final abnormality map M, the abnormality maps of all layers are added together: , where represents the upsampling operation on the i-th abnormality map; to reduce the noise in the abnormality map, the image is smoothed using Gaussian filtering.

[0049] While detecting defects in the image, it is also necessary to judge the detection effect of the model; thus, evaluation metrics are introduced:

[0050] ROC (Receiver Operating Characteristic), the abscissa of ROC is the false positive rate (also called the false positive class rate, False Positive Rate), and the ordinate is the true positive rate (true positive class rate, True Positive Rate); among them, TPR is the ratio of the samples that are actually positive and are correctly judged as positive; FPR is the ratio of the samples that are actually negative and are wrongly judged as positive; given a binary classification model and its threshold, a coordinate point (X = FPR, Y = TPR) can be calculated from the true values and predicted values of all samples.

[0051] AUC (Area Under Curve): When used for classification tasks, the area enclosed by the ROC curve and the coordinate axes ranges from 0.1 to 1. As a numerical value, AUC can intuitively evaluate the quality of a classifier, and the larger the value, the better; when used for localization tasks, generally, the result after binarizing the heat map of the defect probability by setting a threshold is compared with the pixel-level label map for calculation, that is, the pixel-level AUROC index.

[0052] And PRO (Per Region Overlap): It is mostly used for localization tasks. Since the TPR value in the ROC curve is affected by the defect area, if the abnormal area with a large correctly located area will greatly improve the TPR index, but if the defect area with a small mislocated area has little impact, so the first step of PRO is to divide the located defect results and the actual true values into N regions according to the connected components, then find the intersection of the predicted results and the true values in each region, and the PRO value can be obtained by weighted averaging the N regions after dividing the intersection by the true value; similarly, in order to avoid the influence of the threshold on the evaluation results, only the threshold that makes FPN between 0 - 30% is selected.

[0053] In this design, the maximum abnormal value in the abnormal map and the test sample label are used to calculate AUC to obtain the image-level detection result, and the abnormal map and the test sample ground truth label are used to calculate AUC and PRO (FPN = 30%) to obtain the pixel-level localization result and the localization accuracy respectively.

[0054] The method for defect detection of industrial production manufacturing screw images proposed by the present invention is designed as a fully automatic algorithm, including the following algorithm steps:

[0055] Step 1: Initialize the teacher encoder , the dual-branch embedding module and the student decoder , and freeze the parameters.

[0056] Start the loop;

[0057] Step 2: Initial input image Extract three features of different sizes from ; ;

[0058] Step 3: Input to perform compression to obtain the feature ;

[0059] Step 4: and Input to perform feature reconstruction to obtain ;

[0060] Step 5: Calculate the distillation loss between and to update the model loss;

[0061] End the loop;

[0062] Step 6: Calculate the AUC of the test image, pixel-level AUC, and pixel-level PRO using the cosine distance between and ;

[0063] Step 7: The algorithm ends.

[0064] The parts not detailed in the present invention are the prior art. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention; therefore, from any point of view, the above embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention, and any reference signs in the claims should not be regarded as limiting the content of the claims involved.

Claims

1. An image defect detection method based on dual-branch reverse distillation and multiple input images, characterized in that: The following steps are involved: S1. Building a model, including building a teacher network, a student network, and a dual-branch bottleneck structure with global branches and local branches; The teacher network is used to output multiple feature maps e of different scales according to the input original image I, and input the corresponding extracted multi-scale features to the local branch, and input the deepest features in the multi-scale features to the global branch; The local branch and global branch of the dual-branch bottleneck structure are used to compress multi-scale features and the deepest features to obtain local features and global features respectively while retaining normal features, and matrix multiplication is used to fuse local features and global features to obtain dual-branch fusion features; The student network introduces a corresponding number of hierarchically arranged additional image input modules corresponding to multiple feature maps e. The additional image input module at the highest level fuses the dual-branch fusion features with the original image I to obtain a reconstructed feature map d. The additional image input modules at other levels fuse the original image I with the feature map d obtained by the additional image input module at the previous level to obtain a newly reconstructed feature map d. S2, model training; distill the multiple feature maps e output by the teacher network with the multiple feature maps d corresponding to the student network, and use the cosine distance as the distillation loss. For the i-th layer feature: In the formula, represents the output feature of the i-th layer of the teacher network, represents the output feature of the i-th layer of the student network, ∙ represents the dot product; the total distillation loss is the average value of the loss of all layers: Where j represents the number of layers; the goal of model training is to minimize the distillation loss; S3, model testing: by calculating the cosine distance of each pixel, the two-dimensional anomaly map of the i-th layer of the model is obtained: In this formula, (h, w) represents the position of the eigenvector, represents the output feature of the i-th layer of the teacher network, represents the output feature of the i-th layer of the student network, ; To obtain the final anomaly map M, add up the anomaly maps of all layers: , here Indicates upsampling operation on the i-th abnormal graph.

2. The image defect detection method based on dual-branch reverse distillation and multiple input images according to claim 1 is characterized in that: In S1, the teacher network freezes its parameters after determining the initialization parameter method and does not participate in model updating.

3. The image defect detection method based on dual-branch reverse distillation and multiple input images according to claim 1 is characterized in that: In S1, both the local branch and the global branch contain an embedding module and a feature fusion module, in which the 4th residual block of ResNet is used as the embedding module.

4. The image defect detection method based on dual-branch reverse distillation and multiple input images according to claim 1 is characterized in that: In S1, the student network performs average pooling on the original image I before fusing it with the original image I.

5. The image defect detection method based on dual-branch reverse distillation and multiple input images according to claim 1 is characterized in that: In S3, the anomaly map M is smoothed using Gaussian filtering.

6. The image defect detection method based on dual-branch reverse distillation and multiple input images according to claim 1 is characterized in that: In S3, the maximum anomaly value in the anomaly map M and the test sample label are used to calculate the AUC to obtain the image-level detection result. The anomaly map and the test sample ground truth label are used to calculate the AUC and PRO to obtain the pixel-level positioning result and positioning accuracy, respectively.

Citation Information

Patent Citations

  • Industrial image defect detection method based on knowledge distillation

    CN114663392A