Self-supervised industrial defect detection method based on prototype memory guidance
Through the self-supervised method of prototype memory-guided, the defect image is synthesized by significance detection and Berlin noise, combined with the teacher-student network and the average memory module, the problem of logic defect detection difficulties and over-detection in the traditional method is solved, achieving higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202410354812.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-27
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2044-03-27
AI Technical Summary
Traditional methods are difficult to effectively detect logical defects that show local normality, such as target misalignment or missing in transistors, and existing defect synthesis methods can easily lead to over-detection problems, misleading interference factors such as dust in the image background as defects.
Using a self-supervised industrial defect detection method based on prototype memory guidance, a foreground mask is generated through the significance detection network, defect images are synthesized using Berlin noise machines, and features are extracted through the teacher-student network and the average memory module, combining FocalLoss and L1 Loss loss function optimization to enhance the model's detection ability of logical defects.
It improves the detection performance of the model for logical defects, reduces over-detection, enhances the robustness of background interference, saves memory overhead and retrieval time, and improves detection accuracy.
Smart Images

Figure CN118196051B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial defect detection, and in particular to a self-supervised industrial defect detection method based on prototype memory guidance. Background Art
[0002] Industrial defect detection is to detect the appearance defects of various industrial products. It is one of the important technologies to ensure product quality and maintain production stability.
[0003] In recent years, unsupervised industrial defect detection based on knowledge distillation has attracted extensive attention and research. At present, the knowledge distillation-based method can effectively detect surface defects such as scratches, stains or discoloration. However, it is difficult for existing methods to effectively detect logical defects that show local normality, such as normal targets appearing in the wrong position or targets missing, such as Figure 1 As shown in Figure 2, traditional methods can basically detect surface scratch defects on metal nuts, but it is difficult to effectively detect logic defects in transistors. Due to the lack of prior knowledge of normal samples during the testing phase, the defective parts are mistakenly identified as normal. In addition, as Figure 1 As shown in the detection results of the toothbrush test image in , the existing defect synthesis method is prone to over-detection, that is, misleading the model into detecting interference factors such as dust in the image background as defects. Summary of the Invention
[0004] The technical problem solved by the present invention is that traditional methods can basically detect surface scratch defects in metal nuts, but it is difficult to effectively detect logic defects in transistors. This is because there is a lack of prior knowledge of normal samples during the testing phase, which causes the defective parts to be mistakenly identified as normal. The existing defect synthesis method is prone to over-detection problems, that is, misleading the model to detect interference factors such as dust in the image background as defects.
[0005] To solve the above technical problems, the present invention provides the following technical solutions: a self-supervised industrial defect detection method based on prototype memory guidance, comprising step S100, obtaining a foreground mask F of an input image O based on a saliency detection network, generating random noise P of the input image O using a Perlin noise generator, constraining the random noise P to the foreground region using the foreground mask F to obtain a mask image M, and synthesizing a defect image S using the mask image M and an out-of-distribution image A; step S200, extracting features of the input image O and the defect image S using a teacher-student network, and obtaining a pixel-by-pixel similarity F between the features. ts ; Step S300, extracting the prototype feature F of the input image O based on the average memory module AMM mt ; Step S400, pixel-by-pixel similarity F ts and F mtAfter multi-scale discrimination, the FocalLoss and L1 Loss loss functions are used for joint optimization.
[0006] As a preferred embodiment of the self-supervised industrial defect detection method based on prototype memory guidance described in the present invention, the specific method of synthesizing the defect image S in step S100 is:
[0007] The transparency factor is introduced in the defect synthesis process to balance the fusion of the input detection image and the out-of-distribution image. Its mathematical expression is:
[0008]
[0009] Among them, S represents the defect image, O represents the input image, ⊙ represents element-by-element multiplication, and M represents the mask image. represents the inversion of the mask image M, β∈[0,1] represents the transparency factor, and Aug represents three operations randomly selected from equalization, exposure, toner separation, sharpness, and inversion.
[0010] As a preferred embodiment of the self-supervised industrial defect detection method based on prototype memory guidance described in the present invention, in step S200, the input image features extracted by the student network through cosine similarity matching with the teacher network include:
[0011] The synthesized defect image S is input into the student network, and the input image O corresponding to the synthesized defect image S is input into the teacher network. The student network matches the features of the input image O extracted by the teacher network through cosine similarity, reconstructs the features, and repairs the defect features. The mathematical expression is:
[0012]
[0013]
[0014] in, represents the i-th layer feature extracted by the teacher network, represents the i-th layer feature extracted by the student network, h∈[1,H i ]W∈[1,W i ], H i Indicates the height of the i-th layer feature, W i represents the width of the i-th layer feature, ∈[1, 3], represents the two-dimensional defect score map obtained by calculating the cosine similarity along the channel dimension, L cos represents the loss function for optimizing the student network.
[0015] As a preferred embodiment of the self-supervised industrial defect detection method based on prototype memory guidance described in the present invention, the specific process of extracting prototype features in step S300 includes:
[0016] Step S301, using an average memory module (AMM) to extract prototype features, and detecting logical defects in the foreground information using the prototype features;
[0017] Step S302 : Calculate the pixel-by-pixel similarity vector of the difference between the prototype feature and the defect feature.
[0018] As a preferred embodiment of the self-supervised industrial defect detection method based on prototype memory guidance described in the present invention, the mathematical expression of step S301 is:
[0019]
[0020] Among them, F m Represents the prototype feature, represents the i-th layer of prototype features, ∈[1,3] represents the i-th layer of feature extractor, N represents the number of input images 0 in the training set, Represents the i-th layer feature of the n-th input image, n∈[1,N].
[0021] As a preferred embodiment of the self-supervised industrial defect detection method based on prototype memory guidance of the present invention, the mathematical expression of step S302 is:
[0022]
[0023] in, represents the i-th layer of prototype features, Represents the i-th layer feature extracted by the teacher network, upsamples the feature similarity to the same scale, and splices it along the channel dimension to obtain the pixel-by-pixel similarity vector F between the prototype feature and the defect feature mt , the mathematical expression of the prototype feature is:
[0024]
[0025] Among them, F mt represents the prototype feature, and Ψ represents the upsampling function.
[0026] As a preferred embodiment of the self-supervised industrial defect detection method based on prototype memory guidance according to the present invention, the specific process optimized in step S400 is as follows:
[0027] The synthesized defect image is input into the teacher network and the student network. The student network repairs the defect features into normal features, and the teacher network retains the difference output of the defect part. The feature representation of the defect part by the teacher network and the student network will be different. The pixel-by-pixel similarity F of the features extracted by the teacher network and the student network is calculated. ts To express the difference between the two, the mathematical expression is:
[0028]
[0029] in, represents the i-th layer feature extracted by the teacher network, Represents the i-th layer feature extracted by the student network, and calculates the pixel-by-pixel similarity of each layer feature of the teacher network and the student network The feature similarity is upsampled to the same scale and concatenated along the channel dimension to form a fusion feature. Its mathematical expression is:
[0030]
[0031] in,
[0032] The teacher-student network integrates the features F through the segmentation network ts Perform multi-scale discrimination;
[0033] The segmentation network includes a residual block, a void space pyramid pooling module and a segmentation head, and is finally optimized by the FocalLoss and L1 Loss loss functions. The output of the segmentation network is recorded as P 1 , the defect mask Mask is upsampled to the same value as P by bilinear interpolation 1 The mathematical expression for the same scale is:
[0034]
[0035]
[0036]
[0037]
[0038] Among them, h∈[1,H1], w∈[1,W1], H1 represents the height of the segmentation result P1, W1 represents the segmentation result P 1 The width, Represents the segmentation result P 1 A pixel value, L focal Represents the calculation result of FocalLoss, Indicates the calculation result of L1Loss, represents the optimization loss of the teacher-student segmentation network, γ represents the hyperparameter controlling the focusing coefficient, and λ1 and λ2 represent the loss weight hyperparameters.
[0039] The beneficial effects of the present invention are as follows: the prototype memory mechanism is utilized to enhance the model's ability to distinguish the differences between logical defects and normal images, thereby improving the model's detection performance for logical defects. To address the over-detection problem that occurs in the traditional teacher-student network, a saliency detection network and Perlin noise are used to synthesize defect images, thereby improving the distribution consistency between the synthesized images and the real defect images, and the student network is trained through self-supervision using the synthesized defect images, thereby suppressing the over-detection phenomenon. In addition, the average memory module can not only reduce memory overhead, but also eliminate the need to design complex retrieval algorithms, thereby saving the time spent on retrieving the memory library, which is conducive to improving detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 Schematic diagram of the limitations of traditional unsupervised industrial defect detection methods.
[0041] Figure 2 A schematic diagram of the basic flow of a prototype memory-guided self-supervised industrial defect detection method provided by one embodiment of the present invention.
[0042] Figure 3 A schematic diagram of synthetic defects in a prototype memory-guided self-supervised industrial defect detection method provided by one embodiment of the present invention.
[0043] Figure 4 A schematic diagram illustrating the visualization of the detection results of various methods on the MVTec AD dataset for a prototype memory-guided self-supervised industrial defect detection method provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0044] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.
[0045] Example 1
[0046] Reference Figure 2 and Figure 3 , as one embodiment of the present invention, provides a self-supervised industrial defect detection method based on prototype memory guidance, comprising:
[0047] Step S100: obtaining a foreground mask F of the input image O based on the saliency detection network, generating random noise P of the input image O using a Perlin noiser, constraining the random noise P to the foreground region using the foreground mask F to obtain a mask image M, and synthesizing a defect image S using the mask image M and the out-of-distribution image A.
[0048] Step S200: Use the teacher-student network to extract features of the input image O and the defect image S, and obtain the pixel-by-pixel similarity F between the features.ts ;
[0049] Step S300: extracting the prototype feature F of the input image O based on the average memory module AMM mt ;
[0050] Step S400: Calculate the pixel-by-pixel similarity F ts and F mt After multi-scale discrimination, the FocalLoss and L1 Loss loss functions are used for joint optimization.
[0051] The specific method of synthesizing the defect image S in step S100 is:
[0052] The transparency factor is introduced in the defect synthesis process to balance the fusion of the input detection image and the out-of-distribution image. Its mathematical expression is:
[0053]
[0054] Among them, S represents the defect image, O represents the input image, ⊙ represents element-by-element multiplication, and M represents the mask image. represents the inversion of the mask image M, β∈[0,1] represents the transparency factor, and Aug represents three operations randomly selected from equalization, exposure, toner separation, sharpness, and inversion.
[0055] Reference Figure 3 The defects synthesized by this method are concentrated on the foreground objects of the image, which is more consistent with the distribution of most real defects. Using the defect images obtained by this method for training can enhance the robustness of the model to background interference, alleviate the above-mentioned over-detection problem, and thus improve the detection accuracy.
[0056] In step S200, the student network matches the input image features extracted by the teacher network through cosine similarity, including:
[0057] The synthesized defect image S is input into the student network, and the input image O corresponding to the synthesized defect image S is input into the teacher network. The student network matches the features of the input image O extracted by the teacher network through cosine similarity, reconstructs the features, and repairs the defect features. The mathematical expression is:
[0058]
[0059]
[0060] in, represents the i-th layer feature extracted by the teacher network, represents the i-th layer feature extracted by the student network, h∈[1,H i ],w∈[1,W i], H i Indicates the height of the i-th layer feature, W i Represents the width of the i-th layer feature, ∈ [1, 3], represents the two-dimensional defect score map obtained by calculating the cosine similarity along the channel dimension, L cos represents the loss function for optimizing the student network.
[0061] The specific process of extracting prototype features in step S300 includes:
[0062] In step S301, prototype features are extracted using the average memory module (AMM), and logical defects in the foreground information are detected using the prototype features. The introduction of the average memory module not only represents uncommon normal samples but also effectively reduces feature redundancy, thereby obtaining representative prototype features. For each type of industrial product, the memory library only stores one prototype feature. The average memory module not only reduces memory overhead but also eliminates the need for complex retrieval algorithms, thus saving the time spent searching the memory library.
[0063] Step S302 : Calculate the pixel-by-pixel similarity vector of the difference between the prototype feature and the defect feature.
[0064] The mathematical expression of step S301 is:
[0065]
[0066] Among them, F m Represents the prototype feature, represents the i-th layer of prototype features, ∈[1,3] represents the i-th layer of feature extractor, N represents the number of input images 0 in the training set, Represents the i-th layer feature of the n-th input image, n∈[1,N].
[0067] The mathematical expression of step S302 is:
[0068]
[0069] in, represents the i-th layer of prototype features, Represents the i-th layer feature extracted by the teacher network, upsamples the feature similarity to the same scale, and splices it along the channel dimension to obtain the pixel-by-pixel similarity vector F between the prototype feature and the defect feature mt , the mathematical expression of the prototype feature is:
[0070]
[0071] Among them, F mt represents the prototype feature, and Ψ represents the upsampling function.
[0072] The specific process of optimization in step S400 is:
[0073] The synthesized defect image is input into the teacher network and the student network. The student network repairs the defect features into normal features, and the teacher network retains the difference output of the defect part. The feature representation of the defect part by the teacher network and the student network will be different. The pixel-by-pixel similarity F of the features extracted by the teacher network and the student network is calculated. ts To express the difference between the two, the mathematical expression is:
[0074]
[0075] in, represents the i-th layer feature extracted by the teacher network, Represents the i-th layer feature extracted by the student network, and calculates the pixel-by-pixel similarity of each layer feature of the teacher network and the student network The feature similarity is upsampled to the same scale and concatenated along the channel dimension to form a fusion feature. Its mathematical expression is:
[0076]
[0077] in,
[0078] The teacher-student network integrates the features F through the segmentation network ts Perform multi-scale discrimination;
[0079] The segmentation network includes a residual block, a void space pyramid pooling module and a segmentation head, and is finally optimized by the Focal Loss and L1 Loss loss functions. The output of the segmentation network is recorded as P 1 , the defect mask Mask is upsampled to the same value as P by bilinear interpolation 1 The mathematical expression for the same scale is:
[0080]
[0081]
[0082]
[0083]
[0084] Among them, h∈[1,H1], w∈[1,W1], H1 represents the height of the segmentation result P1, W1 represents the segmentation result P 1 The width, Represents the segmentation result P 1 A pixel value, L focalRepresents the calculation result of Focal Loss, Indicates the calculation result of L1 Loss, represents the optimization loss of the teacher-student segmentation network, γ represents the hyperparameter controlling the focusing coefficient, and λ1 and λ2 represent the loss weight hyperparameters.
[0085] The prototype memory mechanism is used to enhance the model's ability to distinguish the differences between logical defects and normal images, thereby improving the model's detection performance for logical defects. To address the over-detection problem that occurs in the traditional teacher-student network, a saliency detection network and Perlin noise are used to synthesize defect images, thereby improving the distribution consistency between the synthesized images and the real defect images. The student network is trained through self-supervision using the synthesized defect images, thereby suppressing the over-detection phenomenon. The average memory module not only reduces memory overhead, but also eliminates the need to design complex retrieval algorithms, thereby saving the time spent on retrieving the memory library, which is conducive to improving detection performance.
[0086] Example 2
[0087] Reference Figure 4 , which is another embodiment of the present invention. Different from the first embodiment, this embodiment provides an experimental verification of a self-supervised industrial defect detection method based on prototype memory guidance. In order to verify and illustrate the technical effects adopted in this method, this embodiment adopts a traditional technical solution and the method of the present invention for comparative testing, and compares the test results by means of scientific demonstration to verify the real effect of this method.
[0088] The proposed method is experimentally compared with mainstream industrial defect detection methods on the MVTec AD benchmark dataset in the field of industrial defect detection. This dataset has a total of 15 categories, including 10 object classes and 5 texture classes. The training set contains 3629 normal images. For each category, hundreds of normal images are used for training, with image sizes ranging from 700×700 to 1024×1024. The test set contains 1725 samples, including normal images and defect images, and all defect images are provided with pixel-level annotations.
[0089] For fair comparison, all defect synthesis methods use images from the DTD dataset as out-of-distribution data; this dataset is a continuously updated texture dataset consisting of 5,640 images, divided into 47 categories based on human perception, with 120 images in each category; in terms of defect synthesis, compared with other texture datasets (such as ImageNet), the DTD dataset can meet the needs of defect detection, and it occupies less memory space, making it easier to apply in practice.
[0090] The input image dimension of the proposed model is set to 256×256 during training; the pre-trained ResNet18 model is used as the backbone network of the average memory module and the teacher network; the atrous spatial pyramid pooling module in the segmentation network contains three atrous convolutional layers, and their expansion rates are set to 6, 12, and 18, respectively; in the training stage, the batch size is set to 12, and the model parameters are updated using the stochastic gradient descent method, with a total of 5000 iterative training times; the first 1000 times are used to train the student network, and the next 4000 times are used to train the two segmentation networks; the focusing coefficient hyperparameter γ is set to 4, the hyperparameter θ that controls the normalization range is set to 0.1, and the hyperparameters λ1 and λ2 used to control the final segmentation loss weight are both set to 0.5.
[0091] The hardware configuration used in the experiment is Inter(R) i5-12500H processor and Nvidia RTX 4060Ti graphics card.
[0092] In the experiment, AUROC, widely used in the field of industrial defect detection, was used as a performance evaluation metric. The ROC curve is plotted with the false positive rate (FPR) as the horizontal axis and the recall rate (TPR) as the vertical axis, describing the relationship between recall rate and false positive rate. AUC, representing the area under the curve, measures the performance described by the ROC curve. The formulas for calculating recall rate and false positive rate are as follows:
[0093]
[0094] Among them, TP represents the number of correctly detected positive examples, FN represents the number of positive examples incorrectly detected as negative examples, FP represents the number of negative examples incorrectly detected as positive examples, and TN represents the number of correctly detected negative examples.
[0095] In the ROC curve, there is a weighted relationship between the TPR value and the defect area. A larger correctly segmented area will greatly compensate for the quantitative indicators of multiple small defect areas with segmentation errors. The predicted results and the true results are divided into N areas according to the connected domain. Then, the intersection of the true result Gn and the predicted result P in each area is calculated in turn, as well as the ratio of the intersection to the true result area Gn. Finally, the weighted average is calculated according to the N connected domains.
[0096]
[0097] Table 1: Comparison of image-level AUROC indicators of various industrial defect detection methods on the MVTec AD dataset.
[0098]
[0099]
[0100] Table 2: Comparison of pixel-level AUROC indicators of various industrial defect detection methods on the MVTec AD dataset.
[0101]
[0102]
[0103] Experimental results analysis
[0104] On the MVTec AD dataset, the proposed method is compared with mainstream unsupervised defect detection methods, including US, P-SVDD, DRAEM, PatchCore, RD, and DeSTSeg. Among them, US introduces knowledge distillation to unsupervised industrial defect detection; P-SVDD encodes image blocks separately through an encoder, improving the overall detection performance; DRAEM uses an autoencoder to reconstruct the synthesized defect image and uses a discriminant network to finely distinguish the reconstructed image from the synthesized defect image, thereby more accurately detecting defects; PatchCore constructs normal patterns by randomly sampling feature maps, achieving efficient feature screening; RD proposes heterogeneous distillation of the teacher network and the student network, thereby alleviating the over-generalization problem of the student network; DeSTSeg, based on the teacher-student network framework, uses a segmentation network to adaptively fuse the similarities of multi-level features, further improving positioning accuracy.
[0105] To analyze the detection performance of various methods for different categories of industrial products, the 15 MVTec AD categories were divided into texture and object categories, and image-level AUROC (AUROC) metrics were compared. As shown in Table 1, all industrial defect detection methods demonstrated strong performance on textured images. Compared to other methods, the proposed method achieved the best defect detection results on textured test images. Compared to the baseline model DeSTSeg, its average image-level AUROC metric improved from 99.3% to 99.8%. The defect detection results for object-class test images show that the proposed method achieved an average image-level AUROC of 99.1%, a significant improvement of 1.6% over the DeSTSeg model. Averaging the detection results across all categories, the proposed method improved the average image-level AUROC metric from 98.1% to 99.3% compared to the DeSTSeg model.
[0106] Unlike image-level detection, pixel-level industrial defect detection requires refined discrimination of each pixel in the image. Therefore, pixel-level defect detection is a more challenging task. As shown in Table 2, the proposed method significantly improves detection performance for both texture and object test images compared to existing mainstream methods. For texture images, the proposed method achieves a best-in-class average pixel-level AUROC of 98.7%, while for object images, the proposed method improves the best-in-class average pixel-level AUROC to 99.1%. Specifically, for transistors containing a large number of logic defects, the proposed method improves the average pixel-level AUROC by 9.1% compared to the DeSTSeg model. The proposed method also achieves the best average pixel-level AUROC on the average detection results of the entire MVTec AD dataset. These experimental results demonstrate the effectiveness of the proposed industrial defect detection framework. This is due to the improved robustness of the model to background interference achieved through the proposed defect synthesis method. The prototype memory-guided branch provides representative prototype information for detection, effectively detecting logic defects.
[0107] To further analyze the performance of each defect detection method, Table 3 compares the pixel-level PRO and AP metrics. Unlike the pixel-by-pixel AUROC metric, the PRO score treats defect regions of all sizes equally, while the Average Precision (AP) metric better evaluates detection results under imbalanced class conditions. As shown in Table 3, the proposed method still outperforms other methods on the more challenging pixel-level PRO and AP metrics. In particular, the proposed method achieves a 2.5% improvement in AP compared to the DeSTSeg model.
[0108] In order to qualitatively evaluate the detection performance of the proposed method, typical test images were selected in the experiment, and the final segmentation results of each detection method were compared by heat map visualization to intuitively show the defect intensity in different areas. Figure 4The paper presents heat maps of test images, true defect labels, and the segmentation results of each method, where Input and GT represent the test image and the corresponding true defect label, respectively. The first column shows the defect detection results of industrial transistors. For such defective images, it is necessary not only to detect the misplaced electronic components but also to mark their locations in the normal image. Therefore, such defects can be considered a combination of two defects: misalignment and missing targets. Most existing unsupervised defect detection methods can only detect the misaligned portion, while the proposed method can accurately detect both defects through a memory mechanism. Furthermore, when detecting small surface defects such as the blurred printing on the pharmaceutical capsules in the fourth column, the scratches on the tablets in the second-to-last column, and the worn screws in the last column, the proposed method effectively alleviates the problem of over-detection and achieves more accurate detection results.
[0109] Table 3: Comparison of pixel-level PRO and AP indicators of various industrial defect detection methods on the MVTecAD dataset.
[0110]
[0111] It should be appreciated that embodiments of the present invention can be implemented or practiced by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The methods can be implemented in a computer program using standard programming techniques, including a non-transitory computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner, according to the methods and figures described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, the program can be run on a programmed application-specific integrated circuit for this purpose.
[0112] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A self-supervised industrial defect detection method based on prototype memory guidance, characterized in that: include: Step S100: obtaining a foreground mask F of the input image O based on the saliency detection network, generating random noise P of the input image O using a Perlin noiser, constraining the random noise P to the foreground region using the foreground mask F to obtain a mask image M, and synthesizing a defect image S using the mask image M and the out-of-distribution image A. Step S200: Use the teacher-student network to extract features of the input image O and the defect image S, and obtain the pixel-by-pixel similarity F between the features. ts ; Step S300: extract the prototype features of the input image O based on the average memory module AMM, and calculate the pixel-by-pixel similarity vector F between the prototype features and the defect features. mt ; Step S400: Calculate the pixel-by-pixel similarity F ts and F mt After multi-scale discrimination, optimization is performed jointly through Focal Loss and L1 Loss loss functions; In step S200, the synthesized defect image S is input into the student network, and the input image O corresponding to the synthesized defect image S is input into the teacher network. The student network matches the features of the input image O extracted by the teacher network through cosine similarity, reconstructs the features, and repairs the defect features, wherein: , , in, represents the i-th layer feature extracted by the teacher network, represents the i-th layer feature extracted by the student network, h∈[1,H i ],w∈[1,W i ] , H i Indicates the height of the i-th layer feature, W i Represents the width of the i-th layer feature, i∈[1,3], C i Indicates the number of channels, represents the two-dimensional defect score map obtained by calculating the cosine similarity along the channel dimension, L cos represents the loss function for optimizing the student network; The specific process of optimization in step S400 is: The teacher-student network splits the network into two parts: ts Perform multi-scale discrimination; The segmentation network includes a residual block, a void space pyramid pooling module and a segmentation head, and is finally optimized by the Focal Loss and L1 Loss loss functions. The output of the segmentation network is recorded as P 1 , the defect mask Mask is upsampled to the same value as P by bilinear interpolation 1 The mathematical expression for the same scale is: , , , , in, , , H1 represents the segmentation result P 1 The height of W1 represents the segmentation result P 1 The width, Represents the segmentation result P 1 A pixel value of Represents the calculation result of Focal Loss, Indicates the calculation result of L1 Loss, represents the optimization loss of the teacher-student segmentation network, represents the hyperparameter that controls the focusing coefficient, and represents the loss weight hyperparameter.
2. The self-supervised industrial defect detection method based on prototype memory guidance according to claim 1, characterized in that: The specific method of synthesizing the defect image S in step S100 is: The transparency factor is introduced in the defect synthesis process to balance the fusion of the input detection image and the out-of-distribution image. Its mathematical expression is: , Among them, S represents the defect image, O represents the input image, ⊙ represents element-by-element multiplication, and M represents the mask image. represents the inversion of the mask image M, β∈[0,1] represents the transparency factor, and Aug represents three operations randomly selected from equalization, exposure, toner separation, sharpness, and inversion.
3. The self-supervised industrial defect detection method based on prototype memory guidance according to claim 2, characterized in that: The specific process of extracting prototype features in step S300 includes: Step S301, using an average memory module (AMM) to extract prototype features, and detecting logical defects in the foreground information using the prototype features; Step S302, calculating a pixel-by-pixel similarity vector of the difference between the prototype feature and the defect feature; The method for extracting prototype features by the average memory module AMM in step S301 is: , Among them, F m Represents the prototype feature, represents the i-th layer of prototype features, i∈[1,3] represents the i-th layer of feature extractor, N represents the number of input images O in the training set, Represents the i-th layer feature of the n-th input image, n∈[1,N]; The mathematical expression of the pixel-by-pixel similarity vector between the prototype feature and the defect feature in step S302 is: , in, represents the i-th layer of prototype features, Represents the i-th layer feature extracted by the teacher network, upsamples the feature similarity to the same scale, and splices it along the channel dimension to obtain the pixel-by-pixel similarity vector F between the prototype feature and the defect feature mt , the mathematical expression is: , Where Ψ represents the upsampling function.
Citation Information
Patent Citations
Construction method and application of OLED novel display device surface defect detection model
CN115619743A
Unsupervised defect detection method of tandem knowledge distillation added with Transform
CN116468667A