An unsupervised defect detection method based on image reconstruction

CN118397373BActive Publication Date: 2026-09-29SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410645832.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-23
Publication Date
2026-09-29
Estimated Expiration
2044-05-23

AI Technical Summary

Technical Problem

[0005]本发明的目的是提供一种基于图像重建的无监督缺陷检测方法,解决了现有技术无法精准检测缺陷,并定位出缺陷的位置,采用一个双分支编码器-解码器结构,利用正常图像和异常图像结构上的差异,分别对图像和结构信息进行编码,共用一个解码器重建图像,该方法对于正常图像和异常图像结构上存在明显缺陷的识别有很好的效果

Benefits of technology

[0039]1、利用结构信息和图像特征之间的关联进行缺陷检测,本发明引入了一个记忆模块记录图像结构特征,在读取时通过计算相似度进行权重平均检索特征,考虑了图像数据结构的多样性,有助于更全面地捕获图像的特征信息。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118397373B_ABST
    Figure CN118397373B_ABST
Patent Text Reader

Abstract

The application discloses an unsupervised defect detection method based on image reconstruction, which comprises the following steps: different structural features of normal image data are stored in different items of a memory module, and a final structural feature is obtained by calculating the similarity between the structural feature and each memory item and then performing weight averaging; the structural similarity (SSIM) is used to measure the similarity between two images; and the application solves the problem that the prior art cannot accurately detect defects and locate the positions of the defects. The application proposes a double-branch encoder structure, which encodes the image and the edge structure respectively, and the reconstructed image is subjected to structure extraction again as a regularization item, so that the image reconstruction error and the structure error are simultaneously constrained, the reconstruction difference is measured by using the structural similarity, the structural difference of the image can be more accurately reflected, and the defect detection effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a defect detection method, specifically an unsupervised defect detection method based on image reconstruction. Background Technology

[0002] Traditional methods primarily rely on known defect characteristics and image processing techniques for defect detection. For example, some defect areas have significantly higher pixel values ​​than normal areas, allowing for defect detection using threshold analysis. However, due to the diversity and unknown nature of defects, different detection methods are often required for different defect types, and these methods may not be applicable to new defects, potentially necessitating complex post-processing. With the application of deep learning, supervised deep learning models can effectively detect defects when the defect type is known and a large number of labeled samples are available. Some researchers have designed models based on the U-Net network, using multi-scale feature fusion to improve detection accuracy. However, training these models typically requires a large amount of labeled data, which is difficult to obtain in practice.

[0003] Therefore, unsupervised defect detection methods have attracted widespread attention. These methods only require readily available normal samples for model training and can achieve defect detection without using real defect samples. The core idea of ​​these methods is to reconstruct an image most similar to the original image and compare it with the original image, detecting defects based on the differences between the two. Some researchers use image reconstruction methods, such as encoder-decoder reconstruction based on image processing (AE), and can locate defects based on the reconstruction error between the input image and the reconstructed image. However, these methods suffer from poor reconstruction quality, leading to poor reconstruction of normal regions in defective images and easily causing false detections of normal pixels. Other researchers introduce memory after the encoder to store normal image features, and retrieve features from memory for reconstruction based on similarity during testing. These methods reconstruct by memorizing features from normal images, consuming a large amount of memory, and do not utilize the correlation between structural information and image texture, making the reconstruction effect difficult to control.

[0004] In summary, the problems with existing technologies are as follows: Traditional methods require designing corresponding defect detection methods based on known defect characteristics, thus only addressing a limited range of defects and failing to tackle novel defects. Supervised deep learning typically requires a large amount of labeled data; however, in reality, due to strict control over defect rates in production processes, most samples are normal, with only a small number being defective, and precisely labeled defective samples are even rarer. Unsupervised methods based on image reconstruction suffer from unpredictable reconstruction results, often resulting in overly good or underlying reconstructions. Previous methods incorporating memory ignore the correlation between structure and texture and consume significant memory, easily leading to missed or false detections of defects. Summary of the Invention

[0005] The purpose of this invention is to provide an unsupervised defect detection method based on image reconstruction, which solves the problem that existing technologies cannot accurately detect defects and locate their positions. It adopts a dual-branch encoder-decoder structure, which utilizes the structural differences between normal and abnormal images to encode image and structural information respectively, and uses a common decoder to reconstruct the image. This method has a good effect on the identification of obvious defects in the structure of normal and abnormal images.

[0006] To achieve the above objectives, this invention provides an unsupervised defect detection method based on image reconstruction, the method comprising:

[0007] (S100) The present invention stores different structural features of normal image data into different items of the memory module, calculates the similarity between the structural features and each memory item, and obtains the final structural features by weighted averaging.

[0008] (S200) This invention uses structural similarity (SSIM) to measure the similarity between two images;

[0009] (S300) The present invention uses a dual-branch encoder-decoder to obtain the reconstructed image;

[0010] In this dual-branch encoder, one encoder extracts features from the normal input image and is called the image encoder, while the other encoder extracts features from the edge structure of the normal input image and is called the structure encoder.

[0011] The image encoder, the structure encoder, and the decoder all adopt the form of the U-Net network architecture.

[0012] Preferably, in step (S300), the present invention uses the Canny edge detection function from the scikit-image library to extract the edge structure of the input image and inputs it into the structural encoder. The original image is input into the image encoder (the image encoder has three layers of encoded features; the bridging connection is the concatenation of channel dimensions, generally referred to as a "bridge" in Unet). The features obtained by the structural encoder are used as query terms and input into the memory module for retrieval. The structural features and the image features from the image encoder are concatenated and fused in the channel dimension and input together into the decoder to obtain the reconstructed image. The channel changes of the image encoder are as follows: C3->C 64 ->C 128 ->C 256 The channel changes of the structure encoder are as follows: C1->C 64 ->C 128 ->C 256 The decoder's channel changes are as follows: C 512 ->C 256 ->C128 ->C 64 ->C3. Here, C3 indicates that the number of channels in the convolution kernel is 3; the other values ​​have similar meanings. The number of channels is only written up to 256 here to reduce the number of model parameters and computational complexity. Furthermore, experiments have shown that increasing the number of channels does not bring a significant performance improvement.

[0013] Preferably, during decoding and reconstruction, this invention only uses the coding features of the last layer of the image encoder. For the structural information in the structural encoder, a "bridge" connection is also used for fusion in the channel dimension. Each layer upsampling corresponds to the addition of features from the structural encoder for splicing and fusion to obtain rich structural features. To avoid over-reconstruction of the image in defective areas, features from each layer of the image encoder are not spliced ​​in here. For the reconstructed image, this invention again uses the scikit-image library to perform Canny edge detection to extract the structure. By calculating the L1 norm of the reconstructed image structure and the original image structure, the image reconstruction error and structural error are constrained at the same time.

[0014] Preferably, in step (S200), the structural similarity SSIM considers the correlation between local pixel regions of the input image and the reconstructed image, and jointly evaluates the differences in brightness, contrast and structure to obtain the defect segmentation result;

[0015] The Structural Similarity SSIM defines the similarity of two K×K images p and q in terms of brightness, contrast, and structure. In image quality assessment, local SSIM calculation is more effective than global SSIM calculation. Therefore, the value of K here does not take the entire image size and can be changed according to the actual assessment situation. Generally, a sliding window with a value of 11 is set. In the experiment, this invention uses a value of 11. Its specific representation is as follows:

[0016] SSIM(p,q)=l(p,q)*c(p,q)*s(p,q) (5)

[0017]

[0018]

[0019]

[0020] Substituting equations (6) to (8) into equation (5), we get:

[0021]

[0022] In equations (6) to (8), u p and u q Let σ represent the mean of image p and image q, respectively. p and σq Let c1 and c2 represent the variances of image p and image q, respectively, and set c1 = 0.01 and c2 = 0.03, and SSIM(p,q) ∈ [-1,1].

[0023] Preferably, when SSIM(p,q) = 1, it means that image p and image q are the same.

[0024] Preferably, in step (S100), the dataset used in this invention is the publicly available MVTec AD and a private PCB board dataset. The PCB board dataset consists of five different structural data, with 50 normal images and 25 abnormal images for each data type. The abnormal data is only used for testing. The input data are all 3-channel color images. When extracting structural information, in order to avoid interference from the background area, this invention uses the Canny edge detection function in the scikit-image library. The obtained single-channel edge image is sent to the structural encoder to obtain structural features, and then sent to the memory module for storage.

[0025] Preferably, the memory module of the present invention includes M items that record the structural features of a normal image, and it can store and retrieve information.

[0026] Preferably, the present invention uses p m ∈R C Let represent the m-th storage item in the memory module, where (m = 1, ..., M). The features obtained from the structure encoder are used as each query item, denoted by q. k ∈R C Let q represent the image height (k = 1, ..., H x W), where H represents the image height, W represents the image width, and C represents the number of channels. k The corresponding features are input into the memory module for storage or retrieval.

[0027] This invention calculates q for each query item. k and all items p in the memory module m The cosine similarity is calculated, and then the corresponding weights are obtained through the softmax function. The retrieved features are obtained by averaging these weights with the weights of each memory item. The weights in the retrieval process are represented as follows:

[0028]

[0029] In equation (1), w k,m Let p represent the attention weight of the k-th query and the m-th item in the memory module, where (m = 1, ..., M). m T Item p m The transpose of p, where p m A dimension can be represented as MXC, which, after transpose, becomes CXM. kThis indicates that each query item is calculated, where (k = 1, ..., HXW), and its dimension can be represented as HXWXC. The dimension obtained by matrix multiplication is HXWXM.

[0030] Final retrieval features It is expressed as follows:

[0031]

[0032] In equation (2), p m ∈R C This represents the m-th storage item in the memory module, where (m = 1, ..., M), w k,m This represents the attention weight of the m-th item in the k-th query and memory module;

[0033] Similarly, when updating items in the memory module, cosine similarity is used to calculate attention weights, and then the closest query item is selected and stored in the memory module. This is to avoid the features stored in the memory module being too similar and lacking distinctiveness. The weights of the stored procedure are represented as follows:

[0034]

[0035] In equation (3) v m,k Let q represent the attention weights of the m-th and k-th query terms in the memory module, where (m = 1, ..., M). k Represents each query term, where (k = 1, ..., HXW), p m T Item p m The transpose, the final update It is expressed as follows:

[0036]

[0037] In equation (4), q k Represents each query item, v m,k This represents the attention weights of the m-th and k-th query terms in the memory module.

[0038] This invention provides an unsupervised defect detection method based on image reconstruction, which solves the problem that existing technologies cannot accurately detect defects and locate their positions, and has the following advantages:

[0039] 1. By utilizing the correlation between structural information and image features for defect detection, this invention introduces a memory module to record image structural features. During reading, features are retrieved by weighted averaging through similarity calculation. This takes into account the diversity of image data structures and helps to capture image feature information more comprehensively.

[0040] 2. Considering the structural differences between normal and abnormal images, this invention proposes a dual-branch encoder structure to better capture image structural information and features. This structure encodes the image and edge structures separately, with the edge structures occupying an additional encoder branch, allowing for the extraction of richer structural features. The reconstructed image undergoes further structural extraction as a regularization term, simultaneously constraining both image reconstruction error and structural error.

[0041] 3. Using structural similarity as a metric to reconstruct differences takes into account the correlation between local image regions and comprehensively evaluates brightness, contrast, and structural information, rather than simply comparing individual pixel values. This method can more accurately reflect the structural differences of the image and improve the effect of defect detection. Without structural similarity, training with only a reconstruction error metric function will lead to over-reconstruction of the image.

[0042] 4. The method of this invention is unsupervised. The memory stores structural information that takes into account structural characteristics. It has good experimental results for images with simple backgrounds, less extracted structural interference information, and structural defects. Attached Figure Description

[0043] Figure 1 This is a neural network architecture diagram provided in Embodiment 1 of the present invention.

[0044] Figure 2 This is a diagram illustrating the reading process of the memory module provided in Embodiment 1 of the present invention.

[0045] Figure 3 This is a storage process diagram of the memory module provided in Embodiment 1 of the present invention.

[0046] Figure 4 This invention provides a diagram showing normal data, abnormal data, and their corresponding structural information for Embodiment 1. Detailed Implementation

[0047] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0048] Example 1

[0049] An unsupervised defect detection method based on image reconstruction, the method comprising:

[0050] (S100) Store different structural features of normal image data into different items of the memory module, calculate the similarity between the structural features and each memory item, and obtain the final structural features by weighted averaging.

[0051] Normal images typically possess a regular structure, and defects often disrupt this structure. This invention utilizes the relationship between image structure and image texture features for defect detection. A structural feature memory module stores different structural features of normal image data into different entries within the memory module. By calculating the similarity between the structural features and each memory entry, a weighted average is applied to obtain the final structural feature. The datasets used in this invention are the publicly available MVTecAD dataset and a proprietary PCB board dataset. The PCB board dataset contains data in five different structural categories, with 50 normal images and 25 abnormal images in each category. Defects include component breakage, component displacement, soldering abnormalities, and scratches. For the MVTecAD dataset, files in the training folder are used for training, and files in the test folder are used for testing. Normal images in the PCB dataset are used for training, while abnormal data is used for testing. The training process employs data augmentation techniques including horizontal and vertical image flipping, rotation (angle values ​​[-20, 20]), translation (distance values ​​[-10, 10]), brightness transformation (values ​​[-0.5, 1.5]), and contrast transformation (values ​​[-0.5, 1.5]), all with a probability of 0.5. The input data consists of 3-channel color images, all resized to 512×512. To avoid interference from background areas when extracting structural information, this invention does not use the edge structure extraction method from the OpenCV library. Instead, it uses the Canny edge detection function from the scikit-image library. The resulting single-channel edge image is fed into a structure encoder to obtain structural features (the input to the structure encoder is the edge image; edge extraction by the scikit-image library is more accurate than other libraries), and then stored in the memory module. The memory module of this invention contains M storage items that record typical features of normal images; it can store and retrieve information.

[0052] This invention uses p m ∈R C Let represent the m-th storage item in the memory module, where (m = 1, ..., M). The features obtained from the structure encoder are used as each query item, denoted by q. k ∈R C Let q represent the image height (k = 1, ..., H x W), where H represents the image height, W represents the image width, and C represents the number of channels. k The corresponding features are input into the memory module for storage or retrieval.

[0053] When retrieving features from the memory module, the most relevant features are not necessarily retrieved. Considering the similarity between different normal image structure features, this invention calculates q for each query item. kand all items p in the memory module m The cosine similarity is calculated, and then the corresponding weights are obtained through the softmax function. The retrieved features are obtained by averaging these weights with the weights of each memory item. The weights in the retrieval process are represented as follows:

[0054]

[0055] In equation (1), w k,m Let p represent the attention weight of the k-th query and the m-th item in the memory module, where (m = 1, ..., M). m T Item p m The transpose of p, where p m A dimension can be represented as MXC, which, after transpose, becomes CXM. k This represents the computation of each query item, where (k = 1, ..., H×W), and its dimension can be represented as H×W×C. The dimension obtained after matrix multiplication is H×W×M. The final retrieval features are... It is expressed as follows:

[0056]

[0057] In equation (2), p m ∈R C This represents the m-th storage item in the memory module, where (m = 1, ..., M), w k,m This represents the attention weight of the m-th item in the k-th query and memory module.

[0058] When storing structural features, considering the diversity of structural features, this invention only records the most typical structural features. When updating items in the memory module, this invention also uses cosine similarity to calculate attention weights, and then selects the closest query item to store in the memory module. This avoids features stored in the memory module being too similar and lacking distinctiveness. The weights in the storage process are represented as follows:

[0059]

[0060] In equation (3) v m,k Let q represent the attention weights of the m-th and k-th query terms in the memory module, where (m = 1, ..., M). k This indicates that each query term is computed, where (k = 1, ..., HXW), p m T Item p m The transpose of .

[0061] Final Update It is expressed as follows:

[0062]

[0063] In equation (4), q k Represents each query item, v m,k This represents the attention weights of the m-th and k-th query terms in the memory module.

[0064] (S200) Use Structural Similarity (SSIM) to measure the similarity between two images.

[0065] Commonly used reconstruction error metrics are the L1 or L2 norm of the corresponding pixels. However, these pixel-level loss functions ignore the correlation between adjacent pixels and cannot effectively segment anomalous regions that differ structurally but have similar pixel values. This invention proposes using Structural Similarity (SSIM) to measure the similarity between two images. This metric considers the correlation between local pixel regions in the input and reconstructed images, jointly evaluating differences in brightness, contrast, and structure to obtain satisfactory defect segmentation results.

[0066] This invention uses SSIM to measure reconstruction similarity. SSIM defines the similarity of two K×K images p and q in terms of brightness, contrast, and structure. In image quality assessment, local SSIM calculation is better than global calculation. Therefore, the value of K here is not the size of the entire image. It can be changed according to the actual assessment situation. Generally, a sliding window with a value of 11 is set. In the experiment, this invention uses a value of 11, which is calculated according to the formula in the existing literature ([1] Wang Z, Pan W, Cuppens-Boulahia N, et al. Image quality assessment: From error visibility to structural similarity[J].2013), and is specifically expressed as follows:

[0067] SSIM(p,q)=l(p,q)*c(p,q)*s(p,q) (5)

[0068]

[0069]

[0070]

[0071] Substituting equations (6) to (8) into equation (5), we get:

[0072]

[0073] In equations (6) to (8), u p and u q Let σ represent the mean of image p and image q, respectively.p and σ q Let c1 and c2 represent the variances of images p and q, respectively, and set c1 = 0.01 and c2 = 0.03. SSIM(p,q) ∈ [-1,1], and when SSIM(p,q) = 1, it indicates that images p and q are the same. To calculate the structural similarity between the entire input image and the reconstructed image, this invention uses a K×K sliding window to calculate the SSIM value on the image. Items in the memory module are selected based on the similarity of the feature vectors.

[0074] (S300) uses a dual-branch encoder-decoder to obtain the reconstructed image.

[0075] This invention employs the Canny edge detection function from the scikit-image library to extract the edge structure of the input image, which is then fed into a structural encoder. The original image is fed into an image encoder. Both encoders and decoders utilize the U-Net network architecture, storing the structural features obtained from the structural encoder in a memory module. The features encoded by the structural encoder are used as query terms and input into the memory module for retrieval. The resulting structural features and image features from the image encoder are concatenated and fused along the channel dimension and then fed into the decoder to obtain the reconstructed image. To avoid reconstructing defective regions in abnormal images, this invention does not use the corresponding branches in the image encoder for "bridging" connections during decoding and reconstruction; instead, it only uses the encoded features from the last layer. However, for the structural information in the structural encoder, this invention uses "bridging" connections for fusion and joint upsampling to reconstruct the image. To regularize the model and constrain both image reconstruction error and structural error, this method enables the detection and localization of defects. Since the parameters of the memory module model are trained only on the structural information of normal samples, the model can reconstruct the structure of normal samples well. For defective samples, the defective regions will lead to larger reconstruction errors. This method is highly effective in identifying obvious defects in the structure of both normal and abnormal images. This invention performs L1 norm calculation on the reconstructed image again through structure extraction, while simultaneously constraining both image reconstruction error and structural error. The loss function used throughout the training process is expressed as follows:

[0076]

[0077] In equation (10), X represents the input image. The image is represented by α = 1 and β = 0.5. The calculation process is shown in equation (9).

[0078] like Figure 1 As shown, this is a neural network architecture diagram provided in Embodiment 1 of the present invention. Figure 1As can be seen, the model has two input branches: an image encoder that extracts image features and a structural encoder that extracts structural features. The output has only one branch, used for decoding and reconstructing the image. During decoding and reconstruction, to avoid over-reconstructing defect areas, only the last layer of features from the image encoder is used. However, to better reconstruct the image structure, richer structural information is needed; therefore, features from each layer of the structural encoder are used. The middle part of the model is a memory module storing structural features, with M terms, each storing one structural feature. Similarity is calculated between the query term and each term in the memory module, and then weighted and averaged to obtain a final output structural feature.

[0079] like Figure 2 The diagram shown illustrates the reading process of the memory module provided in Embodiment 1 of the present invention. Figure 2 It can be seen that during the reading process, the query item q is calculated. k The cosine similarity of each item in the memory module is used to obtain different weight information w after passing through softmax. k,m The final retrieval feature is obtained by multiplying the weights of each item in the memory module.

[0080] like Figure 3 As shown, the storage process of the memory module provided in Embodiment 1 of the present invention is described. Figure 3 It can be seen that during storage, the calculation involves a single item p in the memory module. m The cosine similarity between the query terms and HXW is then processed by softmax to obtain different weight information v. m,k When storing data, considering the differences between different structures, only the query item feature with the highest weight is selected for storage.

[0081] Experiment Example 1: Practical Application

[0082] This invention can be used for defect detection in abnormal images. It has a good detection effect on most defects, but the recognition effect on defects such as small scratches and small area damage is not very good. This may be because the defect area is too small and the structural features of the image cannot reflect the difference well, thus resulting in poor recognition.

[0083] like Figure 4 As shown, Embodiment 1 of the present invention provides normal data and abnormal data and their corresponding structural information diagrams, where A is a normal image; B is an image with defects; C is the structure extracted from the normal image; and D is the structure extracted from the image with defects. Figure 4It is evident that normal images and images with defects exhibit significant structural differences, which aligns with the observations made in this invention. Therefore, this invention extracts the structure of both images using the Canny edge detection function from the scikit-image library. The results clearly show that the structures of the normal and defective images are also distinctly different. This demonstrates that considering the structural information of both normal and defective images in defect detection is meaningful in this invention.

[0084] Although the present invention has been described in detail through the preferred embodiments above, it should be understood that the above description should not be considered as a limitation of the present invention. Various modifications and substitutions to the present invention will be apparent to those skilled in the art after reading the above description. Therefore, the scope of protection of the present invention should be defined by the appended claims.

Claims

1. An unsupervised defect detection method based on image reconstruction, characterized in that, The method includes: (S100) The present invention stores different structural features of normal image data into different items of the memory module, calculates the similarity between the structural features and each memory item, and obtains the final structural features by weighted averaging. (S200) This invention utilizes structural similarity To measure the similarity between two images; (S300) The present invention uses a dual-branch encoder-decoder to obtain the reconstructed image; In the dual-branch encoder, one encoder extracts features from the normal input image and is called the image encoder, while the other encoder extracts features from the edge structure of the normal input image and is called the structure encoder. The image encoder, the structural encoder, and the decoder all adopt the form of the U-Net network architecture; In step (S300), the present invention uses the scikit-image library for execution. The edge detection function extracts the edge structure of the input image and inputs it into the structure encoder. The original image is input into the image encoder. The features obtained by the structure encoder are used as query terms and input into the memory module for retrieval. The structural features and the image features of the image encoder are concatenated and fused in the channel dimension and input into the decoder to obtain the reconstructed image. During decoding and reconstruction, this invention only uses the encoded features of the last layer of the image encoder. For the structural information in the structure encoder, a "bridge" connection is also used for fusion in the channel dimension. Each layer upsampling corresponds to the features added to the structure encoder and then spliced ​​and fused to obtain rich structural features. For the reconstructed image, this invention again executes the scikit-image library. Edge detection extracts structure by calculating the structure of the reconstructed image and the original image. Norms simultaneously constrain image reconstruction errors and structural errors.

2. The unsupervised defect detection method according to claim 1, characterized in that, In step (S200), The structural similarity The correlation between local pixel regions of the input image and the reconstructed image is considered, and the differences are evaluated in terms of brightness, contrast and structure to obtain the defect segmentation result. The structural similarity Two were defined Image of size and images Similarity in brightness, contrast, and structure is crucial in image quality assessment. K The value is a fixed value, 11; its specific representation is as follows: (5) (6) (7) (8) Substituting equations (6) to (8) into equation (5), we get: (9) In equations (6) to (8), and Representing images respectively and images The mean, and Representing images respectively and images The variance, set and , .

3. The unsupervised defect detection method according to claim 2, characterized in that, The When, it represents an image. and images Same.

4. The unsupervised defect detection method according to claim 1, characterized in that, In step (S100), The datasets used in this invention are the publicly available MVTec AD dataset and the proprietary PCB board dataset. The PCB board dataset consists of five different structural data categories, with 50 normal images and 25 abnormal images for each category. Defects include component breakage, component displacement, soldering abnormalities, and scratches. For the MVTec AD dataset, files in the training folder are used for training, and files in the test folder are used for testing. For the PCB dataset, normal images are used for training, and abnormal data is used for testing. The data augmentation used in the training process includes horizontal flipping, vertical flipping, rotation, translation, brightness transformation, and contrast transformation of the image, with a probability of 0.5 for each transformation. The rotation angle is [-20, 20], the translation distance is [-10, 10], the brightness change is [-0.5, 1.5], and the contrast change is [-0.5, 1.5]. The input data are all 3-channel color images, all resized to 512×512. To avoid interference from background areas when extracting structural information, this invention does not use the edge extraction method from the OpenCV library, but instead uses the method from the scikit-image library for performing... The edge detection function generates a single-channel edge image, which is then fed into a structure encoder to obtain structural features, and finally stored in a memory module.

5. The unsupervised defect detection method according to claim 4, characterized in that, The memory module of the present invention includes This item records the structural features of a normal image; it stores and retrieves information.

6. The unsupervised defect detection method according to claim 5, characterized in that, This invention uses Represents the first in the memory module There are 1 storage item, of which The features obtained from the structural encoder are used as each query term, and... It means that among them , Indicates the height of the image. Indicates the width of the image. Indicates the number of channels, The corresponding features are input into the memory module for storage or retrieval. This invention calculates each query item and all items in the memory module The cosine similarity, and then through The function yields corresponding weights, which are then averaged with the weights of each memory item to obtain the retrieved features. The weights in the retrieval process are represented as follows: (1) In formula (1) Let m represent the attention weight of the k-th query and the m-th item in the memory module, where item The transpose of, its Dimensions are represented as MXC, after transpose. The dimension is CXM. ,in Its dimension is represented as HXWXC, and the dimension obtained by matrix multiplication is HXWXM; Final retrieval features It is expressed as follows: (2) Represents the first in the memory module There are 1 storage item, of which , This represents the attention weight of the m-th item in the k-th query and memory module; Similarly, when updating items in the memory module, cosine similarity is used to calculate attention weights, and then the closest query item is selected and stored in the memory module. The weights of the stored procedure are represented as follows: (3) In formula (3) This represents the attention weights of the m-th and k-th query terms in the memory module. item The transpose, the final update It is expressed as follows: (4)。

Citation Information

Patent Citations

  • Mobile phone screen defect segmentation method, device and equipment based on converged network

    CN111553929A

  • Unsupervised defect detection method based on quantization auto-encoder

    CN115375604A