Memory matching industrial defect detection method based on adaptive feature fusion
Through the memory matching method of adaptive feature fusion, defect samples are generated and adaptive hierarchical memory banks are designed. Combined with multi-scale feature fusion, the problem of insufficient detection accuracy and real-time in the existing technology is solved, and efficient defect detection and positioning is achieved to meet industrial application needs.
Patent Information
- Application Number
- CN202510358927.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
AI Technical Summary
The existing industrial defect detection technology is difficult to balance between detection accuracy and inference speed, especially when the defect samples are extremely limited, it is difficult to accurately locate defect areas, and the existing unsupervised methods have shortcomings in detection accuracy and real-time performance.
The memory matching method of adaptive feature fusion is adopted to generate multiple unknown defect samples through enhanced image-level defect simulation, and an adaptive hierarchical memory architecture is designed, combining multi-scale feature fusion and defect positioning optimization technology, and feature extraction and adaptive core central sampling are used to calculate Euclidean distance and defect difference information, and adaptive multi-scale feature fusion and loss function optimization are performed.
It significantly improves the accuracy of defect detection and positioning, realizes efficient real-time detection, and has an inference speed of 89FPS, meeting industrial application needs, and shows excellent detection and positioning performance on MVTecAD and BTAD datasets.
Smart Images

Figure CN120339195A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial defect detection, and particularly relates to a memory matching industrial defect detection method with adaptive feature fusion. Background Art
[0002] From aircraft wings to chip grains, industrial products are ubiquitous in modern society. Industrial defect detection is a key link in improving production quality and efficiency. The core goal of industrial defect detection is to accurately detect and locate defect areas in industrial images.
[0003] In recent years, significant progress has been made in industrial defect detection technologies based on deep learning. However, due to the scarcity of defect samples and the fact that they are far fewer than normal samples, the problem of uneven data distribution has become a major challenge for model learning. In addition, the cost of pixel-level annotation is high, making it difficult to popularize traditional supervised learning methods in practical applications. Therefore, unsupervised defect detection technologies have been widely studied and applied. Currently, the mainstream unsupervised methods can be divided into three categories: methods based on image reconstruction, methods based on synthesis, and methods based on embedding.
[0004] Although existing methods have improved the defect detection ability to a certain extent, there are still many limitations in practical applications. For example, it is difficult to balance detection accuracy and inference speed, especially when the number of defect samples is extremely limited, this problem is particularly prominent. In addition, industrial defects often exhibit subtle and variable morphological features, making it easy for existing methods to miss detections and false detections, and it is difficult to accurately locate the defect area. Therefore, there is an urgent need for an efficient technical solution that can improve the defect detection accuracy and positioning accuracy while meeting the requirements of real-time detection. Summary of the Invention
[0005] Object of the Invention: Aiming at the problems existing in the prior art, the present invention provides a memory matching industrial defect detection method with adaptive feature fusion, which can improve the defect detection accuracy and positioning accuracy while meeting the requirements of real-time detection.
[0006] Technical Solution: The present invention provides a memory matching industrial defect detection method with adaptive feature fusion, including the following steps:
[0007] Step 1: Use normal sample images to generate a variety of unknown defect samples through an enhanced image-level defect simulation method to obtain defect sample images, and the defect sample images retain the structural features of the normal sample images;
[0008] Step 2: Input the image of the sample to be detected, the image of the normal sample, and the defective sample image generated in Step 1 into ResNet34 for multi-scale feature extraction. Input the feature information extracted from the normal sample image into the adaptive hierarchical memory bank, and through adaptive core set sampling, retain the most representative normal sample features. The normal sample features form the "normal pattern".
[0009] Step 3: Calculate the Euclidean distance, that is, calculate the L2 difference, between the feature information extracted from the generated defective sample image and the "normal pattern" in the adaptive hierarchical memory bank to measure the similarity of the feature vectors, and update the "normal pattern". Calculate the L2 difference between the feature information extracted from the image of the sample to be detected and the "normal pattern" in the final adaptive hierarchical memory bank to obtain the defective difference information between the input sample to be detected and the most similar memory sample.
[0010] Step 4: Perform adaptive multi-scale feature fusion on the feature information extracted from the sample to be detected and the defective difference information calculated in Step 3, construct a multi-scale defective area localization map for defective localization optimization, and obtain accurate defective detection and localization results.
[0011] Furthermore, the enhanced image-level defective simulation method in Step 1 is specifically as follows:
[0012] Step 1.1: Randomly generate three binary masks M P , M G , and M V using Perlin noise, Gaussian noise, and cellular noise. Further expand and construct a mask atlas through the union of M P , M G , and M V . Perform binarization processing to extract the foreground mask M I of the normal sample image I, and multiply it pixel by pixel with the masks selected from the mask atlas to obtain the foreground mask M of the defective position:
[0013]
[0014] where the random number P r ∈ [0, 1], and α is set to 0.25.
[0015] Step 1.2: Generate a noise image NI j .
[0016] Step 1.3: Fuse the normal sample I and the noise image NI j to generate a preliminary defective sample P i , and the selection probabilities of the categories of NI j are the same.
[0017] Step 1.4: Invert the foreground mask M to generate the corresponding background mask Utilize the preliminarily generated defective image P i , normal sample image I, foreground mask M and background mask to fuse and generate the final defective sample image I A :
[0018] Furthermore, the noise image NI in the said Step 1.2 j is generated from the following three parts:
[0019] Step 1.2.1: Structural noise image: Randomly adjust the brightness, saturation, hue and contrast of the input normal sample image I, and apply local Gaussian blur or sharpening operations. Divide the image into irregular regions, and apply rotation, flipping, scaling and translation to these regions to create local structural perturbations and misalignments;
[0020] Step 1.2.2: Texture noise image: Apply random transformations to the texture samples to generate texture noise images with rich diversity;
[0021] Step 1.2.3: Morphological noise image: Generate by applying morphological operations to the input normal sample image I, including dilation, erosion, opening and closing operations, to create structural deformations of different scales.
[0022] Furthermore, the specific content of the said Step 2 is as follows:
[0023] Divide the input image into N p regional grids, use blocks 1, 2 and 3 of ResNet34 for feature extraction, and independently store the features of the normal sample images of different layers. At the same time, each layer's memory bank divides the space into N p regional grids to store the features of the mapped positions on the image; Select K feature vectors from each layer for retention, cluster the features of each layer in the adaptive hierarchical memory bank, and the features are divided into C clusters {C1, C2,..., C C}, and cluster centers {u1, u2,..., u C}, where C is the number of clusters for clustering determined by the dynamic optimization method; Within each cluster, evaluate the dispersion of the samples within the cluster, that is, calculate the Euclidean distance between each sample in each cluster and its cluster center sample to obtain the dispersion score d it of each cluster in each layer. According to the magnitude of the dispersion, dynamically adjust the number of samples retained in each cluster, and then independently perform greedy core set sampling within each cluster to retain the most representative normal sample features, and the said normal sample features form the "normal mode".
[0024] Furthermore, the number of clusters of the C clustering is determined by the weighted average of the elbow method and the silhouette coefficient method.
[0025] Furthermore, the specific steps of step 3 are as follows:
[0026] Calculate the L2 difference between the input image feature information FI extracted and the samples in the adaptive hierarchical memory bank B to obtain the difference between the input sample and its most similar memory sample, that is, the defect difference information. The calculation formula is as follows:
[0027]
[0028] Among them, the input image feature information FI is the feature information of the defect sample image after being extracted by ResNet34 in the training stage and the feature information extracted from the sample image to be detected in the inference stage; B j represents the hierarchical memory bank, j ∈ {1, 2, 3} represents the corresponding number of layers of the adaptive hierarchical memory bank, n represents the number of feature information samples in the adaptive hierarchical memory bank, K represents the number of input image feature information samples, and the greater the difference at a certain position, the higher the probability that the corresponding input image area is abnormal.
[0029] Furthermore, in step 4, the adaptive multi-scale feature fusion includes an adaptive feature fusion module and a multi-scale feature information fusion module. The feature information FI extracted from the sample image to be detected j and the calculated defect difference information DDI j are input into the adaptive feature fusion module, j ∈ {1, 2, 3}. A comprehensive feature map is generated through weighted fusion, and residual connections are used to ensure the transmission of information flow and the effective backpropagation of gradients. Further input into the multi-scale feature information fusion module, the resolution of feature maps with different dimensions is aligned by upsampling, then the number of channels is aligned by convolution, and finally an element-wise addition operation is performed to achieve multi-scale feature information fusion, obtaining the finally fused features FF1, FF2, and FF3.
[0030] Furthermore, when optimizing the defect location in step 4, considering that the defect difference information includes features of multiple scales, first calculate the mean value of the defect difference information in the channel dimension, and introduce learnable weight parameters w i on the feature maps of different scales to generate the corresponding defect spatial attention map DM. Subsequently, through upsampling and element-wise calculation, the features of different scales are fully fused, and finally a defect region location optimization map with multi-level features is obtained. The calculation formula is as follows:
[0031]
[0032]
[0033] Among them, H1, H2, and H3 are the number of channels of defect difference information DDI1, DDI2, and DDI3 respectively, and DDI 1i , DDI 2i , DDI 3i are the feature maps with channel i in DDI1, DDI2, and DDI3 respectively. w3, w2, and w1 are the learnable weights of each layer of feature maps. S represents the Sigmoid activation function, and U represents the upsampling operation.
[0034] Furthermore, multiply the defect spatial attention maps DM1, DM2, and DM3 by the corresponding fused features FF1, FF2, and FF3 to obtain the optimized features, and then optimize them using the loss function to generate the final defect detection result; the loss function is jointly optimized by the Focal Loss and L1 Loss functions, and the weights α1 and α2 are introduced to control the relative importance of the two loss functions respectively. The calculation formula is as follows:
[0035] L total =α1L smooth_L1 +α2L focal
[0036] where α1 and α2 are hyperparameters used to adjust the weights of the L1 Loss and Focal loss. Beneficial effects:
[0037] 1. Through the enhanced image-level defect simulation strategy, the present invention helps the model effectively distinguish between normal and abnormal modes; designs an adaptive hierarchical memory bank architecture, hierarchically embeds multi-scale features into the grid memory bank, effectively eliminates redundant calculations, and further improves the inference speed through the adaptive core set sampling strategy.
[0038] 2. The present invention combines multi-scale feature fusion and defect localization optimization technology, makes full use of the differences between multi-scale features and normal modes, and significantly improves the accuracy of defect detection and localization. Experimental results on the MVTecAD industrial dataset show that the I-AUROC and AUCPRO of the memory feature matching industrial defect detection method with adaptive feature fusion reach 99.3% and 97.3% respectively, showing significant advantages compared with the comparison methods. In addition, this method also shows excellent performance on the BTAD unsupervised dataset, further verifying its effectiveness and generality. In terms of real-time performance, the inference speed of the algorithm reaches 89FPS, which is more than 8 times faster than the Patchcore model, meeting the real-time detection requirements in industrial applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a schematic diagram of the overall architecture of the memory matching industrial defect detection method with adaptive feature fusion of the present invention.
[0040] Figure 2 Schematic diagram of the enhanced image-level defect simulation strategy steps of the present invention.
[0041] Figure 3 Schematic diagram of the adaptive multi-scale feature fusion strategy of the present invention.
[0042] Figure 4 Schematic diagram of the adaptive feature fusion module of the present invention.
[0043] Figure 5 Schematic diagram of the parallel attention branch module of the present invention.
[0044] Figure 6 Schematic diagram of the visual comparison of the memory matching industrial defect detection method with adaptive feature fusion of the present invention and other methods in the detection and localization results of MVTecAD object category products.
[0045] Figure 7 Schematic diagram of the visual comparison of the memory matching industrial defect detection method with adaptive feature fusion of the present invention and other methods in the detection and localization results of MVTecAD texture category products.
[0046] Figure 8 Schematic diagram of the comparison of the inference speed of the memory matching industrial defect detection method with adaptive feature fusion of the present invention and other methods. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] Embodiment 1
[0049] Referring to Figures 1-5 , an embodiment of the present invention provides a memory matching industrial defect detection method with adaptive feature fusion, including:
[0050] Step S100, designing an enhanced image-level defect simulation strategy to generate a variety of unknown defect samples. In this process, it is necessary to ensure that the defect samples can retain the structural features of the original normal image while having differences from the normal samples. These simulated defect samples will be used to assist the method in learning and distinguishing the normal mode from the abnormal mode, thereby enhancing the defect detection ability.
[0051] Step S200: Input the sample image to be detected, the normal sample image, and the generated defective sample image into the ResNet34 neural network for multi-scale image feature extraction. The extracted normal sample feature information is input into the adaptive hierarchical memory bank architecture, and the adaptive core set sampling method is used to screen and retain the most representative normal sample feature information to form a stable "normal mode". The establishment of this memory bank helps improve the accuracy of defect detection and significantly optimize the inference speed.
[0052] Step S300: Calculate the L2 difference between the generated defective sample feature information and the "normal mode" of the memory bank, and continuously update the memory bank to improve the accuracy of defect recognition. Input the feature information of the sample to be detected and calculate its L2 difference from the "normal mode" of the final memory bank. In this way, the defect difference information between the input sample and the most similar memory sample can be obtained. The recognition of the defective area is completed by comparing the input image with the normal mode in the memory bank, and the defect difference information plays a crucial role in this process.
[0053] Step S400: Based on the input feature information of the sample to be detected and the defect difference information, perform adaptive multi-scale feature fusion to improve the accuracy of defect detection and localization. Use Focal Loss and L1 Loss as loss functions to optimize the detection results. Finally, obtain the optimized defect detection and localization results, improving the detection reliability for various types of defects.
[0054] As Figure 1 shown, an adaptive feature fusion memory matching industrial defect detection method of the present invention consists of an enhanced image-level defect simulation strategy, an adaptive hierarchical memory bank architecture, and adaptive multi-scale feature fusion and localization optimization. Defective samples are generated from normal samples through the enhanced image-level defect simulation strategy. Based on the U-Net architecture, ResNet34 is sampled as the encoder. In the training stage, only normal samples need to be input. The adaptive feature fusion memory matching industrial defect detection method comprehensively covers the normal mode with the least amount of normal sample features through adaptive core set sampling, and these retained features are regarded as the "memory experience" of the model. In the inference stage, the features of the input sample are compared with the samples in the memory bank to generate defect difference information. Subsequently, an adaptive multi-scale feature fusion module and a localization optimization strategy are adopted to improve the detection accuracy and localization accuracy of the defective area.
[0055] The enhanced image-level defect simulation strategy is as Figure 2 shown to solve problems such as insufficient authenticity of defect generation results, limited types of defect morphologies, or poor alignment accuracy between generated anomalies and masks. The specific method is as follows:
[0056] Defect mask generation: Use Perlin noise, Gaussian noise, and cellular noise to randomly generate three binary masks, denoted as M P , M G , and M V as the basis for generating defect masks. To further increase the diversity of defect regions, through the union of M P , M G , and M V to further expand and construct a mask atlas. To avoid the difference between the data distribution of the generated defect samples and the real defect samples caused by the superposition of noise in the background region, perform binarization to extract the foreground mask M I of the normal sample I. Subsequently, obtain the defect position mask image by multiplying pixel by pixel with the mask selected from the mask atlas.
[0057]
[0058] Among them, the random number P r ∈[0,1], and α is set to 0.25.
[0059] Initial defect generation: To generate the defect sample P i , first input the normal sample I into the initial defect generation module, and generate the initial defect sample by fusing I and the noise image NI j , and the selection probabilities of the categories of NI j are the same. The generation of the noise image NI j comes from the following three parts:
[0060] a. The structural noise image increases randomness by randomly adjusting the brightness, saturation, hue, and contrast of the input image and applying local Gaussian blur or sharpening operations. Then, divide the image into irregular regions and apply geometric transformations such as rotation, flipping, scaling, and translation to these regions to create local structural perturbations and misalignments.
[0061] b. The texture noise image generates a texture noise image with rich diversity by applying random transformations (such as rotation, scaling, color adjustment, etc.) to texture samples (from the DTD texture dataset and the Kylberg texture dataset).
[0062] c. The morphological noise image is generated by applying morphological operations to the input image, including dilation, erosion, opening, and closing. These operations can create structural deformations of different scales to simulate morphological defects such as damaged or contaminated material surfaces.
[0063] Final defect generation: Invert the foreground mask M to generate the corresponding background mask Then, use the initially generated defect image P i, normal sample image I, foreground mask M, and background mask are fused to generate the final defective sample image I A , and its calculation formula is as follows:
[0064]
[0065] In step S200, the multi-scale image information feature extraction, adaptive hierarchical memory bank architecture, and adaptive core set sampling are as Figure 1 shown, and the specific method to improve the defect detection accuracy and inference speed is:
[0066] The input image is divided into N p region grids, and the pre-trained encoder ResNet34 blocks 1, 2, and 3 are used for feature extraction, and their parameters are frozen while the other parts can be trained. The adaptive hierarchical memory bank architecture stores the normal sample image features of different layers independently. At the same time, each layer of the memory bank divides the space into N p region grids to store the features at the mapped positions on the image. K most representative feature vectors are selected and retained from each layer. The features in each layer of the hierarchical memory bank are clustered, and the features are divided into C clusters {C1, C2, …, C C}, and the cluster centers {u1, u2,..., u C}, where the number of clusters C for clustering is determined by dynamic optimization methods (elbow method and silhouette coefficient method). Within each cluster, the dispersion of the samples within the cluster is evaluated, that is, the Euclidean distance between each sample in the i-th cluster and the sample at its cluster center is calculated, so as to obtain the dispersion score d it of each cluster in each layer. Its calculation formula is as follows:
[0067]
[0068] where, represents all samples in the i-th cluster. According to the magnitude of the dispersion, the number of samples retained in each cluster is dynamically adjusted. Clusters with larger dispersion will retain more samples, while dense clusters will retain fewer samples. Then, greedy core set sampling is independently performed within each cluster to avoid imbalance in the retained samples.
[0069] In step S300, the defect difference information calculation is as Figure 1 shown, and the specific method is:
[0070] Defective regions are typically identified by comparing the input image with the normal patterns stored in memory. During the training or inference phase, the feature information FI of the input image is extracted, and the L2 difference between it and the samples in the hierarchical memory bank B is calculated. With the help of the difference, the difference between the input sample and its most similar memory sample can be obtained, that is, the defective difference information. The greater the difference at a certain position, the higher the probability that the corresponding region of the input image is abnormal. The calculation formula is as follows:
[0071]
[0072] Among them, B j represents the hierarchical memory bank, j ∈ {1, 2, 3} represents the corresponding memory bank layer number, n represents the number of feature information samples in the hierarchical memory bank, and K represents the number of input image feature information samples.
[0073] In step S400, adaptive multi-scale feature fusion and defect localization optimization, as well as loss functions, are used to improve the detection and localization accuracy of various defects, such as Figure 3 , Figure 4 and Figure 5 shown. The specific method is as follows:
[0074] Adaptive multi-scale feature fusion includes an adaptive feature fusion module and a multi-scale feature information fusion module. The image information feature FI j and the defective difference information feature DDI j are input into the adaptive feature fusion module, and a comprehensive feature map is generated through weighted fusion. The residual connection is used to ensure the transmission of information flow and the effective backpropagation of gradients. Since different scale information helps to enhance the robustness of the model, multi-scale feature information fusion is further performed, that is, the resolution of feature maps with different dimensions is aligned using upsampling, and then the number of channels is aligned using convolution. Finally, an element-wise addition operation is performed to achieve multi-scale feature information fusion. The fused features are output through the Conv_Norm layer, denoted as FF1, FF2, and FF3 respectively.
[0075] Since small defects often show locality and weak signal features and are easily interfered by textures or noises in complex backgrounds, while large defects may span multiple feature levels, resulting in the problem of inconsistent global and local information. To improve the feature expression ability, a parallel attention mechanism is adopted in the adaptive feature fusion process. Among them, channel attention (CA) is used to allocate weights for different channels, while pixel attention (PA) is used to adjust the spatial distribution. Assuming the input feature map is x, after normalization to reduce the impact of data distribution changes, and then the feature weight allocation is optimized under the action of the attention mechanism, so as to more effectively highlight the defective information.
[0076] The channel attention mechanism can effectively extract global information and adjust the channel dimension of features. The channel attention consists of a CA branch to extract the overall channel features. The calculation formula is as follows:
[0077]
[0078] The pixel attention mechanism enhances the saliency of the defect area and optimizes the expression ability of spatial features by assigning weights to each pixel in the feature map. This mechanism contains a PA branch, which can be used to extract global pixel gating features. The calculation formula is as follows:
[0079]
[0080] During the feature fusion process, in order to give full play to the complementarity of different features y FI and y DDI , first perform global average pooling on the two, and use one-dimensional convolution to extract the global information vectors G FI and G DDI . Stack them and apply the Softmax function to obtain the adjustment weight vectors ω FI and ω DDI . The calculation formula is as follows:
[0081]
[0082] ω DDI = 1 - ω FI
[0083] The feature weight vector is multiplied element-wise with the corresponding feature. Then, the weighted visual information and semantic information features are added together to obtain the fused feature F, and F is sent to the multi-scale feature information fusion module. The calculation formula is as follows:
[0084] F = ω FI × y FI + ω DDI × y DDI
[0085] When optimizing defect localization, considering that DDI contains features of multiple scales, first calculate the mean in the channel dimension. Since the sizes of the defect areas are different, and in order to effectively fuse the feature information of different scales, introduce the learnable weight parameter w i on the feature maps of different scales to generate the corresponding defect spatial attention map (DM). Subsequently, through upsampling and element-wise calculation, the features of different scales are fully fused, and finally, the defect area localization optimization map with multi-level features is obtained. The calculation formula is as follows:
[0086]
[0087] Among them, H1, H2, and H3 are the number of channels of defect difference information DDI1, DDI2, and DDI3 respectively, and DDI 1i , DDI 2i , DDI 3i are the feature maps with channel i in DDI1, DDI2, and DDI3 respectively. w3, w2, and w1 are the learnable weights of each layer of feature maps. S represents the Sigmoid activation function, and U represents the upsampling operation.
[0088] Multiply the defect spatial attention maps DM1, DM2, and DM3 by the corresponding fused features FF1, FF2, and FF3 to obtain the optimized features. Then, use the loss function for optimization to generate the final defect detection result.
[0089] The L1 Loss function helps to retain edge information when predicting the segmentation result, while the Focal loss function can effectively alleviate the area imbalance problem between the normal region and the abnormal region, thereby improving the model's attention to difficult-to-classify regions. For this reason, weights α1 and α2 are introduced to control the relative importance of the two loss functions respectively. The calculation formula is as follows:
[0090] L total = α1L smooth_L1 + α2L focal
[0091] Among them, α1 and α2 are hyperparameters used to adjust the weights of the L1 Loss and Focal loss.
[0092] Referring to Figure 6 , Figure 7 and Figure 8 , the experimental verification of the memory matching industrial defect detection method with adaptive feature fusion of the present invention is provided. To verify the technical effects adopted in this method, in this embodiment, the traditional technical solution is compared with the method of the present invention through a comparative test, and the experimental results are compared by means of scientific demonstration to verify the real effects of this method.
[0093] Datasets: The MVTec AD dataset is used as the main evaluation benchmark. This dataset is widely applied to industrial defect detection tasks and covers 15 industrial product categories, including 10 object categories and 5 texture categories. The dataset contains a total of 3,629 normal (defect-free) training images and 1,725 test images. The test set includes 73 different types of defects, including typical industrial defects such as position deviation, missing parts, scratches, dents, oil stains, etc., which basically cover all possible defect patterns of these category products. The BeanTech dataset (BTAD) is sourced from real industrial scenarios and contains 3 different industrial products. This dataset contains a total of 1,800 normal (defect-free) training images and 1,030 test images, showing product surface and structural defects. The BTAD dataset is often used to supplement the MVTec AD dataset to more comprehensively evaluate the performance of defect detection algorithms.
[0094] Evaluation Metrics: In industrial defect detection tasks, to effectively evaluate the discrimination ability of different models at the image and pixel levels, the area under the receiver operating characteristic curve at the image level (I-AUROC) and the pixel level (P-AUROC) are usually used for evaluation results. However, since the defective area usually only accounts for a very small part of the image, P-AUROC may not accurately reflect the localization accuracy of the model. Recent research has pointed out that P-AUROC may overestimate the model performance and is difficult to truly measure the accuracy of pixel-level defect detection. Therefore, the area under the curve per region overlap at the pixel level (AUCPRO) is adopted as a new evaluation metric. It assigns equal weights when evaluating the performance of all predicted abnormal regions, which can effectively avoid the risk of overestimating the performance by AUROC due to the increase in false positives, and thus more accurately reflects the model's abnormal localization ability at the pixel level. Its calculation formula is as follows:
[0095]
[0096] where N d represents the total number of defective pixels in the image, TP d represents the number of pixels correctly detected as defective among all defective pixels, and FP d represents the number of defective pixels misdetected as normal.
[0097] Therefore, this invention also adopts this method, using I-AUROC to evaluate the model's ability to distinguish normal and defective features at the image level, and using AUCPRO to measure the model's defect localization performance at the pixel level. In addition, to evaluate the real-time inference ability of the model, the frame rate (Frames Per Second, FPS) is adopted as a measure of computational efficiency. The higher the FPS value, the faster the inference speed of the model.
[0098] Experimental Environment and Settings: All experiments were completed on a computing platform equipped with an AMD Ryzen 9 7940H processor (main frequency 4.0 GHz) and an NVIDIA GeForce RTX 4060 Laptop GPU (8GB video memory). The experiments were implemented based on the PyTorch 2.0.1 framework in a Python 3.8 environment and utilized CUDA 11.8 for accelerated computing.
[0099] All images in the dataset were resized to 256×256 as input. During the construction of the memory bank, the number of grid cells N per layer p was set to 4. The optimal number of clusters was determined by the weighted average of the Elbow Method and the Silhouette Coefficient method. During the greedy sampling process, the number of random initializations was set to 10, and the number of retained samples K was set to 15% of the total dataset size, with the specific value adjusted according to the dataset size. In the training stage, the Adam optimizer was used for 4000 iterations, the initial learning rate was set to 0.0008, and the weight decay coefficient was 0.0003. The batch size was set to 8, with normal samples and simulated abnormal samples configured in a 1:1 ratio. The weights of the L1 Loss and Focal loss were set to 0.6 and 0.4 respectively. Finally, the average value of the top 100 abnormal pixels was used as the image-level abnormal score.
[0100] Table 1 Comparative Experiments on the MVTec-AD Dataset
[0101]
[0102]
[0103] Analysis of Experimental Results
[0104] Quantitative analysis: As shown in Table 1 of the experimental results, the memory matching industrial defect detection method with adaptive feature fusion of the present invention shows excellent performance in both I-AUROC and AUCPRO evaluation metrics on the MVTec AD dataset. The units of I-AUROC and AUCPRO shown in the table are %. It can be seen from the data in the table that the memory matching industrial defect detection method with adaptive feature fusion achieves high detection accuracy in multiple categories. Especially in terms of the AUCPRO metric, compared with other methods (such as Patchcore, RD4AD, and CFA), it has achieved better performance. This indicates that the proposed method has strong defect area localization ability in fine-grained defect detection tasks. In addition, in terms of the I-AUROC metric, the present invention also demonstrates stable recognition ability. Further analyzing the average performance, the memory matching industrial defect detection method with adaptive feature fusion reaches 99.3% and 97.3% respectively in terms of I-AUROC and AUCPRO metrics. Compared with the optimal method (Patchcore), the memory matching industrial defect detection method with adaptive feature fusion still shows a certain degree of improvement in these two metrics. This result proves the balance and adaptability of the proposed method in global and local defect area recognition and localization. It should be noted that in some specific categories (such as toothbrush, tile, etc.), the AUCPRO metric of the memory matching industrial defect detection method with adaptive feature fusion is significantly improved compared with other methods, further verifying the effectiveness of the adaptive multi-scale feature fusion method proposed in the present invention in complex background and small defect detection tasks.
[0105] Qualitative analysis: It can be seen from the data in Table 1 that Patchcore (an unsupervised anomaly detection algorithm) performs relatively well in some categories. To more intuitively compare the performance of the memory matching industrial defect detection method with adaptive feature fusion of the present invention and Patchcore on the MVTec AD dataset, Figure 6 and Figure 7Visualize the detection and localization results of the two algorithms, where Input, GT, result1, and result2 represent the input image, the true defect label, the heat map of the defect localization result, and the defect segmentation and localization result map respectively. From the visualization results, it can be observed that the memory matching industrial defect detection method with adaptive feature fusion demonstrates high accuracy in the detection and localization tasks of various types of defects. Especially in dealing with complex defect types such as fine cracks, structural damage, material loss, and color anomalies, it can accurately identify the defect areas. In addition, whether it is texture categories (such as carpet, grid, tile, wood) or object categories (such as bottle, transistor, cable), the memory matching industrial defect detection method with adaptive feature fusion can stably provide high-quality defect detection and localization results. Further comparison with the detection results of Patchcore reveals that although it can accurately identify the defect areas in some categories, there are still certain cases of missed detection or misdetection when dealing with scenarios with complex backgrounds or relatively small defect features (such as metalnut, pill, zipper). In contrast, the memory matching industrial defect detection method with adaptive feature fusion can still maintain stable detection performance in these situations and accurately locate the defect areas, further demonstrating the effectiveness and adaptability of the method of the present invention in industrial defect detection tasks.
[0106] Table 2 Comparative experiments on the BTAD dataset
[0107]
[0108] Analysis of the experimental results of the generalization performance evaluation
[0109] To further verify the defect detection and localization capabilities of the memory matching industrial defect detection method with adaptive feature fusion of the present invention, a benchmark test was conducted on another widely used dataset (i.e., the BTAD dataset). The experimental results are shown in Table 2, where the units of the I-AUROC and AUCPRO metrics are %. The present invention achieved relatively excellent I-AUROC and AUCPRO scores. Especially in terms of the AUCPRO metric, AFFMM-IDD has a significant improvement compared to the Patchcore and RD4AD models, with increases of 12.5% and 13% respectively. This further demonstrates its effectiveness and generality.
[0110] Analysis of the experimental results of the inference time evaluation
[0111] Refer to Figure 8, in addition to defect detection and localization performance, inference speed is also a crucial consideration in the industrial model deployment process. To comprehensively evaluate the real-time detection ability of the memory matching industrial defect detection method with adaptive feature fusion of the present invention, the inference times of different methods on the MVTec AD dataset were compared, and the frame rate (FPS) was used as the performance measurement index. The experimental results show that while ensuring the best detection and localization performance, the present invention also achieves the fastest inference speed, reaching 89 FPS. This result indicates that the present invention has efficient real-time inference ability in actual industrial detection scenarios and can meet the requirements of high-speed production lines.
[0112] Embodiments of the present invention can be implemented by computer hardware, a combination of hardware and software, or computer instructions stored in a non-transitory computer-readable medium. The method can adopt standard programming techniques and use a configured non-transitory computer-readable storage medium in a computer program to enable the computer to perform specified operations in a predetermined manner, as specifically shown in the embodiments and the drawings. The program can use a high-level language, an object-oriented language, or assembly or machine language as required. In some cases, the program can also run on an application-specific integrated circuit (ASIC).
[0113] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable those familiar with this technology to understand the content of the present invention and implement it accordingly, and should not be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. An industrial defect detection method based on memory matching with adaptive feature fusion, characterized in that It includes the following steps: Step 1: Use normal sample images to generate a variety of unknown defect samples through an enhanced image-level defect simulation method to obtain defect sample images, and the defect sample images retain the structural features of the normal sample images; Step 2: Input the sample image to be detected, the normal sample image, and the defect sample images generated in Step 1 into ResNet34 for multi-scale feature extraction. Input the feature information extracted from the normal sample image into the adaptive hierarchical memory bank, and through adaptive core set sampling, retain the most representative normal sample features, and the normal sample features form the "normal mode"; Step 3: Calculate the Euclidean distance, that is, calculate the L2 difference, between the feature information extracted from the generated defect sample images and the "normal mode" in the adaptive hierarchical memory bank to measure the similarity of the feature vectors, and update the "normal mode"; Calculate the L2 difference between the feature information extracted from the sample image to be detected and the "normal mode" in the final adaptive hierarchical memory bank to obtain the defect difference information between the input sample to be detected and the most similar memory sample; Step 4: Perform adaptive multi-scale feature fusion on the feature information extracted from the sample to be detected and the defect difference information calculated in Step 3, construct a multi-scale defect region localization map for defect localization optimization, and obtain accurate defect detection and localization results.
2. The memory matching industrial defect detection method with adaptive feature fusion according to claim 1, characterized in that The enhanced image-level defect simulation method in Step 1 is specifically as follows: Step 1.1: Randomly generate three binary masks M P , M G and M V using Perlin noise, Gaussian noise, and cellular noise. Further expand and construct a mask atlas through the union of M P , M G and M V . Perform binarization processing to extract the foreground mask M I of the normal sample image I, and multiply it pixel by pixel with the masks selected from the mask atlas to obtain the foreground mask M of the defect location: Among them, the random number P r ∈ [0, 1], and α is set to 0.25; Step 1.2: Generate a noise image NI j ; Step 1.3: Fuse the normal sample I and the noise image NI j Generate the preliminary defect sample P i , and NI j has the same category selection probability; Step 1.4: Invert the foreground mask M to generate the corresponding background mask Utilize the preliminarily generated defective image P i , the normal sample image I, the foreground mask M, and the background mask to fuse and generate the final defective sample image I A :
3. The memory matching industrial defect detection method with adaptive feature fusion according to claim 2, wherein, The noise image NI in step 1.2 j is generated from the following three parts: Step 1.2.1: Structural noise image: Randomly adjust the brightness, saturation, hue, and contrast of the input normal sample image I, and apply local Gaussian blur or sharpening operations. Divide the image into irregular regions, and apply rotation, flipping, scaling, and translation to these regions to create local structural perturbations and misalignments; Step 1.2.2: Texture noise image: Apply random transformations to the texture samples to generate texture noise images with rich diversity; Step 1.2.3: Morphological noise image: Generate by applying morphological operations to the input normal sample image I, including dilation, erosion, opening, and closing, to create structural deformations of different scales.
4. An industrial defect detection method based on memory matching with adaptive feature fusion according to claim 1, characterized in that, Step 2 is specifically as follows: Divide the input image into N p regional grids, use blocks 1, 2, and 3 of ResNet34 for feature extraction, independently store the normal sample image features of different layers, and at the same time divide the space of each layer's memory bank into N p regional grids to store the features of the mapped positions on the image; select K feature vectors from each layer for retention, cluster the features of each layer in the adaptive hierarchical memory bank, and the features are divided into C clusters {C1, C2,..., C C}, and cluster centers {u1, u2,..., u C}, where C is the number of clusters for clustering determined by the dynamic optimization method; within each cluster, evaluate the dispersion of the samples within the cluster, that is, calculate the Euclidean distance between the samples in each cluster and the sample of its cluster center to obtain the dispersion score d it of each cluster in each layer. According to the magnitude of the dispersion, dynamically adjust the number of samples retained in each cluster, and then independently perform greedy core set sampling within each cluster to retain the most representative normal sample features, and the normal sample features form a "normal pattern".
5. The memory matching industrial defect detection method with adaptive feature fusion according to claim 4, characterized in that, The number of clusters of the C clustering is determined by the weighted average of the elbow method and the silhouette coefficient method.
6. The memory matching industrial defect detection method with adaptive feature fusion according to claim 1, wherein Step 3 is specifically as follows: Calculate the L2 difference between the extracted input image feature information FI and the samples in the adaptive hierarchical memory bank B to obtain the difference between the input sample and its most similar memory sample, that is, the defect difference information. The calculation formula is as follows: Among them, the input image feature information FI is the feature information of the defective sample image after being extracted by ResNet34 in the training stage, and is the feature information extracted from the sample image to be detected in the inference stage; B j represents the hierarchical memory bank, j ∈ {1, 2, 3} represents the corresponding number of layers of the adaptive hierarchical memory bank, n represents the number of feature information samples of the adaptive hierarchical memory bank, K represents the number of input image feature information samples, and the greater the difference at a certain position, the higher the probability that the corresponding input image area at that position is abnormal.
7. An industrial defect detection method based on memory matching with adaptive feature fusion according to claim 1, characterized in that, In step 4, the adaptive multi-scale feature fusion includes an adaptive feature fusion module and a multi-scale feature information fusion module, and the feature information FI extracted from the sample image to be detected j and the calculated defect difference information DDI j are input into the adaptive feature fusion module, where j ∈ {1, 2, 3}. Through weighted fusion, a comprehensive feature map is generated, and residual connections are used to ensure the transmission of information flow and the effective backpropagation of gradients. Further input into the multi-scale feature information fusion module, the resolution of feature maps of different dimensions is aligned by upsampling, then the number of channels is aligned by convolution, and finally, an element-wise addition operation is performed to achieve multi-scale feature information fusion, obtaining the finally fused features FF1, FF2, and FF3.
8. An industrial defect detection method based on memory matching with adaptive feature fusion according to claim 1, characterized in that When optimizing defect location in step 4, considering that the defect difference information includes features of multiple scales, first calculate the mean value of the defect difference information in the channel dimension, and introduce a learnable weight parameter w on the feature maps of different scales i , generate the corresponding defect spatial attention map DM, and then through upsampling and element-by-element calculation, fully fuse the features of different scales, and finally obtain an optimized defect region location map with multi-level features. Its calculation formula is as follows: where H1, H2, and H3 are the number of channels of defect difference information DDI1, DDI2, and DDI3 respectively, DDI 1i , DDI 2i , DDI 3i are the feature maps with channel i in DDI1, DDI2, and DDI3 respectively, w3, w2, and w1 are the learnable weights of each layer of feature maps, S represents the Sigmoid activation function, and U represents the upsampling operation.
9. The memory matching industrial defect detection method with adaptive feature fusion according to claim 8, wherein Multiply the defect spatial attention maps DM1, DM2, and DM3 by the corresponding fused features FF1, FF2, and FF3 to obtain the optimized features, and then use the loss function for optimization to generate the final defect detection result; The loss function is jointly optimized by the Focal Loss and L1Loss loss functions, and the weights α1 and α2 are introduced to respectively control the relative importance of the two loss functions. The calculation formula is as follows: L total = α1L smooth_L1 + α2L focal Among them, α1 and α2 are hyperparameters used to adjust the weights of the L1Loss loss and the Focal loss loss.
Citation Information
Cited By
CLIP-based double-prompt optimized few-sample industrial anomaly detection method
CN120672752A
A CLIP-based dual-clue optimization method for few-sample industrial anomaly detection
CN120672752B
Defect hotspot detection method based on two-dimensional contour extraction
CN120931568A
Multi-scale aggregation chip surface defect detection method based on comparative learning
CN121190398A
Cable surface anomaly detection method and device based on unsupervised algorithm
CN121458639A