A method for detecting structural anomalies and logical anomalies in an image, a storage medium
Patent Information
- Application Number
- CN202311797132.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-12-25
AI Technical Summary
在实际的工业生产过程中,正常样本数量多,易于大规模收集,而异常样本较少,收集困难,并且在异常图像中,结构异常与逻辑异常经常会同时出现(例如,生产工厂中,包装袋内的螺丝钉数量不符合生产厂商的设定标准,并且部分数量的螺丝钉出现破损等情况,前者属于逻辑异常,后者属于结构异常),导致检测难度大幅提升
[0049]依据上述实施例的图像中结构异常与逻辑异常的检测方法,采用“教师-双学生”架构,通过在预训练的教师网络的知识蒸馏下对结构检测模块和逻辑检测模块进行训练,结构检测模块能够获得教师网络对训练样本的特征提取能力、逻辑检测模块能够重建出训练样本的低级特征图;在检测时,根据结构检测模块对待检测图像提取的特征图与教师网络提取的特征图之间的差异,以及逻辑检测模块重建出的特征图与教师网络提取的低级特征图的差异,通过对比即可以检测出待检测图像中存在结构异常和/或逻辑异常的区域。在训练时可以只用正常样本图像进行训练,检测时通过特征对比同样可以检出结构异常和/或逻辑异常,克服了异常样本缺乏的不足。本方法构建结构检测模块和逻辑检测模块分别负责结构异常以及逻辑异常的检测,使结构异常与逻辑异常的检测在网络模型中解耦,可以更好地解决结构异常与逻辑异常相结合的缺陷分割问题,实现对缺陷的精确检测,大大减小了对正常图像的误检,大幅提高了缺陷检测能力。
Smart Images

Figure CN118014932B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of defect detection technology, specifically to a method and storage medium for detecting structural and logical anomalies in images. Background Technology
[0002] Currently, most industrial product quality inspections utilize automated machine vision, which captures images of products and then performs defect detection on those images. Defects in industrial images can be broadly categorized into two types: structural anomalies and logical anomalies. Structural anomalies refer to surface quality defects in objects within the image, primarily manifested as scratches, dents, dents, contamination, and breakage. Logical anomalies refer to objects in the image violating logical constraints, such as constraints on the quantity or position of objects, including discrepancies in product quantity, incorrect product placement, or missing products. In actual industrial production, normal samples are plentiful and easy to collect on a large scale, while abnormal samples are fewer and more difficult to collect. Furthermore, structural and logical anomalies often coexist in abnormal images (for example, in a production factory, the number of screws in a packaging bag may not meet the manufacturer's standards, and some screws may be damaged; the former is a logical anomaly, and the latter a structural anomaly), significantly increasing the difficulty of detection. Therefore, the visual inspection market has a strong demand for the detection of combined structural and logical anomalies, necessitating improvements to address the shortcomings of existing technologies. Summary of the Invention
[0003] The main technical problem this invention addresses is how to simultaneously detect structural and logical anomalies in an image.
[0004] According to the first aspect, one embodiment provides a method for detecting structural and logical anomalies in an image, comprising:
[0005] Acquire the image to be detected;
[0006] The image to be detected is input into a pre-trained teacher network for feature extraction to obtain multi-layered first feature maps with different resolutions of the image to be detected; the image to be detected is input into a trained structure detection module for feature extraction to obtain multi-layered second feature maps with different resolutions of the image to be detected; the first feature map of the highest layer is input into a trained logic detection module for decoding and reconstruction to obtain a third feature map, the resolution of the third feature map being consistent with the resolution of the first feature map of the lowest layer;
[0007] The second feature map is compared with the first feature map, and the third feature map is compared with the first feature map of the lowest layer, to determine the regions in the image to be detected that have structural and / or logical abnormalities.
[0008] The structure detection module and the logic detection module are obtained through training on the knowledge distillation of the pre-trained teacher network.
[0009] In some embodiments, the structure detection module is the same as the teacher network structure.
[0010] In some embodiments, the first feature map has M layers, with the resolution decreasing and the number of channels increasing in each layer, where M is an integer greater than or equal to 2; the logical detection module includes an M-1 layer of first convolutional layer, as well as a spatial attention submodule and a channel attention submodule;
[0011] The step of inputting the first feature map of the highest layer into the trained logic detection module for decoding and reconstruction to obtain the third feature map includes:
[0012] The first feature map of the highest layer is input into the first convolutional layer of the first layer for convolution processing and then upsampled to obtain an upsampled feature map. The upsampled feature map is then input into the first convolutional layer of the next layer for convolution processing and then upsampled to obtain an upsampled feature map of the next layer. The steps of inputting the upsampled feature map into the first convolutional layer of the next layer for convolution processing and then upsampling to obtain an upsampled feature map of the next layer are repeated until the upsampled feature map of the M-1th layer is obtained. The resolution and number of channels of the upsampled feature map of the kth layer are the same as those of the first feature map of the M-kth layer, where k∈[1,M-1].
[0013] The upsampled feature map of layer M-1 is input into the spatial attention submodule to perform attention enhancement processing in space to obtain a spatial attention map;
[0014] The upsampled feature map of layer M-1 is input into the channel attention submodule to perform attention enhancement processing on the channel to obtain the channel attention map;
[0015] The spatial attention map and the channel attention map are added element-wise to obtain the third feature map.
[0016] In some embodiments, the kernel size of the first convolutional layer is 3×3.
[0017] In some embodiments, the spatial attention submodule includes a second convolutional layer and a batch normalization layer;
[0018] The step of inputting the upsampled feature map of layer M-1 into the spatial attention submodule to perform attention enhancement processing in space to obtain a spatial attention map includes:
[0019] The upsampled feature map of layer M-1 is input into the second convolutional layer for dimensionality reduction to reduce the number of channels of the upsampled feature map of layer M-1 to 1. Then, it is processed by the batch normalization layer to obtain the batch normalized feature map.
[0020] Perform a Sigmoid operation on the batch normalized feature map to obtain a spatial attention score map;
[0021] The spatial attention map is obtained by performing element-wise multiplication of the upsampled feature map of the M-1th layer and the spatial attention score map.
[0022] In some embodiments, the channel attention submodule includes a global average pooling layer and a third convolutional layer;
[0023] The step of inputting the upsampled feature map of layer M-1 into the channel attention submodule to perform attention enhancement processing on the channel to obtain the channel attention map includes:
[0024] The upsampled feature map of the M-1th layer is input into the global average pooling layer to reduce the resolution to 1×1, and then processed by the third convolutional layer to obtain the third convolutional feature map, wherein the number of input and output channels of the third convolutional layer is kept consistent.
[0025] Perform a Sigmoid operation on the third convolutional feature map to obtain a channel attention score map;
[0026] The channel attention map is obtained by performing element-wise multiplication of the upsampled feature map of the M-1th layer and the channel attention score map.
[0027] In some embodiments, the number of second feature maps is the same as the number of first feature maps, and the number of second feature maps is the same as the number of first feature maps. Figure 1 A one-to-one correspondence exists, where the corresponding second feature map and the first feature map have the same resolution; by comparing the difference between the second feature map and the corresponding first feature map, and the difference between the third feature map and the lowest layer first feature map, regions with structural and / or logical anomalies in the image to be detected are determined.
[0028] In some embodiments, the first feature map has M layers, and the resolution of the first feature map in each layer decreases, where M is an integer not less than 2;
[0029] The step of determining regions in the image to be detected that contain structural and / or logical anomalies by comparing the differences between the second feature map and the corresponding first feature map, and the differences between the third feature map and the lowest-level first feature map, includes:
[0030] The squared difference between each second feature map and the corresponding first feature map is calculated to obtain M second difference squared maps. The mean of each second difference squared map is calculated in the channel dimension, and then a sigmoid operation is performed to obtain M structural anomaly fractional sub-maps. The structural anomaly fractional sub-maps other than the structural anomaly fractional sub-map with the largest resolution are upsampled to the size of the structural anomaly fractional sub-map with the largest resolution. Then, element-wise multiplication is performed on the M structural anomaly fractional sub-maps to obtain the structural anomaly fractional map.
[0031] The squared difference between corresponding elements of the third feature map and the first feature map of the lowest layer is calculated to obtain the third difference square map; the mean of the third difference square map is calculated in the channel dimension, and then the Sigmoid operation is performed to obtain the logic anomaly score map.
[0032] The structural anomaly score map and the logical anomaly score map are upsampled to the size of the image to be detected. Then, a preset structural score threshold is used to perform threshold segmentation on the structural anomaly score map to obtain a structural anomaly segmentation map. The structural anomaly segmentation map is used to identify regions in the image to be detected that have structural anomalies. Similarly, a preset logical score threshold is used to perform threshold segmentation on the logical anomaly score map to obtain a logical anomaly segmentation map. The logical anomaly segmentation map is used to identify regions in the image to be detected that have logical anomalies.
[0033] Alternatively, an element-wise multiplication operation is performed on the structural anomaly score map and the logical anomaly score map to obtain a merged anomaly score map. The merged anomaly score map is then upsampled to the size of the image to be detected. Finally, a preset detection threshold is used to perform threshold segmentation on the merged anomaly score map to obtain a defect segmentation map. The defect segmentation map is used to identify regions in the image to be detected where structural anomalies and / or logical anomalies exist.
[0034] In some embodiments, the structure detection module and the logic detection module are trained in the following manner:
[0035] Normal sample images are input into the pre-trained teacher network for feature extraction to obtain M layers of first normal sample feature maps with different resolutions; normal sample images are input into the structure detection module to be trained for feature extraction to obtain M layers of second normal sample feature maps corresponding to the first normal sample feature maps, where M is an integer not less than 2; the first normal sample feature map of the highest layer is input into the logic detection module to be trained for decoding and reconstruction to obtain a third normal sample feature map, the resolution of the third normal sample feature map being consistent with the resolution of the first normal sample feature map of the lowest layer;
[0036] Based on the total loss function Lf The structure detection module and the logic detection module to be trained are trained, and the total loss function L is used. f From the structural loss function L s and logical loss function L l It is determined that the structural loss function L s The logistic loss function L represents the difference between the second normal sample feature map and the corresponding first normal sample feature map. l The difference between the third normal sample feature map and the first normal sample feature map of the lowest layer is represented, and the structure and parameters of the teacher network remain unchanged during training.
[0037] In some embodiments, the structural loss function L s The expression is:
[0038]
[0039] Where for any m∈[1,M], L m The expression is consistent, that is...
[0040]
[0041] Where L represents any L m Let g represent the one-dimensional flattening vector of the first normal sample feature map in the m-th layer, and h represent the one-dimensional flattening vector of the corresponding second normal sample feature map. i Let h represent the i-th element of vector g. i Let represent the i-th element of vector h, and n represent the number of elements in vector g or h. The one-dimensional flattened vector of a feature map refers to a one-dimensional vector formed by arranging the elements of the feature map in order.
[0042] In some embodiments, the logical loss function L l The expression is:
[0043]
[0044] Where q represents the one-dimensional flattening vector of the first normal sample feature map at the lowest layer, and p represents the one-dimensional flattening vector of the third normal sample feature map. j p represents the j-th element of vector q. j Let represent the j-th element of vector p, and o represent the number of elements in vector q or p.
[0045] In some embodiments, the total loss function L f The expression is:
[0046] L f =αL s +βLl ,
[0047] Where α and β are preset weighting coefficients.
[0048] According to a second aspect, one embodiment provides a computer-readable storage medium storing a program that can be executed by a processor to implement the method for detecting structural and logical anomalies in images according to any of the above embodiments.
[0049] The method for detecting structural and logical anomalies in images according to the above embodiments adopts a "teacher-dual-student" architecture. The structural detection module and the logical detection module are trained under the knowledge distillation of a pre-trained teacher network. The structural detection module acquires the feature extraction capability of the teacher network from training samples, and the logical detection module reconstructs low-level feature maps of the training samples. During detection, based on the differences between the feature maps extracted by the structural detection module and those extracted by the teacher network, and the differences between the reconstructed feature maps and the low-level feature maps extracted by the teacher network, regions with structural and / or logical anomalies in the image can be detected through comparison. Training can be performed using only normal sample images, and structural and / or logical anomalies can still be detected through feature comparison during detection, overcoming the deficiency of a lack of abnormal samples. This method constructs a structural detection module and a logical detection module to be responsible for the detection of structural and logical anomalies respectively, decoupling the detection of structural and logical anomalies in the network model. This better solves the defect segmentation problem combining structural and logical anomalies, achieving accurate defect detection, greatly reducing false detections of normal images, and significantly improving defect detection capability. Attached Figure Description
[0050] Figure 1 This is a schematic diagram of the processing flow of network model F provided in one embodiment of the present invention;
[0051] Figure 2 This is a flowchart illustrating a method for detecting structural and logical anomalies in an image, according to one embodiment.
[0052] Figure 3 This is a schematic diagram of the structure of a logic detection module according to one embodiment;
[0053] Figure 4 This is a flowchart of one embodiment in which the first feature map of the highest layer is input into a trained logic detection module for decoding and reconstruction to obtain a third feature map;
[0054] Figure 5 This is a schematic diagram illustrating the structure and processing flow of a spatial attention submodule according to one embodiment.
[0055] Figure 6This is a schematic diagram illustrating the structure and processing flow of a channel attention submodule according to one embodiment.
[0056] Figure 7 This refers to an image to be detected and a defect segmentation image output by the detection method of the present invention, as shown in one embodiment.
[0057] Figure 8 This is a training flowchart for the structure detection module S and the logic detection module L in one embodiment. Detailed Implementation
[0058] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings. Similar elements in different embodiments are referred to by associated similar element reference numerals. In the following embodiments, many details are described to facilitate a better understanding of this application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other elements, materials, or methods. In some cases, certain operations related to this application are not shown or described in the specification. This is to avoid obscuring the core parts of this application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; they can fully understand the related operations based on the description in the specification and general technical knowledge in the art.
[0059] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can be rearranged or adjusted in a manner obvious to those skilled in the art. Therefore, the various orders in the specification and drawings are only for the clear description of a particular embodiment and do not imply a necessary order, unless otherwise stated that a particular order must be followed.
[0060] The serial numbers assigned to components in this document, such as "first" and "second," are used solely to distinguish the described objects and have no sequential or technical meaning. The terms "connection" and "linkage" used in this application, unless otherwise specified, include both direct and indirect connections (linkages). For ease of description, unless otherwise specified, 1×1Conv represents a convolutional layer with a 1×1 kernel, and 3×3Conv represents a convolutional layer with a 3×3 kernel.
[0061] For structural anomalies, some methods can achieve good detection results. For logical anomalies, methods with high detection accuracy cannot meet real-time inference requirements, and real-time solutions have poor detection performance. For anomaly detection involving a small number of samples and a combination of structural and logical anomalies, current methods generally employ pixel-based single-point and multi-element pairing anomaly detection approaches. The process involves extracting structural and logical features from the training images separately, then processing the elements of both features to remove unqualified data, and finally storing them in a feature memory. During inference, the structural and logical features of the image to be tested are extracted again, and the distances are calculated with the corresponding features in the feature memory to obtain structural and logical anomaly score maps. These two anomaly score maps are then fused to obtain the anomaly score map corresponding to the image to be tested. Finally, based on a preset threshold, the final pixel-by-pixel defect segmentation map is obtained. This detection method has two major drawbacks: First, it is complex to use and increases costs: the structural and logical features extracted by the network need to be processed additionally and stored in the corresponding feature memory for inference calculations. This results in a large amount of computation, poor real-time performance, and the inability to use it in a simple "end-to-end" manner, which greatly increases the usage, storage, and computation costs for manufacturers. Second, it has significant limitations on production scenarios: the additional feature memory requires a certain amount of storage space, which makes it impossible to deploy in detection scenarios with strict storage limitations.
[0062] The main objective of this invention is to solve the problem of simultaneously detecting structural and logical anomalies in images when only normal sample data is used for training. Furthermore, in some embodiments, it addresses certain problems existing in the prior art. To this end, this invention proposes a novel method for detecting structural and logical anomalies in images. The designed model can be trained without any defect samples, using only normal product images as training samples, achieving accurate detection of structural and logical anomalies and significantly improving detection performance. Moreover, the detection method of this invention is based on unsupervised deep learning technology, enabling end-to-end model training and inference without complex processing, making it simple to operate. Decoupling the detection of structural and logical anomalies within the network model better solves the defect segmentation problem combining structural and logical anomalies. In some embodiments, the various modules of the network model are designed to be as lightweight as possible while ensuring detection accuracy, improving detection performance and enabling real-time inference, reducing storage space usage, and lowering industrial production costs, thereby achieving cost reduction and efficiency improvement for manufacturing enterprises.
[0063] This invention uses deep learning knowledge transfer technology to build a network structure. The design of the entire network model F is based on a "teacher-dual student" architecture, which mainly consists of three parts: a teacher network T with pre-trained weights, a structure detection module S, and a logic detection module L. The teacher network T is responsible for extracting image features, the structure detection module S is mainly responsible for detecting structural anomalies, and the logic detection module L is mainly responsible for detecting logical anomalies. The structure detection module S and the logic detection module L are collectively referred to as the student network, i.e., the "teacher-dual student" architecture.
[0064] The processing flow of the entire network model F is as follows: Figure 1 As shown, the input is image x, the training phase uses pre-prepared sample images, and the inference phase uses the image of the object to be detected. During the training phase, the structure detection module S and the logic detection module L are trained through knowledge distillation from the pre-trained teacher network T. The teacher network T extracts multi-layer feature maps from the sample image x, for example, three-layer feature maps t1, t2, and t3 (for ease of explanation, this paper and accompanying figures use three layers as an example; those skilled in the art can obtain feature maps of other layers according to actual needs, and this invention is not limited thereto). The structure detection module S also accepts the sample image x as input and outputs multi-layer feature maps s1, s2, and s3. Then, it constructs a loss function L based on s1, s2, and s3 and t1, t2, and t3. s For example, loss functions Loss1, Loss2, and Loss3 are constructed for s1 and t1, s2 and t2, and s3 and t3 respectively, and then combined into loss function L. s The structural detection module S operates under the loss function L. s Distillation training was conducted under control, L s The goal is to guide the structure detection module S to learn the feature extraction capabilities of the teacher network T for sample images. Simultaneously, the logic detection module L accepts the highest-level feature map t3 of the teacher network T as input and outputs feature map l1. l1 and t1 are used to construct the loss function L. l The logic detection module L is in the loss function L l Distillation training was conducted under control, L s The goal is to guide the logic detection module L to reconstruct low-level feature maps from the high-level feature maps extracted by the teacher network T from the sample images. Overall, the structure detection module S and the logic detection module L operate under the corresponding loss function L. s and L l Under their control, we complete the model training together.
[0065] After training with knowledge distillation, based on the differences between the feature map extracted by the structure detection module S and the feature map extracted by the teacher network T, and the differences between the feature map reconstructed by the logic detection module L and the low-level feature map extracted by the teacher network T, regions with structural and / or logical abnormalities in the image to be detected can be identified by comparison.
[0066] This invention employs unsupervised knowledge transfer technology to train the entire network model F. During training, only normal sample images are involved (described below), without any real or artificially created defective samples. After the network model is trained, the image to be detected is input into the network model, and the network model can directly output regions with structural and / or logical anomalies in an end-to-end manner. Here, normal sample images refer to images without structural or logical anomalies.
[0067] The specific structure of the detection method and network model of the present invention will be described below. Please refer to [link / reference needed]. Figure 2 In one embodiment, the detection method includes steps 100 to 300, which are described in detail below.
[0068] Step 100: Obtain the image to be detected.
[0069] The image to be detected can be obtained by taking pictures of the object to be detected using imaging devices such as cameras / video cameras. The object to be detected can be a product obtained from an industrial production line, or an object in other scenarios where structural and logical anomalies need to be detected, such as the detection of object assembly and placement.
[0070] Step 200: Input the image to be detected into the pre-trained teacher network T for feature extraction to obtain multi-layer first feature maps with different resolutions of the image to be detected; input the image to be detected into the trained structure detection module S for feature extraction to obtain multi-layer second feature maps with different resolutions of the image to be detected; input the first feature map of the highest layer into the trained logic detection module L for decoding and reconstruction to obtain the third feature map, the resolution of the third feature map being consistent with the resolution of the first feature map of the lowest layer.
[0071] The teacher network T and the structure detection module S can employ various commonly used feature extraction networks. The first and second feature maps are feature maps output by certain layers (e.g., standard convolutional layers) of the teacher network T and the structure detection module S, respectively. The teacher network T is pre-trained using a public dataset, and its network parameters are loaded with pre-trained weights. In some embodiments of this invention, the teacher network T is a network after removing the fourth feature extraction layer and subsequent layers from ResNet-18, and the first feature map is the feature map output by the first three feature extraction layers of ResNet-18. This design takes into account two points: first, feature maps with too small a resolution are detrimental to semantically intensive tasks such as defect detection; second, it is a lightweight design that prunes unnecessary network layers, ensuring detection accuracy while achieving real-time inference. Therefore, the fourth feature extraction layer of ResNet-18 is removed.
[0072] In some embodiments, the structure detection module S has the same structure as the teacher network T.
[0073] Here, multi-layer first and second feature maps of different resolutions are used simultaneously for distillation and detection, which can improve the detection accuracy of structural anomalies with small areas, thereby completing the task of detecting structural defects in industrial images. Typically, the feature maps output by the network model decrease in resolution from low to high layers, and higher-level feature maps extract more advanced features. For example, for the ResNet-18 that outputs three layers of first feature maps, assuming the data shape of the image to be detected is [1,3,256,256], where 1 is the batch number, 3 is the number of image channels, and [256,256] represents the height and width of the image, ResNet-18 outputs three layers of first feature maps t1, t2, and t3, with data shapes of [1,64,64,64], [1,128,32,32], and [1,256,16,16], respectively.
[0074] The logic detection module L receives the first feature map from the highest layer as input and outputs a third feature map. Structurally, the teacher network T acts as an "encoder," while the logic detection module L is a "decoder," its essential function being the reconstruction of the feature maps. For logical anomalies such as discrepancies in quantity or incorrect positions, the model needs to focus more on the global picture. Higher-level feature maps have larger receptive fields and better global characteristics. At the same time, it is also necessary to preserve positional information. Therefore, the logic detection module L receives the first feature map from the highest layer as input, then decodes it to reconstruct lower-level features, outputting a third feature map. The resolution of the third feature map is consistent with the resolution of the first feature map from the lowest layer, meaning that the first feature map from the lowest layer is used as the reconstruction target.
[0075] The logic detection module L can adopt various decoder structures, and can be composed of multiple convolutional layers and upsampling, etc.
[0076] Step 300: Compare the second feature map with the first feature map, and compare the third feature map with the first feature map of the lowest layer, to determine the regions in the image to be detected that have structural and / or logical abnormalities.
[0077] In some embodiments, the number of second feature maps is the same as the number of first feature maps, and the number of second feature maps is the same as the number of first feature maps. Figure 1 There is a one-to-one correspondence, and the corresponding second feature map and the first feature map have the same resolution. By comparing the difference between the second feature map and the corresponding first feature map, as well as the difference between the third feature map and the first feature map of the lowest layer, it can be determined that there are regions with structural and / or logical abnormalities in the image to be detected.
[0078] Essentially, the structure detection module S learns the knowledge induction of normal samples by the teacher network T during distillation training, but it does not master the knowledge extraction ability of the teacher network T for abnormal samples. Therefore, during inference, the specific location of structural anomalies in the image to be detected can be located based on the difference between the feature maps output by the teacher network T and the structure detection module S for the input image to be detected. For example, regions with a difference greater than a certain threshold are identified as structural anomaly regions.
[0079] Similarly, the logic detection module L, which has only been trained by distillation of normal samples, is unable to reconstruct the low-level features generated by the teacher network T from anomalous samples with logical anomalies. Figure 1 The output is consistent, and logical anomalies can be located based on the difference between the feature map reconstructed by the logic detection module L and the low-level feature map actually generated by the teacher network T. For example, regions with differences greater than a certain threshold are considered as logical anomaly regions.
[0080] In some embodiments, the detection results can be output in the form of a defect segmentation map, where abnormal regions are represented by different pixel values compared to other regions, thus identifying areas in the image to be detected that contain structural and / or logical anomalies. It can be understood that when no structural anomalies exist, the defect segmentation map only identifies logically abnormal regions; when no logical anomalies exist, the defect segmentation map only identifies structurally abnormal regions; and when neither type of anomaly exists, no abnormal regions are displayed on the defect segmentation map. Logically abnormal regions and structurally abnormal regions can be identified using different pixel values, or they can be left undifferentiated, simply indicating the presence of a defect. There can also be two defect segmentation maps: a structural anomaly segmentation map and a logical anomaly segmentation map. The structural anomaly segmentation map identifies regions in the image to be detected that contain structural anomalies, and the logical anomaly segmentation map identifies regions in the image to be detected that contain logical anomalies.
[0081] In some preferred embodiments, the first feature map has M layers, with the resolution decreasing and the number of channels increasing in each layer (i.e., the resolution decreases and the number of channels increases with the number of layers), where M is an integer greater than or equal to 2; the logical detection module L includes an M-1 layer of first convolutional layer Conv1, as well as a spatial attention submodule SA and a channel attention submodule CA, such as Figure 3 As shown. Figure 3 The logical detection module L in the example consists of two first convolutional layers (M=3). The resolution of the first feature map in each layer is halved, and the number of channels is doubled. However, the actual implementation is not limited to this and can be configured according to actual needs. Based on this, please refer to... Figure 4 In step 200, the first feature map of the highest layer is input into the trained logic detection module for decoding and reconstruction to obtain the third feature map, including steps 210 to 240, which are explained in detail below.
[0082] Step 210: Input the first feature map of the highest layer into the first convolutional layer Conv1 of the first layer for convolution processing and then upsample to obtain an upsampled feature map. Input the upsampled feature map into the first convolutional layer Conv1 of the next layer for convolution processing and then upsample to obtain an upsampled feature map of the next layer. Repeat the step of inputting the upsampled feature map into the first convolutional layer Conv1 of the next layer for convolution processing and then upsample to obtain an upsampled feature map of the next layer until the upsampled feature map of the M-1th layer is obtained. The resolution and number of channels of the upsampled feature map of the kth layer are the same as those of the first feature map of the M-kth layer, k∈[1,M-1].
[0083] by Figure 3 For example, assuming the data shape of the third layer first feature map t3 output by the teacher network T is [N,C,H,W], where N is the batch number, C is the number of channels, and [H,W] represents the height and width, then the data shape of the second layer first feature map t2 is [N,C / 2,2H,2W], and the data shape of the first layer first feature map t1 is [N,C / 4,4H,4W]. The third layer first feature map t3 serves as the input of the logic detection module L. First, it passes through the first convolutional layer Conv1 to reduce the number of channels to half of the input, and then performs an upsampling operation to increase the resolution to twice that of the input, the same as the first feature map t2; then it passes through the first convolutional layer Conv1 again to reduce the number of channels to one-quarter of the input, and then performs an upsampling operation to increase the resolution to four times that of the input, the same as the first feature map t1.
[0084] The first convolutional layer, Conv1, has the function of dimensionality reduction. It is preferable to use a convolutional layer with a kernel size of 3×3. The advantages are: first, it can effectively decode and reconstruct features; second, existing technologies generally use 1×1 Conv with a kernel size of 1×1 for channel dimensionality reduction, but using 1×1 Conv here will lead to over-extraction of feature map information and loss of important positional information. Compared with using 1×1 Conv, 3×3 Conv can better preserve positional information, which is beneficial for semantically intensive tasks such as defect segmentation.
[0085] Step 220: Input the upsampled feature map of layer M-1 into the spatial attention submodule SA to perform attention enhancement processing in space to obtain the spatial attention map.
[0086] The spatial attention submodule (SA) primarily aims to enhance the spatial importance of features. One embodiment employs a minimalist yet effective structural design, such as... Figure 5 As shown, it includes a second convolutional layer Conv2 and a batch normalization layer Bn. Step 220 specifically includes: inputting the upsampled feature map inp from the M-1 layer into the second convolutional layer Conv2 for dimensionality reduction to reduce the number of channels inp to 1; then processing it through the batch normalization layer Bn to obtain a batch normalized feature map; performing a Sigmoid operation on the batch normalized feature map to obtain a spatial attention score map; and performing an element-wise multiplication operation (Mul) between the upsampled feature map inp from the M-1 layer and the spatial attention score map to obtain a spatial attention map outs. The second convolutional layer Conv2 can be a 1×1 Conv2.
[0087] Step 230: Input the upsampled feature map of layer M-1 into the channel attention submodule CA to perform attention enhancement processing on the channel to obtain the channel attention map.
[0088] The primary function of the Channel Attention (CA) submodule is to enhance the channel importance of features. One embodiment also employs a very simple yet effective design, such as... Figure 6 As shown, it includes a global average pooling layer (GlobalAvgPool) and a third convolutional layer (Conv3). Step 220 specifically includes: inputting the upsampled feature map inp from the M-1 layer into the global average pooling layer (GlobalAvgPool) to reduce the resolution to 1×1, then processing it through the third convolutional layer (Conv3) to obtain a third convolutional feature map, where the number of input and output channels of the third convolutional layer (Conv3) remains consistent; performing a Sigmoid operation on the third convolutional feature map to obtain a channel attention score map; and performing element-wise multiplication of the upsampled feature map inp from the M-1 layer and the channel attention score map to obtain a channel attention map outc. The third convolutional layer (Conv3) can be a 1x1 Conv3.
[0089] As can be seen, the two attention submodules in the above embodiments are efficient, lightweight, and simple in structure, which helps to ensure detection accuracy and real-time performance.
[0090] Step 240: Perform an element-wise addition operation (Add) on the spatial attention map and the channel attention map to obtain the third feature map.
[0091] like Figure 3 As shown, the data shape of the obtained third feature map output is the same as that of the first feature map t1 of the lowest layer, thus achieving reconstruction.
[0092] For step 300, let's assume that the first feature map has M layers, and the resolution of the first feature map in each layer decreases. The specific process of step 300 is explained below.
[0093] First, the squared differences between corresponding elements of each second feature map and its corresponding first feature map are calculated, resulting in M second difference squared maps. Then, the mean of each second difference squared map is calculated along the channel dimension (i.e., the values of all channels are summed and averaged), followed by a sigmoid operation, resulting in M structural anomaly sub-maps. Next, all structural anomaly sub-maps except the one with the highest resolution are upsampled to the size of the highest-resolution sub-map. Finally, element-wise multiplication is performed on the M structural anomaly sub-maps to obtain the final structural anomaly score.
[0094] The squared difference between corresponding elements of the third feature map and the first feature map of the lowest layer is calculated to obtain the third difference squared map. The mean of the third difference squared map is calculated in the channel dimension, and then the Sigmoid operation is performed to obtain the logic anomaly score map.
[0095] Finally, the structural anomaly score map and the logical anomaly score map are upsampled to the size of the image to be detected. Then, the structural anomaly score map is segmented by a preset structural score threshold to obtain a structural anomaly segmentation map, and the logical anomaly score map is segmented by a preset logical score threshold to obtain a logical anomaly segmentation map.
[0096] Alternatively, an element-wise multiplication operation can be performed on the structural anomaly score map and the logical anomaly score map to obtain a merged anomaly score map. This merged anomaly score map is then upsampled to the size of the image to be detected, and finally, a preset detection threshold is used to perform threshold segmentation on the merged anomaly score map to obtain a defect segmentation map. The above upsampling operations can be implemented using interpolation algorithms.
[0097] by Figure 1For example, firstly, the squares of the differences between corresponding elements of s1 and t1, s2 and t2, and s3 and t3 are calculated to obtain the second difference squared graphs r1, r2, and r3. Then, the mean of the second difference squared graphs r1, r2, and r3 is calculated along the channel dimension, followed by a sigmoid operation to obtain the structural anomaly subgraphs u1, u2, and u3. Next, the structural anomaly subgraphs u2 and u3 are upsampled to the size of u1 to obtain the structural anomaly subgraphs v2 and v3. Finally, element-wise multiplication is performed on u1, v2, and v3 to obtain the structural anomaly score map. s The squared difference between corresponding elements in the third feature map l1 and the lowest-level first feature map t1 is calculated to obtain the third difference squared map y. The mean of the third difference squared map y is calculated along the channel dimension, followed by a Sigmoid operation to obtain the logistic anomaly score map map. l Structural anomaly score map s Logical anomaly score map l Element-level multiplication is performed to obtain the merged anomaly score map result. The merged anomaly score map result is upsampled to the size of the image to be detected. Finally, a preset detection threshold th is used to perform threshold segmentation on the merged anomaly score map result to obtain the defect segmentation map (that is, if the element value of the merged anomaly score map result is greater than or equal to the detection threshold th, it is mapped to 1, otherwise it is 0, 1 represents defect, and 0 represents normal).
[0098] Please refer to Figure 7 , Figure 7 The image on the left is an image to be detected in one embodiment. It can be seen that one light bulb is missing on the left side of the image (the quantity is incorrect), which is a logical anomaly; two light bulbs on the right side are broken, which is a structural anomaly. This image is input into a network model F. Assuming that a merged anomaly score map is obtained, and a threshold segmentation is performed on the merged anomaly score map to obtain a defect segmentation map, with a detection threshold th = 0.5, the defect segmentation map output by network model F is as follows. Figure 7 As shown in the middle right image, two defects were accurately detected.
[0099] The specific training methods for the structure detection module S and the logic detection module L are described below. In some embodiments of this invention, corresponding loss functions are designed for training the structure detection module S and the logic detection module L respectively. The combined design of the structure detection module S and the logic detection module L, as well as the design of the loss function during the training phase, are crucial. Please refer to... Figure 8 The structure detection module S and the logic detection module L are trained through steps 10 to 20, which are explained in detail below.
[0100] Step 10: Input the normal sample image into the pre-trained teacher network for feature extraction to obtain M layers of first normal sample feature maps with different resolutions; input the normal sample image into the structure detection module to be trained for feature extraction to obtain M layers of second normal sample feature maps corresponding to the first normal sample feature maps, where M is an integer not less than 2; input the first normal sample feature map of the highest layer into the logic detection module to be trained for decoding and reconstruction to obtain the third normal sample feature map, the resolution of the third normal sample feature map being consistent with the resolution of the first normal sample feature map of the lowest layer.
[0101] For details on step 10, please refer to the above description of step 200; it will not be repeated here.
[0102] Step 20: Based on the total loss function L f The structure detection module S and the logic detection module L are trained, and the total loss function L is used. f From the structural loss function L s and logical loss function L l Determine the structural loss function L. s The feature map of the second normal sample represents the difference between the feature map of the corresponding first normal sample, thereby enabling the structure detection module S to learn the ability to extract knowledge from normal samples; the logistic loss function L l The difference between the third normal sample feature map and the lowest-level first normal sample feature map is represented, so that the logic detection module L can learn to reconstruct the low-level feature map using the high-level feature map of the normal sample image extracted by the teacher network T.
[0103] It should be noted that the structure and parameters of the teacher network T remain unchanged during the training process. Its network parameters are loaded with pre-trained weights. Throughout the entire network training and inference process, its network structure and parameters remain frozen and are not updated.
[0104] In one embodiment, the structural loss function L s The expression is:
[0105]
[0106] Where L m Let L represent the loss function between the m-th first normal sample feature map and the corresponding second normal sample feature map. For any m∈[1,M], L m The expressions are consistent; to avoid redundancy, let L represent any L. m Its expression is
[0107]
[0108] Where g represents the one-dimensional flattening vector of the first normal sample feature map in the m-th layer, and h represents the one-dimensional flattening vector of the corresponding second normal sample feature map. i Let h represent the i-th element of vector g. i Let represent the i-th element of vector h, and n represent the number of elements in vector g or h. The one-dimensional flattened vector of a feature map refers to a one-dimensional vector formed by arranging the elements of the feature map in order. For example, a feature map with the data shape [1, 64, 64, 64], flattened into a one-dimensional vector, becomes [64×64×64], or [262144], which is a vector containing 262144 elements. Since the first normal sample feature map and the corresponding second normal sample feature map have the same resolution, vectors g and h have the same number of elements.
[0109] In one embodiment, the logical loss function L l The expression is:
[0110]
[0111] Where q represents the one-dimensional flattened vector of the first normal sample feature map of the lowest layer, and p represents the one-dimensional flattened vector of the third normal sample feature map. j p represents the j-th element of vector q. j Let represent the j-th element of vector p, and o represent the number of elements in vector q or p. Since the resolution of the third normal sample feature map and the first normal sample feature map of the lowest layer are the same, the number of elements in vector q and p are the same.
[0112] In one embodiment, the total loss function L f The expression is:
[0113] L f =αL s +βL l ,
[0114] Where α and β are preset weighting coefficients, and an example is α = 0.3 and β = 0.7.
[0115] Those skilled in the art will understand that network models typically require iterative training, meaning that steps 10 to 20 can be repeated multiple times, each time using different normal sample images as input for training.
[0116] The method for detecting structural and logical anomalies in images provided in this invention adopts a "teacher-dual-student" architecture. The structural detection module and the logical detection module are trained using knowledge distillation from a pre-trained teacher network. The structural detection module acquires the feature extraction capability of the teacher network from training samples, and the logical detection module reconstructs low-level feature maps of the training samples. During detection, based on the differences between the feature maps extracted by the structural detection module and those extracted by the teacher network, and the differences between the reconstructed feature maps and the low-level feature maps extracted by the teacher network, regions with structural and / or logical anomalies in the image can be detected through comparison. Compared to existing technologies, this method offers the following advantages:
[0117] (1) During training, only normal sample images can be used for training without any abnormal samples, thus overcoming the lack of abnormal samples.
[0118] (2) Improve detection accuracy. A logic detection module and a structure detection module were designed to detect logical anomalies and structural anomalies, respectively. The efficient combination of the two modules enables accurate detection of defects, which can better solve the defect segmentation problem of combining structural and logical anomalies, greatly reduce false detection of normal sample images, and significantly improve defect detection capability.
[0119] (3) Simple operation and easy deployment, enabling enterprises to reduce costs and increase efficiency. The network model proposed in this invention adopts an end-to-end design, which can realize one-click training and inference. There is no need to specifically save the feature memory related to the image to be tested. The operation is simple, which greatly reduces the operating threshold for users. It is conducive to the detection deployment in various production scenarios, reduces the storage cost and labor cost of production enterprises, and achieves cost reduction and efficiency improvement.
[0120] (4) In some embodiments, the network model is designed with lightweight features, enabling real-time inference. In some embodiments of the present invention, a "teacher-dual-student" architecture is used, and the networks used by each module are designed with lightweight features, which can achieve real-time inference while achieving high detection accuracy.
[0121] Those skilled in the art will understand that all or part of the functions of the various methods in the above embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in a server, another computer, disk, optical disk, flash drive, or external hard drive, etc., and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be achieved.
[0122] The above examples illustrate the present invention only to aid in understanding it and are not intended to limit the scope of the invention. Those skilled in the art can make various simple deductions, modifications, or substitutions based on the principles of this invention.
Claims
1. A method for detecting structural and logical anomalies in an image, characterized in that, include: Acquire the image to be detected; The image to be detected is input into a pre-trained teacher network for feature extraction to obtain a multi-layer first feature map of the image to be detected with different resolutions; The image to be detected is input into a trained structure detection module for feature extraction to obtain multiple layers of second feature maps with different resolutions. The number of second feature maps is the same as the number of first feature maps, and there is a one-to-one correspondence between the second and first feature maps. The corresponding second and first feature maps have the same resolution. The first feature map of the highest layer is input into a trained logic detection module for decoding and reconstruction to obtain a third feature map. The resolution of the third feature map is consistent with the resolution of the first feature map of the lowest layer. The logic detection module adopts a decoder structure. The difference between the second feature map and the corresponding first feature map is compared, and the region in the image to be detected containing structural anomalies is located based on the difference between the second feature map and the corresponding first feature map; the difference between the third feature map and the lowest layer first feature map is compared, and the region in the image to be detected containing logical anomalies is located based on the difference between the third feature map and the lowest layer first feature map; wherein, structural anomalies refer to surface quality defects of objects in the image, and logical anomalies refer to anomalies of objects in the image violating logical constraints. The structure detection module and the logic detection module are trained on normal sample images under the knowledge distillation of the pre-trained teacher network.
2. The detection method as described in claim 1, characterized in that, The structure detection module is the same as the teacher network structure.
3. The detection method as described in claim 1 or 2, characterized in that, The first feature map has M layers, with the resolution decreasing and the number of channels increasing in each layer, where M is an integer greater than or equal to 2; the logical detection module includes an M-1 layer first convolutional layer, a spatial attention submodule, and a channel attention submodule; The step of inputting the first feature map of the highest layer into the trained logic detection module for decoding and reconstruction to obtain the third feature map includes: The first feature map of the highest layer is input into the first convolutional layer of the first layer for convolution processing and then upsampled to obtain an upsampled feature map. The upsampled feature map is then input into the first convolutional layer of the next layer for convolution processing and then upsampled to obtain an upsampled feature map of the next layer. The steps of inputting the upsampled feature map into the first convolutional layer of the next layer for convolution processing and then upsampling to obtain an upsampled feature map of the next layer are repeated until the upsampled feature map of the M-1th layer is obtained. The resolution and number of channels of the upsampled feature map of the kth layer are the same as those of the first feature map of the M-kth layer, where k∈[1, M-1]. The upsampled feature map of layer M-1 is input into the spatial attention submodule to perform attention enhancement processing in space to obtain a spatial attention map; The upsampled feature map of layer M-1 is input into the channel attention submodule to perform attention enhancement processing on the channel to obtain the channel attention map; The spatial attention map and the channel attention map are added element-wise to obtain the third feature map.
4. The detection method as described in claim 3, characterized in that, The kernel size of the first convolutional layer is 3×3.
5. The detection method as described in claim 3, characterized in that, The spatial attention submodule includes a second convolutional layer and a batch normalization layer; The step of inputting the upsampled feature map of layer M-1 into the spatial attention submodule to perform attention enhancement processing in space to obtain a spatial attention map includes: The upsampled feature map of layer M-1 is input into the second convolutional layer for dimensionality reduction to reduce the number of channels of the upsampled feature map of layer M-1 to 1. Then, it is processed by the batch normalization layer to obtain the batch normalized feature map. Perform a Sigmoid operation on the batch normalized feature map to obtain a spatial attention score map; The spatial attention map is obtained by performing element-wise multiplication of the upsampled feature map of the M-1th layer and the spatial attention score map.
6. The detection method as described in claim 3, characterized in that, The channel attention submodule includes a global average pooling layer and a third convolutional layer; The step of inputting the upsampled feature map of layer M-1 into the channel attention submodule to perform attention enhancement processing on the channel to obtain the channel attention map includes: The upsampled feature map of the M-1th layer is input into the global average pooling layer to reduce the resolution to 1×1, and then processed by the third convolutional layer to obtain the third convolutional feature map, wherein the number of input and output channels of the third convolutional layer is kept consistent. Perform a Sigmoid operation on the third convolutional feature map to obtain a channel attention score map; The channel attention map is obtained by performing element-wise multiplication of the upsampled feature map of the M-1th layer and the channel attention score map.
7. The detection method as described in claim 1, characterized in that, The first feature map has M layers, and the resolution of the first feature map in each layer decreases, where M is an integer not less than 2; The step of comparing the difference between the second feature map and the corresponding first feature map, and locating regions with structural anomalies in the image to be detected based on the difference between the second feature map and the corresponding first feature map, and comparing the difference between the third feature map and the lowest layer first feature map, and locating regions with logical anomalies in the image to be detected based on the difference between the third feature map and the lowest layer first feature map, includes: The squared difference between each second feature map and the corresponding first feature map is calculated to obtain M second difference squared maps. The mean of each second difference squared map is calculated in the channel dimension, and then a sigmoid operation is performed to obtain M structural anomaly fractional sub-maps. The structural anomaly fractional sub-maps other than the structural anomaly fractional sub-map with the largest resolution are upsampled to the size of the structural anomaly fractional sub-map with the largest resolution. Then, element-wise multiplication is performed on the M structural anomaly fractional sub-maps to obtain the structural anomaly fractional map. The squared difference between corresponding elements of the third feature map and the first feature map of the lowest layer is calculated to obtain the third difference square map; the mean of the third difference square map is calculated in the channel dimension, and then the Sigmoid operation is performed to obtain the logic anomaly score map. The structural anomaly score map and the logical anomaly score map are upsampled to the size of the image to be detected. Then, a preset structural score threshold is used to perform threshold segmentation on the structural anomaly score map to obtain a structural anomaly segmentation map. The structural anomaly segmentation map is used to identify regions in the image to be detected that have structural anomalies. Similarly, a preset logical score threshold is used to perform threshold segmentation on the logical anomaly score map to obtain a logical anomaly segmentation map. The logical anomaly segmentation map is used to identify regions in the image to be detected that have logical anomalies. Alternatively, an element-wise multiplication operation is performed on the structural anomaly score map and the logical anomaly score map to obtain a merged anomaly score map. The merged anomaly score map is then upsampled to the size of the image to be detected. Finally, a preset detection threshold is used to perform threshold segmentation on the merged anomaly score map to obtain a defect segmentation map. The defect segmentation map is used to identify regions in the image to be detected where structural anomalies and / or logical anomalies exist.
8. The detection method as described in claim 7, characterized in that, The structure detection module and the logic detection module are trained in the following way: Normal sample images are input into the pre-trained teacher network for feature extraction to obtain M layers of first normal sample feature maps with different resolutions; Normal sample images are input into the structure detection module to be trained for feature extraction to obtain an M-layer second normal sample feature map corresponding to the first normal sample feature map, where M is an integer not less than 2; The first normal sample feature map of the highest layer is input into the logic detection module to be trained for decoding and reconstruction to obtain the third normal sample feature map. The resolution of the third normal sample feature map is consistent with the resolution of the first normal sample feature map of the lowest layer. Based on the total loss function L f The structure detection module and the logic detection module to be trained are trained, and the total loss function L is used. f From the structural loss function L s and logical loss function L l It is determined that the structural loss function L s The logistic loss function L represents the difference between the second normal sample feature map and the corresponding first normal sample feature map. l The difference between the third normal sample feature map and the first normal sample feature map of the lowest layer is represented, and the structure and parameters of the teacher network remain unchanged during training.
9. The detection method as described in claim 8, characterized in that, The structural loss function L s The expression is: , Where for any m∈[1,M], L m The expression is consistent, that is... , Where L represents any L m Let g represent the one-dimensional flattening vector of the first normal sample feature map in the m-th layer, and h represent the one-dimensional flattening vector of the corresponding second normal sample feature map. i Let h represent the i-th element of vector g. i Let represent the i-th element of vector h, and n represent the number of elements in vector g or h. The one-dimensional flattened vector of a feature map refers to a one-dimensional vector formed by arranging the elements of the feature map in order.
10. The detection method as described in claim 8, characterized in that, The logical loss function L l The expression is: , Where q represents the one-dimensional flattening vector of the first normal sample feature map at the lowest layer, and p represents the one-dimensional flattening vector of the third normal sample feature map. j p represents the j-th element of vector q. j represents the j-th element of vector p, o represents the number of elements in vector q or p, and the one-dimensional flattened vector of the feature map refers to the one-dimensional vector formed by arranging the elements of the feature map in order.
11. The detection method according to any one of claims 8 to 10, characterized in that, The total loss function L f The expression is: , Where α and β are preset weighting coefficients.
12. A computer-readable storage medium, characterized in that, The medium stores a program that can be executed by a processor to implement the detection method as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Defect detection method and device based on unsupervised learning
CN114862838A
Neural network training methods, devices, and storage media based on knowledge distillation
CN114936605A