A method and system for fine-grained industrial image anomaly detection and localization based on dual-path grid reconstruction
Through the dual-channel grid reconstruction method, the efficiency and accuracy problems in the prior art under the conditions of fine-grained defect detection and limited resources are solved, and efficient and accurate industrial image abnormality detection and positioning are achieved.
Patent Information
- Application Number
- CN202411650783.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-11-19
AI Technical Summary
Existing industrial image anomaly detection methods perform poorly when dealing with fine-grained defects, and have problems with poor generalization ability and high spatial complexity, making it difficult to effectively utilize the knowledge of normal samples and synthetic anomaly samples.
Using a method based on dual-channel grid reconstruction, the final reconstruction features are generated and the exception score is calculated to detect and locate abnormalities through preprocessing and feature extraction, coordinate mapping, dual-channel grid reconstruction and feature refinement.
It improves the accuracy and efficiency of industrial image abnormality detection, can efficiently and accurately perform fine-grained abnormality detection and positioning, and effectively synthesize diversified abnormal data under limited computing resources.
Smart Images

Figure CN119599982B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and image processing, and specifically to a method and system for fine-grained industrial image anomaly detection and localization based on dual-path grid reconstruction. Background Art
[0002] Image anomaly detection and localization aims to identify and precisely segment abnormal regions in images, and is widely applied in fields such as industrial inspection, medical imaging, and video surveillance. However, due to the scarcity of abnormal samples and the diversity of abnormal patterns (such as slight scratches to severe structural damage in industrial production), this task faces great challenges. Under these challenges, more and more research focuses on developing unsupervised methods.
[0003] In industrial image anomaly detection, typical examples of unsupervised methods include PaDiM, SPADE, and PatchCore, which rely on an external vector database to store features extracted from normal samples. During the inference process, anomalies are identified by calculating the Euclidean distance between the test sample and the nearest neighbors in the database. Although these methods are effective, due to the way of storing discrete features, there are problems of poor generalization ability, and a large amount of diverse normal features need to be stored, resulting in high spatial complexity and resource-intensive search operations. In addition, existing industrial image anomaly detection methods often perform poorly when dealing with fine-grained defects, mainly because they lack knowledge of real abnormal data, resulting in a non-compact enough defined boundary of normal data. Although some generative models such as DFMGAN and AnomalyDiffusion can synthesize more realistic anomalies to help refine the normal data boundary, the training of these models requires a large amount of computing resources, limiting their wide use in practical applications.
[0004] Therefore, how to solve the problems of accuracy and efficiency in the above methods, and at the same time utilize the knowledge of normal samples and synthetic abnormal samples to improve the performance of fine-grained anomaly detection has become a research topic. In addition, how to effectively synthesize diverse abnormal data under limited computing resources is also an urgent problem to be solved at present. For this reason, the present application proposes a method and system for fine-grained industrial image anomaly detection and localization based on dual-path grid reconstruction. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for fine-grained industrial image anomaly detection and localization based on dual-path grid reconstruction to solve the problems raised in the above background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for fine-grained industrial image anomaly detection and localization based on dual-path grid reconstruction, including the following steps:
[0007] Step S1: Preprocess the input sample image to be processed by unifying its size and standardizing it. Then, pass it through a pre-trained feature extraction network, take the output of the middle layer to obtain aligned original features, and further synthesize abnormal features on the original features through the FBP module during the abnormal grid training phase. Among them, the sample image to be processed includes industrial product images without label information;
[0008] Step S2: Input the original features in Step S1 into the coordinate mapping module, and map them through the trained convolutional layer of the coordinate mapping module to obtain local and global coordinates;
[0009] Step S3: Use the local and global coordinates obtained in Step S2 to sample from both the normal and abnormal grids, and fuse the sampling results and input them into the CNN module to obtain preliminary reconstruction features. The normal and abnormal grids respectively contain corresponding local and global grids. Sample using the local and global coordinates in the local and global grids to obtain preliminary reconstruction features that take into account both the local details and global information of the image;
[0010] Step S4: Input the original features in Step S1 and the preliminary reconstruction features in Step S3 into the feature refinement module together to obtain the final reconstruction features. Among them, the feature refinement module combines the mean square error and cosine similarity to evaluate the pixel-level similarity between the original features and the preliminary reconstruction features, and guides the weighted fusion of the original features and the preliminary reconstruction features based on this similarity to obtain the final reconstruction features;
[0011] Step S5: Use the difference between the final reconstruction features in Step S4 and the original features in Step S1 as the anomaly score to detect and locate anomalies;
[0012] Step S6: After receiving an inference request sent by the terminal device, map the image to be measured into the feature space, then perform reconstruction at the feature level to obtain the reconstructed features, and compare the reconstructed features with the original features to obtain the inference result.
[0013] Preferably: The pre-trained feature extraction network in Step S1 includes deep convolutional neural networks such as ResNet, VGG, and Inception.
[0014] Preferably: The FBP anomaly synthesis module in Step S1 only synthesizes anomalies with controllable size, shape, intensity, and position for normal samples at the feature level during the abnormal grid training period, and provides corresponding annotation mask images. Specifically, the FBP anomaly synthesis module operates through the following steps:
[0015] Define the block size as B, the block intensity as I, and the block center coordinates (x c , y c ). First, generate the initialization mask M:
[0016] M = zeros(2B + 1, 2B + 1)
[0017] where M is a (2B + 1)×(2B + 1) matrix with an initial value of 0. Then, a random walk mask is generated: 1) Select the initial position (x 0 , y 0 ), where x 0 = B and y 0 = B; 2) Randomly select the number of walk steps N, which is randomly selected from B to 2B; 3) Perform the random walk process, each time randomly select the values of Δx and Δy, whose value range is { -1, 0, 1}, update the position (x k+1 , y k+1 ), where x k+1 = x k + Δx and y k+ = y k + Δy, and mark the corresponding position in the mask M as 1; Finally, generate the abnormal block and paste it onto the feature map: 1) Initialize the block paste tensor P, whose size is 1×1×H×W with an initial value of 0; 2) According to the markings of the mask M, paste the block with intensity I onto the corresponding position of the feature map; 3) Perform Gaussian blur processing on the paste tensor P to obtain P blurred ; 4) Paste the tensor P blurred after Gaussian blur processing onto the feature map.
[0018] Preferably: The coordinate mapping module in step S2 maps the original features to the corresponding coordinate space by learning the mapping relationship in the local and global coordinate spaces. Among them, the local coordinate space is used to capture the local details of the image, and the global coordinate space is used to capture the overall structure of the image.
[0019] Preferably: The process of creating normal and abnormal grids in step S3 is as follows: 1) Create an uninitialized grid tensor; 2) Initialize the grid tensor using the Xavier normal distribution; 3) Convert the initialized tensor to the Parameter type and manage it using ParameterList.
[0020] Preferably: The process of training normal and abnormal grids in step S3 is as follows: 1) First, freeze the normal grid parameters, use normal samples to synthesize abnormalities through the FBP module, and train the abnormal grid; 2) Freeze the abnormal grid parameters and train the normal grid; 3) The coordinate mapping module in S2 and the CNN module in S3 are both trainable modules and participate in the training throughout the process; 4) Use the mean square error as the reconstruction loss for training the normal grid, and its formula is as follows:
[0021]
[0022] Where C represents the number of channels of the feature map, H and W represent the height and width of the feature map respectively, and φ(x) represents the original feature in S1. represents the final reconstructed feature of S4; 5) The training anomaly grid also reconstructs the error in the above manner, where φ(x) is replaced by the feature map after synthesizing the anomaly, and the truncated L1 loss is used as the contrast loss to refine the boundary between the normal feature and the abnormal feature. The formula is as follows:
[0023]
[0024] Where th represents the buffer size around the dividing line, D is a set of sample pair similarities constructed according to the mask, and d + represents the similarity of the positive sample pair, and d - represents the similarity of the negative sample pair.
[0025] Preferably, the feature refinement module in step S4 evaluates the pixel-level similarity between the original feature and the preliminary reconstructed feature by calculating the mean square error (MSE) and cosine similarity. Among them, MSE is used to capture the absolute difference at the pixel level, and cosine similarity is used to capture the directional similarity between feature vectors. Specifically, the feature refinement module operates through the following steps:
[0026] First, calculate the similarity Sim between the original feature φ aligned (x) and the preliminary reconstructed feature The formula is as follows:
[0027]
[0028] Where ∈ is an infinitesimal constant.
[0029]
[0030] Among them, is the indicator function, which is assigned a value of 1 when the condition in the parentheses is satisfied. mse(·, ·) and cosim(·, ·) respectively represent the calculation of the mean square error and cosine similarity. The above similarity Sim is then applied to the reconstructed feature to obtain the final reconstructed feature
[0031]
[0032] Among them, ⊙ represents element-wise multiplication.
[0033] A system for a fine-grained industrial image anomaly detection and localization method based on dual-path grid reconstruction according to any one of the above, the system comprising: a preprocessing module for unifying the size and normalizing the input picture, selecting the output of the intermediate layer after alignment as the original feature through a pre-trained feature extraction network, and during the abnormal grid training, using the FBP anomaly synthesis module. At the normal sample feature level, by defining the size, intensity and position of the abnormal block, using the random walk algorithm to generate a mask, and synthesizing the abnormal block onto the feature map of the normal sample to obtain the original feature;
[0034] A processing module for obtaining the global coordinates and local coordinates of the original feature through a coordinate mapping module; inputting the global coordinates and local coordinates into a dual-path grid reconstruction module, the dual-path grid reconstruction module consisting of a normal grid and an abnormal grid, and the normal and abnormal grids respectively containing corresponding local and global grids. In each path, through the method of grid sampling, global and local feature samplings are obtained, where the global and local feature samplings are fused by concatenating along the channel dimension, and then the features of the two paths of the normal grid and the abnormal grid are weighted and fused; inputting the fused feature sampling into a CNN module to obtain a preliminary reconstructed feature; inputting the preliminary reconstructed feature and the original feature into a feature refinement module, calculating the mean square error (MSE) and cosine similarity to evaluate the pixel-level similarity between the original feature and the preliminary reconstructed feature, and using the similarity result as a weight to fuse the original feature and the preliminary reconstructed feature to obtain the final reconstructed feature; calculating the anomaly score using the difference between the final reconstructed feature and the original feature to detect and localize the anomaly, where the anomaly score is the anomaly score of each pixel, and the image-level anomaly score is the maximum value of all pixel-level anomaly scores;
[0035] A query feedback module for receiving the query image sent by the terminal device, mapping the query image to the feature space and then converting it into a reconstructed feature generated by the model, and then determining the image-level and pixel-level anomaly scores according to the difference between the reconstructed feature and the original feature, and feeding back the query result to the terminal device.
[0036] The beneficial effects of the present invention compared with the prior art are:
[0037] A fine-grained industrial image anomaly detection and localization system based on dual-path grid reconstruction provided by the present invention. First, the system standardizes the input image and extracts the original features through a pre-trained network. In the anomaly grid training stage, an FBP anomaly synthesis module generates a feature map containing anomalies. The processing module maps the original features to the local and global coordinate spaces, and through the dual-path grid reconstruction module, performs feature sampling and splicing to generate preliminary reconstructed features. The feature refinement module calculates the mean square error (MSE) and cosine similarity, fuses the preliminary reconstructed features with the original features to obtain the final reconstructed features, calculates the anomaly score using the difference between the final reconstructed features and the original features, determines the pixel-level and image-level anomalies. The query feedback module receives the query image, maps it to the feature space to generate the reconstructed features, and feedbacks the anomaly detection results according to the difference between the reconstructed features and the original features. This system can efficiently and accurately perform industrial defect detection and localization, improve the detection performance, and perform excellently in terms of inference speed and memory occupancy. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic diagram of one implementation manner of the industrial defect detection and localization model of the present invention;
[0039] Figure 2 It is a schematic diagram of one implementation manner of the preprocessing module of the present invention;
[0040] Figure 3 It is a schematic diagram of the effect of the FBP anomaly synthesis module of the present invention;
[0041] Figure 4 It is a schematic diagram of the method flow of the present invention;
[0042] Figure 5 It is a schematic diagram of the system architecture of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0044] Embodiment
[0045] Please refer to Figure 4 , a method for fine-grained industrial image anomaly detection and localization based on dual-path grid reconstruction in the figure, includes the following steps.
[0046] Step S1: Preprocess the input sample images to be processed by unifying their sizes and standardizing them. Then, through a pre-trained feature extraction network, obtain the original features by taking the output of the intermediate layer for alignment. Further, in the abnormal grid training stage, synthesize abnormal features on the original features through the FBP module. Among them, the sample images to be processed include industrial product images without label information;
[0047] Step S2: Input the original features in Step S1 into the coordinate mapping module, and map them through the trained convolutional layer of the coordinate mapping module to obtain local and global coordinates;
[0048] Step S3: Use the local and global coordinates obtained in Step S2 to sample from two paths of normal and abnormal grids, and fuse the sampling results and input them into the CNN module to obtain preliminary reconstruction features. The normal and abnormal grids respectively contain corresponding local and global grids. Use the local and global coordinates to sample in the local and global grids to obtain preliminary reconstruction features that simultaneously consider the local details and global information of the image;
[0049] Step S4: Input the original features in Step S1 and the preliminary reconstruction features in Step S3 into the feature refinement module together to obtain the final reconstruction features. Among them, the feature refinement module combines the mean square error and cosine similarity to evaluate the pixel-level similarity between the original features and the preliminary reconstruction features, and guides the weighted fusion of the original features and the preliminary reconstruction features according to this similarity to obtain the final reconstruction features;
[0050] Step S5: Use the difference between the final reconstruction features in Step S4 and the original features in Step S1 as the anomaly score to detect and locate anomalies;
[0051] Step S6: After receiving the inference request sent by the terminal device, map the image to be measured into the feature space, then reconstruct it at the feature level to obtain the reconstructed features, and compare the reconstructed features with the original features to obtain the inference result.
[0052] Furthermore, combine the abnormal grid trained using synthetic anomalies and the normal grid trained using normal samples. Under the constraint of the abnormal grid, further refine the normal feature boundaries in the normal grid, thereby improving the detection performance of the model for fine-grained anomalies in complex scenarios. Specifically, as Figure 3 shown, use the FBP anomaly synthesis module to synthesize diverse anomalies at the feature level to train the abnormal grid, obtain the understanding of abnormal samples, and thus guide the normal grid to further refine its normal feature boundaries to reconstruct higher-quality normal features. The main method process is as Figure 3 and Figure 5 shown, including:
[0053] The sample image to be processed is input into the preprocessing module for preprocessing, including: unifying the size of the image, standardizing it, and then inputting it into the pre-trained feature extractor. The intermediate layer features are aligned to the same size and then concatenated to obtain the original features. In the abnormal grid training stage, on the basis of the original features, the FBP abnormal synthesis module synthesizes abnormalities at the feature level as abnormal original features.
[0054] In this embodiment, the preprocessing module mentioned can be implemented as a code program in practical applications. First, the input RGB image is processed to have a unified size, and then the image with the unified size is standardized to follow the standard normal distribution. Then, the intermediate layer features are obtained through the pre-trained feature extractor. The multiple intermediate layer features obtained are aligned to the feature layer with the largest selected resolution and then concatenated together to obtain the original features. In addition, during the abnormal grid training, abnormal blocks are synthesized on the basis of the original features, which includes: generating a mask area by means of random walk according to the position and size parameters, then controlling the intensity of the synthesized abnormal blocks through the intensity parameter and fusing them with the original features, and finally refining the edges of the abnormal blocks through Gaussian blur to obtain the original features with abnormalities. For the specific synthesis effect, please refer to Figure 3 .
[0055] The obtained original features are input into the coordinate mapping module to obtain global and local coordinates for subsequent grid sampling. Among them, the coordinate mapping module includes a global coordinate mapping module and a local coordinate mapping module. The original features are mapped to the global coordinate space through the global coordinate mapping module and to the local coordinate space through the local coordinate mapping module. In the global coordinate mapping module, first, the input original features are mapped to a new feature space through a fully connected layer (Linear), then the Tanh activation function is applied to introduce nonlinear characteristics. Next, the features are further adjusted through another fully connected layer, and finally, the Tanh activation function is applied again. In the local coordinate mapping module, the original features first undergo an inter-channel linear transformation through a 1×1 convolutional layer, then the Tanh activation function is applied to introduce nonlinear characteristics. Then, the features are further adjusted through another 1×1 convolutional layer, and finally, the Tanh activation function is applied again. Through the combination of the global and local coordinate mapping modules, the system can effectively capture the local details and overall structure of the image, thus achieving more accurate feature mapping.
[0056] The obtained global and local coordinates are input into the dual-path grid reconstruction module to obtain preliminary reconstructed features. Among them, the grid is trained to be a coordinate function with infinite resolution, outputting the features corresponding to the coordinates. The output features with infinite resolution are aggregated from the nearby features in the grid according to the coordinates of the input features and the distances to the adjacent features. For example, when taking one-dimensional grid sampling then the output feature of channel C is from the one-dimensional grid Adjacent value interpolation, and its mathematical expression is:
[0057] φ(v; G) = |v - n|G[m] + |v - m|G[n]
[0058]
[0059] Where, is an arbitrary input coordinate normalized to the grid resolution R, G[i] represents the feature of index i from grid G, m and n are the indices to be referenced, where, represent the lower bound and upper bound operations respectively. By interpolating 2 D values of the D-dimensional grid (for example, for a 2D grid, according to the input coordinates, interpolate 2 2 i.e., 4 values) adjacent values, the above equation can be simply extended to grids of higher dimensions. Among them, the dual-path grid reconstruction module includes a normal grid and an abnormal grid, and each grid is further divided into a global grid and a local grid. The global coordinates and local coordinates are respectively input into the dual-path grid reconstruction module, and the grid sampling function is used to perform bilinear interpolation on the input global and local coordinates to generate a sampled feature map. For the grid sampling of normal features, the global coordinates and local coordinates are input into the grid sampling function to respectively obtain the global and local normal feature sampling results; for the grid sampling of abnormal features, the global coordinates and local coordinates are also input into the grid sampling function to respectively obtain the global and local abnormal feature sampling results. Then, the sampling results are normalized to eliminate the dimensional differences of the feature values and ensure the stability and consistency of subsequent processing. The normalized sampling results are input into a predefined convolutional block for processing. The convolutional block consists of multiple convolutional layers, batch normalization layers, and activation functions. The sampled normal features are input into the normal convolutional block for processing to obtain preliminary normal reconstruction features; the sampled abnormal features are input into the abnormal convolutional block for processing to obtain preliminary abnormal reconstruction features. Finally, the preliminary reconstruction features are processed through a 1×1 convolutional layer to further adjust the number of channels and dimensions of the features, and the abnormal reconstruction features and normal reconstruction features are fused to obtain preliminary reconstruction features.
[0060] The obtained preliminary reconstruction features are input into the feature refinement module. By calculating the similarity of the original features as the weight, the original features and the preliminary reconstruction features are fused to obtain the final reconstruction features. Among them, the feature refinement module evaluates the pixel-level similarity between the original features and the preliminary reconstruction features by calculating the mean squared error MSE and cosine similarity. MSE is used to capture the absolute differences at the pixel level, and cosine similarity is used to capture the structural similarity between feature vectors. Specifically, the feature refinement module operates through the following steps:
[0061] First, calculate the original feature φ aligned (x) and the preliminary reconstruction feature The similarity Sim between them is given by the following formula:
[0062]
[0063] where C is the number of channels of the feature map, and ∈ is a very small constant used to avoid division by zero.
[0064]
[0065] where is the indicator function, which is assigned a value of 1 when the condition inside the parentheses is satisfied. mse(·, ·) and cosim(·, ·) represent the calculation of the mean squared error per pixel and the cosine similarity respectively. The above similarity Sim is then applied to the reconstructed features to obtain the final reconstructed features.
[0066]
[0067] where ⊙ represents element-wise multiplication.
[0068] The final reconstructed features are compared with the original features to obtain a pixel-level anomaly localization map, and then an image-level anomaly classification result is obtained based on the pixel-level anomaly localization map.
[0069] Among them, the difference between the final reconstructed feature and the original feature is used as the anomaly score to detect and locate anomalies. The anomaly score is calculated by using the Euclidean distance between the two features. The specific implementation steps are as follows: First, calculate the difference between the reconstructed image feature and the original image feature of the input. This difference is defined as the reconstructed image feature minus the original image feature of the input. Then, square the difference to obtain a squared difference matrix. Next, sum the squared difference matrix in the feature dimension, i.e., the channel dimension, and take the square root to obtain the Euclidean distance matrix at each pixel position. This matrix reflects the degree of difference between the reconstructed image and the input original image in the feature space. To ensure that the spatial resolution of the difference matrix is the same as that of the input image, further perform an upsampling operation on the Euclidean distance matrix. The upsampling operation can adopt bilinear interpolation, nearest neighbor interpolation, or other appropriate interpolation methods to ensure that the upsampled result has the same spatial resolution as the input image. Finally, the Euclidean distance matrix is the anomaly score matrix. By performing threshold processing on this anomaly score matrix, the anomaly region in the input image can be effectively detected and located. The threshold can be preset according to the specific application scenario or determined through experiments. In order to obtain the image-level localization result, in this embodiment, it is achieved by comparing the maximum value in the pixel-level anomaly score matrix with the preset threshold. The specific steps are as follows: Calculate the maximum value in the pixel-level anomaly score matrix, and compare this maximum value with the preset threshold. If the maximum value is greater than or equal to the threshold, it is determined that there is an anomaly in the image; otherwise, it is determined that there is no anomaly in the image.
[0070] The working mode of this embodiment is similar to that of a common object detection engine. When the user inputs the appearance image of an industrial product into the system, the system will convert the image into original features through a preprocessing module, and then input these original features into a processing module to obtain the final reconstructed features. The system calculates the pixel-level anomaly score by comparing the differences between the original features and the final reconstructed features of the product image, and calculates the image-level anomaly classification result based on the maximum value of this sample pixel-level anomaly score. Finally, the corresponding query result is returned to the user. For example, when the user inputs a picture of a screw, the system will return the result: whether there are production defects or anomalies in the screw in this image. If there are, the specific location of the defect will also be given, that is, when the user inputs a picture, the system will inform the user whether there are industrial defects in the item in the picture and the specific location of the defect. The model for detecting anomalies simultaneously considers the ability of the normal grid relying on normal samples and the abnormal grid relying on synthetic anomalies in discovering abnormal patterns. By refining the normal feature boundaries of the normal grid at the feature level through the abnormal grid, the reconstruction performance of the normal grid is significantly enhanced. The present invention adopts a dual-path grid reconstruction and feature refinement technology to obtain a better feature reconstruction result, thereby significantly improving the distinguishability between normal and abnormal.
[0071] In this embodiment, step S1 includes: processing the input RGB image to a unified size, then normalizing the image of the unified size so that it follows the standard normal distribution, and then taking the intermediate layer features through a pre-trained feature extractor. After aligning the multiple taken intermediate layer features to the layer feature with the largest selected resolution, they are spliced together to obtain the original features.
[0072] In addition, during the abnormal grid training, the FBP abnormal synthesis technology is additionally used on the original features to synthesize abnormal blocks with controllable size, shape, intensity, and position at the feature level, and the corresponding annotation mask image is provided. Specifically, the FBP module operates through the following steps:
[0073] Define the block size as B, the block intensity as I, and the block center coordinates as (x c , y c ). First, generate the initialization mask M:
[0074] M = zeroS(2B + 1, 2B + 1)
[0075] where M is a (2B + 1)×(2B + 1) matrix with an initial value of 0. Then, generate the random walk mask:
[0076] 1) Select the initial position (x 0 , y 0 ), where x 0 = B and y 0 = B;
[0077] 2) Randomly select the number of walk steps N, whose value is randomly selected from B to 2B;
[0078] 3) Conduct the random walk process. Each time, randomly select the values of Δx and Δy, whose value range is {-1, 0, 1}, and update the position (x k+1 , y k+1 ), where x k+1 = x k + Δx and y k+1 = y k + Δy, and mark the corresponding position in the mask M as 1.
[0079] Finally, generate the abnormal block and paste it onto the feature map: 1) Initialize the block paste tensor P, whose size is 1×1×H×W with an initial value of 0; 2) According to the marking of the mask M, paste the block with intensity I to the corresponding position of the feature map; 3) Perform Gaussian blur processing on the paste tensor P to obtain P blurred ; 4) Paste the tensor P blurred after Gaussian blur processing onto the feature map.
[0080] In this embodiment, step S2 includes: mapping the original features passing through the pre-trained feature extraction module to the global coordinate space through the global coordinate mapping module and to the local coordinate space through the local coordinate mapping module; in the global coordinate mapping module, for the input original features, first map the input features to a new feature space through a fully connected layer (Linear), then apply the Tanh activation function to introduce non-linearity, then, further adjust the features through another fully connected layer, and finally, apply the Tanh activation function again to ensure the non-linearity of the output features; in the local coordinate mapping module, for the received original features, first perform an inter-channel linear transformation of the input features through a 1×1 convolutional layer, then apply the Tanh activation function to introduce non-linearity, then further adjust the features through another 1×1 convolutional layer, and finally apply the Tanh activation function again to ensure the non-linearity of the output features. Through the combination of the above global and local coordinate mapping modules, the system can effectively capture the local details and overall structure of the image, thereby achieving more accurate feature mapping.
[0081] In this embodiment, step S3 includes: respectively inputting the global coordinates and local coordinates into the dual-path grid reconstruction module: using the grid_sample function provided by torch to perform bilinear interpolation on the feature map according to the input global and local coordinates to obtain the sampled feature map; for the grid sampling of normal features, input the global coordinates and local coordinates into the grid_sample function to respectively obtain the global and local normal feature sampling results; for the grid sampling of abnormal features, also input the global coordinates and local coordinates into the grid_sample function to respectively obtain the global and local abnormal feature sampling results; perform normalization processing on the sampling results to eliminate the dimensional difference of the feature values, then, input the normalized sampling results into a convolutional block for processing: the predefined convolutional block processes the sampled features, and the convolutional block is composed of multiple convolutional layers, batch normalization layers, and activation functions (such as ReLU); input the sampled normal features into the normal convolutional block for processing to obtain the preliminary normal reconstruction features; input the sampled abnormal features into the abnormal convolutional block for processing to obtain the preliminary abnormal reconstruction features, and finally, adjust the number of channels and dimensions of the features through a 1×1 convolutional layer and then use the abnormal reconstruction features to constrain the normal reconstruction features to obtain the preliminary reconstruction features.
[0082] In this embodiment, in step S4, the obtained preliminary reconstruction features are passed through the feature refinement module to obtain the final reconstruction features;
[0083] Among them, the feature refinement module evaluates the pixel-level similarity between the original feature and the preliminary reconstructed feature by calculating the MSE and cosine similarity. The MSE is used to capture the absolute difference at the pixel level, and the cosine similarity is used to capture the structural similarity between feature vectors. Specifically, the feature refinement module operates through the following steps:
[0084] First, calculate the similarity Sim between the original feature φ aligned (x) and the preliminary reconstructed feature The formula is as follows:
[0085]
[0086] where C is the number of channels of the feature map, and ∈ is a very small constant used to avoid division by zero.
[0087]
[0088] where, is the indicator function, which is assigned 1 when the condition in the parentheses is satisfied. mse(·, ·) and cosim(·, ·) respectively represent the calculation of the mean squared error per pixel and the cosine similarity.
[0089] The above similarity Sim is then applied to the reconstructed feature to obtain the final reconstructed feature
[0090]
[0091] where ⊙ represents element-wise multiplication.
[0092] In this embodiment, in step S5, the final reconstructed feature is compared with the original feature to obtain a pixel-level anomaly localization map, and then an image-level anomaly classification result is obtained according to the pixel-level anomaly localization map. Among them, what is compared is the feature space difference between two images, and the feature space difference calculates the anomaly score by using the Euclidean distance, which is achieved through the following steps:
[0093] First, calculate the difference between the reconstructed image feature and the input original image feature φ aligned (x), which is defined as:
[0094]
[0095] Then, square the difference to obtain the squared difference matrix:
[0096] square_diff = diff 2
[0097] Next, sum the squared difference matrix in the feature dimension (i.e., the channel dimension) and take the square root to obtain the Euclidean distance matrix at each pixel position, defined as:
[0098]
[0099] The Euclidean distance matrix pred reflects the degree of difference between the reconstructed image and the input original image in the feature space. To ensure that the spatial resolution of the difference matrix is consistent with the input image, further perform an upsampling operation on the Euclidean distance matrix:
[0100] pred pixel = upsample(pred)
[0101] The upsampling operation can use bilinear interpolation, nearest neighbor interpolation, or other appropriate interpolation methods to ensure that the upsampled matrix has the same spatial resolution as the input image. Finally, the Euclidean distance matrix pred pixel is the anomaly score matrix. By performing threshold processing on this anomaly score matrix, the anomaly regions in the input image can be effectively detected and located. The threshold can be preset according to the specific application scenario or determined through experiments.
[0102] To obtain the image-level localization result, in this embodiment, it is achieved by comparing the maximum value in the pixel-level anomaly score matrix pred pixel with a preset threshold. The specific steps are as follows:
[0103] img_score = max(pred pixel )
[0104] Compare the image-level anomaly score img_score with the preset threshold threshold. If img_score is greater than or equal to the threshold, it is determined that there is an anomaly in the image; otherwise, it is determined that there is no anomaly in the image.
[0105]
[0106] where pred img is a binary variable. When its value is 1, it indicates that there is an anomaly in the image, and when its value is 0, it indicates that there is no anomaly in the image.
[0107] The advantages of the present invention are as follows: By using spatial domain information and frequency domain information to jointly guide the training of the model, the dual-domain information complements each other and jointly discovers abnormal patterns existing in the image. The image preprocessing module highlights the high-frequency information that is easily lost in image reconstruction, greatly avoiding the model seeing abnormalities during inference. The dual-domain feature selection module considers both the spatial domain and the frequency domain when modeling the normal image distribution. The dual-domain reconstruction not only captures the local details and global structure of the image but also effectively utilizes the frequency domain information, thus achieving higher-quality image reconstruction, which directly determines the detection performance of the model.
[0108] It should be noted that this embodiment is not a simple calculation method but can be applied to industrial production and assist in improving the production line of industrial products. For example, in practical applications, the method of this embodiment can be applied to a system as shown in the appendix Figure 1 and includes:
[0109] A preprocessing module, which is used to unify the size of the input pictures and perform standardization processing. Through a pre-trained feature extraction network, the output of the middle layer is selected as the original feature after alignment. During the abnormal grid training, the FBP abnormal synthesis module is used at the feature level of normal samples. By defining the size, intensity, and position of the abnormal block and using the random walk algorithm to generate a mask, the abnormal block is synthesized onto the feature map of the normal sample to obtain the original feature containing abnormalities.
[0110] A processing module maps the above original features to the local and global coordinate spaces through a coordinate mapping module. The coordinate mapping module includes two parts, local and global, which respectively perform local and global coordinate mapping of the features. Then, these coordinates are input into a dual-path grid reconstruction module, and through the method of grid sampling, global and local feature samplings are obtained. The global and local feature samplings are fused by splicing. The normal and abnormal grids respectively contain the corresponding local and global grids. Immediately afterwards, the spliced feature sampling is input into a CNN module to obtain preliminary reconstructed features. Subsequently, the preliminary reconstructed features and the original features are input into a feature refinement module. The pixel-level similarity between the original features and the preliminary reconstructed features is evaluated by calculating the mean square error (MSE) and cosine similarity, and the similarity result is used as a weight to fuse the original features and the preliminary reconstructed features to obtain the final reconstructed features.
[0111] A query feedback module is used to receive the query image sent by the terminal device, map the query image to the feature space and then convert it into the final reconstructed feature generated by the model. The difference between the final reconstructed feature and the original feature is used as an abnormal score to detect and locate abnormalities. The abnormal score is the abnormal score of each pixel, and the image-level abnormal score is determined by selecting the maximum value of all pixel-level abnormal scores, and the query result is fed back to the terminal device.
[0112] Specifically, this embodiment is applicable to the defect detection and localization of industrial images. That is, the product image is converted into a reconstructed image through a pre-training module and the trained model. The model further compares the gradients and color differences between the input product image and the output reconstructed image to detect and locate anomalies, and then returns the corresponding detection and localization results. For example, the working mode is similar to the commonly used image recognition at present. When industrial parts are produced on the assembly line, the machine on the assembly line takes pictures of this part, obtains the image and inputs it into the model. Then the model reconstructs the image and uses the gradient and color anomaly evaluation function to evaluate whether the image is abnormal. Finally, the result is returned. This result can be used to determine whether the assembly line continues to produce. Because when defective industrial products are produced, it indicates that there is a problem with the assembly line and it may be necessary to stop production and repair the machine, otherwise it will cause waste of resources. It can also be used to guide the subsequent repair of the faulty machine. The location where the defect appears may also contain information about the fault.
[0113] This embodiment is applicable to the defect detection and localization of industrial images. Specifically, the product image obtains the original features through a preprocessing module and is reconstructed into the final reconstructed features through a processing module. The model further compares the differences between the original features and the final reconstructed features to detect and locate anomalies and returns the corresponding results. For example, on the production assembly line, the machine takes pictures of industrial parts, obtains the image and inputs it into the model. The model then reconstructs the image at the feature level and uses the MSE-based anomaly evaluation function to determine whether there is a problem with the image. Finally, the system returns the detection result to decide whether to continue production. If the product is defective, this indicates that there may be a fault in the production line and it may be necessary to stop production and perform maintenance to prevent waste of resources. At the same time, the defect information can also provide guidance for the subsequent repair of the machine. The location and form of the defect can reveal the fault information.
[0114] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0115] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for fine-grained industrial image anomaly detection and localization based on dual-path grid reconstruction, characterized in that: The steps include: Step S1, pre-processing the input sample image to be processed to unify the size and standardize it, then aligning the intermediate layer output through the pre-trained feature extraction network to obtain the original features, and then further synthesizing the abnormal features on the original features through the FBP module through the abnormal grid training stage, wherein the sample image to be processed includes an industrial product image without label information; Step S2, inputting the original features in step S1 into the coordinate mapping module, and obtaining the local and global coordinates through the convolution layer mapping trained by the coordinate mapping module; Step S3, using the local and global coordinates obtained in step S2, sampling is performed from the normal and abnormal grids, and the sampling results are fused and input into the CNN module to obtain preliminary reconstruction features, the normal and abnormal grids contain corresponding local and global grids respectively, and the local and global coordinates are used to sample in the local and global grids to obtain preliminary reconstruction features that take into account both local details and global information of the image; Step S4, inputting the original features in step S1 and the preliminary reconstructed features in step S3 into a feature refinement module to obtain final reconstructed features, wherein the feature refinement module combines mean square error and cosine similarity to evaluate the pixel-level similarity between the original features and the preliminary reconstructed features, and guides the weighted fusion of the original features and the preliminary reconstructed features based on this similarity to obtain the final reconstructed features; Step S5, using the difference between the final reconstructed features of step S4 and the original features of step S1 as an anomaly score to detect and locate anomalies; Step S6: after receiving the inference request from the terminal device, the image to be tested is mapped to the feature space, and then reconstructed at the feature level to obtain reconstructed features, and the reconstructed features are compared with the original features to obtain the inference results.
2. According to claim 1, a method for fine-grained industrial image anomaly detection and positioning based on dual-path grid reconstruction is characterized in that: The feature extraction network pre-trained in step S1 includes ResNet, VGG, and Inception deep convolutional neural networks.
3. The method for fine-grained industrial image anomaly detection and location based on dual-path grid reconstruction according to claim 2 is characterized in that: The FBP anomaly synthesis module of step S1 synthesizes anomalies with controllable size, shape, intensity and position for normal samples at the feature level only during the abnormal grid training, and provides corresponding annotated mask images. Specifically, the FBP anomaly synthesis module operates through the following steps: Define the block size as B, the block strength as I, and the block center coordinates (x c ,y c ), first, generate the initialization mask M: M = zeros(2B+1, 2B+1) Where M is a (2B+1)×(2B+1) matrix with an initial value of 0. Then, generate a random walk mask: 1) select the initial position (x0, y0), x0=B, y0=B; 2) randomly select the number of walk steps N, whose value is randomly selected from B to 2B; 3) perform a random walk process, randomly select the value of Δx and Δy each time, and its value range is {-1, 0, 1}, and update the position (x k+1 ,y k+1 ), where x k+1 =x k +Δx,y k+1 =y k +Δy, and mark the corresponding position in the mask M as 1; finally, generate the abnormal block and paste it to the feature map: 1) Initialize the block paste tensor P, whose size is 1×1×H×W and the initial value is 0; 2) According to the mark of the mask M, paste the block with intensity I to the corresponding position of the feature map; 3) Perform Gaussian blur processing on the pasted tensor P to obtain P blurred ; 4) The Gaussian blurred tensor P blurred Paste onto the feature map.
4. The method for fine-grained industrial image anomaly detection and location based on dual-path grid reconstruction according to claim 3 is characterized by: The coordinate mapping module in step S2 maps the original features to the corresponding coordinate space by learning the mapping relationship between the local and global coordinate spaces, wherein the local coordinate space is used to capture the local details of the image, and the global coordinate space is used to capture the overall structure of the image.
5. The method for fine-grained industrial image anomaly detection and location based on dual-path grid reconstruction according to claim 4 is characterized in that: The normal and abnormal grid creation process of step S3 is as follows: 1) creating an uninitialized grid tensor; 2) initializing the grid tensor using Xavier normal distribution; 3) converting the initialized tensor into Parameter type and managing it using ParameterList.
6. The method for fine-grained industrial image anomaly detection and location based on dual-path grid reconstruction according to claim 5 is characterized by: The normal and abnormal grid training process of step S3 is as follows: 1) first freeze the normal grid parameters, use normal samples to synthesize abnormalities through the FBP module, and train the abnormal grid; 2) freeze the abnormal grid parameters and train the normal grid; 3) the coordinate mapping module in S2 and the CNN module in S3 are both trainable modules and participate in the training throughout the process; 4) the mean square error is used as the reconstruction loss for training the normal grid, and the formula is as follows: Where C represents the number of channels of the feature map, H and W represent the height and width of the feature map respectively, φ(x) represents the original feature in S1, represents the final reconstructed features of S4; 5) The training abnormal grid also uses the above method to reconstruct the error, where φ(x) is replaced by the feature map after synthesizing the abnormality, and the truncated L1 loss is used as the contrast loss to refine the boundary between normal features and abnormal features. The formula is as follows: Where th represents the size of the buffer around the dividing line, D is a set of sample pair similarities constructed based on the mask, and d + The table represents the similarity of the positive sample pairs, d - Represents the similarity of the negative sample pair.
7. The method for fine-grained industrial image anomaly detection and location based on dual-path grid reconstruction according to claim 6 is characterized in that: The feature refinement module in step S4 evaluates the pixel-level similarity between the original feature and the preliminary reconstructed feature by calculating the mean square error (MSE) and cosine similarity, wherein MSE is used to capture the absolute difference at the pixel level, and cosine similarity is used to capture the directional similarity between feature vectors. Specifically, the feature refinement module operates through the following steps: First, calculate the original feature φ aligned (x) and preliminary reconstruction features The similarity between them is Sim, and the formula is as follows: where ∈ is an infinitely small constant, in, is an indicator function, and the value is 1 when the conditions in the brackets are met. mse(·,·) and cosim(·,·) represent the calculation of mean square error and cosine similarity, respectively. The above similarity Sim is then applied to the reconstructed features to obtain the final reconstructed features. Among them, ⊙ represents element-by-element multiplication.
8. A system for fine-grained industrial image anomaly detection and positioning method based on dual-path grid reconstruction according to any one of claims 1 to 7, characterized in that: The system includes: a preprocessing module, which is used to unify the size of the input image and standardize it, select the intermediate layer output after alignment as the original feature through the pre-trained feature extraction network, and use the FBP abnormal synthesis module during the abnormal grid training to define the size, strength and position of the abnormal block at the normal sample feature level, use the random walk algorithm to generate a mask, and synthesize the abnormal block onto the feature map of the normal sample to obtain the original feature; A processing module is used to obtain global coordinates and local coordinates of the original features through a coordinate mapping module; the global coordinates and local coordinates are input into a dual-path grid reconstruction module, the dual-path grid reconstruction module is composed of a normal grid and an abnormal grid, the normal grid and the abnormal grid respectively contain corresponding local and global grids, in each path, global and local feature sampling is obtained by grid sampling, wherein the global and local feature sampling are fused by splicing along the channel dimension, and then the features of the normal grid and the abnormal grid are weighted fused; the fused feature sampling is input into a CNN module to obtain preliminary reconstruction features; the preliminary reconstruction features and the original features are input into a feature refinement module, the pixel-level similarity between the original features and the preliminary reconstruction features is evaluated by calculating the mean square error (MSE) and cosine similarity, and the similarity result is used as a weight to fuse the original features and the preliminary reconstruction features to obtain the final reconstruction features; the anomaly score is calculated using the difference between the final reconstruction features and the original features to detect and locate the anomaly, wherein the anomaly score is the anomaly score of each pixel, and the image-level anomaly score is the maximum value of all pixel-level anomaly scores; The query feedback module is used to receive the query image sent by the terminal device, map the query image to the feature space and convert it into a reconstructed feature generated by the model, determine the image-level and pixel-level anomaly scores based on the difference between the reconstructed feature and the original feature, and feed back the query result to the terminal device.
Citation Information
Patent Citations
Foreign matter detection method based on generative adversarial network
CN114220043A
Remote sensing image fine-grained detection method and system based on region proposal enhancement
CN117726940A