Method and System for Industrial Defect Detection and Location Based on Frequency Domain Enhancement

By combining spatial domain and frequency domain information in industrial image abnormality detection, high-pass filters and improved U-Net networks are used for image reconstruction and abnormal detection, the problem of insufficient abnormal positioning accuracy in the prior art is solved, and higher detection accuracy and lower error detection rate are achieved.

CN119515780BActive Publication Date: 2025-07-04NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Patent Information

Application Number
CN202411452542.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2025-07-04
Estimated Expiration
2044-10-17

AI Technical Summary

Technical Problem

In the industrial image abnormality detection, the prior art fails to fully utilize the frequency domain information, resulting in insufficient abnormal positioning accuracy and high error detection rate.

Method used

Using a method based on frequency domain enhancement, high-frequency images are extracted through high-pass filters, and combined with the improved U-Net network of spatial domain and frequency domain feature selection modules, image reconstruction and abnormal detection are performed, and dual-domain feature fusion and multi-scale gradient similarity evaluation function are used to improve abnormal positioning accuracy.

Benefits of technology

While retaining the local details and global frequency characteristics of the image, the accuracy of abnormal detection is significantly improved and the error detection rate is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119515780B_ABST
    Figure CN119515780B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for industrial defect detection and localization based on frequency domain enhancement, which relates to the fields of computer vision and image processing, can discover abnormal patterns in industrial images from multiple perspectives, and also has a relatively fast inference speed. The method includes: processing an input image through a high-pass filter to obtain a high-frequency image; inputting the original image and the high-frequency image into an improved U-Net network integrated with a dual-domain feature selection module; performing image reconstruction simultaneously in the spatial domain and the frequency domain; calculating an anomaly score and a localization map based on the reconstruction results. By combining spatial domain and frequency domain analysis, the present invention realizes a comprehensive characterization of local pixel changes and global frequency patterns in the image, improving the accuracy of anomaly detection and localization. The method has achieved optimal performance on multiple real industrial scenario datasets. The present invention is applicable to anomaly detection and anomaly area localization of industrial product images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image anomaly detection, and in particular, to a method and system for industrial defect detection and localization based on frequency domain enhancement. Background Art

[0002] Anomaly detection of industrial images refers to using image analysis techniques to identify defects in products or components during the manufacturing process. In the manufacturing industry, maintaining the consistency of product quality is crucial, and anomaly detection plays a key role in this process. This technology can detect problems at an early stage, thereby reducing costs, improving efficiency, and ensuring the satisfaction of the end-users of the products.

[0003] In industrial production, since normal conditions dominate the data and abnormal conditions are relatively rare. Therefore, self-supervised image reconstruction methods are widely used in the field of anomaly detection, which use the difference between the original input image and the reconstructed image to detect and locate anomalies. However, there is usually a trade-off between the fidelity of normal reconstruction and the distinguishability of abnormal reconstruction. One way to alleviate the trade-off is to improve the quality of the reconstructed image by means such as adding skip connections or adversarial training, so as to reduce the impact of blurred reconstructed images on the detection results. However, in the process of learning image reconstruction, existing solutions usually only focus on spatial domain information and do not fully exploit the potential of frequency domain information in discovering abnormal patterns. Eventually, it leads to the inability to further improve the anomaly localization accuracy of the image while retaining the local details and global frequency characteristics of the image, thus unable to further reduce the false detection rate. Summary of the Invention

[0004] Embodiments of the present invention provide a method and system for industrial defect detection and localization based on frequency domain enhancement, which can further improve the anomaly localization accuracy of the image while retaining the local details and global frequency characteristics of the image, thereby further reducing the false detection rate.

[0005] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:

[0006] In a first aspect, the method provided by the embodiments of the present invention includes:

[0007] S1. After grayscale processing the image to be processed, a high-frequency image is obtained through a high-pass filter. Among them, the image to be processed includes: a normal industrial product image without label information. Specifically, the sample image to be processed is grayscale processed, and the brightness and texture information of the image are retained after grayscale processing, and a high-frequency image is obtained through a Butterworth high-pass filter. Then, Perlin noise and the abnormal texture source dataset DTD are used to synthesize abnormalities on the high-frequency image. Among them, the sample image to be processed includes: a normal industrial product sample image without label information. Note that the operation of synthesizing abnormalities is only carried out in the training stage, and no abnormalities are synthesized in the testing stage. It should be noted that in the training stage, the operation of synthesizing abnormalities is carried out. The training stage in this embodiment includes S1-S4, which refers to training the model. The model consists of an encoder and a decoder, so the encoder and decoder (the whole model) are trained in the training stage. After the training stage is completed, the model is trained. At this time, the model processes the inference image of the terminal device.

[0008] S2. Input the high-frequency image into the encoder for dual-domain feature selection; specifically, input the high-frequency image with synthesized abnormalities into the improved U-Net encoding network with a dual-domain feature selection module to obtain dual-domain features. Among them, the dual-domain feature selection module selects features in both the spatial domain and the frequency domain at the same time. Among them, after the obtained high-frequency image undergoes operations such as convolution activation, it is respectively input into the spatial domain selection branch and the frequency domain selection branch of the dual-domain feature selection module. In order to maintain the consistency of feature selection, the same activation function is selected for both dual-domain branches. This method of dual-domain feature selection allows the network to simultaneously focus on the local details and global frequency features of the image. Feature selection in the spatial domain is used to capture local structural information such as the shape and texture of objects, while feature selection in the frequency domain is used to capture the overall frequency distribution of the image. After the feature selection operations of the two branches, dual-domain features are obtained.

[0009] S3. Input the dual-domain features obtained in S2 into an encoder for dual-domain feature fusion to obtain a fused image representation. Among them, the dual-domain features include: spatial-domain selected features and frequency-domain selected features. The spatial-domain selected features include the local structure of the image, and the frequency-domain selected features include the global frequency distribution of the image. The fused image representation includes spatial-domain information and frequency-domain information. Specifically, input the dual-domain features into a dual-domain feature fusion module to obtain an image representation containing local details and global frequency information of the image. Among them, the dual-domain feature fusion module performs concatenation and 1×1 convolution on the dual-domain features. The dual-domain feature fusion module accepts dual-domain features, which include: spatial-domain selected features and frequency-domain selected features, which respectively capture the local structure and global frequency distribution of the image. At the beginning, the module will perform preliminary processing on these features, such as normalization or dimension adjustment, to ensure that features in different domains can be effectively fused. Next, the model uses the feature concatenation operation and 1×1 convolution to fuse the dual-domain features. Through the above-mentioned dual-domain feature selection and dual-domain feature fusion operations at multiple scales, an image representation integrating spatial-domain and frequency-domain information is obtained. The S2-S3 in this embodiment implements an unsupervised method, which simultaneously uses spatial-domain information and frequency-domain information to discover abnormal patterns existing in the image, and this embodiment provides a simple and efficient way to fuse dual-domain features. In the prior art, either a supervised method is used to train the model, or only spatial-domain information is mined, without making full use of frequency-domain information, or the model is relatively complex and difficult to deploy. What is implemented in this embodiment is to train the model in an unsupervised manner, without manual annotation, while performing image analysis in the spatial domain and frequency domain, and the model is simple, which is conducive to actual deployment.

[0010] S4. Use the fused image representation for image reconstruction. Specifically, constrain the image reconstruction process through spatial-domain and frequency-domain information. Among them, the decoder simply selects the decoder of the UNet network, which includes a series of convolutional upsampling operations, and finally decodes the image into a reconstructed image, and the reconstructed image does not contain any abnormal regions. The spatial-domain and frequency-domain information includes: l2 loss in the spatial domain, SSIM loss in the spatial domain, and FFL loss in the frequency domain.

[0011] S5. Use the image reconstructed in S4 and the image to be processed for anomaly detection. The anomaly detection result includes the anomaly score at the pixel level. Specifically, the anomaly score calculation method includes the anomaly evaluation functions of gradient similarity and color similarity. Among them, the gradient similarity anomaly evaluation function applies the multi-scale gradient magnitude similarity (MSGMS). The anomaly evaluation function of color similarity converts the image into the CIELAB color space that is more in line with human perception of colors for comparison. The final anomaly score is composed of the gradient similarity and color similarity.

[0012] S6. After receiving the inference image sent by the terminal device, perform the above steps S1 to S5 on the inference image and obtain the anomaly detection result corresponding to the inference image. Then, determine whether the inference image is abnormal according to the anomaly detection result of the inference image. In practical applications, after receiving the query image sent by the terminal device, gray-scale and high-pass filter the query image, then convert it into a reconstructed image through the model. After that, determine the image-level and pixel-level anomaly scores according to the difference between the reconstructed image and the original image, and then determine whether the image is abnormal and the possible abnormal area according to the threshold, and feedback the query result to the terminal device.

[0013] In this embodiment, in S1, obtaining the high-frequency image through the high-pass filter includes: after performing a fast Fourier transform on the image to be processed, using a Butterworth high-pass filter to obtain high-frequency information, and then performing an inverse fast Fourier transform on the high-frequency information to obtain the high-frequency image. Specifically, the Butterworth high-pass filter attenuates the low-frequency components and passes the high-frequency components by setting the spectral center point and cut-off frequency, etc. Specifically, the sample image to be processed is first gray-scaled, then the frequency-domain information is obtained through a fast Fourier transform, the frequency-domain information is high-pass filtered using a Butterworth high-pass filter, and finally the high-frequency image is obtained through an inverse fast Fourier transform. Adjusting the parameters of the Butterworth high-pass filter can obtain a more suitable high-frequency image.

[0014] In this embodiment, in S1, it further includes: if in the training stage, the image to be processed is a sample image, and anomalies are synthesized on the high-frequency image of the sample image. The synthesized anomalies include: generating an anomaly mask using Perlin noise, and then generating an abnormal area in the high-frequency image using the anomaly texture source dataset DTD and the mask generated by Perlin noise.

[0015] In this embodiment, in S2, it includes: in the spatial domain selection branch, spatial domain feature selection is performed through convolution 3×3 convolution - BatchNorm - ReLU - 3×3 convolution - BatchNorm; in the frequency domain selection branch, the convolved image is converted into frequency domain information through fast Fourier transform, then frequency domain feature selection is performed through 1×1 convolution - ReLU - 1×1 convolution, and finally the frequency domain information is converted back into spatial domain information through inverse fast Fourier transform. Specifically, the reconstruction network includes: an improved U-Net encoder with a dual-domain feature selection module and a UNet decoder; the dual-domain feature selection module includes a spatial domain feature selection branch and a frequency domain feature selection branch; during the feature selection process, the feature maps obtained after convolution are respectively input into the spatial domain feature selection branch and the frequency domain feature selection branch. Spatial domain features are obtained through spatial domain feature selection, and frequency domain features are obtained through frequency domain feature selection. The spatial domain features and the frequency domain features are called dual-domain features; in the spatial domain feature selection branch, 3×3 convolution - BatchNorm - ReLU - 3×3 convolution - BatchNorm is used for spatial domain feature selection. Through the superposition of two layers of convolution and the intermediate activation function, the network can effectively extract useful features from the input data and maintain the stability of training through batch normalization; in the frequency domain feature selection branch, the convolved image samples are converted from the spatial domain to the frequency domain through fast Fourier transform, then 1×1 convolution - ReLU - 1×1 convolution is used to achieve cross-channel information aggregation and adjust and optimize the frequency domain features, and finally through inverse fast Fourier transform, the processed frequency domain image is converted back to the spatial domain.

[0016] In this embodiment, in S3, it includes: in the dual-domain feature fusion module, 1×1 convolution is performed after splicing the dual-domain features, where dual-domain feature selection and dual-domain feature fusion are performed in each encoding layer of the encoder; in each encoding layer, the features output by the dual-domain feature fusion module are spliced with the features input to the dual-domain feature selection module, and then downsampling is performed to obtain the output of one encoding layer, and the output of this encoding layer is used as the input of the next encoding layer, and the features obtained from the last encoding layer are used as the image representation. Specifically, in S3, it includes: inputting the obtained dual-domain features into the dual-domain feature fusion module; where the dual-domain feature fusion module splices and performs 1×1 convolution on the dual-domain features. Through the above-mentioned dual-domain feature selection and dual-domain feature fusion operations at multiple scales, an image representation that combines spatial domain and frequency domain information is obtained; specifically, dual-domain feature selection and dual-domain feature fusion are performed in each encoding layer. The features output by the dual-domain feature fusion module are spliced with the features input to the dual-domain feature selection module, and then downsampling is performed to obtain the output of each encoding layer, and this output is used as the input of the next encoding layer, and the features obtained from the last encoding layer are called the image representation.

[0017] In this embodiment, in S4, the obtained image representation is decoded by a decoder into a reconstructed image, and the image reconstruction process is constrained by spatial domain and frequency domain information; wherein, the decoder is composed of convolutional upsampling operations, and the decoder layers correspond one-to-one with the encoder layers, and the output of each layer of the encoder will be concatenated with the decoder layer through skip connections. The spatial domain information for constraining the image reconstruction process includes: the l2 loss in the spatial domain and the structural similarity (SSIM) loss in the spatial domain. Among them, the l2 loss in the spatial domain is: l2 represents that the loss emphasizes pixel-level accuracy, and I and I r are the quantized representations of the original image and the reconstructed image respectively; the SSIM loss is: μ x and μ y are the average luminances of I and I r respectively, and are the variances of I and I r respectively, σ xy is the covariance of two I and I r respectively, C1 and C2 are small constants added to avoid the denominator being zero, H and W are the height and width of the image respectively, N p is equal to the number of pixels in the image, i and j respectively represent the row and column indices of the pixel in the upper left corner of the current image window, and SSIM(I, I r ) i,j represents the structural similarity index, which is designed to evaluate the similarity of two images in terms of structure, luminance, and contrast. Specifically, the SSIM loss is the structural similarity index loss, which is designed to evaluate the similarity of two images in terms of structure, luminance, and contrast. To achieve this, instead of calculating a single SSIM value for the entire image, SSIM is calculated separately on multiple small windows of the image. These windows are usually of a fixed size, such as square windows of 8x8 or 11x11 pixels. Here, i and j refer to the row and column indices of the pixel in the upper left corner of the current window. And SSIM(I, I_r)_{i,j} represents the structural similarity index between the original image I and the reconstructed image I_r; the frequency domain information for constraining the image reconstruction process includes the frequency domain reconstruction focus loss FFL between the reconstructed image and the original input image, wherein, M represents the height of the spectrum, N represents the width of the spectrum, F r (u, v) is the frequency value at the spectrum coordinate (u, v) of the real image, and the corresponding F f (u, v) represents the corresponding value of the spectrum of the reconstructed image, w(u, v) is an element of the spectrum weight matrix, representing the weight assigned to the frequency component at the position (u, v), and w(u, v) = |F r(u, v)-F f (u, v)| α , where α is the scaling factor of flexibility, and u and v respectively represent the indices of the horizontal frequency and the vertical frequency in the frequency domain. Specifically, they represent the indices of the horizontal frequency and the vertical frequency in the frequency space after the Fourier transform of the image.

[0018] Furthermore, a reconstruction model for image reconstruction is established, and the reconstruction model is optimized by the Adam algorithm. During the optimization process, the learning rate is adjusted by the MultiStepLR learning rate scheduler, and the MultiStepLR learning rate scheduler reduces the learning rate according to preset milestones. The optimization objective is: minL rec , L rec represents the loss information for constraining the image reconstruction process, and L rec = l2(I, I r ) + L SSIM (I, I r ) + L FFL .

[0019] In this embodiment, in S5, the process of anomaly detection includes: where MSGMS(x, y) represents the multi-scale gradient magnitude similarity function, which is used to calculate the multi-scale gradient magnitude similarity between two images x and y. S represents the number of scales to be considered, and GMS s (x, y) is the gradient magnitude similarity at scale i,

[0020] The anomaly score A of the gradient similarity anomaly assessment gradient = 1 M×N - MSGMS(I, I r ), and the anomaly score of the color similarity anomaly assessment L represents brightness, and a and b represent colors; the anomaly score A of the pixels in the image score = (kA color + A gradient ) * f mean , f mean is the average smoothing filter, and k is a scaling factor used to make A color and A gradient of the same order of magnitude.

[0021] In the second aspect, the system provided by the embodiment of the present invention includes:

[0022] A preprocessing module for grayscale processing the image to be processed and obtaining a high-frequency image through a high-pass filter, where the image to be processed includes: a normal industrial product image without label information;

[0023] A processing module for inputting a high-frequency image into an encoder for dual-domain feature selection. In the encoder, there is a dual-domain feature selection module and a dual-domain feature fusion module. The dual-domain feature selection module includes a spatial-domain selection branch and a frequency-domain selection branch. Input the obtained dual-domain features into the encoder for dual-domain feature fusion to obtain a fused image representation. The dual-domain features include spatial-domain selection features and frequency-domain selection features. The spatial-domain selection features include the local structure of the image, and the frequency-domain selection features include the global frequency distribution of the image. The fused image representation includes spatial-domain information and frequency-domain information. Use the fused image representation for image reconstruction. Use the reconstructed image and the image to be processed for anomaly detection. The anomaly detection result includes a pixel-level anomaly score.

[0024] A query feedback module for, after receiving an inference image sent by a terminal device, obtaining the corresponding anomaly detection result for the inference image, then determining whether the inference image is abnormal according to the anomaly detection result of the inference image, and feeding back a query result to the terminal device.

[0025] The method and system for industrial defect detection and localization based on frequency-domain enhancement provided by the embodiments of the present invention input a sample image to be processed into a preprocessing module for grayscale processing and Butterworth high-pass filtering to obtain a high-frequency image. Obtain a high-frequency image containing anomalies through an anomaly texture synthesis module for the high-frequency image. Input the high-frequency image containing anomalies into an improved U-Net encoding network with a dual-domain feature selection module to obtain dual-domain features. Obtain an image representation that combines spatial-domain and frequency-domain information through a dual-domain feature fusion module for the dual-domain features. Decode the image representation into a reconstructed image through a decoder, and use the difference between the reconstructed image and the original image as an anomaly score to detect and localize anomalies. The present invention is applicable to the field of industrial image anomaly detection and localization, and can accurately locate image anomalies while retaining the local details and global frequency features of the image, reducing the false detection rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 It is a schematic diagram of a possible implementation manner of an industrial defect detection and localization model provided by the embodiments of the present invention.

[0028] Figure 2 Schematic diagram of a possible implementation manner of the preprocessing module provided by an embodiment of the present invention.

[0029] Figure 3 Schematic diagram of a possible implementation manner of the dual-domain feature selection module provided by an embodiment of the present invention.

[0030] Figure 4 Schematic diagram of the method flow provided by an embodiment of the present invention.

[0031] Figure 5 Schematic diagram of the system architecture provided by an embodiment of the present invention. Detailed implementation manners

[0032] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the drawings and specific implementation manners. The implementation manners of the present invention will be described in detail below. Examples of the implementation manners are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements having the same or similar functions from beginning to end. The implementation manners described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention. Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any and all combinations of one or more of the associated listed items. Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless defined as herein.

[0033] An embodiment of the present invention provides an industrial defect detection and localization method using spatial domain and frequency domain analysis. Specifically, it is an improvement of a solution belonging to the anomaly detection technology and is a variant of the autoencoder network. The main design idea is as follows: By combining spatial domain feature selection and frequency domain feature selection, normal and abnormal patterns are considered from a dual-domain perspective, and the features selected in the dual domain are fused to generate an image representation that contains both local details and global frequency information of the image. During the pixel reconstruction process, both spatial domain loss and frequency domain loss are used to constrain the reconstruction process to ensure that normal regions and abnormal regions can be correctly reconstructed, which is beneficial for subsequent anomaly detection and localization. Specifically, as Figure 1 shown, this method uses dual-domain feature selection and fusion to obtain a better image representation and uses dual-domain loss to guide image reconstruction, more comprehensively identifying abnormal patterns and reconstructing high-quality normal images. The main method flow is as Figure 1 、 4 shown, including:

[0034] Input the sample image to be processed into a preprocessing module for preprocessing, including: grayscale the image to retain the brightness and texture information of the image, and obtain a high-frequency image through a high-pass filter. Then, use Perlin noise and the abnormal texture source dataset DTD to synthesize anomalies on the high-frequency image.

[0035] The preprocessing module mentioned in this embodiment can be implemented as a code program in practical applications. As shown in the appendix Figure 2 shown, we first grayscale the input RGB image, then perform a fast Fourier transform on the grayscale image to convert it from the spatial domain to the frequency domain, and then obtain a high-frequency image through a Butterworth high-pass filter. A more suitable high-frequency image can be obtained by setting the parameters of the Butterworth high-pass filter. After high-pass filtering, perform an inverse fast Fourier transform to convert the image from the frequency domain to the spatial domain. Finally, synthesize abnormal regions on the high-frequency image, which includes: generating random mask regions using Perlin noise and synthesizing anomalies in these mask regions using the abnormal texture source dataset DTD.

[0036] Input the obtained high-frequency image containing synthesized anomalies into the dual-domain feature selection module after convolution.

[0037] Among them, the dual-domain feature selection module refers to a module that performs feature selection simultaneously in the spatial domain and the frequency domain, including a spatial domain branch and a frequency domain branch. The convolved image is input into these two branches for processing respectively. Specifically, it can be seen as Figure 3Among them, the spatial domain branch is derived from the encoding layer Block of a neural network ResNet specifically used to extract feature information in images. The frequency domain branch is similar to the spatial domain branch. To maintain the consistency of dual-domain feature selection, the frequency domain branch adds Fourier transform and inverse Fourier transform, and replaces all 3×3 convolutions in the spatial domain branch with 1×1 convolutions. After dual-domain feature selection, dual-domain features are obtained.

[0038] The obtained dual-domain features are input into a dual-domain feature fusion module to obtain an image representation containing local details and global frequency information of the image.

[0039] Among them, the dual-domain feature fusion module refers to a module that splices and fuses the obtained spatial domain features and frequency domain features. The dual-domain feature fusion module directly splices the dual-domain features and then performs 1×1 convolution to fuse the dual-domain features. The fused dual-domain features are simply spliced with the input of the dual-domain feature selection module to obtain the output of the corresponding encoding layer. It should be noted that dual-domain feature selection and dual-domain feature fusion operations are performed for each encoding layer. Specifically, the feature map output from the previous encoding layer serves as the input of the dual-domain feature selection module. After being processed by the dual-domain feature selection module, dual-domain features are obtained. Then, the dual-domain features are fused through the dual-domain feature fusion module. Finally, the fused dual-domain features are simply spliced with the feature map output from one encoding layer and then pooled to obtain the output of the corresponding encoding layer. The output of the last encoding layer serves as the image representation containing local details and global frequency information of the image. The encoder has 5 encoding layers, indicating that 5 times of dual-domain feature selection and dual-domain feature fusion operations are performed during image encoding.

[0040] The obtained image representation containing local details and global frequency information of the image is decoded into a reconstructed image through a decoder.

[0041] Among them, the decoder uses the decoder of the popular image segmentation network UNet in deep learning, which consists of convolutional operations and upsampling operations. The decoder decodes the image representation into a reconstructed image, and skip connections are applied during the encoding and decoding process. Specifically, the input of each decoding layer of the decoder includes: the output of the previous decoding layer and the output of the corresponding encoder layer. Specifically, the first decoding layer receives the output of the last encoding layer as input, the second decoding layer receives the output of the penultimate encoding layer and the output of the first decoding layer as input, and so on. During the reconstruction process where the sample image is encoded by the encoder into an image representation and the image representation is decoded by the decoder into a reconstructed image, spatial domain and frequency domain information are used to constrain the reconstruction process. The spatial domain information that constrains the image reconstruction process is the spatial domain reconstruction loss between the reconstructed image and the original input image, including the l2 loss in the spatial domain and the structural similarity (SSIM) loss in the spatial domain. The frequency domain information that constrains the image reconstruction process is the frequency domain reconstruction loss between the reconstructed image and the original input image.

[0042] The reconstructed image and the original input image are compared for gradient and color differences through an anomaly evaluation function of gradient similarity and color similarity to calculate the anomaly score.

[0043] Among them, the gradient similarity applies the multi-scale gradient magnitude similarity (MSGMS). MSGMS utilizes the gradient information of the image at different scales and can provide a more comprehensive description of image features than a single-scale gradient. The anomaly evaluation function of color similarity converts the image into the CIELAB color space that is more in line with human perception of color for comparison. The CIELAB space consists of three components: L represents brightness, and a and b represent colors. The anomaly evaluation function of color similarity ignores L to reduce the influence of illumination.

[0044] The working mode of this embodiment is similar to that of a common image recognition engine. When the user inputs the appearance image of an industrial product into the system, the system will convert the model into a high-frequency image through a preprocessing module, and then the high-frequency image is input into the reconstruction model to obtain a reconstructed image. The system will calculate the abnormal score to determine the abnormal area by comparing the difference between the input product image and the reconstructed image, and then return the corresponding query result to the user. For example, when the user inputs a picture of a nail, the system will return the result to the user: whether there are production defects (abnormalities) in the nails in this image, and if so, the location of the defects will also be given. That is, when the user inputs a picture, the system will tell the user whether there are industrial defects in the item in the picture and the location of the industrial defects. Among them, the model for detecting abnormalities simultaneously considers the capabilities of spatial domain information and frequency domain information in discovering abnormal patterns, and increases the attention to frequency domain information at the pixel level (image reconstruction) and the feature level (feature selection). The spatial domain information and the frequency domain information complement each other and jointly guide the model training. The present invention adopts dual-domain feature selection and dual-domain reconstruction to obtain a better image representation and a higher-quality reconstructed image, increasing the distinguishability between normal and abnormal.

[0045] In this embodiment, in step S1, it includes: performing high-pass filtering on the grayscale image data through a Butterworth high-pass filter. Among them, the cut-off frequency of the Butterworth high-pass filter is taken as 30, which means that the frequency components in the frequency domain with a distance from the center frequency less than or equal to 30 will be significantly attenuated, and the order n of the filter is set to 2. The order determines the attenuation speed of the filter and the width of the transition band. A stability factor is also used to avoid division-by-zero errors in the calculation and ensure numerical stability. These three parameters jointly define the behavior and effect of the Butterworth high-pass filter. By adjusting these parameters, the influence degree and characteristics of the filter on the image can be controlled, such as edge enhancement, detail extraction, or noise suppression.

[0046] Subsequently, Perlin noise and the anomaly texture source dataset DTD are used to synthesize anomalies. Specifically, first, random noise images are generated by Perlin noise to capture various abnormal shapes, and are binarized into an anomaly mask map M through a randomly uniformly sampled threshold. a , the anomaly texture source image A is sampled from an anomaly source image dataset independent of the distribution of the input image. Subsequently, data augmentation is performed on the image A (such as: degree change, color change, automatic contrast, etc.). The enhanced texture image A is a masked by the anomaly mask map M h and mixed with the high-frequency image I a to create anomalies close to the normal distribution, thereby helping to tighten the decision boundary in the training network. Therefore, the synthesized anomaly image I

[0047]

[0048] in, It is M a The inverse of , ⊙ is the element-wise multiplication operation, and β is the opacity parameter in the blending process. This parameter is uniformly sampled from an interval [0.1, 1.0]. Random blending and enhancement can generate different abnormal images from a single texture.

[0049] In this embodiment, step S2 includes:

[0050] Specifically, in each encoding layer of the encoder, the input feature map can be input into the two branches of the dual-domain feature fusion module after operations such as convolution activation to obtain spatial domain features and frequency domain features. Specifically, the spatial domain feature selection branch includes: 3×3 convolution-BatchNorm-ReLU-3×3 convolution-BatchNorm. The frequency domain feature selection branch includes: FFT-1×1 convolution-ReLU-1×1 convolution-IFFT. Here, in order to maintain the consistency of the two selection branches, ReLU is selected as the activation function. FFT and IFFT represent fast Fourier transform and the corresponding inverse transform respectively. 1×1 convolution is mainly to achieve cross-channel information aggregation. The features obtained by the dual-domain feature selection module are called dual-domain features.

[0051] In this embodiment, step S3 includes:

[0052] The obtained dual-domain features are concatenated and fused through a simple dual-domain feature fusion module to obtain image features containing local details and global frequency information of the image. Among them, feature selection in the spatial domain helps to capture local structural information such as the shape and texture of the object, while feature selection in the frequency domain is conducive to capturing the overall frequency distribution of the image. In addition, dual-domain feature selection and dual-domain feature fusion are performed in each coding layer. Specifically, the input of each layer is the feature map output by the previous coding layer. These feature maps first enter the dual-domain feature selection module and generate dual-domain features through processing. Then, these dual-domain features are fused through the dual-domain feature fusion module. The fused dual-domain features are simply concatenated with the feature map of the layer, and then pooling operations are performed to form the output of the coding layer. Finally, the output of the last coding layer will represent the local details and global frequency information of the entire image, also known as image representation. The entire encoder contains 5 such coding layers, so dual-domain feature selection and dual-domain feature fusion will be performed 5 times during the entire image coding process.

[0053] In this embodiment, in step S4, the obtained image representation is decoded into a reconstructed image by a decoder, and the image reconstruction process is constrained by spatial domain and frequency domain information;

[0054] Among them, the decoder consists of transposed convolutional operations. The decoder layers correspond one-to-one with the encoder layers, and the output of each layer of the encoder will pass through a skip connection and be concatenated with the decoder layer.

[0055] Specifically, based on traditional image reconstruction methods, the system employs an innovative twin reconstruction strategy that performs image reconstruction simultaneously in the spatial domain and the frequency domain. This method not only captures the local details and global structure of the image but also effectively utilizes frequency domain information, thereby achieving higher-quality image reconstruction. The spatial domain information that constrains the image reconstruction process is the spatial domain reconstruction loss between the reconstructed image and the original input image, including the l2 loss in the spatial domain and the structural similarity (SSIM) loss in the spatial domain. The l2 loss in the spatial domain is: The l2 loss emphasizes pixel-level accuracy, where I and I r are the original image and the reconstructed image, respectively. The SSIM loss in the spatial domain pays more attention to the overall structure of the image, and the definition of SSIM is as follows:

[0056]

[0057] where μ x and μ y are the average luminances of images I and I r respectively, and are the variances of the two images, σ xy is the covariance of the two images, and C1, C2 are small constants added to avoid the denominator being zero. The SSIM loss is defined as:

[0058]

[0059] where H and W are the height and width of the image respectively. N p equals the number of pixels in the image. The frequency domain information that constrains the image reconstruction process is the focus frequency loss (FFL) between the reconstructed image and the original input image. FFL is a frequency domain-based loss function designed to address some details that traditional spatial domain loss functions (such as mean square error or cross entropy) may overlook. Specifically, it defines the scaled Euclidean distance of these vectors by reducing the weight of simple frequencies using a dynamic spectral weight matrix. Intuitively, the matrix is dynamically updated during training according to the non-uniform distribution of the current loss for each frequency. Then, the model will quickly focus on hard frequencies and gradually refine the generated frequencies to improve the image quality. FFL is defined as:

[0060]

[0061] where M represents the height of the spectrum, N represents the width of the spectrum, F r(u, v) is the frequency value at the spectral coordinates (u, v) of the real image, corresponding to F f (u, v) represents the corresponding value of the spectrum of the reconstructed image. w(u, v) is an element of the spectral weight matrix, representing the weight assigned to the frequency component at position (u, v), and is defined as:

[0062] w(u, v) = |F r (u, v) - F f (u, v)| α

[0063] where α is the scaling factor of flexibility, and L FFL can be regarded as the weighted average of the frequency distances between the real image and the fake image. It focuses on synthesizing hard frequencies by reducing the weights of simple frequencies. Therefore, the spatial domain and frequency domain information used to constrain the image reconstruction process is:

[0064] L rec = l2(I, Ir) + L SSIM (I, I r ) + L FFL

[0065] Using the loss information L rec to guide the model to reconstruct the image, and treating the losses in the spatial domain and frequency domain equally during the training process.

[0066] In this embodiment, in step S5, the difference between the reconstructed image and the input original image is used as the anomaly score to detect and locate anomalies, where the gradients and colors of the two images are compared. To accurately evaluate these differences, the system adopts an anomaly evaluation function of gradient similarity and color similarity. The gradient similarity uses the multi-scale gradient magnitude similarity (MSGMS). The MSGMS method uses the image pyramid technology to calculate the gradient magnitude similarity (GMS) between the original image and the reconstructed image at multiple scales. The specific operation is as follows: First, create multiple versions of the image at different scales, and then calculate the gradients of the two corresponding images at each scale respectively. By comparing these gradient magnitudes, the similarity score at each scale can be obtained, and finally these scores are averaged to obtain the final MSGMS value. The advantage of this multi-scale method is that it can capture and evaluate the details and structural differences of the image at different resolutions. MSGMS is defined as:

[0067]

[0068] where MSGMS(x, y) is the multi-scale gradient magnitude similarity between two images x and y, S represents the number of scales to be considered, and GMS s(x, y) is the gradient magnitude similarity at scale i, and the gradient similarity anomaly evaluation function can be defined as:

[0069] A gradien t = 1 M×N -MSGMS(I, I r )

[0070] The anomaly evaluation function of color similarity converts the image into the CIELAB color space that is more in line with human perception of colors for comparison. The CIELAB space consists of three components: L represents luminance, and a and b represent colors. To reduce the influence of illumination changes in anomaly detection, this evaluation function deliberately ignores the L channel and only uses the two color channels a and b for comparison. This is because the luminance L component is very sensitive to illumination conditions and may mask or distort important anomaly signals caused by color differences. The anomaly evaluation function of color similarity is defined as:

[0071]

[0072] where the superscripts a and b represent the color components in the CIELAB space, and I and I r represent the original image and the reconstructed image respectively. The calculation formula for the difference score between images is defined as:

[0073] A score =(kA color +A gradient )*f mean

[0074] where f mean is the average smoothing filter. k is a scaling factor that makes A color and A gradient of the same order of magnitude. The A score calculates the anomaly score for each pixel in the image. For the image-level detection score, the maximum value among all pixel-level anomaly scores is selected. In addition, the color similarity anomaly evaluation function using the CIELAB color space and ignoring the luminance component can provide a more stable and accurate way to analyze and compare color differences between images. This method strengthens the ability of the anomaly detection model and enables it to more effectively identify and locate color-related anomalies in complex visual environments.

[0075] ​The advantages of this embodiment are as follows: By using spatial domain information and frequency domain information to jointly guide the training of the model, the two-domain information complements each other and jointly discovers abnormal patterns existing in the image. The image preprocessing module highlights the high-frequency information that is easily lost in image reconstruction, greatly avoiding the model seeing abnormalities during inference. The two-domain feature selection module considers both the spatial domain and the frequency domain when modeling the normal image distribution. The two-domain reconstruction not only captures the local details and global structure of the image, but also effectively utilizes the frequency domain information, thus achieving higher-quality image reconstruction, which directly determines the detection performance of the model.

[0076] It should be noted that this embodiment is not a simple calculation method, but can be applied in industrial production and assist in improving the production line of industrial products. For example, in practical applications, the method of this embodiment can be applied to a system as Figure 5 shown, including:

[0077] A preprocessing module, configured to perform a fast Fourier transform on the image data to be processed, obtain high-frequency information by using a Butterworth high-pass filter, and perform an inverse fast Fourier transform on the high-frequency information to obtain a high-frequency image;

[0078] A processing module, configured to input the high-frequency image into an improved U-Net encoding network with a two-domain feature selection module to obtain two-domain features, wherein the two-domain feature selection module selects features in both the spatial domain and the frequency domain; input the two-domain features into a two-domain feature fusion module to obtain an image representation containing local details and global frequency information of the image, wherein the two-domain feature fusion module splices the two-domain features and performs 1×1 convolution; decode the image representation into a reconstructed image through a UNet decoding network, wherein the reconstructed image is constrained by a reconstruction loss in the spatial domain and a reconstruction loss in the frequency domain; use the difference between the reconstructed image and the original image as an anomaly score to detect and locate anomalies, wherein the difference includes the color difference and gradient difference of the image, the anomaly score is the anomaly score of each pixel, and the image-level anomaly score selects the maximum value of all pixel-level anomaly scores;

[0079] A query feedback module, configured to receive a query image sent by a terminal device, convert the query image into a reconstructed image generated by the model, and then determine the image-level and pixel-level anomaly scores according to the difference between the reconstructed image and the original image, and feedback the query result to the terminal device.

[0080] Specifically, this embodiment is applicable to the defect detection and localization of industrial images. That is, the product image is converted into a reconstructed image through a pre-training module and the trained model. The model will further compare the gradient and color differences between the input product image and the output reconstructed image to detect and locate abnormalities, and then return the corresponding detection and localization results. For example, the working mode is similar to the commonly used image recognition at present. When industrial parts are produced on the assembly line, the machine on the assembly line takes pictures of this part, obtains the image and inputs it into the model. Then the model reconstructs the image and uses the gradient and color anomaly evaluation function to evaluate whether the image is abnormal. Finally, the result is returned. This result can be used to determine whether the assembly line continues to produce. Because when defective industrial products are produced, it indicates that there is a problem with the assembly line, and it may be necessary to stop production and repair the machine, otherwise it will cause waste of resources. It can also be used to guide the subsequent repair of faulty machines. The location where the defect appears may also contain fault information.

[0081] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment. As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for industrial defect detection and localization based on frequency domain enhancement, characterized in that, Including: S1. After grayscaling the image to be processed, obtain a high-frequency image through a high-pass filter, where the image to be processed includes: a normal industrial product image without label information; S2. Input the high-frequency image into an encoder for dual-domain feature selection. Among them, the encoder includes a dual-domain feature selection module and a dual-domain feature fusion module, and the dual-domain feature selection module includes a spatial domain selection branch and a frequency domain selection branch; S3. Input the dual-domain features obtained in S2 into the encoder for dual-domain feature fusion to obtain a fused image representation. Among them, the dual-domain features include: spatial domain selection features and frequency domain selection features. The spatial domain selection features include the local structure of the image, and the frequency domain selection features include the global frequency distribution of the image. The fused image representation includes spatial domain information and frequency domain information; S4. Use the fused image representation for image reconstruction; S5. Use the image reconstructed in S4 and the image to be processed for anomaly detection. The anomaly detection result includes a pixel-level anomaly score; S6. After receiving the inference image sent by the terminal device, perform the above steps S1 to S5 on the inference image and obtain the anomaly detection result corresponding to the inference image. Then, determine whether the inference image is abnormal according to the anomaly detection result of the inference image. In S2, it includes: In the spatial domain selection branch, perform spatial domain feature selection through convolution 3×3 convolution - BatchNorm - ReLU - 3×3 convolution - BatchNorm; In the frequency domain selection branch, convert the convolved image to frequency domain information through the fast Fourier transform, then perform frequency domain feature selection through 1×1 convolution - ReLU - 1×1 convolution, and finally convert the frequency domain information to spatial domain information through the inverse fast Fourier transform; In S3, it includes: In the dual-domain feature fusion module, perform 1×1 convolution after splicing the dual-domain features. Among them, dual-domain feature selection and dual-domain feature fusion are performed on each encoding layer of the encoder; In each encoding layer, splice the features output by the dual-domain feature fusion module with the features input to the dual-domain feature selection module, then perform downsampling to obtain the output of one encoding layer. The output of this encoding layer is used as the input of the next encoding layer, and the features obtained from the last encoding layer are used as the image representation.

2. The method according to claim 1, characterized in that, In S1, the obtaining of the high-frequency image through the high-pass filter includes: After performing the fast Fourier transform on the image to be processed, obtain high-frequency information using a Butterworth high-pass filter, and then perform the inverse fast Fourier transform on the high-frequency information to obtain the high-frequency image.

3. The method according to claim 1, wherein In S1, it also includes: If in the training stage, the image to be processed is a sample image, and anomalies are synthesized on the high-frequency image of the sample image. The synthesized anomalies include: generating an anomaly mask using Perlin noise, and then generating an anomaly area in the high-frequency image using the anomaly texture source dataset DTD and the mask generated by Perlin noise.

4. The method according to claim 1, wherein In S4, the spatial domain information used to constrain the image reconstruction process includes: The l2 loss in the spatial domain and the structural similarity (SSIM) loss in the spatial domain, where the l2 loss in the spatial domain is: The l2 loss emphasizes pixel-level accuracy, and I and I r are the quantized representations of the original image and the reconstructed image, respectively; The SSIM loss is as follows: μ x and μ y are the average luminances of I and I r respectively, and are the variances of I and I r respectively, σ xy is the covariance of two I and I r respectively, C1 and C2 are small constants added to avoid the denominator being zero, H and W are the height and width of the image respectively, N p equals the number of pixels in the image, i and j represent the row and column indices of the top - left pixel of the current image window respectively, SSIM(I, I r ) i,j represents the structural similarity index, which is designed to evaluate the similarity of two images in terms of structure, luminance, and contrast; The frequency domain information used to constrain the image reconstruction process includes the frequency domain reconstruction focus loss FFL between the reconstructed image and the original input image. where M represents the height of the spectrum, N represents the width of the spectrum, and F r (u, v) is the frequency value at the spectrum coordinates (u, v) of the real image, and the corresponding F f (u, v) represents the corresponding value of the spectrum of the reconstructed image. w(u, v) is an element of the spectrum weight matrix, representing the weight assigned to the frequency component at position (u, v). w(u, v) = |F r (u, v) - F f (u, v)| α , where α is the proportionality factor of flexibility, and u and v respectively represent the indices of the horizontal frequency and vertical frequency in the frequency domain.

5. The method according to claim 4, characterized in that, It also includes: Build a reconstruction model for image reconstruction and optimize the reconstruction model through the Adam algorithm. During the optimization process, adjust the learning rate through the MultiStepLR learning rate scheduler, and the MultiStepLR learning rate scheduler reduces the learning rate according to preset milestones. The optimization objective is: minL rec , L rec represents the loss information used to constrain the image reconstruction process, and L rec = l2(I, I r ) + L SSIM (I, I r ) + L FFL .

6. The method according to claim 4, wherein In S5, the process of anomaly detection includes: Among them, MSGMS(x,y) represents the multi-scale gradient magnitude similarity function, which is used to calculate the multi-scale gradient magnitude similarity between two images x and y. S represents the number of scales to be considered, and GMS i (x,y) is the gradient magnitude similarity at scale i. Anomaly score A for gradient similarity anomaly assessment gradient = 1 M×N - MSGMS(I, I r ), Anomaly score for color similarity anomaly assessment The superscripts a and b represent the color components in the CIELAB space; Abnormal score A of pixels in the image score =(kA color +A gradient )*f mean , f mean is an average smoothing filter, and k is a scaling factor used to make A color and A gradient be of the same order of magnitude.

7. A system for industrial defect detection and localization based on frequency domain enhancement, characterized in that, It includes: A preprocessing module for grayscale processing the image to be processed and then obtaining a high-frequency image through a high-pass filter, where the image to be processed includes: a normal industrial product image without label information; A processing module for inputting the high-frequency image into an encoder for dual-domain feature selection. In the encoder, there is a dual-domain feature selection module and a dual-domain feature fusion module. The dual-domain feature selection module includes a spatial domain selection branch and a frequency domain selection branch; inputting the obtained dual-domain features into the encoder for dual-domain feature fusion to obtain a fused image representation. The dual-domain features include: spatial domain selection features and frequency domain selection features. The spatial domain selection features include the local structure of the image, and the frequency domain selection features include the global frequency distribution of the image. The fused image representation includes spatial domain information and frequency domain information; using the fused image representation for image reconstruction; using the reconstructed image and the image to be processed for anomaly detection, and the anomaly detection result includes a pixel-level anomaly score; A query feedback module for obtaining the corresponding anomaly detection result after receiving the inference image sent by the terminal device, then determining whether the inference image is abnormal according to the anomaly detection result of the inference image, and feeding back the query result to the terminal device; The processing module is specifically configured to perform spatial domain feature selection in the spatial domain selection branch through convolution 3×3 convolution - BatchNorm - ReLU - 3×3 convolution - BatchNorm; in the frequency domain selection branch, convert the convolved image to frequency domain information through the fast Fourier transform, then perform frequency domain feature selection through 1×1 convolution - ReLU - 1×1 convolution, and finally convert the frequency domain information to spatial domain information through the inverse fast Fourier transform; Perform 1×1 convolution after splicing the dual-domain features in the dual-domain feature fusion module, where each encoding layer of the encoder performs dual-domain feature selection and dual-domain feature fusion; in each encoding layer, splice the features output by the dual-domain feature fusion module with the features input to the dual-domain feature selection module, then perform downsampling to obtain the output of one encoding layer, and use the output of this encoding layer as the input of the next encoding layer. The features obtained from the last encoding layer are used as the image representation.

Citation Information

Patent Citations

  • Defect detection CT image reconstruction method and system fusing FBP algorithm and IR algorithm

    CN117218220A

  • SAR (Synthetic Aperture Radar) image change detection method based on double-domain network

    CN117690019A

Cited By

  • Spring surface defect detection method based on machine vision

    CN122312634A

  • A spring surface defect detection method based on machine vision

    CN122312634B