A method, device and storage medium for detecting metal surface defects

The defect-synthesis image is generated through Q-Shift dual-tree complex wavelet transformation and convolutional network, and wavelet domain processing and image reconstruction are carried out, which solves the problems of low efficiency and low accuracy of metal surface defect detection in the prior art, and achieves efficient and accurate detection without real samples.

CN115439408BActive Publication Date: 2025-07-08SOUTH CHINA UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210920393.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-02
Publication Date
2025-07-08
Estimated Expiration
2042-08-02

AI Technical Summary

Technical Problem

The prior art relies on manual visual inspection in the detection of surface defects of arc-shaped metal workpieces, which has low efficiency and low accuracy, and requires a large number of defect samples based on supervised learning methods. The distortion of synthetic samples leads to a decrease in detection ability and it is difficult to adapt to complex surface textures.

Method used

Using a method based on Q-Shift dual-tree complex wavelet transformation and convolutional network, by generating defect synthesis images and performing wavelet domain processing, the reconstruction network and the decision network are used to distinguish normal areas and defect areas, and the image prediction model is trained for detection.

Benefits of technology

It realizes high-precision metal surface defect detection without real defect sample training. It is suitable for surfaces with random texture characteristics, with fast detection speed, high accuracy and good adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115439408B_ABST
    Figure CN115439408B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus and storage medium for detecting metal surface defects. The method includes: obtaining a defect synthesis image according to a normal sample image; performing dual-tree complex wavelet transform on the defect synthesis image to transform the image features from the pixel domain to the wavelet domain, obtaining a low-frequency component and high-frequency components of multiple scales; modifying the obtained low-frequency component and high-frequency components, and performing inverse dual-tree complex wavelet transform to obtain a reconstructed image; training an image prediction model by using the defect synthesis image and the reconstructed image; obtaining an image to be detected, inputting the image to be detected into the trained image prediction model, and outputting a detection result. The present invention does not need to collect real defect samples, and only needs normal samples to have good defect detection and positioning capabilities, and can be widely used for automatic online detection of metal surfaces with a certain metallic luster. The present invention can be widely applied to the technical field of metal surface defect detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of defect detection on metal surfaces, and particularly to a method, device and storage medium for detecting defects on metal surfaces. Background Art

[0002] In the production process of arc-shaped metal workpieces (such as electronic commutators and white vehicle bodies), defects on the product surface directly affect the product quality. Due to problems such as poor material imaging characteristics, diverse defect types, various shapes, and low defect contrast on the surfaces of electronic commutators and white vehicle bodies, the defect detection on the surfaces of such products mainly relies on experienced technical workers through manual visual inspection at present. Its disadvantages such as low efficiency, low precision, low reproducibility, and boring work have become the key bottlenecks restricting enterprises from improving product quality and enhancing market competitiveness. For example, internal data of a certain automobile company shows that in 2019, defects in the body before painting (i.e., white vehicle body) accounted for 46% of the total defects, but 65% of these defects were only discovered after the vehicle was painted, greatly increasing the repair cost of the automobile.

[0003] Currently, surface defect detection mainly adopts supervised learning-based methods. Although these methods are relatively mature, model training requires collecting a large number of defect samples, which is very difficult in many applications. Therefore, some studies use artificial synthesis to generate training samples. However, due to the distortion of the synthesized samples, they cannot well fit the true distribution of defects. Therefore, once deployed to real production scenarios, the detection ability of the model will seriously decline.

[0004] Unsupervised methods refer to methods that only use normal (i.e., defect-free) samples to train a neural network model for detection, which can greatly reduce the training cost of the model. Its basic principle is: reconstruct a "normal" sample image for the test sample, and determine the area with differences between the two images as the defect area. Therefore, the core module of the algorithm based on this principle is the image reconstruction module in the network model. An ideal reconstruction module should "unchangedly" retain the defect-free area in the input sample, and at the same time fuse context information to "seamlessly" repair the defect area, so that the overall reconstruction result is highly faithful. However, how to only use normal samples to guide the network model to distinguish the normal area and the defect area and perform adaptive image repair is the main technical difficulty of this type of method. On the one hand, current methods generally perform constraints and reconstructions on images in the pixel domain. These methods have good effects on regular texture images, but it is difficult to effectively constrain the network to effectively describe the main features of random texture images. On the other hand, due to the strong random surface texture on the arc-shaped metal surfaces such as electronic commutators, the current mainstream reconstruction network modules are difficult to accurately reconstruct the normal area of the image. Therefore, there is a large amount of noise caused by reconstruction errors in the detection results, which limits the detection ability of this method. Summary of the Invention

[0005] To at least partly solve one of the technical problems existing in the prior art, an object of the present invention is to provide a method, a device and a storage medium for detecting metal surface defects.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A method for detecting metal surface defects includes the following steps:

[0008] Obtain a defect synthesis image based on a normal sample image; wherein, a defect area is provided on the defect synthesis image;

[0009] Perform dual-tree complex wavelet transform on the defect synthesis image to transform the image features from the pixel domain to the wavelet domain, and obtain a low-frequency component and high-frequency components of multiple scales;

[0010] Modify the obtained low-frequency component and high-frequency components, and perform inverse dual-tree complex wavelet transform to obtain a reconstructed image;

[0011] Use the defect synthesis image and the reconstructed image to train an image prediction model;

[0012] Obtain an image to be detected, input the image to be detected into the trained image prediction model, and output a detection result.

[0013] Further, the obtaining of the defect synthesis image based on the normal sample image includes:

[0014] Obtain a defect candidate region, and obtain a mask image according to the defect candidate region;

[0015] Generate an abnormal image according to the defect candidate region; wherein, the abnormal image includes a salt-and-pepper noise image, a Gaussian noise image, a faded image or a heterologous dataset image;

[0016] Fuse the normal sample image and the abnormal image according to the mask image to generate a defect area, and obtain a defect synthesis image.

[0017] Further, the expression of the defect synthesis image is as follows:

[0018]

[0019] In the formula, is the synthesis image, is the normal sample image, is the abnormal image, β is a fusion factor, is a random floating point number obeying a uniform distribution, ⊙ represents pixel multiplication, represents the mask image, is the binary image after taking the inverse of the mask image of the mask image.

[0020] Furthermore, modifying the obtained low-frequency component and high-frequency component and performing inverse dual-tree complex wavelet transform to obtain a reconstructed image includes:

[0021] For the low-frequency wavelet coefficient map, use a reconstruction network to reconstruct the low-frequency component as the modified low-frequency component. The reconstructed wavelet coefficient map is consistent with the wavelet coefficients of the normal sample image. Figure 1 consistent;

[0022] For the high-frequency wavelet coefficient map, each scale of the high-frequency component corresponds to a decision module; the decision module takes the modulus of the real and imaginary parts of the wavelet coefficient map in each direction and outputs a fractional value for each local part in the image; if the output fractional value is greater than a preset threshold, retain the wavelet coefficients of that local part; otherwise, set the wavelet coefficients of that local part to 0; obtain the modified high-frequency component.

[0023] Perform inverse dual-tree complex wavelet transform on the modified low-frequency component and high-frequency component to obtain a reconstructed image.

[0024] Furthermore, the reconstruction network is an autoencoder network, and the expression of the loss function of the reconstruction network is:

[0025]

[0026] where is the wavelet coefficient map output by the reconstruction network, is the low-frequency wavelet coefficient map obtained by performing the same dual-tree complex wavelet decomposition on the normal sample image; represents the structural similarity loss value between the two images.

[0027] Furthermore, the expression of the loss function corresponding to the high-frequency component is:

[0028]

[0029] where is the high-frequency component image processed by the decision module, is the high-frequency component image of the normal sample image, ⊙ is pixel multiplication, is the mask image, is the mask image the binary image after taking the inverse.

[0030] Furthermore, the expression of the loss function in the pixel domain corresponding to the reconstructed image is:

[0031]

[0032] where is the reconstructed image, is a normal sample image, is a mask image, is a mask image is a binary image after inversion.

[0033] Further, the training of the image prediction model using the defect synthesis image and the reconstructed image includes:

[0034] Obtain the defect synthesis image and the reconstructed image, as well as the residual image between the two images, and stack the defect synthesis image, the reconstructed image, and the residual image along the channel number direction as the input of the image prediction model;

[0035] Use the Focal loss function to train the image prediction model so that the defect area detected by the image prediction model is consistent with the defect candidate area on the defect synthesis image.

[0036] Another technical solution adopted by the present invention is:

[0037] A metal surface defect detection device includes:

[0038] At least one processor;

[0039] At least one memory for storing at least one program;

[0040] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0041] Another technical solution adopted by the present invention is:

[0042] A computer-readable storage medium stores a program executable by a processor, and the program executable by the processor is used to execute the above method when executed by the processor.

[0043] The beneficial effects of the present invention are: The present invention does not need to collect real defect samples, and only needs normal samples to have good defect detection and positioning capabilities, and can be widely used for automatic on-line detection of metal surfaces with a certain metallic luster. Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the relevant technical solution drawings in the embodiments of the present invention or the prior art. It should be understood that the drawings introduced below are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention, and those skilled in the art can also obtain other drawings based on these drawings without creative efforts.

[0045] Figure 1 is the network framework flow chart in the embodiment of the present invention;

[0046] Figure 2 These are schematic diagrams of some synthetic image examples in the embodiments of the present invention; among them, Figure 2 (a) is a synthetic image of salt-and-pepper noise, Figure 2 (b) is a synthetic image of Gaussian noise, Figure 2 (c) is a faded synthetic image, Figure 2 (d) is a synthetic image of a heterogeneous dataset;

[0047] Figure 3 This is a schematic diagram of the reconstruction network model structure in the embodiments of the present invention;

[0048] Figure 4 This is a schematic diagram of the decision module structure in the embodiments of the present invention;

[0049] Figure 5 These are the original images and reconstructed images of normal examples and defective examples in the KSDD2 dataset by the reconstruction network in the embodiments of the present invention; among them, Figure 5 (a) is the original image of a normal example, Figure 5 (b) is the reconstructed image of the corresponding example; Figure 5 (c) is the original image of a defective example, Figure 5 (d) is the reconstructed image of the defective example;

[0050] Figure 6 This is a schematic diagram of the image prediction module structure in the embodiments of the present invention;

[0051] Figure 7 These are the detection result graphs on the KSDD2 open-source dataset and the phosphating plate dataset simulating the surface of a white vehicle body in the embodiments of the present invention;

[0052] Figure 8 This is a flowchart of the steps of a method for detecting metal surface defects in the embodiments of the present invention. Detailed implementation manners

[0053] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0054] In the description of the present invention, it should be understood that regarding the orientation description, for example, the orientation or positional relationship indicated by up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.

[0055] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is more than two, greater than, less than, exceeding, etc. are understood as not including the present number, above, below, within, etc. are understood as including the present number. If the first and second are described, it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0056] In the description of the present invention, unless otherwise clearly defined, words such as setting, installing, connecting, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.

[0057] As shown in Figure 1 and 8 In this embodiment, a method for detecting metal surface defects is based on Q-Shift dual-tree complex wavelet transform and convolutional network. First, an artificial defect image is generated by an image enhancement module, then the image is subjected to dual-tree complex wavelet transform to transform the image features from the pixel domain to the wavelet domain. Subsequently, the reconstruction network and the decision network are used to process the low-frequency component and the high-frequency component respectively. Immediately afterwards, the inverse dual-tree complex wavelet transform is performed to transform the image from the wavelet domain back to the pixel domain to obtain a reconstructed image. Finally, the image prediction module outputs the detection result. Before the detection of this method, it is not necessary to collect any real defect samples. Only a certain number of normal samples of this kind need to be collected to train the network model, and end-to-end defect detection can be achieved. The method specifically includes the following steps:

[0058] S1. Obtain a defect synthesis image according to the normal sample image; wherein, a defect candidate area is provided on the defect synthesis image.

[0059] Image enhancement module: By using image processing technology, the normal image is randomly damaged to obtain a combined defect synthesis image, providing "defect samples" for subsequent modules.

[0060] Specifically, the image enhancement module is composed of a mask generation module and an abnormal image generation module. First, the mask generation module is used to generate a defect candidate area, and then the abnormal image generation module is used to fill the content into the defect candidate area to obtain a synthesis image, so as to guide the network model to distinguish the normal area and the defect area.

[0061] As an alternative implementation, the mask generation module is a random algorithm for generating defect candidate regions. The mask generation module can generate three major types of mask regions with different shapes and sizes: regular regions, irregular regions, and Perlin noise regions.

[0062] The random generation algorithm for regular region masks can randomly generate rectangular regions of different sizes. The random generation algorithm for irregular region masks can generate various irregular regions, which are composed of three primitives: line segments, circles, and squares. The random generation algorithm for Perlin noise region masks is a random algorithm based on Perlin noise images, and candidate regions can be obtained by thresholding the Perlin noise images.

[0063] Using the above mask generation module, a series of various candidate regions can be obtained. Then, the abnormal image generation module is used to fill the content within the candidate regions. The abnormal image generation module can generate four types of abnormal images: salt-and-pepper noise images, Gaussian noise images, faded images, and heterologous dataset images.

[0064] After obtaining the mask image and the abnormal image, image synthesis can be performed to finally obtain a synthesized image. The synthesized image can be represented by the following formula:

[0065]

[0066] In the formula, is the synthesized image, is the normal sample image, is the abnormal image, β is the fusion factor, is a random floating-point number subject to a uniform distribution, ⊙ represents pixel multiplication, represents the mask image.

[0067] S2. Perform a dual-tree complex wavelet transform on the defect synthesized image to transform the image features from the pixel domain to the wavelet domain, and obtain the low-frequency component and high-frequency components at multiple scales.

[0068] S3. Modify the obtained low-frequency component and high-frequency components, and perform an inverse dual-tree complex wavelet transform to obtain a reconstructed image.

[0069] Image reconstruction module: Adopt a network model composed of Q-Shift dual-tree complex wavelet transform and convolutional neural network. This module is composed of a reconstruction network and a decision module. In the wavelet domain, the reconstruction network reconstructs the low-frequency component, and the decision module classifies and reconstructs the high-frequency components. In the pixel domain, additional constraints are imposed on the normal regions in the image.

[0070] Specifically, the image reconstruction module consists of a reconstruction network and multiple decision-making modules. First, the defect synthesis image obtained in step S1 is used as the input. The image reconstruction module first performs a dual-tree complex wavelet transform on the image to obtain multi-level high-frequency components and low-frequency components respectively. Then, the reconstruction network reconstructs the low-frequency components, and the decision-making module makes a local decision for each high-frequency component map, so as to only retain the high-frequency components in the normal area. Then, an inverse dual-tree complex wavelet transform is performed to obtain the reconstructed image.

[0071] As an optional implementation manner, first, a multi-level Q-Shift dual-tree complex wavelet transform is performed on the image to obtain high-frequency components and low-frequency components at multiple scales. For the low-frequency wavelet coefficient map, the reconstruction network reconstructs it, and it is required that the reconstructed wavelet coefficient map is consistent with the wavelet coefficients of the image before damage. The reconstruction network adopted in this embodiment is an autoencoder network. For the high-frequency wavelet coefficient map, each scale of high-frequency component corresponds to a decision-making module. The decision-making module first calculates the modulus of the real part and the imaginary part of the wavelet coefficient map in each direction, and then the convolutional network outputs a fractional value for each small local area in the image. If the fractional value is greater than 0.5, the wavelet coefficients of this local area are retained, otherwise the wavelet coefficients of this local area are set to 0. After the above processing, the modified low-frequency components and high-frequency components can be obtained, and then an inverse dual-tree complex wavelet transform is performed on them to restore the image from the wavelet domain to the pixel domain to obtain the reconstructed image. Figure 1 S4. Use the defect synthesis image and the reconstructed image to train the image prediction model.

[0072] The image prediction module: Stack the synthesis image, the reconstructed image, and the residual image between them along the channel number direction as the input of the model. The network structure adopts a network model similar to the Unet structure. The model outputs an abnormal score value for each pixel position in the image, and takes the maximum abnormal score value in the image as the abnormal score of the image.

[0073] Specifically, for defect segmentation and localization: Stack the synthesis image obtained in step S2, the reconstructed image obtained in step S3, and the residual map between the two along the channel number direction, and use it as the input of the image prediction module. The network model directly outputs the abnormal score value of each pixel position to obtain an abnormal score map, and takes the maximum abnormal score value in the image as the abnormal score of the image.

[0074] As an optional implementation manner, first, use the defect synthesis image in step S1, the reconstructed image in step S2, and the residual map between the two images as the input. Then, during the training process, use the Focal loss to require the image prediction module to output the same as the defect candidate area in step S1. Finally, take the maximum value in the detection result map output by the image prediction module as the abnormal score of the image.

[0075] ​

[0076] S5. Obtain the image to be detected, input the image to be detected into the trained image prediction model, and output the detection result.

[0077] The above method will be explained in detail with reference to the accompanying drawings and specific embodiments.

[0078] In this embodiment, the KolektorSDD2 open-source dataset (abbreviated as KSDD2) is used as the detection object, and a method for defect detection of industrial metal products is introduced, including the following steps:

[0079] S101. The image sizes of the original KSDD2 dataset are not fixed, and directly adjusting the resolution has a certain negative impact on the performance of the present invention. Therefore, first obtain the minimum side length (184 pixels in length) in the training set and the test set, and then randomly crop square image patches with a side length of 184 pixels for each image in the training set and the test set, and uniformly adjust the size to 256 pixels × 256 pixels. The training set contains 2085 normal samples, and the test set contains 51 defect samples and 99 normal samples.

[0080] S102. The image enhancement module consists of two sub-modules: a mask generation module and an abnormal image generation module. The mask generation module can generate 3 types of mask regions with different shapes and sizes: regular regions, irregular regions, and Perlin noise regions. The random generation algorithm for the regular region mask is to randomly generate rectangular regions of different sizes. It randomly generates rectangular regions that occupy one-tenth to one-fourth of the image area according to the size of the input image. The random generation algorithm for the irregular region mask generates irregular regions composed of three primitives: line segments, circles, and squares. The line width, radius, and side length have a value range of [1, 25] pixels, the line segment length has a value range of [10, 70] pixels, the included angle between adjacent sides has a range of [0°, 130°], and each connected region is randomly composed of 1 to 5 similar primitives. The maximum number of enhancement times for the same image is 10 times. The random generation algorithm for the Perlin noise region mask is a random algorithm based on Perlin noise images. First, the random algorithm generates a Perlin noise image. To further improve the randomness of the region, the random algorithm randomly rotates the noise image, and the angle range is [-45°, 45°]. Then, the rotated image is thresholded, and the region greater than the threshold is the defect candidate region. The threshold selected in this example is 0.5.

[0081] Then, the abnormal image generation module can generate four types of abnormal images: salt-and-pepper noise images, Gaussian noise images, faded images, and heterologous dataset images. For salt-and-pepper noise images, the signal-to-noise ratio is a random number following a uniform distribution with a value range of [0.1, 0.9]. Gaussian noise images are truncated normal distributions, which follow a normal distribution with a mean of 0 and a standard deviation of 0.5, with an upper limit of 1.0 and a lower limit of 0. For faded images, the random algorithm first inputs the image of the corresponding normal sample and normalizes it to 0 to 1, and then randomly adds a floating-point number to the entire image, causing the brightness of the entire image to deviate. Therefore, when this image is synthesized with the original image, local gray value changes will occur. The floating-point number used in this example is a random number following a uniform distribution with a value range of -[0.1, 0.5] ∪ [0.1, 0.5]. For heterologous dataset images, the DTD texture dataset is selected in this example.

[0082] After obtaining the mask image and abnormal images, image synthesis can be performed to finally obtain the synthetic image. The synthetic image can be expressed by the following formula:

[0083]

[0084] In the formula, is the synthetic image, is the normal sample image, is the abnormal image (salt-and-pepper noise image, Gaussian noise image, faded image, and heterologous dataset image), β is the fusion factor, a random floating-point number following a uniform distribution with a value range of [0, 0.5], ⊙ represents pixel multiplication, represents the mask image.

[0085] Figure 2 Four synthetic image examples are shown. Figure 2 (a) is the salt-and-pepper noise synthetic image, Figure 2 (b) is the Gaussian noise synthetic image, Figure 2 (c) is the faded synthetic image, Figure 2 (d) is the heterologous dataset synthetic image. The proportion of different mask types is: {regular region: line segment irregular region: circular irregular region: square irregular region: Berlin noise region: defect-free region = 3:1:1:1:3:1}. In the images with artificial defects, the proportion of different defect types is: {salt-and-pepper noise enhancement: Gaussian noise enhancement: faded enhancement: heterologous image enhancement = 1:1:4:4}.

[0086] Therefore, two images are obtained in this step, the synthetic image and the mask image.

[0087] S103. After obtaining the synthetic image from the previous step, as the input of the image reconstruction module, first perform a 3-level Q-Shifit dual-tree complex wavelet decomposition on the image. The length of the mother wave at the first level is 13 / 19 tap, and the length at the second level and above is 10 tap. Each time decomposition is performed, the image is downsampled by 2. After 3-level decomposition, three scales of high-frequency components and one scale of low-frequency component will be obtained.

[0088] S104. For the low-frequency component, use a reconstruction network to reconstruct the wavelet coefficient map. In this example, a network model with an autoencoder structure is used. Figure 3 The schematic diagram of the network structure of the reconstruction network for performing J-level dual-tree complex wavelet decomposition is shown. In the figure, J represents the number of decomposition levels, h is the height of the image, w is the width of the image, and the number above the feature map represents the number of channels of the feature map; Conv 5×5, ReLU represents a convolutional layer with a convolutional kernel size of 5, a stride of 1, and a ReLU activation function; Conv 3×3, ReLU represents a convolutional layer with a convolutional kernel size of 3, a stride of 1, and a ReLU activation function; Maxpooling, 2×2 represents a max pooling layer with a convolutional kernel size of 2 and a stride of 2; Upsample, 2 represents bilinear interpolation upsampling with a magnification factor of 2; Conv 3×3 represents a convolutional layer with a convolutional kernel size of 3, a stride of 1, and no activation function.

[0089] The loss function of the reconstruction network is:

[0090]

[0091] In the formula, is the wavelet coefficient map output by the reconstruction network, is the low-frequency wavelet coefficient map obtained by performing the same dual-tree complex wavelet decomposition on a normal image (the image before artificial damage). The structural similarity loss, the expression is:

[0092]

[0093] In the formula, represents the structural similarity loss value of two images within a certain window, μ x , μ y , σ x , σ y are the mean and standard deviation of images x and y respectively; σ xy is the covariance of x and y; the stability factors C1 and C2 are constants, which are 0.0001 and 0.0009 respectively.

[0094] S105. For high-frequency components, each level of wavelet decomposition generates wavelet coefficients in six directions of ±15°, ±45°, and ±75° corresponding to the scale, and the wavelet coefficients in each direction consist of a real part and an imaginary part. In this example, decision modules with the same network structure are used for high-frequency components of different scales. Here, the decision module of the high-frequency component at the Jth level is taken as an example for introduction.

[0095] The classification network in the decision module uses image block-level classification to determine whether to retain the high-frequency subband in any small region in each direction. The decision module first performs a modulus operation on the real part and the imaginary part:

[0096]

[0097] In the formula, is the modulus of the real part and the imaginary part of the high-frequency subband coefficient, is the real part of the high-frequency subband coefficient, is the imaginary part of the high-frequency subband coefficient.

[0098] The modulus of the high-frequency subband coefficient maps in each direction of the Jth-level dual-tree complex wavelet decomposition is input into the classification network. The input size of the classification network is The output size is Therefore, the classification score at each pixel position in the result of the classification network corresponds to a small region of 2 J+2 pixels × 2 J+2 pixels in the original image. After obtaining the classification score, the classification score is adjusted to the same shape as the input by nearest-neighbor linear interpolation. In the training stage, the adjusted classification score is pixel-multiplied with the original high-frequency wavelet coefficient map to obtain the "reconstructed" wavelet coefficient map. In the testing stage, the classification score is binarized with a threshold of 0.5. The classification score greater than 0.5 is set to 1, and the classification score less than 0.5 is set to 0, thus realizing decision classification at the small image block level:

[0099]

[0100] In the formula, represents the original high-frequency wavelet coefficient, and d represents the classification score at the current pixel position.

[0101] Figure 4In it, the numbers above the feature maps indicate the shapes of the feature maps. Modulo operation represents the modulo operation. Conv 5×5,ReLU represents a convolutional layer with a convolutional kernel size of 5, a stride of 1, and a ReLU activation function; Conv 3×3,ReLU represents a convolutional layer with a convolutional kernel size of 3, a stride of 1, and a ReLU activation function; Maxpooling,2×2 represents a max pooling layer with a convolutional kernel size of 2 and a stride of 2; Upsample,4 represents nearest neighbor interpolation upsampling with an amplification factor of 4; Conv 3×3,Sigmoid represents an unbiased convolutional layer with a convolutional kernel size of 3, a stride of 1, and a Sigmoid activation function.

[0102] Performing a 3-level complex wavelet transform will generate 3 high-frequency components of different scales and 1 low-frequency component of one scale. Set the low-frequency component to 0, and then perform the inverse wavelet transform to obtain an image with only high-frequency components. The calculation formula for the loss function of the high-frequency component is:

[0103]

[0104] In the formula, is the high-frequency component image of the decision module "reconstruction", is the high-frequency component image before the image is damaged, ⊙ is pixel multiplication, is the mask image obtained in step S2.

[0105] S106. After S104 and S105, the reconstructed low-frequency component and high-frequency component can be obtained. Perform the inverse dual-tree complex wavelet transform on them to obtain the reconstructed image. The calculation formula for the loss function in the pixel domain is:

[0106]

[0107] In the formula, is the reconstructed image, is the normal image.

[0108] Figure 5 Shows the reconstructed images of the normal images and defective images in the KSDD2 dataset by the reconstruction module, where Figure 5 (a) is a normal sample, Figure 5 (b) is the reconstructed image corresponding to the normal sample, Figure 5 (c) is a defective sample (the red curve area is the defective area), Figure 5 (d) is the reconstructed image corresponding to the defective sample. Compare the synthesized image and the reconstructed image to obtain the L1 residual map.

[0109] S107. Stack the synthesized image, the reconstructed image, and the residual image as the input of the image prediction module.

[0110] The network structure of the image prediction module is shown in Figure 6 , where the numbers above the feature maps represent the number of channels of the features, and the numbers in the lower left corner of the feature maps represent the sizes of the feature maps. Conv 5×5, ReLU represents a convolutional layer with a convolutional kernel size of 5, a stride of 1, and a ReLU activation function; Maxpooling, 2×2 represents a max pooling layer with a convolutional kernel size of 2 and a stride of 2; Conv3×3, ReLU represents a convolutional layer with a convolutional kernel size of 3, a stride of 1, and a ReLU activation function; Conv 1×1, Sigmoid represents an unbiassed convolutional layer with a convolutional kernel size of 1, a stride of 1, and a Sigmoid activation function; dConv 3×3, ReLU represents a deconvolutional layer with a convolutional kernel size of 3, a stride of 2, and a ReLU activation function; Concatenation represents stacking along the number of channels; Skip, Concatenation represents a skip connection layer; Maximum represents the operation of taking the maximum value.

[0111] The loss function of the image prediction module adopts the Focal loss, and the formula is as follows:

[0112]

[0113] In the formula, is the anomaly score map output by the image prediction module, is the mask image. Among them, the γ value of the Focal loss is 2.

[0114] The maximum anomaly score value in the anomaly score map is used as the anomaly score of the image. In addition, in this example, the optimizer selects the Adam optimizer with a learning rate of 0.0001, and the number of iterations is 100.

[0115] Figure 7 Some detection results in this case and some detection result diagrams of the method in this paper on the phosphating plate with a certain curvature change similar to the white body metal are shown. In the figure, the first row is the original image, the second row is the reconstructed image, the third row is the GT image, and the fourth row is the detection result, that is, the anomaly score map. The closer the color is to dark red, the higher the anomaly score, and the closer it is to dark blue, the lower the anomaly score. The results show that the method in this paper has good detection performance for metal surfaces.

[0116] To sum up, compared with the prior art, the method of this embodiment has the following advantages and beneficial effects:

[0117] 1. The method provided by the present invention does not require collecting any real defect samples before detection, nor does it need standard images as reference templates. In particular, it can be applied to surfaces with strongly random texture features and also has good detection performance for regular texture surfaces. It has the advantages of fast detection speed, high accuracy, stable detection results, and good adaptability.

[0118] 2. By combining a convolutional neural network and a dual-tree complex wavelet transform, the present invention makes the texture information of the normal region in the reconstructed image richer, can effectively suppress the reconstruction error. In addition, by stacking the test image, the reconstructed image, and their residual maps as the input of the image prediction module, it can better guide the network to segment and locate the defect region.

[0119] 3. The image enhancement module in the present invention provides a defect random generation algorithm with various shapes and types for artificial defects on the texture surface, which has certain reference significance for other defect detection algorithms.

[0120] This embodiment also provides a metal surface defect detection device, including:

[0121] At least one processor;

[0122] At least one memory for storing at least one program;

[0123] When the at least one program is executed by the at least one processor, the at least one processor implements the method as Figure 8 shown.

[0124] The metal surface defect detection device of this embodiment can execute a metal surface defect detection method provided by the method embodiment of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0125] This application embodiment also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 8 the method shown.

[0126] This embodiment also provides a storage medium, storing instructions or programs that can execute a metal surface defect detection method provided by the method embodiment of the present invention. When the instructions or programs are run, any combination of implementation steps of the method embodiment can be executed, and the corresponding functions and beneficial effects of the method are possessed.

[0127] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks may sometimes be executed in the reverse order. Additionally, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.

[0128] Furthermore, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those skilled in the art will be able to implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0129] If the described functions are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, or part of the technical solution, may be embodied in the form of a software product stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0130] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.

[0131] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.

[0132] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0133] In the above description of this specification, the descriptions referring to the terms "one embodiment / Example", "another embodiment / Example", or "certain embodiments / Examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0134] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0135] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without violating the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A method for detecting metal surface defects, characterized in that, It includes the following steps: Obtain a defect synthesis image based on a normal sample image; wherein, a defect area is provided on the defect synthesis image; Perform dual-tree complex wavelet transform on the defect synthesis image to transform the image features from the pixel domain to the wavelet domain, obtaining a low-frequency component and high-frequency components of multiple scales; Modify the obtained low-frequency component and high-frequency components, and perform inverse dual-tree complex wavelet transform to obtain a reconstructed image; Train an image prediction model using the defect synthesis image and the reconstructed image; Obtain an image to be detected, input the image to be detected into the trained image prediction model, and output a detection result; The step of modifying the obtained low-frequency component and high-frequency components, and performing inverse dual-tree complex wavelet transform to obtain a reconstructed image includes: For the low-frequency wavelet coefficient map, use a reconstruction network to reconstruct the low-frequency component as the modified low-frequency component, and the reconstructed wavelet coefficient map is consistent with the wavelet coefficient map of the normal sample image; For the high-frequency wavelet coefficient map, each scale of high-frequency component corresponds to a decision module; the decision module takes the modulus of the real part and the imaginary part of the wavelet coefficient map in each direction, and outputs a score value for each local part in the image; if the output score value is greater than a preset threshold, retain the wavelet coefficients of this local part; otherwise, set the wavelet coefficients of this local part to 0; obtain the modified high-frequency component; Perform inverse dual-tree complex wavelet transform on the modified low-frequency component and high-frequency components to obtain a reconstructed image; The step of training the image prediction model using the defect synthesis image and the reconstructed image includes: Obtain the defect synthesis image and the reconstructed image, as well as the residual map between the two images, and stack the defect synthesis image, the reconstructed image, and the residual image along the channel number direction as the input of the image prediction model; Use the Focal loss function to train the image prediction model so that the defect area detected by the image prediction model is consistent with the defect candidate area on the defect synthesis image.

2. The method for detecting metal surface defects according to claim 1, wherein The step of obtaining a defect synthesis image based on a normal sample image includes: Obtain a defect candidate area, and obtain a mask image according to the defect candidate area; Generate an abnormal image according to the defect candidate area; wherein, the abnormal image includes a salt-and-pepper noise image, a Gaussian noise image, a faded image, or an image from a heterogeneous dataset; According to the mask image, fuse the normal sample image and the abnormal image to generate a defect area, and obtain a defect synthesis image.

3. A method for detecting metal surface defects according to claim 2, characterized in that, The expression of the defect synthesis image is as follows: Wherein, is the synthetic image, is the normal sample image, is the abnormal image; is the fusion factor, which is a random floating point number obeying the uniform distribution; represents pixel dot multiplication, represents the mask image, is the mask image the binary image after taking the inverse.

4. A method for detecting metal surface defects according to claim 1, characterized in that, The reconstruction network is an autoencoder network, and the expression of the loss function of the reconstruction network is: In the formula, is the wavelet coefficient map output by the reconstruction network, is the low-frequency wavelet coefficient map obtained by performing the same dual-tree complex wavelet decomposition on the normal sample image; represents the structural similarity loss value of the two images.

5. A method for detecting metal surface defects according to claim 4, characterized in that, The expression of the loss function corresponding to the high-frequency component is: In the formula, is the high-frequency component image processed by the decision-making module, is the high-frequency component image of the normal sample image, is pixel multiplication, is the mask image, is the mask image which is the binary image after inversion.

6. A method for detecting metal surface defects according to claim 4, characterized in that, The expression of the loss function in the pixel domain corresponding to the reconstructed image is: In the formula, is the reconstructed image, is the normal sample image, is the mask image, is the binary image after taking the inverse of the mask image.

7. A metal surface defect detection device, characterized in that, It includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1-6.

8. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor is used to execute the method according to any one of claims 1-6 when executed by the processor.

Citation Information

Patent Citations

  • Power equipment defect identification method based on image fusion deep learning model

    CN112184661A