A detection method for defective areas of colored fabrics based on generative adversarial networks
By constructing a color fabric defect detection method based on a generative adversarial network, and reconstructing and repairing the color fabric image using the EFFGAN model, the problem of difficulty in detecting complex textures and tiny defects in the prior art is solved, and a fast and accurate defect detection effect is achieved.
Patent Information
- Application Number
- CN202111305800.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-05
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-11-05
AI Technical Summary
The prior art is difficult to effectively detect complex textures and tiny defects colored fabrics. Traditional texture defect visual detection methods cannot perform good detection of multiple texture types at the same time, and the unsupervised deep learning model has poor detection effect on complex textures.
The color fabric defect detection method based on the Generative Adversarial Network (GAN) is adopted to construct an image reconstruction and repair model EFFGAN model with unsupervised learning. Through the collaborative training of generator G and discriminator D, the generator encodes and decodes the input image, and the discriminator adjusts the gradient feedback, guides the generator's training, and generates the reconstructed image to detect defect areas.
It realizes rapid and accurate reconstruction and repair of colored fabric images, can effectively detect defect areas of excellent fabrics, improve detection efficiency and accuracy, and is suitable for complex textures and high-resolution colored fabric images.
Smart Images

Figure CN114119500B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of textile appearance detection methods, and relates to a method for detecting defective areas of colored fabrics based on a generative adversarial network. Background Art
[0002] Colored fabrics are made of dyed yarns, combined with changes in organizational structure, color matching and finishing processes. With beautiful flower shapes, unique styles and a wide variety of varieties, they have become indispensable materials in people's daily lives and industrial applications, and have high added value. However, due to its long production process and complex process, various forms of defects often occur in actual production. While color-dyed fabrics are gaining wider application, people have also put forward higher requirements for their quality. The generation of defects will greatly reduce the value of fabric products and have a significant impact on the benefits and image of color-dyed fabric production companies. As the texture types of color-dyed fabrics become more and more complex, the types of color-dyed fabric defects are gradually increasing, and the difficulty of defect detection is constantly increasing. The original manual detection method is for workers to visually inspect the fabrics according to personal experience and evaluation standards to find defects. This method has low detection speed, high missed detection rate and high cost. Under the huge commercial pressure of high standards, strict requirements and low costs faced by major textile companies, it cannot meet the requirements of real-time and accurate detection of enterprises. Therefore, there is an urgent need for an automated detection method to replace the low-precision and low-efficiency manual visual inspection.
[0003] Among the automated defect detection methods, using computer vision technology to detect color fabric defects has become a research hotspot for many scholars at home and abroad. Most traditional visual texture defect detection methods use artificially designed features extracted from each small texture image block to distinguish defect areas from non-defective areas. The detection performance of these methods is largely limited by the recognition ability of the artificially designed features extracted from each texture image block. They have certain effects in the detection of fabric defects with simple textures, but are not ideal in the detection of color fabrics with complex textures and tiny defects. Therefore, traditional texture defect visual detection methods cannot perform good detection of multiple texture types at the same time. In recent years, with the rapid development of deep learning, deep neural networks have also been widely used in color fabric defect detection and classification. Compared with traditional feature extraction methods, deep neural networks can extract deep abstract features that are difficult to extract with traditional visual methods. Therefore, supervised learning methods in deep learning have been widely used in fabric defect detection due to their powerful image recognition capabilities. However, in actual production, the marking process of defect samples is complicated, the time and labor costs are high, and the number of defect samples is far less than that of normal samples, resulting in sample imbalance problems, which brings many difficulties to supervised color fabric defect detection methods. In response to the above problems, some scholars have begun to explore the use of unsupervised deep learning models to achieve color fabric defect detection methods. The key to unsupervised methods is whether the texture features of fabric images can be effectively extracted, thereby effectively reconstructing the test samples into defect-free images.
[0004] Zhang et al. proposed a color pattern fabric defect detection algorithm called unsupervised denoising convolutional autoencoder (DCAE), which realizes the detection and location of color pattern fabric defects by processing the residual image of the test image and its reconstructed image, but this method is only applicable to fabrics with simple background texture. Mei et al. proposed a defect detection framework based on multi-scale convolutional denoising autoencoder (MSCDAE), which combines the image pyramid hierarchy and the idea of convolutional denoising autoencoder to detect fabric defects. However, for color pattern fabrics with complex textures, this method is prone to over-detection. Zhang et al. proposed a U-shaped convolutional denoising autoencoder model (UDCAE) based on traditional autoencoders and classic U-Net, but this model only encodes and decodes fabric images at a single scale, so information loss is inevitable. Wei et al. used the VAE model and introduced the mean structure similarity (MSSIM) as the loss function to realize real-time fabric defect detection. This method is suitable for defect detection scenarios of small-resolution simple texture fabrics, but cannot effectively reconstruct irregular and complex texture fabrics, and thus cannot detect and locate fabric defects. Hu et al. proposed an unsupervised fabric defect detection model based on deep convolutional generative adversarial network (DCGAN). The model constructs a reconstructed image of the image to be tested, and then performs residual analysis on the likelihood map of the reconstructed image and the original image to detect the defective area. However, this method is suitable for fabric images with low resolution and high background texture similarity. It is difficult to obtain effective detection results for complex fabric images with high resolution.
[0005] Although the above unsupervised method does not require a large number of labeled defect data sets and only needs to use easily available defect-free samples as input, the autoencoder and variational autoencoder structures are too simple to learn complex texture features, so the detection effect on colored fabrics is not good. The existing GAN model for unsupervised fabric defect detection is also difficult to obtain stable detection results due to problems such as high training difficulty, and cannot solve the problem of colored fabric defect detection in the production process. Summary of the invention
[0006] The purpose of the present invention is to provide a method for detecting defective areas of colored fabrics based on a generative adversarial network, which can effectively reconstruct and repair colored fabric images, thereby quickly and accurately detecting colored fabric defects.
[0007] The technical solution adopted by the present invention is a method for detecting defective areas of colored fabrics based on a generative adversarial network, which is specifically implemented according to the following steps:
[0008] Step 1: Construct an image reconstruction and restoration model EFFGAN based on unsupervised learning. The model consists of two parts: the generator G and the discriminator D.
[0009] Step 2, construct a training set of defect-free colored fabric images including defect-free colored fabric images, then superimpose Gaussian noise on the defect-free colored fabric images, and send the defect-free colored fabric images after superimposing noise into the EFFGAN model constructed in step 1, the generator G extracts and restores features of the input image through encoding and decoding operations, the discriminator D continuously adjusts the gradient feedback to the generator G, guides the training of the generator G, and when the number of training times reaches the set number of iterations, the trained EFFGAN model is obtained;
[0010] Step 3: Input the color fabric image to be detected into the EFFGAN model trained in step 2 to output the corresponding reconstructed image, and then perform detection to determine the defective area.
[0011] The present invention is also characterized in that
[0012] In step 1, the input layer and output layer of the generator G are both three-channel image structures, and the generator uses the feature pyramid structure FPN as the feature extractor.
[0013] The feature pyramid structure FPN includes five backbone modules for feature extraction connected from bottom to top, denoted as C0, C1, C2, C3, and C4, and three feature fusion modules connected from top to bottom, denoted as A0, A1, and A2. The backbone modules C4 and C3 are also connected to the feature fusion module A0 through the lateral connection part, and the backbone modules C2 and C1 are also connected to A1 and A2 one by one through the lateral connection part. The outputs of the three feature fusion modules and the outputs of the top backbone module C4 and the bottom backbone module C0 after lateral connection are respectively obtained through channel series operation to obtain feature maps, and then the feature maps are up-sampled to obtain reconstructed images.
[0014] The backbone module is composed of the MBConv and Fused-MBConv structures of the EfficientNetV2 network, or the Bottleneck residual block structure of the MobileNetV2 network;
[0015] When the EfficientNetV2 network is used, the five backbone modules connected from bottom to top are as follows: the C0 part includes a convolution layer with a convolution kernel size of 3×3 and a stride of 2 and two Fused-MBConv modules with a convolution kernel size of 3×3 and a stride of 1 connected thereto; the C1 part includes four Fused-MBConv modules with a convolution kernel size of 3×3 and a stride of 2 connected thereto; the C2 part includes four Fused-MBConv modules with a convolution kernel size of 3×3 and a stride of 2 connected thereto; the C3 part includes six MBConv modules with a convolution kernel size of 3×3 and a stride of 2 connected thereto; the C4 part includes five MBConv modules with a convolution kernel size of 3×3 and a stride of 1 connected thereto; the five parts correspond to feature maps with output sizes of 128×128, 64×64, 32×32, 16×16, and 8×8, respectively, and the output of the previous backbone module is used as the input of the next backbone module;
[0016] When the MobileNetV2 network is used, the five backbone modules connected from bottom to top are as follows: the C0 part includes a convolution layer with a convolution kernel size of 3×3 and a stride of 2, a convolution layer with a convolution kernel size of 3×3 and a stride of 1, and a Bottleneck residual block module with a convolution kernel size of 3×3 and a stride of 1; the C1 part includes two Bottleneck residualblock modules with a convolution kernel size of 3×3, a stride of 2 and a stride of 1 connected to it; the C2 part includes a Bottleneck residual block module with a convolution kernel size of 3×3 and a stride of 2 connected to it, and two Bottleneck residual block modules with a convolution kernel size of 3×3 and a stride of 1; the C3 part includes a MBConv module with a convolution kernel size of 3×3 and a stride of 2 connected to it, and three Bottleneck residual block module; the C4 part includes a MBConv module with a convolution kernel size of 3×3 and a stride of 2 and four Bottleneck residual block modules with a convolution kernel size of 3×3 and a stride of 1 connected in sequence; the five parts correspond to feature maps with output sizes of 128×128, 64×64, 32×32, 16×16, and 8×8 respectively; and the output of the previous backbone module is used as the input of the next backbone module;
[0017] The lateral connections after the five backbone modules are:
[0018] The feature map output by the backbone module C0 part passes through a convolution layer with a convolution kernel size of 1×1 and a stride of 1, and the output size is a feature map of 128×128 and 64 channels;
[0019] The feature map output by the backbone module C1 passes through a convolution layer with a convolution kernel size of 1×1 and a stride of 1, and the output size is 64×64 and the number of channels is 128.
[0020] The feature map output by the backbone module C2 passes through a convolution layer with a convolution kernel size of 1×1 and a stride of 1, and the output size is 32×32 and the number of channels is 128.
[0021] The feature map output by the backbone module C3 passes through a convolution layer with a convolution kernel size of 1×1 and a stride of 1, and the output size is 16×16 with 128 channels.
[0022] The feature map output by the backbone module C4 passes through a convolution layer with a convolution kernel size of 1×1 and a stride of 1, and the output size is 8×8 and the number of channels is 128.
[0023] The operation of feature fusion module A0 is as follows:
[0024] The 8×8 feature map output by the backbone module C4 through the horizontal connection is upsampled by 2 times, and then the corresponding elements are added to the 16×16 feature map output by the backbone module C3 through the horizontal connection, and then a convolution layer with a convolution kernel size of 1×1 and a stride of 1 is passed to obtain the output feature map;
[0025] The specific operation of feature fusion module A1 is as follows:
[0026] The feature map output by the feature fusion module A0 is upsampled by 2 times, and then the corresponding elements are added to the 16×16 feature map output by the backbone module C2 through the horizontal connection, and then passed through a convolution layer with a convolution kernel size of 1×1 and a step size of 1 to obtain the output feature map;
[0027] The specific operation of feature fusion module A2 is as follows:
[0028] The feature map output by the feature fusion module A1 is upsampled by 2 times, and then the corresponding pixels are added to the 16×16 feature map output by the C3 part after horizontal connection, and then a convolution layer with a convolution kernel size of 1×1 and a step size of 1 is passed to obtain the output feature map;
[0029] The channel series operation is as follows:
[0030] The feature map output by the backbone module C4 through the lateral connection part is sequentially passed through two convolution layers with a convolution kernel size of 3×3 and a step size of 1, and a 2-fold upsampling layer to obtain a feature map output; the feature map output by the feature fusion module A0 is sequentially passed through two convolution layers with a convolution kernel size of 3×3 and a step size of 1, and a 4-fold upsampling layer to obtain a feature map output; the feature map output by the feature fusion module A1 is sequentially passed through two convolution layers with a convolution kernel size of 3×3 and a step size of 1, and a 6-fold upsampling layer to obtain a feature map output; the feature map output by the feature fusion module A2 is sequentially passed through two convolution layers with a convolution kernel size of 3×3 and a step size of 1, and an 8-fold upsampling layer to obtain a feature map output; the above four feature map outputs are all output feature maps with a size of 64×64 and a channel number of 128, and then these four feature maps of the same size are channel-connected to finally generate a feature map with a size of 64×64 and a channel number of 256;
[0031] The reconstructed image obtained by upsampling the feature map is as follows:
[0032] The feature map obtained by the channel series operation passes through a convolution layer with a convolution kernel size of 3×3 and a stride of 1, a 2x upsampling layer, and then performs a corresponding pixel addition operation on the feature map output by the C0 part through a horizontal connection. Then it passes through a convolution layer with a convolution kernel size of 3×3 and a stride of 1, a 2x upsampling layer, two convolution layers with a convolution kernel size of 3×3 and a stride of 1, and a tanh activation layer, and finally restores to the resolution size of the original image to obtain the reconstructed image.
[0033] Except for the backbone module, all other parts of the generator use Instance Normalization as the normalization method, and the activation function is ReLU.
[0034] The discriminator D is a patch-level discriminator, which includes three convolution layers with a kernel size of 4×4 and a stride of 2 connected in sequence to compress the image size of the input discriminator D to 32×32, and the number of channels gradually changes from 3 to 256. Then, two convolution layers with a kernel size of 4×4 and a stride of 1 are used to change the number of feature channels from 256 to 1, and the final discrimination result is output. The first four convolution layers of the discriminator D are connected with a LeakyRelu activation function, and there is an Instance Normalization between the second to fourth convolution layers and their activation functions.
[0035] In step 2, Gaussian noise is added to the defect-free colored fabric image according to formula (1):
[0036]
[0037] Where X is a defect-free colored fabric image, N(0,1) is a Gaussian noise with a standard normal distribution of mean 0 and standard deviation 1, c is the ratio of superimposed noise, c is 0.2, This is the image of the defect-free colored fabric with noise superimposed on it.
[0038] The total loss function of the EFFGAN model training process in step 2 is as follows:
[0039] L total =0.5·L x +0.01 L content +0.01 L adv (5)
[0040] Among them, L x is the pixel-level loss, L content is the content loss, L adv To combat losses;
[0041] Pixel level loss L x , content loss L content , against the loss L adv As shown in formula (2), (3), and (4) respectively:
[0042]
[0043]
[0044]
[0045] Where X(i) is the image of defect-free colored fabric with noise added. and Respectively represent the results after being processed by the generator and the discriminator, is the color fabric reconstruction image output by the EFFGAN model, n is the number of training samples, and are the weights and biases during training; It is the feature map obtained after the jth convolution activation before the i-th maximum pooling layer when the VGG19 network is pre-trained on ImageNet. W and H represent the length and width of the feature map respectively.
[0046] The training in step 2 is to minimize L total To optimize the model parameters, the Adam optimizer was used with a learning rate of 0.0001, and the maximum number of training iterations was set to be no less than the number of samples in the training set of defect-free images of colored fabrics to obtain the trained EFFGAN model.
[0047] In step 3, the defective area is determined to be:
[0048] Step 3.1, grayscale the color fabric image to be detected and its corresponding reconstructed image. The specific operation is shown in formula (6):
[0049] X gray =0.2125X r +0.7154X g +0.0721X b (6)
[0050] Where: X gray is the grayscale image of the color fabric to be detected or its corresponding reconstructed image; X r , X g , X b are the pixel values of the three different color channels of the color fabric image to be detected or its corresponding reconstructed image, and the pixel value range of the grayscale image is 0 to 255;
[0051] Step 3.2: The grayscale image of the color fabric image to be detected or its corresponding reconstructed image is subjected to Gaussian filtering by sliding window convolution operation using a Gaussian kernel of size 3×3 to obtain a filtered image. The specific operation is shown in formula (9):
[0052]
[0053] in, is the image of the color fabric to be detected or the corresponding reconstructed image after Gaussian filtering, G(x,y) is the Gaussian kernel function, (x, y) is the pixel coordinate of the color fabric image to be detected or the grayscale image of the reconstructed image, σ x , σ y are the pixel standard deviations of the color fabric image to be detected or the grayscale image of the reconstructed image in the x-axis and y-axis directions respectively;
[0054] Step 3.3, calculate the difference between the color fabric image to be detected and the corresponding reconstructed image after Gaussian filtering in step 3.2, and obtain the residual image. The residual image is obtained specifically according to formula (8):
[0055]
[0056] Where, X res is the residual image, X gray&Gaussian , are respectively an image of the color fabric image to be detected after being subjected to Gaussian filtering and an image of the reconstructed image after being subjected to Gaussian filtering;
[0057] Step 3.4: Binarize the residual image obtained in step 3.3 using the adaptive threshold method. The binarization operation is shown in formula (9):
[0058]
[0059] Where f(p) is the binarized value, p is the pixel value of the residual image, T is the adaptive threshold of the residual image, μ is the mean of the residual image, σ is the standard deviation of the residual image, and γ is the coefficient of the standard deviation;
[0060] Step 3.5, the binarized residual image is closed, and the closing operation is shown in formula (10):
[0061]
[0062] Where, X binary is the binary image obtained after the residual image is binarized, E is the 3×3 closed operation structure element, is the image dilation operation, ! is the image erosion operation, X closing is the final detection result image;
[0063] Step 3.6, analyze the value of each pixel in the final test result image to determine whether there is a defective area. If there is no difference in the test result image, that is, the pixel values in the image are all 0, it means that the input color fabric has no defects; if there are two pixel values 0 and 1 in the test result image, it means that the input color fabric image has defects, and the defective area is the area with a pixel value of 1.
[0064] The beneficial effects of the present invention are:
[0065] The present invention constructs an unsupervised color fabric image reconstruction and repair model EFFGAN model, and uses the constructed database to train the model. The trained model obtains the color fabric image reconstruction and repair capability, so that when detecting the color fabric image to be tested, by analyzing the residual image between the original color fabric image to be tested and the reconstructed image, the generator combines the efficient fusion framework FPN with two advanced lightweight backbones, the EfficientNetV2 network and the MobileNetV2 network, to construct a lightweight model to improve the detection efficiency. The FPN structure fuses the features containing more position and color information output by the shallow convolution with the features containing rich semantic information output by the deep convolution, thereby avoiding the loss of information and making the reconstruction result more realistic. Therefore, the present invention can quickly and accurately detect the defects of the color fabric. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is a flow chart of step 3 in a method for detecting defective areas of colored fabrics based on a generative adversarial network according to the present invention;
[0067] Figure 2It is a structural diagram of an EFFGAN model generator G in a method for detecting defective areas of colored fabrics based on a generative adversarial network according to the present invention;
[0068] Figure 3 It is a structural diagram of the MB Conv layer, the Fused-MB Conv layer and the Bottleneck residual block of the EFFGAN model in a method for detecting defective areas of colored fabrics based on a generative adversarial network of the present invention;
[0069] Figure 4 It is a structural diagram of the EFFGAN model discriminator D in a method for detecting defective areas of colored fabrics based on a generative adversarial network according to the present invention;
[0070] Figure 5 It is a part of non-defective samples in the experimental samples in the method for detecting defective areas of colored fabrics based on a generative adversarial network of the present invention;
[0071] Figure 6 It is a partial defect sample in the experimental sample in the method for detecting defective areas of colored fabrics based on a generative adversarial network of the present invention;
[0072] Figure 7 This is a comparison chart of the detection results of the EFFGAN model used in the experiment in the detection method of the defective area of the colored fabric based on the generative adversarial network of the present invention and the DCGAN, DCAE, MSDCAE, UDCAE and VAE_L2SSIM models. DETAILED DESCRIPTION
[0073] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0074] The present invention provides a method for detecting defective areas of colored fabrics based on a generative adversarial network, which is specifically implemented according to the following steps:
[0075] Step 1: Construct an image reconstruction and restoration model EFFGAN based on unsupervised learning. The model consists of two parts: the generator G and the discriminator D.
[0076] like Figure 2-3 As shown, the input layer and output layer of the generator G are both three-channel image structures, and the generator uses the feature pyramid structure FPN as the feature extractor.
[0077] The feature pyramid structure FPN includes five backbone modules for feature extraction connected from bottom to top, denoted as C0, C1, C2, C3, and C4, and three feature fusion modules connected from top to bottom, denoted as A0, A1, and A2. The backbone modules C4 and C3 are also connected to the feature fusion module A0 through the lateral connection part, and the backbone modules C2 and C1 are also connected to A1 and A2 one by one through the lateral connection part. The outputs of the three feature fusion modules and the outputs of the top backbone module C4 and the bottom backbone module C0 after lateral connection are respectively obtained through channel series operation to obtain feature maps, and then the feature maps are up-sampled to obtain reconstructed images.
[0078] The backbone module is composed of the MBConv and Fused-MBConv structures of the EfficientNetV2 network, or the Bottleneck residual block structure of the MobileNetV2 network;
[0079] When the EfficientNetV2 network is used, the five backbone modules connected from bottom to top are as follows: the C0 part includes a convolution layer with a convolution kernel size of 3×3 and a stride of 2 and two Fused-MBConv modules with a convolution kernel size of 3×3 and a stride of 1 connected thereto; the C1 part includes four Fused-MBConv modules with a convolution kernel size of 3×3 and a stride of 2 connected thereto; the C2 part includes four Fused-MBConv modules with a convolution kernel size of 3×3 and a stride of 2 connected thereto; the C3 part includes six MBConv modules with a convolution kernel size of 3×3 and a stride of 2 connected thereto; the C4 part includes five MBConv modules with a convolution kernel size of 3×3 and a stride of 1 connected thereto; the five parts correspond to feature maps with output sizes of 128×128, 64×64, 32×32, 16×16, and 8×8, respectively, and the output of the previous backbone module is used as the input of the next backbone module;
[0080] When the MobileNetV2 network is used, the five backbone modules connected from bottom to top are as follows: the C0 part includes a convolution layer with a convolution kernel size of 3×3 and a stride of 2, a convolution layer with a convolution kernel size of 3×3 and a stride of 1, and a Bottleneck residual block module with a convolution kernel size of 3×3 and a stride of 1; the C1 part includes two Bottleneck residualblock modules with a convolution kernel size of 3×3, a stride of 2 and a stride of 1 connected to it; the C2 part includes a Bottleneck residual block module with a convolution kernel size of 3×3 and a stride of 2 connected to it, and two Bottleneck residual block modules with a convolution kernel size of 3×3 and a stride of 1; the C3 part includes a MBConv module with a convolution kernel size of 3×3 and a stride of 2 connected to it, and three Bottleneck residual block module; the C4 part includes a MBConv module with a convolution kernel size of 3×3 and a stride of 2 and four Bottleneck residual block modules with a convolution kernel size of 3×3 and a stride of 1 connected in sequence; the five parts correspond to feature maps with output sizes of 128×128, 64×64, 32×32, 16×16, and 8×8 respectively; and the output of the previous backbone module is used as the input of the next backbone module;
[0081] The lateral connections after the five backbone modules are:
[0082] The feature map output by the backbone module C0 passes through a convolution layer with a convolution kernel size of 1×1 and a stride of 1, and the output size is 128×128 and the number of channels is 64.
[0083] The feature map output by the backbone module C1 passes through a convolution layer with a convolution kernel size of 1×1 and a stride of 1, and the output size is 64×64 and the number of channels is 128.
[0084] The feature map output by the backbone module C2 passes through a convolution layer with a convolution kernel size of 1×1 and a stride of 1, and the output size is 32×32 and the number of channels is 128.
[0085] The feature map output by the backbone module C3 passes through a convolution layer with a convolution kernel size of 1×1 and a stride of 1, and the output size is 16×16 with 128 channels.
[0086] The feature map output by the backbone module C4 passes through a convolution layer with a convolution kernel size of 1×1 and a stride of 1, and the output size is 8×8 and the number of channels is 128.
[0087] The operation of feature fusion module A0 is as follows:
[0088] The 8×8 feature map output by the backbone module C4 through the horizontal connection is upsampled by 2 times, and then the corresponding elements are added to the 16×16 feature map output by the backbone module C3 through the horizontal connection, and then a convolution layer with a convolution kernel size of 1×1 and a stride of 1 is passed to obtain the output feature map;
[0089] The specific operation of feature fusion module A1 is as follows:
[0090] The feature map output by the feature fusion module A0 is upsampled by 2 times, and then the corresponding elements are added to the 16×16 feature map output by the backbone module C2 through the horizontal connection, and then passed through a convolution layer with a convolution kernel size of 1×1 and a step size of 1 to obtain the output feature map;
[0091] The specific operation of feature fusion module A2 is as follows:
[0092] The feature map output by the feature fusion module A1 is upsampled by 2 times, and then the corresponding pixels are added to the 16×16 feature map output by the C3 part after horizontal connection, and then a convolution layer with a convolution kernel size of 1×1 and a step size of 1 is passed to obtain the output feature map;
[0093] The channel series operation is as follows:
[0094] The feature map output by the backbone module C4 through the lateral connection part is sequentially passed through two convolution layers with a convolution kernel size of 3×3 and a step size of 1, and a 2-fold upsampling layer to obtain a feature map output; the feature map output by the feature fusion module A0 is sequentially passed through two convolution layers with a convolution kernel size of 3×3 and a step size of 1, and a 4-fold upsampling layer to obtain a feature map output; the feature map output by the feature fusion module A1 is sequentially passed through two convolution layers with a convolution kernel size of 3×3 and a step size of 1, and a 6-fold upsampling layer to obtain a feature map output; the feature map output by the feature fusion module A2 is sequentially passed through two convolution layers with a convolution kernel size of 3×3 and a step size of 1, and an 8-fold upsampling layer to obtain a feature map output; the above four feature map outputs are all output feature maps with a size of 64×64 and a channel number of 128, and then these four feature maps of the same size are channel-connected to finally generate a feature map with a size of 64×64 and a channel number of 256;
[0095] The reconstructed image obtained by upsampling the feature map is as follows:
[0096] The feature map obtained by the channel series operation passes through a convolution layer with a convolution kernel size of 3×3 and a stride of 1, a 2x upsampling layer, and then performs a corresponding pixel addition operation on the feature map output by the C0 part through a horizontal connection. Then it passes through a convolution layer with a convolution kernel size of 3×3 and a stride of 1, a 2x upsampling layer, two convolution layers with a convolution kernel size of 3×3 and a stride of 1, and a tanh activation layer, and finally restores to the resolution size of the original image to obtain the reconstructed image.
[0097] Except for the backbone module, all other parts of the generator use Instance Normalization as the normalization method, and the activation function is ReLU.
[0098] like Figure 4 As shown in the figure, the discriminator D is a patch-level discriminator, including three sequentially connected convolution layers with a kernel size of 4×4 and a stride of 2, which compress the image size of the input discriminator D to 32×32 and gradually change the number of channels from 3 to 256. Then, two convolution layers with a kernel size of 4×4 and a stride of 1 are used to change the number of feature channels from 256 to 1, and the final discrimination result is output. The first four convolution layers of the discriminator D are all connected with a LeakyRelu activation function, and there is an Instance Normalization between the second to fourth convolution layers and their activation functions. The input of the discriminator has two parts: the reconstructed image output by the generator and the original image of the defect-free colored fabric without noise for training.
[0099] The discriminator compares the difference between the image generated by the generator and the original image, and converts the difference between them into an adversarial loss function L adv Feedback to the generator.
[0100] Step 2, construct a training set of defect-free colored fabric images including defect-free colored fabric images, then superimpose Gaussian noise on the defect-free colored fabric images, and send the defect-free colored fabric images after superimposing noise into the EFFGAN model constructed in step 1, the generator G extracts and restores features of the input image through encoding and decoding operations, the discriminator D continuously adjusts the gradient feedback to the generator G, guides the training of the generator G, and when the number of training times reaches the set number of iterations, the trained EFFGAN model is obtained;
[0101] Among them, the Gaussian noise is added to the defect-free colored fabric image according to formula (1):
[0102]
[0103] Where X is a defect-free colored fabric image, N(0,1) is a Gaussian noise with a standard normal distribution of mean 0 and standard deviation 1, c is the ratio of superimposed noise, c is 0.2, This is the image of the defect-free colored fabric with noise superimposed on it.
[0104] The total loss function during the training process of the EFFGAN model is as follows:
[0105] L total =0.5·L x +0.01 L content +0.01 L adv (5)
[0106] Among them, L x is the pixel-level loss, L content is the content loss, L adv To combat losses;
[0107] Pixel level loss L x , content loss L content , against the loss L adv As shown in formula (2), (3), and (4) respectively:
[0108]
[0109]
[0110]
[0111] Where X(i) is the image of defect-free colored fabric with noise added. and Respectively represent the results after being processed by the generator and the discriminator, is the color fabric reconstruction image output by the EFFGAN model, n is the number of training samples, and are the weights and biases during training; It is the feature map obtained after the jth convolution activation before the i-th maximum pooling layer when the VGG19 network is pre-trained on ImageNet. W and H represent the length and width of the feature map respectively.
[0112] Train to minimize L total To optimize the model parameters, the Adam optimizer was used with a learning rate of 0.0001, and the maximum number of training iterations was set to be no less than the number of samples in the training set of defect-free images of colored fabrics to obtain the trained EFFGAN model.
[0113] Step 3: Input the color fabric image to be tested into the EFFGAN model trained in step 2 to output the corresponding reconstructed image, and then perform the test to determine the defect area. The process is as follows: Figure 1 As shown, specifically:
[0114] Step 3.1, grayscale the color fabric image to be detected and its corresponding reconstructed image. The specific operation is shown in formula (6):
[0115] X gray =0.2125X r +0.7154X g +0.0721X b (6)
[0116] Where: X gray is the grayscale image of the color fabric to be detected or its corresponding reconstructed image; X r , X g , X b are the pixel values of the three different color channels of the color fabric image to be detected or its corresponding reconstructed image, and the pixel value range of the grayscale image is 0 to 255;
[0117] Step 3.2: The grayscale image of the color fabric image to be detected or its corresponding reconstructed image is subjected to Gaussian filtering by sliding window convolution operation using a Gaussian kernel of size 3×3 to obtain a filtered image. The specific operation is shown in formula (9):
[0118]
[0119] in, is the image of the color fabric to be detected or the corresponding reconstructed image after Gaussian filtering, G(x,y) is the Gaussian kernel function, (x, y) is the pixel coordinate of the color fabric image to be detected or the grayscale image of the reconstructed image, σ x , σ y are the pixel standard deviations of the color fabric image to be detected or the grayscale image of the reconstructed image in the x-axis and y-axis directions respectively;
[0120] Step 3.3, calculate the difference between the color fabric image to be detected and the corresponding reconstructed image after Gaussian filtering in step 3.2, and obtain the residual image. The residual image is obtained specifically according to formula (8):
[0121]
[0122] Where, X res is the residual image, X gray&Gaussian , are respectively an image of the color fabric image to be detected after being subjected to Gaussian filtering and an image of the reconstructed image after being subjected to Gaussian filtering;
[0123] Step 3.4: Binarize the residual image obtained in step 3.3 using the adaptive threshold method. The binarization operation is shown in formula (9):
[0124]
[0125] Where f(p) is the binarized value, p is the pixel value of the residual image, T is the adaptive threshold of the residual image, μ is the mean of the residual image, σ is the standard deviation of the residual image, and γ is the coefficient of the standard deviation. In the experiment, γ=3 is selected;
[0126] Step 3.5, the binarized residual image is closed, and the closing operation is shown in formula (10):
[0127]
[0128] In the formula, X binary is the binary image obtained after the residual image is binarized, E is the 3×3 closed operation structure element, is the image dilation operation, ! is the image erosion operation, X closing is the final detection result image;
[0129] Step 3.6, analyze the value of each pixel in the final test result image to determine whether there is a defective area. If there is no difference in the test result image, that is, the pixel values in the image are all 0, it means that the input color fabric has no defects; if there are two pixel values 0 and 1 in the test result image, it means that the input color fabric image has defects, and the defective area is the area with a pixel value of 1.
[0130] The following is a specific embodiment of a method for detecting defective areas of a colored fabric based on a generative adversarial network according to the present invention:
[0131] Experimental device preparation: The hardware configuration of the modeling, training and defect detection experiments of the EFFGAN model in the present invention is as follows: the hardware environment is Intel(R) Core(TM) i7-6850K CPU@3.60GHz; GeForce RTX 3090(24G) GPU; memory 128G. The software configuration is as follows: the operating system is Ubuntu 18.04, CUDA11.2, cuDNN8.2.0, Python3.8.5, Pytorch1.7.0.
[0132] Samples to be tested in the experiment: The data set used in the experiment comes from Guangdong Esquel Textile Co., Ltd. It can be divided into three types according to the complexity of the pattern: Simple Lattices (SL1~SL19), Stripe Patterns (SP1~SP26), and Complex Lattices (CL1~CL21), containing a total of 66 samples of colored fabrics with different patterns. This paper selects ten data sets with different patterns from the three types of data sets for training and testing, namely SL8, SL9, SL10, SL11, SP3, SP5, SP19, SP24, CL2, and CL3. Some samples of each data set are as follows Figure 5 , 6 As shown, Figure 5 This is a sample of the colored fabric without defects, attached Figure 6 This is a sample of some defects in the colored fabric.
[0133] Experimental evaluation indicators: The detection result images were analyzed qualitatively and quantitatively. Qualitative analysis is an intuitive illustration of the defect detection area. Quantitative analysis uses five indicators to evaluate the model: precision (Precision, P), recall (Recall, R), F1-measure (F1), accuracy (Accuracy, Acc) and average intersection over union (IoU). Among them, the definitions of precision, F1-measure, recall, accuracy, and average intersection over union are shown in formulas (11), (12), (13), (14), and (15), respectively:
[0134]
[0135]
[0136]
[0137]
[0138]
[0139] In the formula, TP means that the positive sample is predicted to be positive, TN means that the positive sample is predicted to be negative, FP means that the negative sample is predicted to be positive, and FN means that the negative sample is predicted to be negative.
[0140] Experimental process: First, an unsupervised color fabric reconstruction and repair model EFFGAN model is established; then, the model is trained using defect-free color fabric samples, and the trained model has the ability to reconstruct and repair; finally, when detecting the color fabric image to be tested, the residual image between the original color fabric image to be tested and the reconstructed and repaired color woven shirt piece image is analyzed to achieve rapid detection of color fabric defect areas.
[0141] Qualitative analysis of experimental results: In this experiment, the EFFGAN model is trained with defect-free colored fabric images. The trained EFFGAN model has the ability to reconstruct and repair colored fabric images. Finally, the residual image between the colored fabric image to be tested and the reconstructed image is calculated, and the defect area is detected and located by residual analysis. In order to more intuitively compare the detection results of different unsupervised detection methods, the EFFGAN proposed in this application is experimentally compared with five unsupervised detection methods including DCGAN, DCAE, MSDCAE, UDCAE and VAE_L2SSIM. The experimental results are shown in the attached figure. Figure 7 As shown, through the attached Figure 7 It can be seen that the EFFGAN model of the present application can well repair the defective areas in the color fabric image on the basis of accurately restoring the color fabric images of different patterns. By comparing with the ground truth of the defective area, it can be seen that the EFFGAN model of the present application has good detection results for various patterns.
[0142] Quantitative analysis of experimental results: Through experiments, the precision (P), recall (R), F1-measure (F1), accuracy (Acc) and average intersection over union (IoU) of the defect image detection results of the EFFGAN model for ten color fabric data sets are compared. The larger the value of the evaluation index, the better the detection result. The results are shown in Table 1:
[0143] Table 1. EFFGAN model detection results under five evaluation indicators
[0144]
[0145] Experimental summary: The present invention is essentially an unsupervised modeling method based on the EFFGAN model. By calculating the residual between the fabric image to be tested and the model reconstructed image, mathematical morphological analysis is performed to achieve defect detection and positioning of colored fabrics. This method uses defect-free samples to establish an unsupervised EFFGAN model, which can effectively avoid practical problems such as the scarcity of defective samples, the high cost of annotating large-scale data, and the poor generalization ability of artificially designed defect features. At the same time, the method proposed in the present invention can meet the process requirements of colored fabric production in terms of detection accuracy, and provides an automated defect detection solution that is easy to practice in engineering for the defect detection process of the colored fabric manufacturing industry.
Claims
1. A detection method for defective areas of colored woven fabrics based on a generative adversarial network, characterized in that, it is specifically implemented according to the following steps: Step 1, construct an image reconstruction and repair model EFFGAN model based on unsupervised learning. This model consists of two parts: a generator G and a discriminator D. The input layer and output layer of the generator G are both three-channel image structures, and the generator uses a Feature Pyramid Network (FPN) as a feature extractor; The Feature Pyramid Network (FPN) includes 5 backbone modules for feature extraction connected sequentially from bottom to top, denoted as C0, C1, C2, C3, C4 respectively, and 3 feature fusion modules connected sequentially from top to bottom, denoted as A0, A1, A2 respectively. The backbone modules C4 and C3 are also respectively connected to the feature fusion module A0 through lateral connection parts. The backbone modules C2 and C1 are also respectively connected to A1 and A2 in one-to-one correspondence through lateral connection parts. The outputs of the three feature fusion modules, as well as the outputs of the topmost backbone module C4 and the bottommost backbone module C0 after lateral connection, are obtained through channel concatenation operation to get a feature map, and then the feature map is upsampled to obtain a reconstructed image; The backbone module is composed of the MBConv and Fused-MBConv structures of the EfficientNetV2 network, or is composed of the Bottleneck residual block structure of the MobileNetV2 network; When using the EfficientNetV2 network, the 5 backbone modules connected sequentially from bottom to top are specifically: The C0 part includes a convolutional layer with a kernel size of 3×3 and a stride of 2 connected sequentially, and two Fused-MBConv modules with a kernel size of 3×3 and a stride of 1; The C1 part includes four Fused-MBConv modules with a kernel size of 3×3 and a stride of 2 connected sequentially; The C2 part includes four Fused-MBConv modules with a kernel size of 3×3 and a stride of 2 connected sequentially; The C3 part includes six MBConv modules with a kernel size of 3×3 and a stride of 2 connected sequentially; The C4 part includes five MBConv modules with a kernel size of 3×3 and a stride of 1 connected sequentially; The five parts respectively output feature maps with sizes of 128×128, 64×64, 32×32, 16×16, 8×8, and the output of the previous backbone module is used as the input of the next backbone module; When the MobileNetV2 network is used, the five backbone modules connected from bottom to top are as follows: the C0 part includes a convolution layer with a convolution kernel size of 3×3 and a stride of 2, a convolution layer with a convolution kernel size of 3×3 and a stride of 1, and a Bottleneck residual block module with a convolution kernel size of 3×3 and a stride of 1; the C1 part includes two Bottleneck residual block modules with a convolution kernel size of 3×3, a stride of 2 and a stride of 1 connected to it; the C2 part includes a Bottleneck residualblock module with a convolution kernel size of 3×3 and a stride of 2 connected to it, and two Bottleneck residual block modules with a convolution kernel size of 3×3 and a stride of 1; the C3 part includes a MBConv module with a convolution kernel size of 3×3 and a stride of 2, and three Bottleneck residual block module; the C4 part includes a MBConv module with a convolution kernel size of 3×3 and a stride of 2 and four Bottleneck residualblock modules with a convolution kernel size of 3×3 and a stride of 1 connected in sequence; the five parts correspond to feature maps with output sizes of 128×128, 64×64, 32×32, 16×16, and 8×8 respectively; and the output of the previous backbone module is used as the input of the next backbone module; The lateral connection parts after the five backbone modules are respectively: The feature map output by the backbone module C0 part passes through a convolution layer with a convolution kernel size of 1×1 and a step size of 1, and the output size is a feature map of 128×128 and 64 channels; The feature map output by the backbone module C1 part passes through a convolution layer with a convolution kernel size of 1×1 and a step size of 1, and the output size is a feature map of 64×64 and 128 channels; The feature map output by the backbone module C2 part passes through a convolution layer with a convolution kernel size of 1×1 and a step size of 1, and the output size is a feature map of 32×32 and 128 channels; The feature map output by the backbone module C3 part passes through a convolution layer with a convolution kernel size of 1×1 and a step size of 1, and the output size is a feature map of 16×16 and 128 channels; The feature map output by the backbone module C4 part passes through a convolution layer with a convolution kernel size of 1×1 and a step size of 1, and the output size is 8×8 and the number of channels is 128. The operation of the feature fusion module A0 is specifically as follows: The 8×8 feature map output by the backbone module C4 through the horizontal connection is upsampled by 2 times, and then the corresponding elements are added to the 16×16 feature map output by the backbone module C3 through the horizontal connection, and then a convolution layer with a convolution kernel size of 1×1 and a stride of 1 is passed to obtain the output feature map; The operation of the feature fusion module A1 is specifically as follows: The feature map output by the feature fusion module A0 is upsampled by 2 times, and then the corresponding elements are added to the 16×16 feature map output by the backbone module C2 through the horizontal connection, and then passed through a convolution layer with a convolution kernel size of 1×1 and a step size of 1 to obtain the output feature map; The operation of the feature fusion module A2 is specifically as follows: The feature map output by the feature fusion module A1 is upsampled by 2 times, and then the corresponding pixels are added to the 16×16 feature map output by the C3 part after horizontal connection, and then a convolution layer with a convolution kernel size of 1×1 and a step size of 1 is passed to obtain the output feature map; The channel series operation is as follows: The feature map output by the backbone module C4 through the lateral connection part is sequentially passed through two convolution layers with a convolution kernel size of 3×3 and a step size of 1, and a 2-fold upsampling layer to obtain a feature map output; the feature map output by the feature fusion module A0 is sequentially passed through two convolution layers with a convolution kernel size of 3×3 and a step size of 1, and a 4-fold upsampling layer to obtain a feature map output; the feature map output by the feature fusion module A1 is sequentially passed through two convolution layers with a convolution kernel size of 3×3 and a step size of 1, and a 6-fold upsampling layer to obtain a feature map output; the feature map output by the feature fusion module A2 is sequentially passed through two convolution layers with a convolution kernel size of 3×3 and a step size of 1, and an 8-fold upsampling layer to obtain a feature map output; the above four feature map outputs are all output feature maps with a size of 64×64 and a channel number of 128, and then these four feature maps of the same size are channel-connected to finally generate a feature map with a size of 64×64 and a channel number of 256; The reconstructed image obtained by upsampling the feature map is as follows: The feature map obtained by the channel series operation is sequentially passed through a convolution layer with a convolution kernel size of 3×3 and a stride of 1, a 2x upsampling layer, and then the corresponding pixels are added with the feature map output by the C0 part through the horizontal connection. Then, it is sequentially passed through a convolution layer with a convolution kernel size of 3×3 and a stride of 1, a 2x upsampling layer, two convolution layers with a convolution kernel size of 3×3 and a stride of 1, and a tanh activation layer, and finally restored to the resolution size of the original image to obtain the reconstructed image; The generator uses Instance Normalization as the normalization method except for the backbone module, and the activation function is ReLU; The discriminator D is a patch-level discriminator, including three sequentially connected convolution layers with a convolution kernel size of 4×4 and a step size of 2, which compress the image size of the input discriminator D to 32×32, and gradually change the number of channels from 3 to 256, and then change the number of feature channels from 256 to 1 through two convolution layers with a convolution kernel size of 4×4 and a step size of 1, and output the final discrimination result. The first four convolution layers of the discriminator D are all connected with a LeakyRelu activation function, and there is an Instance Normalization between the second to fourth convolution layers and their activation functions; Step 2, construct a training set of defect-free colored fabric images including defect-free colored fabric images, then superimpose Gaussian noise on the defect-free colored fabric images, and send the defect-free colored fabric images after superimposing noise into the EFFGAN model constructed in step 1, the generator G extracts and restores features of the input image through encoding and decoding operations, the discriminator D continuously adjusts the gradient feedback to the generator G, guides the training of the generator G, and when the number of training times reaches the set number of iterations, the trained EFFGAN model is obtained; Step 3: Input the color fabric image to be detected into the EFFGAN model trained in step 2 to output the corresponding reconstructed image, and then perform detection to determine the defective area.
2. A method for detecting defective areas of colored fabrics based on a generative adversarial network according to claim 1, It is characterized in that In step 2, Gaussian noise is added to the defect-free colored fabric image according to formula (1): (1) In the formula, For defect-free colored fabric images, is Gaussian noise that follows a standard normal distribution with a mean of 0 and a standard deviation of 1. To express the ratio of superimposed noise, is 0.2, This is the image of the defect-free colored fabric with noise superimposed on it.
3. A method for detecting defective areas of colored fabrics based on a generative adversarial network according to claim 2, It is characterized in that The total loss function of the EFFGAN model training process in step 2 is as follows: (5) in, is the pixel-level loss, For content loss, To combat losses; Pixel level loss , content loss , Fighting Losses As shown in formula (2), (3), and (4) respectively: (2) (3) (4) In the formula, This is the image of defect-free colored fabric with noise superimposed on it. ( )and ( ) represent the results obtained after being processed by the generator and the discriminator, respectively. This is the reconstruction image of the colored fabric output by the EFFGAN model. is the number of training samples, and are the weights and biases during training; When the VGG19 network is pre-trained on ImageNet, Before the maximum pooling layer, The feature map obtained after the convolution activation is and Represent the length and width of the feature map respectively.
4. The method for detecting defective areas of colored fabrics based on a generative adversarial network according to claim 3, It is characterized in that The training in step 2 is to minimize To optimize the model parameters, the Adam optimizer was used with a learning rate of 0.0001, and the maximum number of training iterations was set to be no less than the number of samples in the training set of defect-free images of colored fabrics to obtain the trained EFFGAN model.
5. The method for detecting defective areas of colored fabrics based on a generative adversarial network according to claim 3, It is characterized in that The defective area determined by the detection in step 3 is specifically: Step 3.1: grayscale the color fabric image to be detected and its corresponding reconstructed image. The specific operation is shown in formula (6): (6) Where: is the grayscale image of the color fabric to be detected or its corresponding reconstructed image; , , are the pixel values of the three different color channels of the color fabric image to be detected or its corresponding reconstructed image, and the pixel value range of the grayscale image is 0 to 255; Step 3.2: The grayscale image of the color fabric to be detected or its corresponding reconstructed image is converted into The Gaussian kernel of the size is subjected to sliding window convolution operation to perform Gaussian filtering to obtain the filtered image. The specific operation is shown in formula (7): = (7) in, is the image of the colored fabric to be detected or the corresponding reconstructed image after Gaussian filtering, is the Gaussian kernel function, , is the pixel coordinate of the color fabric image to be detected or the grayscale image of the reconstructed image, , are the color fabric image to be detected or the grayscale image of the reconstructed image axis, The pixel standard deviation in the axis direction; Step 3.3, calculate the difference between the color fabric image to be detected and the corresponding reconstructed image after Gaussian filtering in step 3.2, and obtain the residual image. The residual image is obtained according to formula (8): (8) In the formula, is the residual image, , are respectively an image of the color fabric image to be detected after being subjected to Gaussian filtering and an image of the reconstructed image after being subjected to Gaussian filtering; Step 3.4: Binarize the residual image obtained in step 3.3 using the adaptive threshold method. The binarization operation is shown in formula (9): (9) In the formula, is the value after binarization, is the pixel value of the residual image, is the adaptive threshold of the residual image, is the mean of the residual image, is the standard deviation of the residual image, is the coefficient of standard deviation; Step 3.5, the binarized residual image is closed, and the closing operation is shown in formula (10): (10) In the formula, is the binary image obtained after the residual image is binarized, for The closed operation structural element of is the image dilation operation, is the image erosion operation, is the final detection result image; Step 3.6, analyze the value of each pixel in the final test result image to determine whether there is a defective area. If there is no difference in the test result image, that is, the pixel values in the image are all 0, it means that the input color fabric has no defects; if there are two pixel values 0 and 1 in the test result image, it means that the input color fabric image has defects, and the defective area is the area with a pixel value of 1.
Citation Information
Patent Citations
Photovoltaic module unsupervised defect detection method based on GAN improved algorithm
CN111340791A
Cross-domain small sample image classification model method focusing on fine-grained recognition
CN112766378A