A method for enhancing fabric defect data based on generative adversarial networks
By generating an adversarial network, the fabric defect data is enhanced, and the problem of unnatural superposition of defect samples is solved, the performance and robustness of the defect detector are improved, and the data set is expanded.
Patent Information
- Application Number
- CN202210595194.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-05-28
AI Technical Summary
Existing data enhancement methods cannot effectively superimpose defect samples onto new cloth quickly. The small number of defects on the cloth leads to small sample learning problems, and the undefective superposition of defects will naturally reduce detector performance.
Generative adversarial networks are used for data augmentation, multi-scale feature extraction is performed through multi-layer convolution and instance regularization, upsampling module is built in combination with spatial adaptive normalization, parallel defect prospects and transparency generation branches are designed, transparency control and spatial constraints are introduced, and PatchGAN discriminator is used for supervision.
The natural synthesis of defect samples on new cloth is achieved, which improves the performance and robustness of the cloth defect detector and expands the defect data set.
Smart Images

Figure CN115205616B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of general image data processing or generation, belonging to the field of machine vision, and particularly relates to a method for enhancing fabric defect data based on a generative adversarial network. Background Art
[0002] Fabric defect detection is an important link in the fabric production process. Early defect detection was mainly completed by manual screening. In recent years, with the rapid rise of machine learning technology and its large-scale application in the field of computer vision, fabric defect detection has gradually developed towards automation and intelligence.
[0003] Fabric defect detection based on machine learning usually uses object detection networks such as YOLO, SSD, Cascade-RCNN, etc. The deployment of these networks requires a sufficient number of datasets as support. However, in the textile production field, due to the variety of fabric types, the existing datasets cannot well cover all fabric types. Coupled with the difficulty in collecting fabric defect samples and low efficiency, the above factors limit the improvement of fabric defect detection performance and are prone to cause the problem of network overfitting. Therefore, effectively enhancing the collected fabric datasets to improve the generalization and robustness of neural networks is a work with great practical application value.
[0004] Traditional data augmentation methods can be divided into spatial geometric transformation methods, color transformation methods, and multi-sample synthesis methods. Spatial geometric transformations include flipping, rotation, cropping, deformation, and scaling, etc.; color transformations include adding noise, blurring, erasing, color-changing, and filling, etc.; in the multi-sample synthesis method, Chawla et al. (see Chawla N V, Bowyer K W, Hall L O, et al. SMOTE: Synthetic Minority Over-sampling Technique [J]. Journal of Artificial Intelligence Research, 2002, 16(1): 321-357.) aimed at the phenomenon of sample class imbalance and used a method based on feature space interpolation to synthesize new samples for the small-sample class to balance the number of samples. Inoue proposed a simple multi-sample synthesis method SamplePairing (see Inoue H. Data Augmentation by Pairing Samples for Images Classification [J]. arXiv preprint arXiv:1801.02929, 2018.), that is, two images in the training set are first subjected to basic data augmentation and then superimposed and synthesized into a new sample in the form of pixel average values, and the label is any one of the original sample classes. The scale of the dataset processed by this method expands from N to N×N times, and the improvement in performance is quite remarkable. Zhang et al. proposed the Mixup method (see Zhang H, Cisse M, Dauphin Y N, et al. mixup: Beyond Empirical Risk Minimization [J]. 2017.), which uses linear interpolation to obtain new sample data. Experiments show that this method can improve the generalization error of the network in the dataset and enhance the robustness and stability of the network model.
[0005] With the development of deep learning technology, in 2014, Goodfellow et al. proposed the generative adversarial network architecture GAN (see Goodfellow I J, Pouget-Abadie J, Mirza M, et al. Generative Adversarial Networks[J]. Advances in Neural Information Processing Systems, 2014, 3: 2672-2680.), which has shown amazing effects and great potential in the field of image generation. Many researchers have continuously improved the GAN architecture, and many methods for applying GAN to data augmentation have been successively proposed. Tanaka et al. used GAN to synthesize medical data for cancer detection training and obtained better performance than the original small dataset (see Tanaka F, Aranha C. Data Augmentation Using GANs[J]. arXiv preprint, arXiv:1904.09135, 2019.). Liu et al. designed an end-to-end data augmentation scheme for the Cityspace dataset based on the Pix2pixHD model (see Wang T C, Liu M Y, Zhu J Y, et al. High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs[J]. IEEE International Conference on Computer Vision and Pattern Recognition, 2018, 8798-8807.) (see Liu S, Zhang J, Chen Y, et al. Pixel Level Data Augmentation for Semantic Image Segmentation using Generative Adversarial Networks[J]. IEEE International Conference on Acoustics, Speech and Signal Processing, 2019: 1902-1906.). Compared with the traditional data augmentation scheme, this scheme further improved the mIoU of PSPNet in the 19-class semantic segmentation task of Cityspace by 2.1%.Zhang et al. proposed a defect synthesis framework Defect-GAN (see Zhang G, Cui K, Hung TY, et al. Defect-GAN: High-Fidelity Defect Synthesis for Automated Defect Inspection[J]. IEEE Winter Conference on Applications of Computer, 2021: 2523-2533), which achieved good visual effects on the CODEBRIM concrete defect dataset. At the same time, ResNet34 and DenseNet152 were used as concrete defect detectors, and the accuracy of both increased significantly after data augmentation using Defect GAN.
[0006] Existing data augmentation algorithms, whether based on traditional methods or deep learning methods, are not suitable for directly augmenting fabric defect data, mainly having several problems: (1) Traditional data augmentation methods cannot directly and quickly overlay defects from an existing defect sample library onto a new fabric. (2) The number of fabric defects is small, which belongs to the problem of small sample learning. Training a GAN network itself requires a large sample library for support. (3) Data augmentation often overlays defects on a certain area of the fabric. Therefore, it is required that the transition between the augmented area and the entire fabric image is smooth. If the defect overlay is not natural enough, it will instead reduce the performance of the detector. Summary of the Invention
[0007] The present invention solves the problems existing in the prior art and provides an optimized fabric defect data augmentation method based on a generative adversarial network.
[0008] The technical solution adopted by the present invention is a fabric defect data augmentation method based on a generative adversarial network. After building and training the generative adversarial network, for the input defective fabric image to be data-augmented, multi-scale feature extraction and style encoding are performed with multiple convolutional layers and instance normalization, an upsampling convolutional module is constructed with spatial adaptive normalization, and the network tasks are decoupled with two parallel convolutional branches, corresponding to the generation of defect foreground and defect transparency respectively. With the cooperation of transparency control, defect morphology and spatial constraints, a defective fabric picture after data augmentation is finally obtained.
[0009] Preferably, the built generative adversarial network includes:
[0010] A feature encoding module for extracting feature maps of different resolutions during downsampling and encoding the image features to obtain the mean μ and variance σ representing the image style 2 ;
[0011] A flaw generation module for upsampling the style noise sampled from the image style mean μ and variance σ 2 and continuously introducing the feature maps extracted by the feature encoding module during the upsampling process to obtain synthetic flaw samples;
[0012] A discriminator module for determining whether the input sample is a real flaw sample or a synthetic flaw sample, and guiding the flaw generation module to synthesize flaw samples during the model iterative training process, so that the synthetic flaw samples are close to the real flaw samples in terms of image authenticity and clarity.
[0013] Preferably, the feature encoding module includes 5 convolutional blocks and 2 fully connected layers;
[0014] Any convolutional block includes a standard 3×3 convolutional layer, an instance normalization layer, and a leaky linear activation function layer;
[0015] The 5 convolutional blocks respectively output feature maps F0, F1, F2, F3, and F4 at five different resolutions. At the same time, F4 is input into two parallel 256-dimensional fully connected layers to respectively calculate the image style mean vector μ and variance vector σ 2 , and perform single Gaussian sampling on μ and σ 2 to obtain the style noise z.
[0016] In the present invention, the feature encoding module takes a flawless cloth sample as the input and consists of five convolutional blocks and two fully connected layers. Among them, the convolutional blocks are used to extract feature information and downsample. Each convolutional block consists of a standard 3×3 convolution Conv, an instance normalization InstanceNorm, and a leaky linear activation function LeakyReLU. The stride of the first Conv layer is 1 and the zero-padding is 1. The strides of the remaining four Conv layers are 2 and the zero-padding is 2 to achieve the downsampling function and obtain feature maps at different resolutions. Instance Normalize can perform normalization operations for single image instances, ensuring the independence between individual image instances and retaining the style of the image instances, which is more suitable than BatchNormalize in image synthesis tasks. LeakyReLU is similar to ReLU, introducing sparsity into the network and improving computational performance. In addition, since LeakyReLU still has a small gradient when the activation value is negative, this can avoid the problem of neuron "death" and accelerate network convergence.
[0017] In the present invention, after passing through the feature encoding module, five feature maps F0, F1, F2, F3, and F4 at different resolutions can be obtained from the output ends of the five convolutional blocks. Among them, the resolution of F0 is the same as that of the input image and has the richest texture feature information. The length and width of F1, F2, F3, and F4 are 1 / 2, 1 / 4, 1 / 8, and 1 / 16 of the input image respectively. In addition, since F4 contains high-dimensional semantic information, it is necessary to input it into two parallel fully connected layers with 256 dimensions to calculate the image style mean μ and variance σ 2 , and sample the mean μ and variance σ 2 to obtain the style noise z.
[0018] Preferably, the defect generation module includes a parallel defect foreground generation branch and a defect transparency generation branch;
[0019] The defect foreground generation branch includes a fully connected layer and 6 upsampling modules. A residual connection layer is provided between the input and output of each upsampling module. The projected noise z obtained after the input style noise z passes through the fully connected layer project , after passing through 6 upsampling modules, a high-resolution feature map is obtained, and the foreground map O is output after passing through the Tanh activation function forge ;
[0020] The defect transparency generation branch includes a fully connected layer and 6 upsampling modules. A residual connection layer is provided between the input and output of each upsampling module. The input standard normal distribution noise z noise , after passing through 6 upsampling modules, a high-resolution feature map is obtained, and the transparency image O is output after passing through the Sigmoid activation function alpha ;
[0021] Get S defect =S normal ·(1 - O alpha ) + O forge ·O alpha , where S defect is the synthesized defect map, S normal is the defect-free sample input to the feature encoding module, O forge is the output of the defect foreground generation branch, and O alpha is the output of the defect transparency generation branch.
[0022] Preferably, the upsampling module consists of two serially connected convolutional blocks, and each convolutional block includes a SPADE layer, a non-linear activation function ReLU, and a standard 3×3 convolution arranged in sequence.
[0023] Preferably, the inputs of the defect generation module include the 256-dimensional style noise z obtained by the feature encoding module, the feature maps F0, F1, F2, F3, and F4 output by the intermediate layer of the feature encoding module, and the defect mask map M randomly sampled from the preset defect sample library;
[0024] The projection noise z obtained by inputting the style noise z into the corresponding fully connected layer project and the standard normal distribution noise z noise are respectively input into 6 upsampling modules in the corresponding generation branches for decoding. Meanwhile, the defect mask map M is input at the upsampling modules to guide the decoding. Sequentially, the feature maps F4, F3, F2, F1, and F0 are concatenated at the output ends of the first 5 upsampling modules respectively as the inputs of the next-level upsampling modules. For the defect foreground generation branch, the feature map is compressed to 3 dimensions through the last upsampling module and output through the corresponding activation function. For the defect transparency generation branch, the feature map is compressed to 1 dimension through the last upsampling module and output through the corresponding activation function.
[0025] In the present invention, the defect foreground generation branch takes the random style noise z as the input and continuously improves the image resolution through a series of convolutions and bilinear interpolation upsamplings. During the defect foreground generation process, in order to make the generated defect shape and position controllable, it is necessary to introduce a defect semantic mask map for constraint. Currently, the mainstream methods for translating semantic maps into actual images include: the pix2pix series of models directly use the mask as the input for image generation. In the case of the same mask input, since it is required to approach the target image under any input noise, it will cause the problem of blurred generated images. In addition, due to BatchNormalize normalizing the data distribution, the styles of the generated images tend to be the same. To address the above problems, Nvidia researchers proposed a SPADE Normalize method (see Park T, Liu M Y, Wang T C, et al. Semantic Image Synthesis With Spatially-Adaptive Normalization[C]. Conference on Computer Vision and Pattern Recognition, IEEE, 2019: 2332-2341.). This method does not directly use the mask map as the input, but introduces the semantic mask map on the basis of Batch Normalize to guide the calculation of the affine coefficients, obtaining very excellent visual effects. Under the same semantic mask input, different styles of synthetic images can be obtained according to different input noises.
[0026] In the present invention, the defective foreground generation branch uses SPADE Normalize to introduce the semantic map. The SPADEResblock in GauGAN is used as the basic upsampling module. This module is composed of two cascaded convolutional modules. Each convolutional module is sequentially composed of a SPADE layer, a non-linear activation function ReLU, and a standard 3×3 convolution. The convolution stride is 1, the zero-padding is 1, and a residual connection is established between the input and the output. This module has good visual performance in actual experiments; at the same time, the image features F0 to F4 extracted in the feature encoding module are used as the common input, so that during the upsampling process, the texture, shape, etc. of the defects are the same as those of the defect-free cloth original Figure 1 consistent.
[0027] In the present invention, the defective transparency generation branch has the same structure as the defective foreground branch. The middle upsampling layer also receives the image features F0 to F4 extracted in the feature encoding module. The difference is that the input it receives is the standard normal distribution noise with a mean of 0 and a variance of 1, which can increase the diversity of the images generated by the model to a certain extent.
[0028] The final defective image synthesis formula is as follows:
[0029] S defect = S normal ·(1 - O alpha ) + O forge ·O alpha (1)
[0030] where S defect is the synthesized defective map, S normal is the defect-free sample input to the feature encoding module, O forge is the output of the defective foreground generation branch, and O alpha is the output of the defective transparency generation branch.
[0031] Preferably, the discriminator module includes three parallel branches B0, B1, and B2. Any branch includes an input layer and several modules CIL. The module CIL includes a standard 3×3 convolutional layer, an instance normalization layer, and a leaky non-linear activation function layer arranged in sequence;
[0032] In branch B0, there are 3 modules CIL, which are used to receive the image of the original size W×H for discrimination;
[0033] In branch B1, there are 2 modules CIL, which are used to receive the image of size (W / 2)×(H / 2) for discrimination;
[0034] In branch B2, there is 1 module CIL, which is used to receive the image of size (W / 4)×(H / 4) for discrimination.
[0035] In the present invention, PatchGAN proposes a block-based discriminator: after the image passes through various convolutional layers, it is not directly output to the activation function or the fully connected layer. Instead, the output of the convolution is mapped to an N×N matrix, and each value in the matrix represents the evaluation value of a region in the original image. This designed discriminator can pay attention to the details of more regions and improve the quality of the images generated by the GAN. pix2pixHD introduces a multi-scale design based on the discriminator of PatchGAN. The image is sent to the block discriminators with the same structure at different resolutions respectively, and finally an evaluation value integrating the multi-scale images is output. Such a design enables the discriminator to not only pay attention to local details, but also further expand the receptive field, enabling the discriminator to evaluate the overall image and further improving the quality and clarity of the GAN images. The method uses the multi-scale Patch discriminator of pix2pixHD to evaluate the scores of the original image, the original image with half the length and width, and the original image with a quarter of the length and width.
[0036] Preferably, the training of the generative adversarial network includes the following steps:
[0037] Step 1.1: Design the loss function;
[0038] Step 1.2: Produce and enhance the data set to obtain the defective data set Ω1 and the defect-free cloth data set Ω2;
[0039] Step 1.3: Use the Adam optimizer to optimize the network model parameters, and use the loss function L gan Train the defective foreground generation branch, and use the loss function L alpha Train the defective transparency generation branch, and optimize the defective transparency generation branch during the backpropagation process.
[0040] Preferably,
[0041] The final loss function L all = L gan + λ alpha × L alpha ;
[0042] where x represents the real defective sample sampled from the real defective data distribution p r , z represents the random noise sampled from the standard normal distribution N(0,1), represents the synthetic defective sample sampled from the defective generator data distribution , D represents the discriminator, G represents the generator, m represents the real defective mask image sampled from the real defective mask data distribution p ralpha , and a represents the synthetic defective transparency map sampled from the defective transparency generator data distribution p galpha . λdisc is a hyperparameter for controlling the authenticity of synthetic defects, λ disc is a hyperparameter for controlling the intensity of weight penalty, λ alpha is a hyperparameter for controlling the similarity between the synthetic defect transparency map and the real defect mask map.
[0043] In the present invention, the discriminator of the common GAN needs to map the output to a certain numerical range through an activation function, which will cause that when the activation value of one of the generator or the discriminator is too large, the other party cannot effectively learn the gradient. It is necessary to carefully coordinate the training degree of the generator and the discriminator to ensure the stability of GAN training. WGAN proposes a new loss function, which makes three modifications to the traditional GAN and solves the common training instability and mode collapse in GAN. The modifications of WGAN are as follows: the activation function in the discriminator is removed; the logarithmic loss function is not used in the loss function; the Weight clipping technique is adopted to limit the parameter range of the discriminator between -0.01 and 0.01 to meet the 1-Lipschitz constraint. The specific loss function of WGAN is as follows:
[0044]
[0045] WGAN-GP further improves on the basis of WGAN. Since the simple and rough Weight clipping in WGAN will weaken the model modeling ability and cause the phenomenon of gradient disappearance or gradient explosion, WGAN-GP uses gradient penalty (GP) to replace Weight clipping, and finally obtains the loss function as follows:
[0046]
[0047] In addition, the method of the present invention introduces the defect morphology and spatial constraints by calculating the similarity between the transparent channel and the mask map. Considering that the transparency within the mask area should be determined by the network itself and is not restricted by this constraint, only the similarity outside the mask area is calculated, and the morphological space constraint loss function is obtained as follows:
[0048]
[0049] In summary, the method of the present invention combines the WGAN-GP loss function and the morphological space constraint loss function to obtain the final loss function as:
[0050] L all = L gan + λ alpha × L alpha (5)
[0051] In the present invention, due to the large number of network model parameters, the gradient penalty value of WGAN-GP is several orders of magnitude larger than other loss function values, and it is easy for the network to fall into a mediocre solution where no images are generated during training. Therefore, λ grad should not be selected too large, and the range is [1e-3, 1e-5].
[0052] In the present invention, since the defect generation module is a multi-output network, it is necessary to train different branches separately to accelerate model convergence and improve model performance. During the training process, the present invention uses the Adam optimizer to optimize the network, and the learning rate is adjusted using a strategy of decaying every epoch. The learning rate decay formula is as follows:
[0053]
[0054] where lr new is the changed learning rate, lr old is the changed learning rate, curepoch is the current training epoch, maxepoch is the maximum training epoch, and power is the decay factor.
[0055] In the present invention, first, semantic annotation is performed on different types of defects on different cloths photographed on the production line to obtain a defect dataset Ω1; then, data augmentation of vertical and horizontal random flipping is performed on the data in Ω1 to increase the diversity of defect data; finally, defect-free cloth images are photographed from the production line to form a defect-free dataset Ω2.
[0056] In the present invention, the process of training the network model using the defect dataset Ω1 and the defect-free dataset Ω2 is divided into three stages: generator content training, generator transparency training, and discriminator training.
[0057] In the generator content training stage, first, a defect region mask image is randomly sampled from the defect dataset Ω1, and a defect-free cloth image is randomly sampled from the defect-free dataset Ω2. The two are used as the input data of the generator to obtain a synthetic sample, which is sent to the discriminator for evaluation. The loss function L in Equation (3) is used to gan train the defect foreground generation branch.
[0058] The generator transparency training is the same as the generator content training stage, except that the loss function L in Equation (4) is used in this stage for alpha training, and only the defect transparency generation branch is optimized during the backpropagation process.
[0059] In the discriminator training phase, the discriminator receives defective samples and corresponding defective mask images as inputs, and the output is whether the defective samples belong to real images or synthetic images generated by the generator. The real defective samples and defective mask images are obtained from the defective dataset Ω1, and the synthetic defective samples are generated by the generator. The loss function L in Equation (3) is used in this phase. gan Train the discriminator's ability to distinguish between true and false.
[0060] Preferably, using the trained generative adversarial network to enhance the cloth defect data includes the following steps:
[0061] Step 2.1: Perform N r random bounding box selections on each image in the defect-free cloth dataset Ω2 to obtain a number of background images of 256×256 pixels, where N r ranges from [2, 10];
[0062] Step 2.2: For each background image, randomly sample a mask image from the defective dataset Ω1, and input the background image and the mask image into the defective generation network model to obtain a synthetic defective image;
[0063] Step 2.3: Paste the synthetic defective image back to the bounding box area of the corresponding defect-free cloth image in Step 2.1 to obtain a defective cloth image generated by data augmentation.
[0064] The present invention relates to an optimized method for enhancing cloth defect data based on a generative adversarial network. After building and training the generative adversarial network, for the input defective cloth image to be data-enhanced, multi-scale feature extraction and style encoding are performed with multi-layer convolution and instance regularization, an upsampling convolution module is constructed with spatial adaptive normalization, and the network tasks are decoupled with 2 parallel convolution branches, corresponding to the generation of defective foreground and defective transparency respectively. With the cooperation of transparency control, defective morphology and spatial constraints, a defective cloth image after data augmentation is finally obtained.
[0065] The present invention uses multi-layer convolution and Instance Normalize to perform multi-scale feature extraction and style encoding on the input image; on the basis of referring to GauGAN, SPADE Normalize is used to construct an upsampling convolution module. SPADE Normalize can ensure that when introducing semantic segmentation information, the style characteristics of the feature map are still retained, and finally high-resolution features are obtained; for the high-resolution features obtained by upsampling, two parallel convolution branches are designed to decouple the network tasks. The two branches are respectively responsible for generating the defective foreground and defective transparency, and the concept of transparency control is introduced to make the defect synthesis more natural and smooth; at the same time, defect morphology and spatial constraints are introduced to limit the coordinates and morphology of defect generation, so that the data augmentation effect is controllable. In addition, the discriminator of the present invention is designed by referring to the block discriminator of PatchGAN to supervise the defect synthesis effect.
[0066] Using the method of the present invention, data augmentation can be performed on the existing data set, and the performance of the cloth defect detector can be improved by balancing the number of sample categories and expanding the defective samples. It can also perform defect overlay on the new cloth without defects to quickly generate a sufficient amount of defective data set under the new background. Brief Description of the Drawings
[0067] Figure 1 is the content block diagram of the present invention;
[0068] Figure 2 is the overall network structure of the present invention;
[0069] Figure 3 is the network structure of the feature encoding module;
[0070] Figure 4 is the network structure of the defect generation module;
[0071] Figure 5 is the SPADEResblock structure;
[0072] Figure 6 is the multi-scale block discriminator structure;
[0073] Figure 7 is the data augmentation effect diagram of the knot defect of the present invention;
[0074] Figure 8 is the data augmentation effect diagram of the dirt defect of the present invention;
[0075] Figure 9 is the data augmentation effect diagram of the colored yarn defect of the present invention;
[0076] Figure 10 is the data augmentation effect diagram of the drawn yarn defect of the present invention. Detailed Embodiment
[0077] The present invention will be further described in detail below in conjunction with embodiments, but the protection scope of the present invention is not limited thereto.
[0078] The computer hardware configuration selected for the present invention is as follows: CPU Intel i7-11700K@3.6GHz, GPU RTX3090, 24GB video memory, and 16GB memory. The software platform is Ubuntu18.04 64-bit system and is implemented based on Pytorch1.8.
[0079] As Figure 1 shown, the fabric defect data enhancement method based on the generative adversarial network includes three parts:
[0080] (1) Building a generative adversarial network;
[0081] (2) Training and optimizing the neural network;
[0082] (3) Using the trained neural network model to perform data enhancement on the existing data set;
[0083] The specific construction of the generative adversarial network includes:
[0084] (1-1) Feature encoding module
[0085] The feature encoding module takes the flawless fabric sample S normal as input, and the resolution of the sample image is width W×height H. It consists of five convolutional blocks and two fully connected layers, and its structure is as Figure 3 shown. Among them, the convolutional block is used to extract feature information and downsample. Each convolutional block consists of a standard 3×3 convolution Conv, instance normalization InstanceNorm, and a leaky rectified linear activation function LeakyReLU. The stride of the first Conv layer is 1, and the zero-padding is 1; the strides of the remaining four Conv layers are 2, and the zero-padding is 2 to achieve the downsampling function and obtain feature maps at different resolutions; the slope of LeakyReLU is 0.2 when the negative activation value is input. After passing through the feature encoding module, five feature maps F0, F1, F2, F3, and F4 at different resolutions can be obtained from the output ends of the five convolutional modules, which are used as the input for the subsequent network. Among them, the resolution of F0 is the same as that of the input image, and the length and width of F1, F2, F3, and F4 are 1 / 2, 1 / 4, 1 / 8, and 1 / 16 of the input image respectively.
[0086] The fully connected layer is used to calculate the style mean and variance of the feature map. After expanding the highest-dimensional feature map F4 into a one-dimensional sequence, it is input into two parallel 256-dimensional fully connected layers, and the 256-dimensional style mean μ and variance σ of the image are respectively output and calculated 2 , and for the mean μ and variance σ2 Sample 256 - dimensional style noise z.
[0087] (1 - 2) Defect generation module
[0088] The defect generation module consists of two parallel branches with the same structure, namely the defect foreground generation branch and the defect transparency generation branch. The overall framework is as Figure 4 shown. Both are decoding networks used to decode noise into defect images and defect transparencies. The specific structure is as follows:
[0089] (1 - 2 - 1) Defect foreground generation branch
[0090] The defect foreground generation branch uses SPADE Normalize to introduce the semantic map and takes the SPADEResblock in GauGAN as the basic module. The structure of the SPADEResblock module is as Figure 5 shown. This module consists of two serially connected convolutional modules. Each convolutional module is composed of a SPADE layer, a non - linear activation function ReLU, and a standard 3×3 convolution in sequence. The convolution stride is 1, the zero - padding is 1, and a residual connection is established between the input and the output.
[0091] The defect foreground generation branch receives the input of three groups of data, namely: the 256 - dimensional style noise z obtained by the feature encoding module, the feature maps F0, F1, F2, F3, and F4 output by the middle layer of the feature encoding module, and the defect mask map M randomly sampled from the existing defect sample library. The style noise z passes through a fully - connected layer to obtain the projected noise z project , where the output dimension of the fully - connected layer is an integer multiple of the input dimension, and the multiple range is [128, 512]. Here, 256 is taken. Adjust z project to a two - dimensional feature map with a resolution of (W / 32)×(H / 32). The basic decoding module consists of a SPADEResblock module and bilinear interpolation 2 - fold upsampling. The two - dimensional feature map is input into 5 serially connected basic decoding modules for decoding, and at the same time, the defect mask map M is input at the SPADEResblock end to guide the decoding.
[0092] At the output end of each basic decoding module, the feature maps F1 / F2 / F3 / F4 with the corresponding resolution in the feature encoding module are concatenated in dimension as the input of the next - level basic decoding module. After passing through 5 basic decoding modules, a high - resolution feature map F with a resolution of W×H is finally obtained out , and F out and F0 are concatenated in dimension and compressed to 3 dimensions through a layer of SPADEResblock, and then output the foreground map O after passing through the Tanh activation function forge .
[0093] (1-2-2) Defect transparency generation branch
[0094] The defect transparency generation branch has the same structure as the defect foreground branch, and the upsampling intermediate layer also receives the image features F0 to F4 extracted from the feature encoding module. The main difference is that its input is the standard normal distribution noise z with a mean of 0 and a variance of 1 noise ; the activation function at the output end is changed from Tanh to Sigmoid; the final output dimension is a 1D transparency image O alpha , and finally the synthetic defect map is obtained by Equation (1).
[0095] (1-3) Discriminator module
[0096] The present invention uses a multi-scale block discriminator to discriminate the generated cloth defect map. The main structure is as Figure 6 shown, and there are three parallel branches B0, B1, and B2. The three branches have input layers with the same parameters, all of which are standard 3×3 convolutions plus the leaky non-linear activation function LeakyReLU. Among them, the convolution stride is 2, the zero-padding is 2, and the negative slope of LeakyReLU is 0.2. In addition, the standard 3×3 convolution, instance normalization InstanceNorm, and leaky non-linear activation function LeakyReLU are regarded as the basic module CIL, where the convolution stride is 2, the zero-padding is 2, and the negative slope of LeakyReLU is 0.2. Branch B0 contains 3 cascaded CILs, B1 contains 2 cascaded CILs, and B2 contains 1 CIL.
[0097] The multi-scale block discriminator receives four-channel input image data, where channels 0 to 2 are the cloth defect images, and channel 3 is the defect mask image. The original size W×H image is input into branch B0 for discrimination, the (W / 2)×(H / 2) image is input into branch B1 for discrimination, and the (W / 4)×(H / 4) image is input into branch B2 for discrimination.
[0098] The neural network training optimization specifically includes:
[0099] (2-1) Loss function design
[0100] The present invention uses the loss function in Equation (5) to train the model. Since the number of network model parameters is large, the gradient penalty value of WGAN-GP is several orders of magnitude larger than that of other loss function values, and it is easy to make the network fall into a mediocre solution that does not generate images during training. Therefore, λ grad should not be selected too large, and the range is [1e-3, 1e-5], and 1e-4 is taken here; λ disc range is [0.1, 5], and 1 is taken here; λ alpha range is [1, 10], and 5 is taken here.
[0101] (2-2) Dataset Production and Enhancement
[0102] Due to the lack of a publicly available semantic annotation dataset for cloth defects, the present invention performs semantic annotation on different cloths and different types of defects captured on the production line to obtain dataset Ω1, which contains a total of 460 images and 652 defect instances, including 4 common defects: knots, dirt spots, colored yarns, and drawn threads. Since the number of datasets is insufficient, in order to ensure that the model avoids mode collapse and overfitting, data augmentation by randomly performing vertical and horizontal flips on the data is carried out during the training process, where the probabilities of vertical and horizontal flips are both 0.5. The defect-free cloth dataset Ω2 is captured from the factory production line, without the need for additional annotation work and with no restrictions on the quantity.
[0103] (2-3) Model Training Process
[0104] The present invention uses the Adam optimizer to optimize the network model parameters. The initial learning rate ranges from [1e-6, 5e-4], and here 1e-5 is taken. Learning rate decay uses Equation (6), where the power ranges from [0.1, 10], and here 4 is taken.
[0105] During the discriminator training stage, the loss function L is used gan Train the discriminator's ability to judge true and false.
[0106] During the generator content training stage, the defect mask image and the defect-free sample are fed into the generator, and the generated defect sample is sent to the discriminator for evaluation, using the loss function L gan Train the defect foreground generation branch. The generator transparency training is the same as the generator content training stage, except that the loss function L is used in this stage alpha for training, and only the defect transparency generation branch is optimized during the backpropagation process.
[0107] The process of using the trained neural network model to perform data augmentation on the existing dataset specifically includes:
[0108] Use the trained feature encoding module and defect generation module to perform data augmentation on the existing defect-free cloth images. First, randomly select N r times from the defect-free cloths in dataset Ω2 to obtain multiple background images of 256×256 pixels. The range of N r is [2, 10], and here 5 is taken. For each background image, randomly sample a mask image from dataset Ω1, input both of them into the feature encoding module for feature extraction, input the obtained feature map into the defect generation module to obtain a synthetic defect image, and finally paste the synthetic defect back to the corresponding selected area of the cloth to complete the data augmentation.
[0109] Figures 7 to 10This is the data augmentation effect diagram of the method of the present invention. Data augmentation is performed on four types of defects: knots, dirty spots, colored yarns, and drawn threads. As shown in the figure, the cloth defect data augmentation algorithm of the present invention can superimpose realistic defect patterns on a defect-free cloth to achieve the expansion of cloth defect data.
[0110] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0111] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a machine for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.
[0112] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.
[0113] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.
[0114] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0115] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for enhancing fabric defect data based on a generative adversarial network, characterized in that: The method constructs and trains a generative adversarial network; The generative adversarial network includes: A feature encoding module, which is used to extract feature maps of different resolutions during downsampling, and encode the image features to obtain the mean representing the image style and variance ; Perform single Gaussian sampling on and to obtain style noise ; A defect generation module for upsampling the style noise sampled from the image style mean and variance and continuously introducing the feature maps extracted by the feature encoding module during the upsampling process to obtain synthetic defect samples; The defect generation module includes a parallel defect foreground generation branch and a defect transparency generation branch; The defective foreground generation branch includes a fully connected layer and six upsampling modules. A residual connection layer is provided between the input and output of each upsampling module, and the input style noise The projected noise obtained after passing through the fully connected layer , after passing through six upsampling modules, a high-resolution feature map is obtained, and the foreground map is output after passing through the Tanh activation function ; The defect transparency generation branch includes a fully connected layer and six upsampling modules. A residual connection layer is provided between the input and output of each upsampling module, and standard normal distribution noise is input , and a high-resolution feature map is obtained after passing through six upsampling modules, and a transparency image is output through a Sigmoid activation function ; Obtained , where is the synthetic defect map, is the input defect-free sample of the feature encoding module, is the output of the defect foreground generation branch, is the output of the defect transparency generation branch; A discriminator module, which is used to judge whether the input sample is a real defect sample or a synthetic defect sample, and guides the defect generation module to synthesize defect samples during the iterative training process of the model; For the input defective cloth image to be data-augmented, multi-scale feature extraction and style encoding are performed with multi-layer convolution and instance regularization, an upsampling convolution module is constructed with spatially adaptive normalization, and the network tasks are decoupled with 2 parallel convolution branches, corresponding to the generation of defect foreground and defect transparency respectively. With the cooperation of transparency control, defect morphology and spatial constraints, a defective cloth image after data augmentation is finally obtained.
2. The method for enhancing fabric defect data based on a generative adversarial network according to claim 1, wherein: The feature encoding module includes 5 convolutional blocks and 2 fully connected layers; Any convolutional block includes a standard 3×3 convolutional layer, an instance normalization layer and a leaky linear activation function layer; The five convolutional blocks respectively output feature maps at five different resolutions , , , and . At the same time, is input into two parallel fully-connected layers with 256 dimensions to respectively calculate the mean vector of the image style and the variance vector .
3. A method for enhancing fabric defect data based on a generative adversarial network according to claim 1, characterized in that: The upsampling module is composed of two cascaded convolutional blocks, and each convolutional block includes a SPADE layer, a non-linear activation function ReLU and a standard 3×3 convolution arranged in sequence.
4. A method for enhancing fabric defect data based on a generative adversarial network according to claim 2, characterized in that: The input of the defect generation module includes the 256-dimensional style noise obtained by the feature encoding module , the feature map output by the intermediate layer of the feature encoding module , , , and and the defect mask map randomly sampled from the preset defect sample library ; The projection noise obtained by inputting the style noise into the corresponding fully connected layer and the standard normal distribution noise are respectively input into 6 upsampling modules in the corresponding generation branches for decoding. At the same time, a defective mask image is input into the upsampling modules to guide the decoding. Feature maps , , , , are sequentially concatenated at the output ends of the first 5 upsampling modules and used as the input for the next-level upsampling module. For the defective foreground generation branch, the feature map is compressed to 3 dimensions by the last upsampling module and output after passing through the corresponding activation function. For the defective transparency generation branch, the feature map is compressed to 1 dimension by the last upsampling module and output after passing through the corresponding activation function.
5. A method for enhancing fabric defect data based on a generative adversarial network according to claim 1, characterized in that: The discriminator module includes three parallel branches , and . Any one of the branches includes an input layer and several modules . The module includes a standard 3×3 convolutional layer, an instance normalization layer, and a leaky non-linear activation function layer that are sequentially arranged; Branch Among them, the modules are three, used to receive images of the original size W×H for discrimination; Branch In which, the modules are two and used to receive images with the size of (W / 2)×(H / 2) for discrimination; Branch In the there is 1 module, which is used to receive an image with a size of (W / 4)×(H / 4) and perform discrimination.
6. A method for enhancing fabric defect data based on a generative adversarial network according to claim 1, characterized in that: The training of the generative adversarial network includes the following steps: Step 1.1: Design a loss function; Step 1.2: Dataset production and enhancement to obtain a defective dataset and a non-defective cloth dataset ; Step 1.3: Optimize the network model parameters using the Adam optimizer, and use the loss function to train the defective foreground generation branch, and use the loss function to train the defective transparency generation branch, and optimize the defective transparency generation branch during the backpropagation process.
7. According to the method for augmenting cloth defect data based on a generative adversarial network as described in claim 6, characterized in that: , , Final loss function ; where represents the real defect samples sampled from the real defect data distribution ; represents the random noise sampled from the standard normal distribution ; represents the synthetic defect samples sampled from the defect generator data distribution ; represents the discriminator ; represents the generator ; represents the real defect mask map sampled from the real defect mask data distribution ; is a hyperparameter for controlling the authenticity of the synthetic defect is a hyperparameter for controlling the weight penalty intensity is a hyperparameter for controlling the similarity between the synthetic defect transparency map and the real defect mask map 8. A method for enhancing fabric defect data based on a generative adversarial network according to claim 6, characterized in that: Data augmentation of cloth defect data with the trained generative adversarial network includes the following steps: Step 2.1: For the dataset of flawless cloth each image in is randomly boxed to obtain a number of background images of 256×256 pixels, where the range is [2, 10]; Step 2.2: For each background image, randomly sample a mask image from the defect dataset and input the background image and the mask image into the defect generation network model to obtain a synthetic defect image; Step 2.3: Paste the synthetic defect map back to the boxed area of the corresponding defect-free cloth image in step 2.1 to obtain a defective cloth image generated by data augmentation.
Citation Information
Patent Citations
Training method and device of image conversion model generator
CN111709873A
Generative adversarial network training method and device and image style migration method and device
CN111862274A