Cloth defect image generation method based on weight map guidance

By introducing a weight map guidance mechanism and the ConvNeXt V2 module in the fabric defect image generation, the problems of insufficient diversity and poor realism in the generated image in the prior art are solved, and higher quality and diversity of fabric defect image generation are achieved.

CN120219874APending Publication Date: 2025-06-27ZHEJIANG SCI-TECH UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510275888.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has pattern crashes when generating cloth defect images, resulting in insufficient diversity of generated images, and the role of the discriminator in the network is ignored, affecting the reality and quality of the generated images.

Method used

Using a cloth defect image generation method based on weight map guidance, a more realistic and diverse image is generated by introducing an attention mechanism into the generator of the CycleGAN network, and a more realistic and diverse image is generated by combining the weighted parts of the original input image. At the same time, the discriminator structure is improved and the ConvNeXt V2 module is added to improve its ability to provide meticulous feedback.

Benefits of technology

It realizes the generation of multiple cloth defect pictures without pairing samples, which improves the accuracy and realism of image generation and enhances the quality and diversity of generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219874A_ABST
    Figure CN120219874A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image generation, and discloses a cloth flaw image generation method based on weight map guidance, which comprises the following steps: collecting a cloth image with flaws, preprocessing the image into a group of sub-images, and then sequentially inputting the sub-images into a WMG-GAN network trained offline to generate a new cloth flaw image. The WMG-GAN network is based on a CyclGAN network, an attention mechanism thought is added to an output result image of a CyclGAN network generator, and the output result image is divided into a foreground weight map and a feature weight map; and a ConvNeXtV2 module is additionally arranged behind the Leaky ReLU function of each middle level of the CycleGAN network discriminator. According to the image generation method, the cloth defect picture can be generated without matching samples, the foreground can be selectively changed and the background can be reserved by using the foreground and the feature weight map, and thus more real image generation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cloth defect detection and image generation, and specifically relates to a method for generating cloth defect images guided by a weight map. Background Technique

[0002] Cloth defects refer to the defect problems that appear on the surface of cloth during the weaving process. These defects are mainly caused by equipment failures, yarn problems, foreign objects being caught in, poor processing, excessive stretching, etc. With the development of computer vision and machine learning technologies, automatic defect detection systems based on image analysis have gradually become a research hotspot. However, obtaining cloth defect images faces many challenges. First of all, there are a wide variety of cloth defects, which increases the difficulty of image acquisition. Secondly, the number of cloth defect images is relatively small, and the scale of the dataset severely limits the training and generalization capabilities of the model. In addition, due to the complexity of the production environment, it further affects the quality and usability of the images. Currently, there are many methods for generating images, among which Generative Adversarial Networks (GANs) is one of the most representative. However, GAN has some deficiencies in the training process, such as the mode collapse phenomenon, that is, the generator can only generate a limited number of types of images, resulting in insufficient diversity of the generated images.

[0003] To solve the above problems, many scholars have conducted extensive research and optimization on the GAN model. Deep Convolutional GAN (DCGAN) replaces the fully connected layer with a convolutional layer and applies batch normalization technology; Auxiliary Classifier GAN (ACGAN) makes modifications in the discriminator part, not only to judge the authenticity of the input image, but also to try to accurately predict the class label of the image. These two models have achieved relatively obvious success in the field of image translation, but most tasks are troubled by having few or no paired input-output samples. When paired training data cannot be obtained, image-to-image conversion becomes a difficult problem. The proposed CycleGAN solves the problem of difficult acquisition of paired data: it introduces a cycle consistency loss, that is, the converted image should be as close as possible to the original image when converted back, thus ensuring the stability and authenticity of the conversion process. However, it still has obvious defects: when the generator tries to restore the image under the constraint of the cycle consistency loss, it may ignore some important structural details and pay more attention to those common and easily replicated features, resulting in the generated image being unable to accurately restore the original input image. Therefore, although the generated image may seem similar, the details may be distorted. If these details are directly mixed together, the colors and textures of the background and foreground content of the generated image are prone to entanglement, making the image look unrealistic and also resulting in problems of low quality and poor diversity of the generated image.

[0004] In addition, most existing generative models only focus on the ability of the generator to generate images and ignore the role of the discriminator in the network. The training process of GAN is essentially a game process, in which the generator and the discriminator jointly optimize through mutual confrontation. During training, the generator may overfit the common patterns in the training data and ignore the high-level structure and detailed features of the images. This makes the generated images appear blurred, unnatural and lack a sense of reality visually. In addition, the generator relies on the feedback provided by the discriminator to learn the data distribution during the training process. If the discriminator cannot provide accurate feedback, it is difficult for the generator to generate high-quality images on unseen data. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method for generating fabric defect images guided by a weight map, so as to combine a feature weight map and a foreground weight map to realize the generation of fabric defect images with higher accuracy and realism.

[0006] To solve the above technical problem, the present invention provides a method for generating fabric defect images guided by a weight map, including: collecting fabric pictures with defects and preprocessing the images into a group of sub-images, and then sequentially inputting the sub-images into an offline trained WMG-GAN network to generate fabric defect pictures with defect type labels;

[0007] The WMG-GAN network is an improvement based on the CycleGAN network. An idea of adopting an attention mechanism is added to the output result image of the CycleGAN network generator. The output result image is divided into a foreground weight map and a feature weight map. After the foreground weight map is processed by Sigmoid activation, it is combined with the feature weight map, and then added with the weighted part of the original input image to generate the generated image G y ; The discriminator adopts a PatchGan structure, and a ConvNeXtV2 module is added after the LeakyReLU activation function of each intermediate layer of the CycleGAN network.

[0008] As an improvement of the method for generating fabric defect images guided by a weight map according to the present invention:

[0009] The generated image G y is:

[0010] G y = x×(1 - F y2 ) + F y1 ×F y2 (1)

[0011] where F y1 is the feature weight map of image y; F y2is the foreground weight map of image y, where x represents the input image.

[0012] As a further improvement of a method for generating fabric defect images guided by a weight map according to the present invention:

[0013] The process of offline training of the WMG-GAN network is as follows:

[0014] After collecting fabric images and performing image preprocessing and data augmentation, the samples are extended to include weft defect images, warp defect images, and normal images and labeled, and then divided into a training set and a test set. The labeled images in the training set are input into the WMG-GAN network for training. During the training process, the loss function value is calculated and the model parameters are iteratively optimized by backpropagation. After reaching the preset number of epochs, the training is completed and the training ends;

[0015] The labeled images in the test set are input into the trained model to generate new images, and various evaluation indicators of the images are calculated to meet the preset standards, thereby obtaining an offline-trained WMG-GAN network.

[0016] As a further improvement of a method for generating fabric defect images guided by a weight map according to the present invention:

[0017] The data augmentation technology includes highlighting rotation, noise addition, and brightness adjustment.

[0018] As a further improvement of a method for generating fabric defect images guided by a weight map according to the present invention:

[0019] The image preprocessing is to perform image segmentation based on a sliding window.

[0020] As a further improvement of a method for generating fabric defect images guided by a weight map according to the present invention:

[0021] The loss function is:

[0022] L = L GAN + L F + L LPIPS (7)

[0023] Wherein,

[0024] The adversarial loss of mapping G X→Y is:

[0025]

[0026] The adversarial loss of mapping G Y→X is:

[0027]

[0028] Mapping G X→Y The weight map-guided adversarial loss of is:

[0029]

[0030] Where D Y 's task is to distinguish and compare the generated image pairs [F y1 , G X→Y (x)] and the real image pairs [F y1 , y];

[0031] L LPIPS is the similarity loss function:

[0032] Send a set of images x and x0 into the network F for feature extraction, then calculate the distance between the features of x and x0 in different channels, then extract the feature stacks in different convolutional layers, and normalize the feature stacks in different channels. The normalized result is written as Subsequently, scale and activate the channels through ; finally, calculate the distance using the L2 norm, and take the average in the space and sum over the channels to calculate d0:

[0033]

[0034] Where wl is the result calculated by the cosine distance formula, h is the height of the input image, w is the width of the input image, L refers to the selected network layer, and are the feature responses of images x and x0 on this channel respectively, H L is the number of feature maps of this layer, and W L is the weight of the feature map of this channel;

[0035] Pass d0 and d1 into a model containing a RELU fully connected layer with two 32 channels, a single-channel fully connected layer, and a sigmoid layer for training. The similarity loss function is:

[0036] L LPIPS (x, x0, x1, h) = -h log G(d(x, x0), d(x, x1)) - (1 - h) log(1 - G(d(x, x0), d(x, x1))) (3)

[0038] Where x1 represents the image predicted by the pre-trained model; d1 is the true value calculated by the pre-trained model.

[0039] The beneficial effects of the present invention are mainly reflected in:

[0040] 1. The present invention is a cutting-edge generative adversarial network that can generate various fabric defect images without paired samples. By using foreground and feature weight maps, the model can selectively change the foreground while preserving the background, thus achieving more realistic image generation.

[0041] 2. The generator of the present invention combines an autoencoder with skip connections to maintain the data structure and mitigate the vanishing gradient problem. The discriminator is enhanced by integrating the ConvNeXt V2 module, improving its ability to provide detailed feedback and thus enhancing the quality of image generation.

[0042] 3. The present invention proposes a novel loss function that combines LPIPS and cycle consistency loss to enhance the visual quality of the generated images. LPIPS relies on high-level feature extraction and is consistent with human visual perception, enabling the network to generate images with higher structural accuracy and realism. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The following further elaborates on the specific implementation manners of the present invention in conjunction with the accompanying drawings.

[0044] Figure 1 Schematic diagram of a method for generating fabric defect images guided by a weight map according to the present invention;

[0045] Figure 2 Schematic diagram of the structure of the CycleGAN model;

[0046] Figure 3 Schematic diagram of the structure of the generator of the WMG-GAN network according to the present invention;

[0047] Figure 4 Schematic diagram of the structure of the ConvNeXt V2 module;

[0048] Figure 5 Schematic diagram of the structure of the discriminator of the WMG-GAN network according to the present invention;

[0049] Figure 6 Schematic diagram of the similarity loss function LPIPS;

[0050] Figure 7 Examples of fabric defect images in domains X, Y, and Z;

[0051] Figure 8 Comparison example diagrams of the generated images of the GAN, DCGAN, ACGAN, CycleGAN, and WMG-GAN models in domains X, Y, and Z. DETAILED DESCRIPTION OF THE INVENTION

[0052] The present invention will be further described below in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto:

[0053] Embodiment 1. A method for generating fabric defect images guided by a weight map. The WMG-GAN network of the present invention is constructed based on the CycleGAN network for generating fabric defect images. The generator of the WMG-GAN network can generate a feature weight map and a foreground weight map. These two weight maps help the model focus more on modifying the foreground part, thereby reducing the interference of the background on the generation process and improving the overall quality of the generated images. The discriminator integrates the ConvNeXt V2 module and also takes the weight map as an input, improving the recognition accuracy of the difference between real and generated images, thereby promoting the generator to generate more realistic and delicate images. Finally, the LPIPS loss is used as the cycle consistency loss during offline training, and this is used as a new loss function to improve the visual quality of the images, as Figure 1 shown below:

[0054] 1. Improve the CycleGAN network

[0055] 1.1. Construct the basic CycleGAN network

[0056] The traditional GAN model mainly consists of a generator and a discriminator. During training, paired training data is usually required. However, in actual image acquisition, it is difficult to obtain image pairs with different styles. The CycleGAN network solves this problem by introducing a cycle consistency loss. The input X-domain image is converted into a Y-domain image through the first generator; then, through the second generator, the generated Y-domain image is converted back into an X-domain image, and the same operation is performed on the image in the Y domain. The input image is converted twice and then compared with the original input image to obtain the cycle consistency loss.

[0057] The structural diagram of the CycleGAN network is as Figure 2 shown below.

[0058] The generator of the CycleGAN network is based on the ResNet architecture and includes an encoder (input layer and two downsampling layers), residual blocks, and a decoder (two upsampling layers and an output layer). The input RGB image first undergoes reflection padding in the input layer

[0059] (Reflection Padding) reduces boundary artifacts, followed by operations of a 7×7 convolutional layer, instance normalization, and ReLU activation. Each downsampling layer consists of operations of a 3×3 convolutional layer, instance normalization, and ReLU activation. After each downsampling, the spatial size is halved and the number of channels is doubled. Each residual block contains a two-layer structure. One layer is for operations of a 3x3 convolutional layer, instance normalization, and ReLU activation, and the other layer is for operations of a 3x3 convolutional layer and instance normalization, aiming to preserve the key features of the image while performing style transformation. Each upsampling layer includes transposed convolution, instance normalization, and ReLU activation operations. After each upsampling, the spatial size is doubled. The output layer uses a 7×7 convolutional layer to restore the resolution consistent with the input image.

[0060] The discriminator of CycleGAN adopts the PatchGan structure, which provides local discriminative ability by splitting the input image into multiple small patches and making independent authenticity judgments for each patch. The discriminator usually adopts a convolutional neural network structure, including an input layer, four convolutional layers, and an output layer. Each convolutional layer includes convolution operations, a normalization layer (such as spectral normalization), and a LeakyReLU activation function. The output layer includes convolution operations and a Sigmoid activation function to generate a single-channel output map and probability values.

[0061] The CycleGAN network adopts a structure that combines an AutoEncoder and SkipConnection. The AutoEncoder is an unsupervised learning method that encodes and decodes data to preserve its structural information. By integrating the AutoEncoder structure within the generator, it can effectively extract the core features of the input image during the encoding stage and reconstruct these features during the decoding stage to generate images in the target domain, thereby enhancing the quality and detail retention ability of the generated images. Additionally, combining the dimensionality reduction ability of the AutoEncoder enables the model to more precisely map the transformation relationship between different domains while maintaining the original style consistency, thereby improving the realism and diversity of the generated results. It is applicable to application scenarios that require high-fidelity image conversion. SkipConnection, also known as a skip connection, is a structural design used in deep neural networks. It directly connects the output of a certain layer to the input of subsequent layers, bypassing some intermediate layers. By adding SkipConnection inside the generator, low-level features (such as edges, textures, etc.) can be directly passed to the high-level output, ensuring that detail information is effectively retained, thereby enhancing the realism and detail accuracy of the generated images. Additionally, SkipConnection enables CycleGAN to more precisely reconstruct the key features of the original image during the cross-domain conversion process, reducing artifacts and distortions, and ultimately achieving a more realistic and high-quality image conversion effect.

[0062] 1.2. Generator of the WMG-GAN Network

[0063] To address the problem that CycleGAN ignores important structural details and focuses on common and easily replicable features, resulting in the generated images being unable to accurately restore the original input images, the present invention optimizes the generator of CycleGAN. Specifically, the generator of the WMG-GAN network of the present invention further processes the output result of the generator of the CycleGAN network using the idea of an attention mechanism: the output result is divided into a foreground weight map and a feature weight map. The foreground weight map is combined with the feature weight map after being processed by the Sigmoid activation, and then added to the weighted part of the original input image to produce the final output result. This structure allows the model to flexibly modify or retain different regions of the input image.

[0064] To learn the most discriminative semantic objects in the X domain and the Y domain while minimizing the modification of the irrelevant background part, the attention mechanism method is adopted inside the generator of the WMG-GAN network to generate the foreground weight map and the feature weight map. In the defect generation task, it is crucial to be able to retain the background of the source domain and the foreground of the target domain. The purpose of the present invention is to learn two mappings between the X domain and the Y domain, namely:

[0065] G X→Y : x → [F y1 , F y2 → G y ,

[0066] G Y→X : y → [F x1 , F x2 → G x ,

[0067] Among them, G x and G y are the generated images; F x1 and F y1 are the feature weight maps of image x and image y respectively; F x2 and F y2 are the foreground weight maps of image x and image y respectively. The foreground weight map F x2 and the foreground weight map F y2 identify the most discriminative foreground regions; the feature weight maps F x1 and the feature weight map F y1 are used to specify the importance degree of each pixel value, which represents the specific part that needs to be converted into the target domain style. The foreground weight map and the feature weight map act on different parts of the image together to ensure that only important semantic objects are converted while the background remains unchanged. Finally, we fuse the input image x, the feature weight map F y1 , and the foreground weight map F x1 to obtain the generated image G y . The final formula can be written as:

[0068] G y = x × (1 - F y2 ) + F y1 × F y2 (1)

[0069] Among them, x × (1 - F y2 ) represents retaining the background of the source domain image; F y1 × F y2 represents focusing on the foreground of the target domain image, that is, the specific region that needs to be changed. Finally, we combine the two to represent the generated image G y that retains the background information and makes changes in the foreground part.

[0070] 1.3. Discriminator of the WMG - GAN network

[0071] The discriminator of CycleGAN mainly focuses on the details of local regions, resulting in the generated images may look realistic locally, but may lack consistency and coherence globally. Therefore, we introduce a ConvNeXt V2 module after the LeakyReLU activation function in the four convolutional layers of the discriminator of CycleGAN as the discriminator of the WMG-GAN network of the present invention to capture complex features. The last special layer (output layer) of the discriminator keeps the spatial size unchanged and continues to increase the depth, outputting a single-channel prediction map for judging whether the input is real or fake.

[0072] 1.3.1, ConvNeXt V2

[0073] ConvNeXt V2 is an advanced convolutional neural network architecture, which is an upgraded version of ConvNeXt. ConvNeXt itself is a fully convolutional network inspired by Transformer, aiming to combine the advantages of convolutional networks and Transformer. ConvNeXt V2 is further optimized on this basis, introducing new designs and technologies to improve the performance and efficiency of the model. Specifically, ConvNeXt V2 introduces a GRN (Global Response Normalization) layer on the basis of ConvNeXt, enhancing the feature competition between channels and improving the selectivity and contrast of features. The main purpose of the GRN layer is to improve the performance of the model by enhancing the key channels in the feature map. It normalizes the global response of each channel, thereby enhancing those feature channels that are more important for the task while suppressing those unimportant channels. The operations of the GRN layer can be divided into the following steps:

[0074] Calculate the global response: For the input feature map X ∈ R H×W×C (where H, W, and C are the height, width, and number of channels of the feature map respectively), calculate the global response of each channel That is, the sum of the squares of all pixel values in this channel.

[0075] Calculate the normalization factor:

[0076] Normalize the global response to obtain the normalization factor of each channel where σ is a very small constant used to prevent division by zero errors.

[0077] Apply the normalization factor:

[0078] Apply the normalization factor to the original feature map X to obtain the normalized feature map Y, that is, Y h,w,i = X h,w,i × s i .

[0079] The GRN layer is similar to an attention mechanism that recalibrates features. ConvNeXt V2 can thus extract high-level features in images more effectively, including textures, edges, and other subtle visual information. The ConvNeXt V2 structure diagram is as Figure 4 shown.

[0080] 1.3.2, Discriminator Guided by Weight Map

[0081] Different from the traditional CycleGAN model, our discriminator is guided by a weight map. Specifically, it is similar to PatchGan in structure, but it receives not only real images and generated images, but also the weight maps F y1 and F y2 as inputs. For example, the input of D y is the pair of generated images [F y1 , G X→Y (x)] and the pair of real images [F y1 , y], and it distinguishes between them.

[0082] In addition, our discriminator also integrates the ConvNeXt V2 module, which improves the selectivity and contrast of features by enhancing competition between features. When there is noise or irregular textures in the background, the traditional PatchGAN may over-focus on these areas, resulting in the generated images being affected by background noise. If the ConvNeXt V2 module is added, the discriminator can better focus on key areas, reduce the interference of background noise, enabling the discriminator to more accurately identify important regions in the image, thereby improving the accuracy of evaluating the quality of generated images and driving the generator to generate finer images.

[0083] 2. Model Offline Training

[0084] 2.1. Dataset Establishment

[0085] The dataset used for the model's offline training was collected from a textile production enterprise. The cloth images were taken by an area array CCD camera facing the exit of the finished cloth on the production line. The images captured by the camera have an extreme aspect ratio of 3072×96 pixels. To address this issue, an image preprocessing method based on sliding window segmentation was proposed, where the sliding window size is 96×96 pixels. The preprocessed sub-images are 96×96 pixels, matching the height of the original 3072×96 image. A 3-pixel overlap is applied in the width direction. Using the sliding window, the image is divided into 96×96 sub-images, and finally the window has an overlap of 93 pixels. A total of 33 sliding windows are used to completely segment the original image.

[0086] Then, using data augmentation techniques such as salient rotation, noise addition, brightness adjustment, etc., the obtained sub-images are expanded into 1459 sub-graphs, including weft defect images (X domain), warp defect images (Y domain), and normal images (Z domain). The images in each domain are divided into a training set and a test set: the X domain has 887 training images and 70 test images; the Y domain has 278 training images and 49 test images; the Z domain has 145 training images and 30 test images. Examples of warp defects, weft defects, and defect-free pictures are shown as Figure 7 shown below.

[0087] 2.2. Loss Function

[0088] 2.2.1. Similarity Loss LPIPS

[0089] The cycle consistency loss is usually calculated by comparing the difference between the image after two transformations and the original image. Common measurement methods include L1 and L2 norms. To further improve the cycle transformation effect of the WMG-GAN network and enhance the model's perception ability of image quality changes, we adopt the similarity loss LPIPS (Learned Perceptual Image Patch Similarity) as the cycle consistency loss in WMG-GAN. LPIPS measures the similarity between images through the high-level features extracted by a pre-trained feature extraction network. Different from traditional methods based on pixel-level differences (such as L1 or L2 loss), the LPIPS loss can better capture the perception of image changes by the human visual system. The LPIPS calculation method is as follows: A set of images x and x0 are fed into the network F for feature extraction, and then the distance between the features of x and x0 in different channels is calculated. Then, feature stacks are extracted in different convolutional layers, and the feature stacks in different channels are normalized. The normalized result is written as Subsequently, through scale and activate the channels. Finally, the L2 norm is used to calculate the distance, and the average value is taken in the space and the sum is taken over the channels to calculate d0. The calculation formula can be written as:

[0090]

[0091] where wl is the result calculated by the cosine distance formula, h is the height of the input image, w is the width of the input image, L refers to the selected network layer, and are the feature responses of images x and x0 on this channel respectively, H L is the number of feature maps of this layer, and W L is the weight of the feature map of this channel.

[0092] Finally, d0 and d1 are passed into a model with a RELU fully connected layer containing two 32-channel layers, a single-channel fully connected layer, and a sigmoid layer for training. The similarity loss is as follows:

[0093] L LPIPS (x, x0, x1, h) = -h log G(d(x, x0), d(x, x1)) - (1 - h) log(1 - G(d(x, x0), d(x, x1))) (3)

[0095] Among them, x1 represents the image predicted by the pre-trained model; d1 is the true value calculated by the pre-trained model.

[0096] 2.2.2. Final loss function

[0097] In addition to the LPIPS loss, we also use adversarial loss and weight map-guided adversarial loss.

[0098] Adversarial loss means that for a generator G from domain X to domain Y and a discriminator D, the adversarial loss function encourages the generator G to generate an image for which the discriminator D gives a probability close to 1 (i.e., an image considered to be real). The goal of this method is to correctly distinguish between real images generated by the generator G and fake images. For the discriminator D, it is necessary to maximize the probability that real images are judged to be real and the probability that fake images (images generated by the generator) are judged to be fake. For the mapping G X→Y , the adversarial loss can be expressed by the following formula:

[0099]

[0100] The generator G tries to minimize this objective, while the discriminator D tries to maximize it. For another mapping G Y→X , the adversarial loss can be defined as:

[0101]

[0102] Among them, the symbol E represents the mathematical expectation operator.

[0103] The role of the weight map-guided adversarial loss is to keep the background unchanged as much as possible. This can make the feature weight maps generated by real images and generated images as consistent as possible, so that key information can be retained. For the mapping G X→Y , the weight map-guided adversarial loss can be expressed by the following formula:

[0104]

[0105] Among them, D Y 's task is to distinguish and compare the generated image pair pFy1 , G X→Y (x)] and the real image pair [F y1 , y].

[0106] The final complete loss function is defined as follows:

[0107] L = L GAN + L F + L LPIPS (7)

[0108] Among them, L GAN represents the adversarial loss, L F represents the weight map-guided adversarial loss, and L LPIPS represents the similarity loss.

[0109] 2.3. Offline training and testing process

[0110] The labeled images in the training set are input into the WMG-GAN network for training. The initial learning rate of the Adam optimizer is 0.0002 and is dynamically adjusted using an exponential decay strategy after 100 epochs. The batch size is set to 16, and the number of epochs is set to 400. During the training process, the loss function value is calculated and the model parameters are iteratively optimized through backpropagation. The training ends after the 400th epoch is completed.

[0111] The testing process is to put the test set images into the trained model to generate new images and calculate various evaluation metrics of the images. Among them, for the generation of weft direction defect images, FID = 15.6935, SSIM = 0.4784, PSNR = 31.1031, and LPIPS = 0.0799; for the generation of warp direction defect images, FID = 20.9138, SSIM = 0.2489, PSNR = 20.8625, and LPIPS = 0.0849; for the generation of normal images, FID = 20.8715, SSIM = 0.1722, PSNR = 20.2126, and LPIPS = 0.1209. All meet the preset index requirements for online production use, thus obtaining the offline trained WMG-GAN network.

[0112] 3. Using the offline trained WMG-GAN network

[0113] Use a CCD camera to collect pictures of the cloth with defects, and cut them into a group of sub-images of 96x96 pixels using a moving window preprocessing method. Then, the sub-images are sequentially input into the offline trained WMG-GAN network obtained in step 2 to generate new images, thereby obtaining cloth defect pictures with different defect type markings, achieving the purpose of expanding the cloth defect images.

[0114] 4. Experiments

[0115] 4.1. The experiments adopted the dataset of Example 1

[0116] 4.2. Evaluation Metrics

[0117] The experimental evaluation metrics include FID, SSIM, PSNR, and LPIPS, and the details are introduced as follows:

[0118] FID: FID calculates the mean and covariance matrix of the feature vectors of the generated images and the real images, and then calculates the distance between the generated distribution and the real distribution. The lower the FID value, the smaller the semantic difference between the generated image and the real image, and the higher the quality of the generated image. The calculation formula can be written as:

[0119]

[0120] where is the L2 norm of the square of the difference between the mean vectors, Tr represents the trace of the matrix (i.e., the sum of the diagonal elements of the matrix), is the covariance matrix Σ g and Σ r is the square root of the product of the two matrices, representing the matrix obtained by taking the square root of the eigenvalues of the product of the two matrices.

[0121] SSIM: SSIM is a metric used to measure the structural similarity between two images. It calculates the similarity by comparing the brightness, contrast, and structural information of the images. It takes into account the perception of the human visual system for images and performs particularly well when dealing with images with complex textures and structures. Its calculation formula can be written as:

[0122]

[0123] where μ x and μ y are the averages of images x and y respectively; σ x 2 and σ y 2 are the variances of images x and y respectively; σ xy is the covariance of images x and y; c1 and c2 are small constants used to stabilize the denominator.

[0124] PSNR: PSNR measures the degree of image distortion by calculating the mean square error (MSE) between the original image and the processed image. Although PSNR is simple and intuitive to calculate, it mainly focuses on pixel-level differences and may not fully reflect the perceptual quality of the human visual system. Therefore, combining other more complex perceptual quality evaluation metrics (such as SSIM or LPIPS) may provide a more comprehensive evaluation. Its calculation formula can be written as:

[0125]

[0126] Among them, MSE is the mean square error of two images; MaxValue is the maximum value that image pixels can take.

[0127] LPIPS: LPIPS obtains the final similarity score by weighted summation of features at different layers. This method can better reflect the perception of image differences by the human visual system.

[0128] 4.3. Ablation Experiment

[0129] To systematically evaluate the impact of each component on the model performance, we conducted ablation experiments. The basic CycleGAN model was successively improved in the generator (G-CycleGAN), improved in the generator + discriminator (GD-CycleGAN), and improved in the generator + discriminator + loss function (WMG-GAN). Then, the dataset of Example 1 was respectively passed through each comparison network to generate extended defective images, and evaluation metrics FID, SSIM, PSNR, and LPIPS were calculated. The experimental results are listed in Tables 1, 2, and 3 respectively.

[0130] Table 1 Comparison Results of Evaluation Metrics for Extended Images of Latitudinal Defects (X Domain)

[0131]

[0132] Table 2 Comparison Results of Evaluation Metrics for Extended Images of Longitudinal Defects (Y Domain)

[0133]

[0134] Table 3 Comparison Results of Evaluation Metrics for Extended Images of Normal Images (Z Domain)

[0135]

[0136] The above experimental results show that: when only the generator is modified, that is, only the weight map is used to guide the generation of images, this will lead to a decrease in the quality of some image generations and the loss of overall structural information of the images. This is because the weak feature extraction ability of the traditional PatchGAN discriminator makes it difficult for the model to capture the subtle differences and high-level features in the images, thus unable to provide accurate feedback to the generator. Moreover, our model relies on the weight map to distinguish foreground and background regions. If the discriminator cannot accurately identify these regions, the generated images may have problems of background and foreground confusion. This will result in the generated images by the generator lacking rich details and poor realism. However, when the discriminator is improved, the quality of the generated images will be significantly improved. When LPIPS is further introduced as the cyclic consistency loss, the performance of the model will be further improved.

[0137] 4.4. Comparative Experiments

[0138] When evaluating the performance of WMG-GAN for image generation, a group of models including GAN, DCGAN, ACGAN, CycleGAN, and WMG-GAN were compared. All models were trained for 400 epochs on the same dataset.

[0139] Tables 4, 5, and 6 respectively present the evaluation metrics of various defects in the fabric images generated by different GAN models.

[0140] Comparison Results of Evaluation Metrics for Extended Images of Weft Defects (X Domain) in Table 4

[0141]

[0142] Comparison Results of Evaluation Metrics for Extended Images of Warp Defects (Y Domain) in Table 5

[0143]

[0144] Comparison Results of Evaluation Metrics for Extended Images of Normal Images (Z Domain) in Table 6

[0145]

[0146] Based on the comparative experiments, it can be seen that our model demonstrates excellent performance in generating various defect images. Specifically, our model achieves a lower FID value, indicating that the generated defect images are highly similar to real defect images in terms of quality and diversity. In terms of SSIM, our model also scores highly, which indicates a very high structural similarity between the generated defect images and the original defect images, thus confirming the effectiveness of the model in retaining structural information. In addition, our model shows excellent performance in terms of PSNR, with a high score, reflecting the superior reconstruction quality and low distortion of the generated defect images, demonstrating the accuracy and stability of the model in the image generation process. Finally, our model obtains a lower LPIPS score, indicating a high perceptual similarity between the generated defect images and the original images, further enhancing the model's ability to generate high-quality and perceptually realistic defect images.

[0147] 4.5. Intuitive Comparison of Images Generated by Different Networks

[0148] In order to more directly and clearly demonstrate that the improved model can more stably retain the main features of the defect in the process of generating images, and to intuitively demonstrate and compare the generation capabilities of defect images of different GAN models, we conducted image generation experiments in the X, Y and Z domains using GAN, DCGAN, ACGAN, CycleGAN and WMG-GAN models. The experimental results are shown in Figure 2. Figure 8 The WMG-GAN network of the present invention can complete accurate and clear expansion generation tasks for normal images, warp defect images, or weft defect images.

[0149] Finally, it should be noted that the above examples are only some specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments, and there are many variations. All variations that can be directly derived or associated with the content disclosed by a person skilled in the art should be considered as the protection scope of the present invention.

Claims

1. A method for generating cloth defect images based on weight map guidance, characterized in that: The method includes collecting defective cloth images and preprocessing them into a group of sub-images, and then inputting the sub-images into the offline trained WMG-GAN network in sequence to generate cloth defect images with defect type labels; The WMG-GAN network is an improvement based on the CycleGAN network. The output image of the CycleGAN network generator is added with the idea of ​​using the attention mechanism. The output image is divided into a foreground weight map and a feature weight map. The foreground weight map is combined with the feature weight map after Sigmoid activation processing, and then the weighted part of the original input image is added to generate the generated image G. y ; The discriminator adopts the PatchGan structure, and adds a ConvNeXtV2 module after the LeakyReLU activation function of each intermediate layer of the CycleGAN network.

2. The method for generating cloth defect images based on weight map guidance according to claim 1, characterized in that: The generated image G y for: G y =x×(1-F y2 )+F y1 ×F y2 (1) Among them, F y1 is the feature weight map of image y; F y2 is the foreground weight map of image y, and x represents the input image.

3. The method for generating cloth defect images based on weight map guidance according to claim 2, characterized in that: The process of offline training of the WMG-GAN network is as follows: After collecting cloth images and undergoing image preprocessing and data enhancement, the samples are expanded to include weft defect images, warp defect images, and normal images and annotated, and then divided into training sets and test sets. The annotated images in the training set are input into the WMG-GAN network for training. During the training process, the loss function value is calculated and the model parameters are optimized through back propagation iteration. The training ends when the preset number of epochs is reached; The labeled images in the test set are input into the trained model to generate new images, and the various evaluation indicators of the images are calculated to meet the preset standards, thereby obtaining the offline trained WMG-GAN network.

4. The method for generating cloth defect images based on weight map guidance according to claim 3, characterized in that: The data enhancement techniques include salient rotation, noise addition, and brightness adjustment.

5. The method for generating cloth defect images based on weight map guidance according to claim 4, characterized in that: The image preprocessing is to perform image segmentation based on a sliding window.

6. The method for generating cloth defect images based on weight map guidance according to claim 5, characterized in that: The loss function is: L=L GAN +L F +L LPIPS (7) in, Mapping G X→Y The adversarial loss is: Mapping G Y→X The adversarial loss is: Mapping G X→Y The adversarial loss guided by the weight graph of is: Where D Y The task is to distinguish and compare the generated image pairs [F y1 ,G X→Y (x)] and the real image pair [F y1 ,y]; L LPIPS is the similarity loss function: A set of images x and x0 are fed into the network F for feature extraction, and then the distance between the features of x and x0 in different channels is calculated. Then, the feature stacks are extracted in different convolutional layers and the feature stacks in different channels are normalized. The normalized result is written as Then through w Scaling activation channels are performed; finally, the distance is calculated using the L2 norm, and the average is taken in space, and d0 is calculated by summing over the channels: Among them, wl is the result calculated by the cosine distance formula, h is the height of the input image, w is the width of the input image, and L refers to the selected network layer. and are the feature responses of images x and x0 on this channel, H L is the number of feature maps in this layer, W L is the weight of the channel feature map; Pass d0 and d1 to a model containing two 32-channel RELU fully connected layers, a single-channel fully connected layer, and a sigmoid layer for training. The similarity loss function is: L LPIPS (x,x0,x1,h)=-h log G(d(x,x0),d(x,x1))-(1-h)log(1-G(d(x,x0),d(x,x1))) (3) Among them, x1 represents the image predicted by the pre-trained model; d1 is the true value calculated by the pre-trained model.