A Defect Image Generation Algorithm, Storage Medium and Device Based on Deep Learning

Through the feature processing framework and background maintenance network combined with multi-scale isometric convolution network and bidirectional LSTM, the problem of insufficient feature extraction and insufficient background effect of traditional GANs in multi-scale defects is solved, and the generated image quality and downstream recognition performance are significantly improved.

CN120013905BActive Publication Date: 2025-07-08ANHUI POLYTECHNIC UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510099667.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-07-08
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

In the prior art, traditional GANs have problems of insufficient extraction or loss of features when generating multi-scale defect features, and it is difficult to take into account both defect feature generation and background effects, resulting in insufficient generated image quality, especially in complex backgrounds, which affects the training of detection models.

Method used

A feature processing framework combining multi-scale isometric convolution networks and bidirectional LSTMs is adopted, and parallel feature extraction is performed through convolution kernels of different scales of 1×1, 3×3 and 5×5. The forward and backward propagation mechanisms of bidirectional LSTMs are adaptively allocated feature weights, and a background maintenance network and Gram matrix constraints are introduced to design a complete loss function system.

Benefits of technology

It effectively improves the quality of generated images, ensures that the comprehensive representation of defect features and background consistency is maintained, and improves the performance of downstream recognition tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013905B_ABST
    Figure CN120013905B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image generation, and particularly relates to a defect image generation algorithm, a storage medium and a device based on deep learning. The algorithm includes the following steps: Step S1, constructing a defect image generation model; Step S2, training and optimizing the defect image generation model; Step S3, inputting a number of real defect-free images based on the trained defect image generation model, and the defect image generation model outputs corresponding generated defect images to achieve the expansion of the data set. Among them, the defect image generation model includes two generator modules and two discriminators. The first generator module sequentially includes a mask module, a generator one and a background preservation network. The second generator module sequentially includes a generator two and a background preservation network. The present invention adopts a processing strategy of using a mask to separate the defect area and the non-defect area, and innovatively introduces a background preservation network, effectively solving the problem of background texture distortion that easily occurs in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image generation, and particularly relates to a defect image generation algorithm, a storage medium, and a device based on deep learning. Background Art

[0002] Defect detection based on deep learning is an important link for the industrial sector to improve product quality, and the quality of the defect image dataset has an important impact on the deep learning network. In the actual industrial production process, since obtaining real defect samples requires a large amount of human cost and the naturally generated defect data is limited, there is generally a problem of insufficient quantity in the defect dataset, which severely limits the performance of subsequent deep learning models.

[0003] In the prior art, there are already methods for generating defect images through a generative adversarial network (GAN). However, these methods have the following problems: Traditional GANs have limitations in the transferability of generated images. When extracting multi-scale defect features of different sizes, problems such as insufficient extraction or feature loss are likely to occur; it is difficult to simultaneously balance defect feature generation and background effect preservation, resulting in limitations in the quality of generated images; the generation ability is limited under complex backgrounds, and there are perceptible differences between the generated samples and real samples, affecting the training of the detection model. Summary of the Invention

[0004] The purpose of the present invention is to provide a defect image generation algorithm based on deep learning, which is used to solve the technical problems in the prior art that it is impossible to extract and reproduce multi-scale defect features, and the image quality is insufficient when generating sample images containing defects, and it is impossible to balance defect features and background effects.

[0005] The described defect image generation algorithm based on deep learning includes the following steps.

[0006] Step S1, construct a defect image generation model;

[0007] Step S2, train and optimize the defect image generation model;

[0008] Step S3, based on the trained defect image generation model, input a number of real defect-free images, and the defect image generation model outputs corresponding generated defect images to expand the dataset;

[0009] Among them, the defect image generation model includes two generator modules and two discriminators. The two generator modules are the first generator module and the second generator module, corresponding to generator one and generator two respectively; the two discriminators are discriminator one and discriminator two respectively. The first generator module is used to convert the input real defect-free image into a generated defect image, and to convert the input generated defect-free image into a reconstructed defect image; the second generator module is used to convert the input real defect image into a generated defect-free image, and to convert the input generated defect image into a reconstructed defect-free image; discriminator one is used to judge whether the generated defect image is real, and discriminator two is used to judge whether the generated defect-free image is real; the first generator module sequentially includes a mask module, generator one and a background preservation network, and the second generator module sequentially includes generator two and a background preservation network.

[0010] Preferably, the mask module is used to divide the input image into a defect area and a non-defect area using a mask. Generator one includes a multi-scale feature extraction and bidirectional LSTM fusion module. The multi-scale feature extraction and bidirectional LSTM fusion module is used to capture defect features at different scales and adaptively fuse the defect features at different scales using bidirectional LSTM technology for temporal modeling.

[0011] Preferably, the features at different scales extracted are concatenated to form the result F of multi-scale feature extraction out , and the corresponding arithmetic expression is:

[0012] F out = Concat[F1, F2, F3],

[0013] where, F i represents the feature map of the i-th scale branch, i = 1, 2, 3; in the feature fusion stage, first a reshaping operation is performed on the feature map of each scale branch, and the corresponding operation process is expressed as the following arithmetic expression:

[0014]

[0015] where, F i view represents the result of the i-th scale branch obtained by the view operation, and x i represents the result of the i-th scale branch obtained by the permute operation. B / C / H / W are the batch number / channel number / height / width of the feature map. The results x i obtained from each scale branch after processing together form multi-scale defect features;

[0016] After that, a bidirectional LSTM is used to perform temporal modeling on the multi-scale defect features, and the bidirectional LSTM is utilized to determine the adaptive fusion weights. The defect image generation model dynamically adjusts the weights of each scale according to the importance of the extracted defect features.

[0017] Preferably, during the fusion process of the bidirectional LSTM, the hidden state is updated in the forward and backward propagation processes, and the features in the two directions are concatenated to obtain the bidirectional feature representation. Then, an attention mechanism is introduced to calculate the fusion weights, and the softmax function is used to convert the original scores into the weight distribution w:

[0018] where Pool(·) represents the temporal pooling operation, W is the fusion weight of the two hidden state outputs and b is the corresponding bias term; then the fusion weight coefficients w of each layer are calculated i , and the calculation formula is:

[0019]

[0020] where z i represents the original score of the i-th layer obtained by temporal pooling, and z j represents the original score of the j-th layer obtained by temporal pooling; the output feature representation obtained by multi-scale feature fusion is:

[0021]

[0022] where F' out represents the output feature, w i represents the fusion weight coefficients of each layer, W i is the fusion weight of the corresponding layer, and b i is the bias term of the corresponding layer.

[0023] Preferably, the LSTM calculates the hidden state of the current moment of a single LSTM through the collaborative work of multiple gating units, and the calculation formula is:

[0024]

[0025] where represents the concat operation of the feature maps F i of the three scale branches, and P(V(F i )) represents performing the View operation and the Permute operation on the feature map F i in sequence, σ[W f (·)+b f represents the forget gate, and W f and b f represent the corresponding weights and bias terms respectively, and tanh[Wc (·) + b c represents the update of the cell state, W c and b c represent the corresponding weight and bias terms respectively, C i represents the cell state, ⊙ represents element-wise multiplication, σ[W o (·) + b o represents the output gate processing, W o and b o represent the corresponding weight and bias terms respectively, h(t) represents the final hidden state;

[0026] The fusion process of the bidirectional LSTM includes updating the hidden state h(t) to the forward and backward propagation processes through its own calculation formula, including:

[0027] Forward propagation process: Backward propagation process: Feature concatenation in two directions: Among them, f t represents the input feature at time t, and represent the hidden state outputs of the forward LSTM and the backward LSTM respectively.

[0028] Preferably, the background retention network includes a VGG19 feature extraction network and a Gram matrix. First, the multi-scale deep features are extracted by the VGG19 feature extraction network, which is represented by the following formula: φ l (I) = Networks l (I), l ∈ {relu11, relu21, relu31, relu41},

[0029] Among them, φ l represents the feature map of the l-th layer of the VGG19 feature extraction network, relu11, relu21, relu31, relu41 represent the four layers of the VGG19 feature extraction network, and Networks represents the VGG19 feature extraction network;

[0030] Then, the statistical information is compared through the Gram matrix to guide the training of the subsequent generator; the defect image generation model introduces a background retention loss, and the background retention loss combines two-level constraints of the background texture loss and the background style loss; in the background retention network, the loss function of the background texture loss corresponding to the non-defect area is:

[0031]

[0032] Among them, M is the mask information of the defect area in the image, λ l is the weight coefficient of the l-th layer feature, Igen and I real are the image generated by the generation module and the real image respectively; and the loss function of the corresponding background style loss is as follows:

[0033]

[0034] Among them, is the Gram matrix, and the corresponding calculation formula is:

[0035]

[0036] Among them, C l / H l / W l is the number of channels / height / width of the feature map, is the normalization factor.

[0037] Preferably, the mask module uses a mask to divide the training image into a defect area and a non-defect area based on the defect image, introduces an identity loss, and the corresponding loss function is as follows:

[0038]

[0039] Among them, M is the mask information of the defect area in the image, and G(·) represents Generator 1.

[0040] Preferably, the total loss function of the defect image generation model is:

[0041]

[0042] Among them, is the adversarial loss function, and λ GAN is the corresponding weight; is the cycle consistency loss function, and λ cyc is the corresponding weight; is the identity loss, used to optimize the mask division result, and λ idt is the corresponding weight; is the background texture loss, and λ background is the corresponding weight; is the background style loss, and λ style is the corresponding weight.

[0043] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and is characterized in that: when the computer program is executed by a processor, it implements the steps of a defect image generation algorithm based on deep learning as described above.

[0044] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor, wherein: when the processor executes the computer program, the steps of a defect image generation algorithm based on deep learning as described above are implemented.

[0045] The present invention has the following advantages:

[0046] 1. The algorithm provided by the present invention innovatively introduces a feature processing framework that combines a multi-scale equidistant convolutional network and a bidirectional LSTM. By using three different scales of convolutional kernels, namely 1×1, 3×3, and 5×5, for parallel feature extraction, local details, medium-scale structures, and large-scale context information of defects can be captured simultaneously. The forward and backward propagation mechanisms of the bidirectional LSTM can fully explore the temporal correlations between these multi-scale features, and adaptively allocate the importance weights of different features through an attention mechanism, thereby obtaining a more comprehensive and accurate defect feature representation and effectively improving the quality of the generated images.

[0047] 2. The present invention adopts a processing strategy of using a mask to separate the defect area and the non-defect area, and innovatively introduces a background preservation network and a Gram matrix constraint to maintain the features of the non-defect area. The background preservation network represents the background information by extracting feature maps at different levels, and the Gram matrix constrains the consistency of the background style by calculating the internal correlations of the feature maps. This dual constraint mechanism ensures that the generated images can maintain background features highly consistent with the original images in the non-defect area, effectively solving the problem of background texture distortion easily occurring in the prior art.

[0048] 3. The present invention designs a complete loss function system, including multiple components such as adversarial loss, cycle consistency loss, identity loss, background preservation loss, and style consistency loss. This multiple loss constraint ensures that the generated images can achieve good effects in multiple aspects such as defect feature performance, background texture preservation, and overall authenticity. Experimental results show that the images generated by the method of the present invention not only have higher structural similarity and peak signal-to-noise ratio, but also exhibit significantly better performance in downstream recognition tasks, verifying the effectiveness of this method in defect image generation and data augmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is the basic flowchart of a defect image generation algorithm based on deep learning of the present invention.

[0050] Figure 2 is the structural diagram of the defect image generation model of the present invention.

[0051] Figure 3 is the flowchart of the first generator module in the present invention.

[0052] Figure 4 It is the structural diagram of the adaptive multi-scale feature bidirectional LSTM in the present invention.

[0053] Figure 5 It is the structural diagram of Generator 1 in the present invention.

[0054] Figure 6 It is the structural diagram of Generator 2 in the present invention.

[0055] Figure 7 It is the structural diagram of the discriminator in the present invention.

[0056] Figure 8 It is the structural diagram of the feature extraction network used in the background preservation module of the present invention.

[0057] Figure 9 It is the comparison chart of the defect generation results between the present invention and the prior art. Among them, the original image represents the real defect-free image, and color, cut, Hole, Metal, and Thread represent five types of defects. Cycle-GAN represents the effects of five types of defects generated by the cycle_gan model. DFM-GAN represents the effects of five types of defects generated by the DFM-GAN model. Dual-GAN represents the effects of five types of defects generated by the Dual-GAN model. SD-GAN represents the effects of five types of defects generated by the SD-GAN model. SD-GAN represents the effects of five types of defects generated by the SD-GAN model.

[0058] Figure 10 It is the comparison effect diagram of the defect generation results between the present invention and the prior art in terms of indicators. Among them, in each evaluation index, the SSIM index evaluates the structural similarity of the image, and the higher the score, the more similar the structure of the generated image is to the reference image; the PSNR index evaluates the image quality, and the higher the value, the smaller the image distortion; the MSE and MAE indexes are used to evaluate the pixel-level error of the image, and the lower the value, the closer the generated image is to the reference image; LPIPS evaluates the perceptual similarity of the image, and the lower the value, the better the perceptual quality of the generated image; the FID index evaluates the authenticity and diversity of the generated image, and the lower the value, the closer the distribution of the generated image is to the real image.

[0059] Figure 11 It is the comparison effect diagram of the background preservation of the defect generation results between the present invention and the prior art in terms of indicators.

[0060] Figure 12 It is the comparison effect diagram of the background of the defect generation results between the present invention and the prior art.

[0061] Figure 13This is the experimental flowchart for training the defect generation results of the present invention and the prior art by feeding them into the recognition network. To verify the effectiveness of the defect images generated by different models for downstream tasks, a dataset authenticity experiment and a dataset augmentation experiment were designed. The authenticity experiment used the Datasets1 dataset. It consisted of 1000 images generated by each model. These images were input into the recognition networks of ResNet50 and DenseNet121 for training and then directly used to recognize real fabric defect images. The recognition network served as a discriminator to judge the authenticity of the generated dataset. The augmentation experiment used Datasets2, which was a dataset composed of 1000 images generated by each model combined with 234 real defect images to simulate the effect of using generated images for training when the dataset was too small. These augmented datasets were also input into the above recognition network to obtain weights. To comprehensively evaluate the performance of the dataset for deep learning, pre-trained weights were used for training and testing at the same time.

[0062] Figure 14 This is the experimental result graph of training the defect generation results of the present invention and the prior art by feeding them into the recognition network. The generated images are highly similar to the real images. On the Datasets2 test set containing real images, the performance of FDIE-GAN is further improved, indicating that the images generated by FDIE-GAN can effectively improve the performance of the recognition network. Detailed implementation manners

[0063] The following is a more detailed description of the specific implementation manners of the present invention by referring to the accompanying drawings and describing the embodiments, so as to help those skilled in the art have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.

[0064] Embodiment 1.

[0065] As Figures 1 - 8 shown, the present invention provides a defect image generation algorithm based on deep learning, including the following steps.

[0066] Step S1, constructing a defect image generation model based on the CycleGAN network.

[0067] The defect image generation model includes two generator modules each containing a corresponding generator and two discriminators. The two generator modules are the first generator module and the second generator module, corresponding to generator one (G) and generator two (F) respectively. The two discriminators are discriminator one (D y ) and discriminator two (D x) The first generator module is used to convert the input real defect-free image into a generated defect image and to convert the input generated defect-free image into a reconstructed defect image; the second generator module is used to convert the input real defect image into a generated defect-free image and to convert the input generated defect image into a reconstructed defect-free image; discriminator one is used to determine whether the generated defect image is real, and discriminator two is used to determine whether the generated defect-free image is real; the first generator module sequentially includes a mask module, generator one, and a background preservation network, and the second generator module sequentially includes generator two and a background preservation network.

[0068] The mask module is used to divide the input image into a defect area and a non-defect area using a mask. During training, the mask module is obtained by manually marking defect images through LabelImg, which improves training efficiency and accuracy and ensures accurate segmentation of the defect area.

[0069] Generator one includes an encoder, a multi-scale feature extraction and bidirectional LSTM (Long Short-Term Memory Network) fusion module, a decoder, and an output layer. The encoder is used to perform feature extraction and downsampling on the image segmented by the mask. The multi-scale feature extraction and bidirectional LSTM fusion module is used to capture defect features at different scales and uses bidirectional LSTM technology for temporal modeling to adaptively fuse defect features at different scales. In this way, the encoder of generator one can dynamically adjust the weights of defect feature information at each scale according to the extracted defect features and adaptively select the most suitable feature representation for different defect features. For example, different scales of convolutional kernels are required to extract small-area hole-like defects and large-area cut-and-twist-like defects, so this solution can avoid insufficient feature extraction or feature loss.

[0070] Generator two uses a U-Net architecture as the backbone network and is optimized for feature extraction and reconstruction of fabric defect images.

[0071] Both discriminators are composed of five convolutional layers. In the discriminator, the first four convolutional layers use the ReLU activation function, and a batch normalization layer is set after each convolutional layer except the first one. The last layer uses the Tanh activation function. The size of the convolutional kernel for each layer is 4×4, the stride of the first three layers is 2, and the stride of the last two layers is 1. The number of channels of the features gradually changes from 64 in the first layer to 1 in the final output layer. The output of the discriminator is a discrimination matrix, where each element corresponds to a local receptive field in the input image, and the receptive field size is 63×63 pixels. This enables the discriminator to perform true / false discrimination on local areas of the fabric image, forcing the generator to not only focus on overall authenticity during the adversarial learning process but also pay attention to maintaining the fineness of local structures and defect features.

[0072] The background preservation network includes a VGG19 feature extraction network and a Gram matrix. First, the VGG19 feature extraction network is used to extract deep features at multiple scales, and then the Gram matrix is used to compare statistical information, thereby guiding the training of the subsequent generator, especially to achieve the preservation of image background information.

[0073] Step S2: Train and optimize the defect image generation model through adversarial games.

[0074] In step S2, the mask module is obtained by manually marking the defect image through LabelImg, which improves the training efficiency and accuracy and ensures the accurate segmentation of the defect area.

[0075] The input image of the first generator module is first divided into regions through the mask module, and then a multi-scale equidistant convolutional network is used in generator one (G) to extract features of the defect area. The multi-scale equidistant convolutional network contains three parallel convolutional branches, which use 1×1 convolution, 3×3 convolution, and 5×5 convolution kernels respectively, and cooperate with different dilation rates for feature extraction. Among them, the 1×1 convolution mainly extracts local detail features, the 3×3 convolution extracts medium-scale structural features, and the 5×5 convolution extracts large-range context information. This step expands the receptive field without increasing the number of parameters by setting different dilation rates, enabling the network to comprehensively capture defect features at different scales. The features extracted at different scales are connected to form the result F of multi-scale feature extraction. out , and the corresponding arithmetic expression is:

[0076] F out = Concat[F1, F2, F3],

[0077] where F i represents the feature map of the i-th scale branch, i = 1, 2, 3. In the feature fusion stage, the present invention first performs a reshaping operation on the feature map of each scale branch, that is, the result F of multi-scale feature extraction out is obtained through the view operation to get results at different scales, and then the feature recombination of each scale branch is realized through the corresponding permute operation; the corresponding operation process is expressed as the following arithmetic expression:

[0078]

[0079] where F i view represents the result of the i-th scale branch obtained through the view operation, and x i represents the result of the i-th scale branch obtained through the permute operation, and is the number of channels / height / width of the feature map. Such a reshaping operation flattens the spatial dimension and adjusts the dimension order, so that the results x obtained by each scale branch after processingi Meet the input requirements of the LSTM, and these results x i Together, they form multi-scale defect features.

[0080] After that, a bidirectional LSTM is used to perform temporal modeling on the multi-scale defect features. Among them, the individual LSTM is as Figure 3 shown. Through the collaborative work of multiple gated units, the hidden state of the individual LSTM at the current moment is finally calculated by h(t). The calculation formula is:

[0081]

[0082] Among them, represents the concat operation of the feature maps F i of the three-scale branches, and P(V(F i )) represents performing the View operation and the Permute operation on the feature map F i in sequence. σ[W f (·)+b f represents the forget gate, W f and b f represent the corresponding weight and bias terms respectively. tanh[W c (·)+b c represents the cell state update, W c and b c represent the corresponding weight and bias terms respectively. C i represents the cell state, ⊙ represents element-wise multiplication, and σ[W o (·)+b o represents the output gate processing. W o and b o represent the corresponding weight and bias terms respectively. h(t) represents the final hidden state.

[0083] The fusion process of the bidirectional LSTM includes updating the hidden state h(t) to the forward and backward propagation processes through its own calculation formula, including:

[0084] Forward propagation process: Backward propagation process: Feature concatenation in two directions: Among them, f t represents the input feature at time t, and represent the hidden state outputs of the forward LSTM and the backward LSTM respectively.

[0085] After obtaining the bidirectional feature representation, in order to achieve the adaptive fusion of multi-scale features, this step introduces an attention mechanism to calculate the fusion weights and uses the softmax function to convert the original scores into a weight distribution w:

[0086] Among them, Pool(·) represents the temporal pooling operation, W is the fusion weight of the two hidden state outputs and b is the corresponding bias term. Then, the fusion weight coefficients w of each layer are calculated i , and the calculation formula is:

[0087]

[0088] where, z i represents the original score of the i-th layer obtained by temporal pooling, and z j represents the original score of the j-th layer obtained by temporal pooling. Combining the above formulas, the output feature obtained by multi-scale feature fusion is expressed as:

[0089]

[0090] where, F' out represents the output feature, w i represents the fusion weight coefficients of each layer, W i is the fusion weight of the corresponding layer, and b i is the bias term of the corresponding layer. The bidirectional LSTM is used to determine the adaptive fusion weight, so that the defect image generation model (FDIE-GAN model) can dynamically adjust the weights of each scale according to the importance of the extracted defect features, and adaptively select the most suitable feature representation for different defects.

[0091] After that, the decoder adopts an upsampling structure symmetric to the encoder to reconstruct the fused feature map into a defect generation image. Each layer of the decoder is equipped with instance normalization and ReLU activation function after upsampling to ensure the quality of feature reconstruction, and finally the final generated defect image is output through a 7×7 convolutional layer.

[0092] Similarly, the first generator module also converts the input generated defect-free image into a reconstructed defect image.

[0093] Meanwhile, the second generator module is used to convert the input real defect image into a generated defect-free image, specifically including: in the feature encoding stage, Generator Two gradually compresses the spatial dimension of the input image through successive downsampling convolutional modules, while capturing background features and defect features. At the position with the lowest feature dimension, the network sets a bottleneck convolutional layer to effectively compress and reorganize the key feature information of the fabric defects. In the feature decoding stage, the network adopts a structure symmetric to the encoder, and gradually restores the spatial resolution of the feature map through a series of transposed convolutional modules. Generator Two also introduces a skip connection mechanism between the corresponding encoding and decoding layers to directly transfer the background features and defect features on the fabric surface to the deeper layers, ensuring that key defect information will not be lost during the image generation process. Finally, Generator Two normalizes the feature map through a Softmax layer and outputs the generated defect-free image. This module is also used to convert the input generated defect image into a reconstructed defect-free image.

[0094] During training, Discriminator One is used to judge whether the generated defect image is real, and Discriminator Two is used to judge whether the generated defect-free image is real; and a loss function is used to optimize the model, so as to make the generated image increasingly approximate the real image. The specific relevant loss functions are as follows.

[0095] The total loss function of the defect image generation model is:

[0096]

[0097] Among them, is the adversarial loss function, and λ GAN is the corresponding weight; is the cycle consistency loss function, and λ cyc is the corresponding weight; is the identity loss, which is used to optimize the mask division result, and λ idt is the corresponding weight; is the background texture loss, and λ background is the corresponding weight; is the background style loss, and λ style is the corresponding weight.

[0098] The adversarial loss function is used to evaluate the authenticity of the generated samples. Generally speaking, the adversarial loss functions for the generator G and the discriminator D are shown as follows:

[0099]

[0100] Among them, P data (x) and P data(y) is the data distribution of the two domains, and G(·) and D(·) represent the generator G and the discriminator D. Based on the above formula, the adversarial generation loss of generator one and discriminator one and the adversarial generation loss of generator two and discriminator two can be calculated, thereby calculating the adversarial loss function of the entire model.

[0101] Due to the reversibility of the transformation in the CycleGAN network, the cycle consistency loss function is:

[0102]

[0103] where G(·) represents generator one, F(·) represents generator two, and ||·||1 represents the L1 norm, and the cycle consistency loss function is calculated therefrom. This loss function ensures that the source domain samples, as input images, can reconstruct the original samples (i.e., reconstruct the defective image and the defect-free image) after cyclic transformation through the generator modules where the two generators are located.

[0104] Inside generator one, based on the defective image, a mask (mask) is used to divide the training image into a defective area and a non-defective area; therefore, the defective image generation model also introduces an identity loss, and the corresponding loss function is as follows:

[0105]

[0106] where M is the mask information of the defective area in the image, and G(·) represents generator one. This loss is used to optimize the division of the image into a defective area and a non-defective area.

[0107] Considering that the texture in the fabric image needs to be kept intact, a background preservation loss is introduced. The background preservation loss combines two levels of constraints: the background texture loss and the background style loss. In the background preservation network, the image first uses the VGG19 feature extraction network to extract feature maps at different levels, which is represented by the following formula: φ l (I) = Networks l (I), l ∈ {relu11, relu21, relu31, relu41},

[0108] where φ l represents the feature map of the l-th layer of the VGG19 feature extraction network, relu11, relu21, relu31, relu41 represent the four layers of the VGG19 feature extraction network, and Networks represents the VGG19 feature extraction network. Thus, the loss function for the background texture loss corresponding to the non-defective area is:

[0109]

[0110] Among them, M is the mask information of the defect area in the image, and λ l is the weight coefficient of the l-th layer feature, I gen and I real are the images generated by the generation module (the first generation module and the second generation module) and the real image respectively.

[0111] And the loss function of the corresponding background style loss is:

[0112]

[0113] Among them, is the Gram matrix, and the corresponding calculation formula is:

[0114]

[0115] Among them, C l / H l / W l is the number of channels / height / width of the feature map, is the normalization factor.

[0116] Step S3, based on the trained defect image generation model, input several real defect-free images, and the defect image generation model outputs the corresponding generated defect images to realize the expansion of the dataset.

[0117] The real defect-free images pass through the first generator module to generate generated defect images with different scales and different categories.

[0118] Next, the process of a defect image generation algorithm based on deep learning will be described in combination with specific experiments.

[0119] As Figures 9 - 14 shown, for the experiment of verifying the semantic descriptors in similar scenarios, in order to verify the performance of the proposed FDIE-GAN model, an objective and subjective evaluation scheme was designed. In the objective evaluation stage, four representative benchmark models were selected for comparison: Cycle-GAN, DFM-GAN, Dual-GAN, and SDGAN. And six metrics, SSIM (Structural similarity index measure), PSNR (Peak signal-to-noise ratio), MSE (Mean squared error), MAE (Mean absolute error), LPIPS (Learned perceptual image patch similarity), and FID (Frechet inception distance), were used to comprehensively evaluate the images generated by each model. Figure 9 、10 The results show that this method is superior to the comparative methods in terms of overall performance. It is worth noting that there are obvious differences in the processing difficulties of different defect types. Although Metal - type defects generally achieve good processing results in terms of indicators, their actual performance is poor. This reflects that Metal defects are too small and the background is similar, and the network model often ignores this defect. In addition, although SDGAN performs outstandingly in single - item indicators such as PSNR, its relatively high FID value indicates that there is still room for improvement in the diversity and authenticity of the generated images.

[0120] In summary, while maintaining a low pixel - level error (MSE, MAE), this method can obtain a high structural similarity (SSIM) and peak signal - to - noise ratio (PSNR), which indicates that this method can well maintain the overall structure of the image while keeping the detail accuracy. In contrast, other comparative methods such as DFM - GAN and SDGAN may perform outstandingly in some indicators, but it is difficult to maintain excellent performance in all evaluation dimensions. These findings not only verify the effectiveness of this method but also provide important references for subsequent research on optimizing the processing strategies for different defect types and evenly improving the model performance. The results in Tables 11 and 12 show that the texture of the images generated by the FDIE - GAN model in the non - defect area is highly maintained. The experimental results show that this method is superior to the comparative methods in terms of overall performance.

[0121] As Figure 13 shown, in another group of experiments, in order to verify the effectiveness of the defect images generated by different models for downstream tasks, a dataset augmentation experiment was designed. As Figure 11 shown, for each generation model, 1000 generated defect images are combined with 200 real defect images to form the augmented dataset Datasets1, and a dataset Datasets2 of 1000 pure generated fabric defect images. These augmented datasets are respectively input into the recognition networks based on ResNet50 and DenseNet121 for training to obtain the recognition network weights. In order to evaluate the quality of the generated images and their contribution to downstream tasks, the pre - trained weights of ResNet50 and DenseNet121 are used for training to obtain weights for testing.

[0122] The results show that on the D1 test set that only uses the generated images, the classification network trained by FDIE-GAN demonstrates excellent generalization ability. The accuracies of ResNet50 and DenseNet121 reach 83.18% and 80.37% respectively, which are significantly better than other comparison methods. And it has an obvious advantage over other transfer models in the recognition task of pure generated images. On the D2 test set that contains real images, the performance of FDIE-GAN is further improved. The accuracies of ResNet50 and DenseNet121 reach 92.52% and 86.91% respectively, indicating that the images generated by FDIE-GAN can effectively improve the performance of the recognition network.

[0123] Embodiment 2.

[0124] Corresponding to Embodiment 1 of the present invention, Embodiment 2 of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the following steps are implemented according to the method of Embodiment 1.

[0125] Step S1, construct a defect image generation model based on the CycleGAN network.

[0126] Step S2, train and optimize the defect image generation model through adversarial games.

[0127] Step S3, input a number of real defect-free images based on the trained defect image generation model, and the defect image generation model outputs the corresponding generated defect images to achieve the expansion of the data set.

[0128] The above storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), and optical discs that can store program codes.

[0129] For the specific limitations of the steps implemented after the program in the above computer-readable storage medium, reference can be made to Embodiment 1, and details will not be described here again.

[0130] Embodiment 3.

[0131] Corresponding to Embodiment 1 of the present invention, Embodiment 3 of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor executes the program, the following steps are implemented according to the method of Embodiment 1.

[0132] Step S1, construct a defect image generation model based on the CycleGAN network.

[0133] Step S2, train and optimize the defect image generation model through adversarial games.

[0134] Step S3: Based on the trained defect image generation model, several real defect-free images are input, and the defect image generation model outputs corresponding generated defect images to implement the expansion of the dataset.

[0135] For the specific limitations on the implementation steps of the computer device, reference can be made to Embodiment 1, and details are not described herein again.

[0136] It should be noted that each block in the block diagram and / or flowchart in the accompanying drawings of the present invention, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and machine instructions obtained.

[0137] The present invention has been described exemplarily above with reference to the accompanying drawings. Obviously, the specific implementation of the present invention is not limited by the above methods. As long as various non-substantive improvements are made by adopting the inventive concept and technical solution of the present invention, or the inventive concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.

Claims

1. A method for generating defect images based on deep learning, characterized in that: It includes the following steps: Step S1, constructing a defect image generation model; Step S2, training and optimizing the defect image generation model; Step S3, based on the trained defect image generation model, inputting a number of real defect-free images, and the defect image generation model outputs corresponding generated defect images to realize the expansion of the dataset; Among them, the defect image generation model includes two generator modules and two discriminators. The two generator modules are the first generator module and the second generator module, corresponding to generator one and generator two respectively; the two discriminators are discriminator one and discriminator two respectively. The first generator module is used to convert the input real defect-free image into a generated defect image, and to convert the input generated defect-free image into a reconstructed defect image; the second generator module is used to convert the input real defect image into a generated defect-free image, and to convert the input generated defect image into a reconstructed defect-free image; discriminator one is used to judge whether the generated defect image is real, and discriminator two is used to judge whether the generated defect-free image is real; the first generator module sequentially includes a mask module, generator one, and a background preservation network, and the second generator module sequentially includes generator two and a background preservation network; The mask module is used to divide the input image into a defect area and a non-defect area using a mask. Generator one includes a multi-scale feature extraction and bidirectional LSTM fusion module. The multi-scale feature extraction and bidirectional LSTM fusion module is used to capture defect features at different scales, and uses bidirectional LSTM technology for temporal modeling to adaptively fuse defect features at different scales; The features extracted at different scales are concatenated to form the result F of multi-scale feature extraction out , and the corresponding formula is expressed as: F out = Concat[F1, F2, F3], Among them, F i represents the feature map of the i-th scale branch, where i = 1, 2, 3; in the feature fusion stage, first perform a reshaping operation on the feature maps of each scale branch, and the corresponding operation process is expressed by the following formula: Among them, F i view represents the result of the i-th scale branch obtained by the view operation, and x i represents the result of the i-th scale branch obtained by the permute operation. B / C / H / W are the batch number / channel number / height / width of the feature map. The results x i obtained from each scale branch after processing together form multi-scale defect features; After that, bidirectional LSTM is used for temporal modeling of multi-scale defect features, and bidirectional LSTM is used to determine the adaptive fusion weights. The defect image generation model dynamically adjusts the weights of each scale according to the importance of the extracted defect features.

2. The method for generating a defect image based on deep learning according to claim 1, characterized in that: During the fusion process of the bidirectional LSTM, the hidden state is updated for the forward and backward propagation processes, and the features in both directions are concatenated to obtain a bidirectional feature representation; then, an attention mechanism is introduced to calculate the fusion weights, and the softmax function is used to convert the original scores into a weight distribution w: where Pool(·) represents the temporal pooling operation, W is the fusion weight of the outputs of the two hidden states and b is the corresponding bias term; then, the fusion weight coefficients w i for each layer are calculated, and the calculation formula is: Among them, z i represents the original score of the i-th layer obtained by temporal pooling, and z j represents the original score of the j-th layer obtained by temporal pooling; the output feature obtained by multi-scale feature fusion is expressed as: Among them, F' out represents the output feature, w i represents the fusion weight coefficient of each layer, W i is the fusion weight of the corresponding layer, and b i is the bias term of the corresponding layer.

3. The method for generating a defect image based on deep learning according to claim 2, wherein: LSTM calculates the hidden state of the individual LSTM at the current moment through the collaborative work of multiple gating units. The calculation formula is: in, The feature map F representing the three scale branches i The concat operation, P(V(F i )) represents the feature map F i Execute the View operation and the Permute operation in sequence, σ[W f (·)+b f ] represents the forget gate, W f and b f Represent the corresponding weight and bias terms respectively, tanh[W c (·)+b c ] represents the unit status update, W c and b c Represent the corresponding weight and bias terms, C i represents the cell state, ⊙ represents element-wise multiplication, σ[W o (·)+b o ] indicates output gate processing, W o and b o They represent the corresponding weights and bias terms respectively, and h(t) represents the final hidden state. The fusion process of the bidirectional LSTM includes updating the hidden state h(t) through its own calculation formula into two propagation processes, including: Forward propagation process: Backward propagation process: Feature concatenation in both directions: Among them, f t represents the input feature at time t, and respectively represent the hidden state outputs of the forward LSTM and the backward LSTM.

4. A method for generating defect images based on deep learning according to claim 1, characterized in that: The background preservation network includes a VGG19 feature extraction network and a Gram matrix. First, the VGG19 feature extraction network extracts deep features at multiple scales, which is represented by the following formula: φ l (I) = Networks l (I), l ∈ {relu11, relu21, relu31, relu41}, where φ l represents the feature map of the l-th layer of the VGG19 feature extraction network, relu11, relu21, relu31, relu41 represent four layers of the VGG19 feature extraction network, and Networks represents the VGG19 feature extraction network; Then, statistical information comparison is carried out through the Gram matrix to guide the training of the subsequent generator; the defect image generation model introduces a background preservation loss, and the background preservation loss combines two-level constraints of background texture loss and background style loss; in the background preservation network, the loss function of the background texture loss corresponding to the non-defect area is: Among them, M is the mask information of the defect area in the image, and λ l is the weight coefficient of the l-th layer feature, I gen and I real are the image generated by the generation module and the real image respectively; and the loss function of the corresponding background style loss is as follows: Among them, is the Gram matrix, and the corresponding calculation formula is: Among them, C l / H l / W l is the number of channels / height / width of the feature map, is the normalization factor.

5. A method for generating defect images based on deep learning according to claim 4, characterized in that: The mask module divides the training image into a defect area and a non-defect area using a mask based on the defect image, and introduces an identity loss. The corresponding loss function is as follows: Among them, M is the mask information of the defect area in the image, and G(·) represents generator one.

6. A method for generating defect images based on deep learning according to claim 5, characterized in that: The total loss function of the defect image generation model is: Among them, is the adversarial loss function, and λ GAN is the corresponding weight; is the cycle consistency loss function, and λ cyc is the corresponding weight; is the identity loss, which is used to optimize the mask segmentation result, and λ idt is the corresponding weight; is the background texture loss, and λ background is the corresponding weight; is the background style loss, and λ style is the corresponding weight.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of a defect image generation method based on deep learning as described in any one of claims 1-6.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of a method for generating a defect image based on deep learning as described in any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Fabric defect detection method, fabric defect detection apparatus, computer equipment and computer readable medium

    CN109187579A

  • CTPN-based cloth defect detection method

    CN115239615A