Defect image generation algorithm based on deep learning, storage medium and equipment

By combining the feature processing framework of multi-scale isometric convolutional network and bidirectional LSTM, the limitations of traditional GANs in multi-scale defect feature extraction and background maintenance are solved, high-quality defect image generation and data set expansion are achieved, and the performance of downstream recognition tasks is significantly improved.

CN120013905AActive Publication Date: 2025-05-16ANHUI POLYTECHNIC UNIV

Patent Information

Application Number
CN202510099667.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

In the prior art, traditional GANs have limitations in the mobility of generated images and the extraction of multi-scale defect feature, and it is difficult to take into account both the generation of defect feature and the maintenance of background effects, resulting in insufficient quality of generated images.

Method used

A feature processing framework combining multi-scale isometric convolution networks and bidirectional LSTMs is adopted to perform parallel feature extraction through convolution kernels of three different scales: 1×1, 3×3 and 5×5, timing modeling and feature fusion are combined with bidirectional LSTMs, and background maintenance network and Gram matrix constraints are introduced to ensure that the generated image can maintain background features that are highly consistent with the original image in non-defective areas.

Benefits of technology

It effectively improves the quality of generated images, ensures comprehensive and accurate representation of defect features, solves the problem of background texture distortion, and shows significantly better performance in downstream recognition tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013905A_ABST
    Figure CN120013905A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image generation, and particularly relates to a defect image generation algorithm based on deep learning, a storage medium and equipment, and the algorithm comprises the following steps: S1, constructing a defect image generation model; s2, performing training optimization on the defect image generation model; s3, inputting a plurality of real defect-free images based on the trained defect image generation model, and outputting corresponding defect generation images by the defect image generation model to realize data set expansion; wherein the defect image generation model comprises two generator modules and two discriminators, each first generator module sequentially comprises a mask module, a first generator and a background preserving network, and each second generator module sequentially comprises a second generator and a background preserving network. According to the method, the processing strategy that the defect area and the non-defect area are separated through mask is adopted, the background preserving network is innovatively introduced, and the problem that in the prior art, background texture distortion is likely to occur is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image generation, and specifically relates to a defect image generation algorithm, storage medium and device based on deep learning. Background Art

[0002] Defect detection based on deep learning is an important part of improving product quality in the industry, and the quality of defect image datasets has an important impact on deep learning networks. In actual industrial production processes, obtaining real defect samples requires a lot of manpower costs and naturally generated defect data is limited, resulting in a general problem of insufficient defect datasets, which seriously limits the performance of subsequent deep learning models.

[0003] There are methods for generating defect images through generative adversarial networks (GAN) in the prior art. However, these methods have the following problems: traditional GAN ​​has limitations in the transferability of generated images. When extracting multi-scale defect features of different sizes, it is easy to have problems of insufficient extraction or feature loss; it is difficult to simultaneously take into account defect feature generation and background effect preservation, resulting in limitations in the quality of generated images; the generation ability is limited under complex backgrounds, and there are perceptible differences between generated samples and real samples, which affects the training of detection models. Summary of the invention

[0004] The purpose of the present invention is to provide a defect image generation algorithm based on deep learning, which is used to solve the technical problems in the prior art that it is impossible to extract and reproduce multi-scale defect features, and the image quality is insufficient when generating sample images containing defects, and it is impossible to take into account both defect features and background effects.

[0005] The defect image generation algorithm based on deep learning includes the following steps.

[0006] Step S1, constructing a defect image generation model;

[0007] Step S2, training and optimizing the defect image generation model;

[0008] Step S3: Based on the trained defect image generation model, a number of real defect-free images are input, and the defect image generation model outputs corresponding generated defect images to expand the data set;

[0009] Among them, the defect image generation model includes two generator modules and two discriminators. The two generator modules are respectively a first generator module and a second generator module, corresponding to generator one and generator two respectively; the two discriminators are respectively discriminator one and discriminator two, the first generator module is used to convert the input real defect-free image into a generated defect image, and to convert the input generated defect-free image into a reconstructed defect image; the second generator module is used to convert the input real defect image into a generated defect-free image, and to convert the input generated defect image into a reconstructed defect-free image; discriminator one is used to judge whether the generated defect image is real, and discriminator two is used to judge whether the generated defect-free image is real; the first generator modules all include a mask module, generator one and a background preservation network in sequence, and the second generator modules all include generator two and a background preservation network in sequence.

[0010] Preferably, the mask module is used to use a mask to divide the input image into defect areas and non-defect areas, and the generator includes a multi-scale feature extraction and bidirectional LSTM fusion module, which is used to capture defect features of different scales and use bidirectional LSTM technology to perform time series modeling to adaptively fuse defect features of different scales.

[0011] Preferably, the extracted features of different scales are connected to form the result of multi-scale feature extraction F out , the corresponding formula is expressed as:

[0012] F out =Concat[F1,F2,F3],

[0013] Among them, F i represents the feature map of the i-th scale branch, i = 1, 2, 3; in the feature fusion stage, the feature map of each scale branch is first reshaped, and the corresponding operation process is expressed as the following formula:

[0014]

[0015] Among them, F i view represents the result of the i-th scale branch obtained by the view operation, x i represents the result of the i-th scale branch obtained by the permute operation, B / C / H / W are the batch number / channel number / height / width of the feature map, and the result x obtained by each scale branch after processing i Together they form multi-scale defect features;

[0016] Afterwards, bidirectional LSTM is used to perform time series modeling on multi-scale defect features, and bidirectional LSTM is used to determine the adaptive fusion weights. The defect image generation model dynamically adjusts the weights of each scale according to the importance of each defect feature extracted.

[0017] Preferably, in the fusion process of the bidirectional LSTM, the hidden state is updated as two forward and backward propagation processes, and the features in the two directions are concatenated to obtain a bidirectional feature representation; then the attention mechanism is introduced to calculate the fusion weight, and the softmax function is used to convert the original score into a weight distribution w:

[0018] Among them, Pool(·) represents the temporal pooling operation, and W is the output of two hidden states. and The fusion weight of each layer is calculated, and b is the corresponding bias term; then the fusion weight coefficient w of each layer is calculated i , the calculation formula is:

[0019]

[0020] Among them, z i represents the original score of the i-th layer obtained by temporal pooling, z j represents the j-th layer original score obtained by temporal pooling; the output feature obtained by multi-scale feature fusion is expressed as:

[0021]

[0022] Among them, F' out represents the output feature, w i Represents the fusion weight coefficient of each layer, W i is the fusion weight of the corresponding layer, b i is the bias term of the corresponding layer.

[0023] Preferably, LSTM calculates the hidden state of a single LSTM at the current moment through the collaborative work of multiple gating units, and the calculation formula is:

[0024]

[0025] in, The feature map F representing the three scale branches i The concat operation, P(V(F i )) represents the feature map F i Execute the View operation and the Permute operation in sequence, σ[W f (·)+b f ] represents the forget gate, W f and b f Represent the corresponding weight and bias terms respectively, tanh[Wc (·)+b c ] represents the unit status update, W c and b c Represent the corresponding weight and bias terms, C i represents the cell state, ⊙ represents element-wise multiplication, σ[W o (·)+b o ] indicates output gate processing, W o and b o They represent the corresponding weights and bias terms respectively, and h(t) represents the final hidden state;

[0026] The fusion process of the bidirectional LSTM includes updating the hidden state h(t) through its own calculation formula into two propagation processes, including:

[0027] Forward propagation process: Back propagation process: Feature stitching in two directions: Among them, f t represents the input features at time t, and Represent the hidden state outputs of the forward LSTM and the backward LSTM respectively.

[0028] Preferably, the background preservation network includes a VGG19 feature extraction network and a Gram matrix. The multi-scale deep features are first extracted by the VGG19 feature extraction network, which is expressed as follows: l (I)=Networks l (I),l∈{relu11,relu21,relu31,relu41},

[0029] Among them, φ l Represents the feature map of the lth layer of the VGG19 feature extraction network, relu11, relu21, relu31, relu41 represent the four layers of the VGG19 feature extraction network, and Networks represents the VGG19 feature extraction network;

[0030] Then, the statistical information is compared through the Gram matrix to guide the subsequent training of the generator; the defect image generation model introduces the background preservation loss, which combines the two-level constraints of background texture loss and background style loss; in the background preservation network, the loss function corresponding to the background texture loss of the non-defect area is:

[0031]

[0032] Among them, M is the mask information of the defect area in the image, λ l is the weight coefficient of the l-th layer feature, Igen and I real are the images generated by the generation module and the real images respectively; and the corresponding background style loss function is:

[0033]

[0034] in, is the Gram matrix, and the corresponding calculation formula is:

[0035]

[0036] Among them, C l / H l / W l is the number of channels / height / width of the feature map, is the normalization factor.

[0037] Preferably, the mask module uses a mask based on the defect image to divide the training image into defect areas and non-defect areas, introducing identity loss, and the corresponding loss function is as follows:

[0038]

[0039] Among them, M is the mask information of the defect area in the image, and G(·) represents generator one.

[0040] Preferably, the total loss function of the defect image generation model is:

[0041]

[0042] in, To combat the loss function, λ GAN is the corresponding weight; is the cycle consistency loss function, λ cyc is the corresponding weight; is the identity loss, used to optimize the mask segmentation result, λ idt is the corresponding weight; is the background texture loss, λ background is the corresponding weight; is the background style loss, λ style is the corresponding weight.

[0043] The present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of a defect image generation algorithm based on deep learning as described above are implemented.

[0044] The present invention also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that when the processor executes the computer program, the steps of a defect image generation algorithm based on deep learning as described above are implemented.

[0045] The present invention has the following advantages:

[0046] 1. The algorithm provided by the present invention innovatively introduces a feature processing framework that combines a multi-scale equidistant convolutional network with a bidirectional LSTM. By performing parallel feature extraction with convolution kernels of three different scales, 1×1, 3×3, and 5×5, the local details, medium-scale structures, and large-scale contextual information of the defects can be captured simultaneously. The forward and backward propagation mechanisms of the bidirectional LSTM can fully exploit the temporal correlation between these multi-scale features, and adaptively allocate the importance weights of different features through the attention mechanism, thereby obtaining a more comprehensive and accurate representation of defect features, effectively improving the quality of the generated image.

[0047] 2. The present invention adopts a processing strategy of separating defect areas from non-defect areas by using masks, and innovatively introduces background preservation network and Gram matrix constraints to preserve the features of non-defect areas. The background preservation network represents background information by extracting feature maps at different levels, and the Gram matrix constrains the consistency of background style by calculating the internal correlation of feature maps. This dual constraint mechanism ensures that the generated image can maintain background features that are highly consistent with the original image in non-defect areas, effectively solving the problem of background texture distortion that is prone to occur in the prior art.

[0048] 3. The present invention designs a complete loss function system, which includes multiple components such as adversarial loss, cycle consistency loss, identity loss, background preservation loss and style consistency loss. This multiple loss constraint ensures that the generated image can achieve good results in multiple aspects such as defect feature expression, background texture preservation and overall authenticity. Experimental results show that the images generated by the method of the present invention not only have higher structural similarity and peak signal-to-noise ratio, but also show significantly better performance in downstream recognition tasks, verifying the effectiveness of this method in defect image generation and data expansion. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a basic flow chart of a defect image generation algorithm based on deep learning in the present invention.

[0050] Figure 2 It is a structural diagram of the defect image generation model of the present invention.

[0051] Figure 3 This is a flow chart of the first generator module in the present invention.

[0052] Figure 4 This is a structural diagram of the adaptive multi-scale feature bidirectional LSTM in the present invention.

[0053] Figure 5 This is a structural diagram of generator one in the present invention.

[0054] Figure 6 This is a structural diagram of generator 2 in the present invention.

[0055] Figure 7 It is a structural diagram of the discriminator in the present invention.

[0056] Figure 8 This is a structural diagram of the feature extraction network used in the background preservation module of the present invention.

[0057] Fig. 9 This is a comparison chart of the defect generation results of the present invention and the prior art. The original image represents a true defect-free image, and color, cut, hole, metal, and thread represent five types of defects. Cycle-GAN represents the five types of defect effects generated by the cycle_gan model. DFM-GAN represents the five types of defect effects generated by the DFM-GAN model. Dual-GAN represents the five types of defect effects generated by the Dual-GAN model. SD-GAN represents the five types of defect effects generated by the SD-GAN model. SD-GAN represents the five types of defect effects generated by the SD-GAN model.

[0058] Fig.10 The figure is a comparison effect diagram of the defect generation results of the present invention and the prior art in terms of indicators. Among them, the SSIM indicator evaluates the structural similarity of the image, and the higher the score, the more similar the structure of the generated image is to the reference image; the PSNR indicator evaluates the image quality, and the higher the value, the smaller the image distortion; the MSE and MAE indicators evaluate the error at the pixel level of the image, and the lower the value, the closer the generated image is to the reference image; LPIPS evaluates the perceptual similarity of the image, and the lower the value, the better the perceptual quality of the generated image; the FID indicator evaluates the authenticity and diversity of the generated image, and the lower the value, the closer the distribution of the generated image is to the real image.

[0059] Fig.11 This is a comparison diagram of the defect generation results of the present invention and the prior art based on the background retention index.

[0060] Fig.12 This is a comparison diagram of the background of defect generation results of the present invention and the prior art.

[0061] Fig.13The experimental flow chart of the defect generation results of the present invention and the prior art being fed into the recognition network for training. In order to verify the effectiveness of defect images generated by different models for downstream tasks, a dataset authenticity experiment and a dataset expansion experiment were designed. The authenticity experiment used the Datasets1 dataset. It consists of 1,000 images generated by each model. After being input into the recognition networks of ResNet50 and DenseNet121 for training, the real fabric defect images can be directly identified. The recognition network acts as a discriminator to judge the authenticity of the generated dataset. The expansion experiment uses Datasets2, which is a dataset composed of 1,000 images generated by each model and 234 real defect images to simulate the effect of using generated images for training when the dataset is too small. These expanded datasets are also input into the above-mentioned recognition network for training to obtain weights. In order to comprehensively evaluate the performance of the dataset for deep learning, pre-trained weights are used for training and testing at the same time.

[0062] Fig.14 The experimental results of the defect generation results of the present invention and the prior art are fed into the recognition network training. The generated images are highly similar to the real images. On the Datasets2 test set containing real images, the performance of FDIE-GAN is further improved, indicating that the images generated by FDIE-GAN can effectively improve the performance of the recognition network. DETAILED DESCRIPTION

[0063] The specific implementation modes of the present invention will be further explained in detail below by describing the embodiments with reference to the accompanying drawings, so as to help those skilled in the art to have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.

[0064] Embodiment 1.

[0065] like Figure 1-Figure 8 As shown, the present invention provides a defect image generation algorithm based on deep learning, comprising the following steps.

[0066] Step S1, constructing a defect image generation model based on the CycleGAN network.

[0067] The defect image generation model includes two generator modules including corresponding generators and two discriminators. The two generator modules are the first generator module and the second generator module, corresponding to generator one (G) and generator two (F), respectively. The two discriminators are discriminator one (D y ) and Discriminator 2 (D x); the first generator module is used to convert the input real defect-free image into a generated defect image, and is used to convert the input generated defect-free image into a reconstructed defect image; the second generator module is used to convert the input real defect image into a generated defect-free image, and is used to convert the input generated defect image into a reconstructed defect-free image; the discriminator one is used to judge whether the generated defect image is real, and the discriminator two is used to judge whether the generated defect-free image is real; the first generator module includes a mask module, a generator one and a background preservation network in sequence, and the second generator module includes a generator two and a background preservation network in sequence.

[0068] The mask module is used to divide the input image into defect areas and non-defect areas using a mask. During training, the mask module is obtained by manually labeling the defect image through LabelImg, which improves the training efficiency and accuracy and ensures that the defect area is accurately segmented.

[0069] The generator 1 includes an encoder, a multi-scale feature extraction and bidirectional LSTM (long short-term memory network) fusion module, a decoder and an output layer. The encoder is used to extract features and downsample the image after mask segmentation. The multi-scale feature extraction and bidirectional LSTM fusion module is used to capture defect features of different scales, and use bidirectional LSTM technology to perform time series modeling to adaptively fuse defect features of different scales. In this way, the encoder of generator 1 can dynamically adjust the weights of defect feature information of each scale according to the extracted defect features, and adaptively select the most appropriate feature representation for different defect features. For example, for small area defects such as holes and large area defects such as cuts, convolution kernels of different scales are required for extraction, so this solution can avoid insufficient feature extraction or feature loss.

[0070] The generator 2 uses the U-Net architecture as the backbone network and is optimized for feature extraction and reconstruction of fabric defect images.

[0071] Both discriminators consist of five convolutional layers. In the discriminator, the first four convolutional layers use the ReLU activation function. Except for the first convolutional layer, a batch normalization layer is set after each convolutional layer, and the last layer uses the Tanh activation function. The convolution kernel size of each layer is 4×4, with a stride of 2 for the first three layers and 1 for the last two layers. The number of feature channels gradually changes from 64 in the first layer to 1 in the last output layer. The output of the discriminator is a discriminant matrix, in which each element corresponds to a local receptive field in the input image, and the receptive field size is 63×63 pixels. This enables the discriminator to distinguish true from false in local areas of the fabric image, forcing the generator to focus not only on the overall authenticity in the adversarial learning process, but also on maintaining the fineness of local structure and defect features.

[0072] The background preservation network includes a VGG19 feature extraction network and a Gram matrix. The multi-scale deep features are first extracted through the VGG19 feature extraction network, and then the statistical information is compared through the Gram matrix, so as to guide the subsequent training of the generator, especially to achieve the preservation of image background information.

[0073] Step S2: training and optimizing the defect image generation model through adversarial game.

[0074] In step S2, the mask module is obtained by manually labeling the defect image through LabelImg, thereby improving training efficiency and accuracy and ensuring accurate segmentation of the defect area.

[0075] The input image of the first generator module is first divided into regions through the mask module, and then a multi-scale equidistant convolutional network is used in the generator (G) to extract features from the defect area. The multi-scale equidistant convolutional network contains three parallel convolution branches, which use 1×1 convolution, 3×3 convolution and 5×5 convolution kernels respectively, and extract features with different voiding rates. Among them, 1×1 convolution mainly extracts local detail features, 3×3 convolution extracts medium-scale structural features, and 5×5 convolution extracts large-scale context information. This step expands the receptive field without increasing the number of parameters by setting different voiding rates, so that the network can fully capture defect features of different scales. The extracted features of different scales are connected to form the result of multi-scale feature extraction F. out , the corresponding formula is expressed as:

[0076] F out =Concat[F1,F2,F3],

[0077] Among them, F i represents the feature map of the i-th scale branch, i = 1, 2, 3. In the feature fusion stage, the present invention first reshapes the feature map of each scale branch, that is, the result F of multi-scale feature extraction out Through the view operation, we get the results of different scales, and then through the corresponding permute operation, we realize the feature reorganization of each scale branch. The corresponding operation process is expressed as the following formula:

[0078]

[0079] Among them, F i view represents the result of the i-th scale branch obtained by the view operation, x i The result of the i-th scale branch obtained by the permute operation is the number of channels / height / width of the feature map. This reshaping operation flattens the spatial dimensions and adjusts the dimension order so that the result x obtained by each scale branch after processingi These results x meet the LSTM input requirements. i Together they form multi-scale defect features.

[0080] Afterwards, bidirectional LSTM is used to perform time series modeling on multi-scale defect features, where a single LSTM is Figure 3 As shown in the figure, through the collaborative work of multiple gating units, h(t) is finally used to calculate the hidden state of a single LSTM at the current moment. The calculation formula is:

[0081]

[0082] in, The feature map F representing the three scale branches i The concat operation, P(V(F i )) represents the feature map F i Execute the View operation and the Permute operation in sequence, σ[W f (·)+b f ] represents the forget gate, W f and b f Represent the corresponding weight and bias terms respectively, tanh[W c (·)+b c ] represents the unit status update, W c and b c Represent the corresponding weight and bias terms, C i represents the cell state, ⊙ represents element-wise multiplication, σ[W o (·)+b o ] indicates output gate processing, W o and b o They represent the corresponding weights and bias terms respectively, and h(t) represents the final hidden state.

[0083] The fusion process of the bidirectional LSTM includes updating the hidden state h(t) through its own calculation formula into two propagation processes, including:

[0084] Forward propagation process: Back propagation process: Feature stitching in two directions: Among them, f t represents the input features at time t, and Represent the hidden state outputs of the forward LSTM and the backward LSTM respectively.

[0085] After obtaining the bidirectional feature representation, in order to achieve adaptive fusion of multi-scale features, this step introduces the attention mechanism to calculate the fusion weights, and uses the softmax function to convert the original scores into weight distribution w:

[0086] Among them, Pool(·) represents the temporal pooling operation, and W is the output of two hidden states. and The fusion weight of each layer is calculated as follows: i , the calculation formula is:

[0087]

[0088] Among them, z i represents the original score of the i-th layer obtained by temporal pooling, z j Represents the j-th layer original score obtained by temporal pooling. Combining the above formula, the output feature obtained by multi-scale feature fusion is expressed as:

[0089]

[0090] Among them, F' out represents the output feature, w i Represents the fusion weight coefficient of each layer, W i is the fusion weight of the corresponding layer, b i The bidirectional LSTM is used to determine the adaptive fusion weight, so that the defect image generation model (FDIE-GAN model) can dynamically adjust the weights of each scale according to the importance of each defect feature extracted, and adaptively select the most appropriate feature representation for different defects.

[0091] The decoder then uses an upsampling structure symmetrical to the encoder to reconstruct the fused feature map into a defect generation image. Each layer of the decoder is equipped with instance normalization and ReLU activation function after upsampling to ensure the quality of feature reconstruction, and finally outputs the final generated defect image through a 7×7 convolution layer.

[0092] Similarly, the first generator module also converts the input generated defect-free image into a reconstructed defect image.

[0093] At the same time, the second generator module is used to convert the input real defect image into a generated defect-free image, specifically including: in the feature encoding stage, the generator 2 gradually compresses the spatial dimension of the input image through continuous downsampling convolution modules, while capturing background features and defect features. At the lowest feature dimension, the network sets a bottleneck convolution layer to effectively compress and reorganize the key feature information of fabric defects. In the feature decoding stage, the network adopts a symmetrical structure with the encoder, and gradually restores the spatial resolution of the feature map through a series of deconvolution modules. Generator 2 also introduces a jump connection mechanism between the corresponding encoding and decoding layers to directly pass the background features and defect features of the fabric surface to the deep layer, ensuring that the key defect information is not lost during the image generation process. Finally, Generator 2 normalizes the feature map through the Softmax layer and outputs a generated defect-free image. This module is also used to convert the input generated defect image into a reconstructed defect-free image.

[0094] During training, discriminator 1 is used to determine whether the generated defect image is real, and discriminator 2 is used to determine whether the generated defect-free image is real; and the loss function is used to optimize the model, so that the generated image is closer and closer to the real image. The relevant loss function is as follows.

[0095] The total loss function of the defect image generation model is:

[0096]

[0097] in, To combat the loss function, λ GAN is the corresponding weight; is the cycle consistency loss function, λ cyc is the corresponding weight; is the identity loss, used to optimize the mask segmentation result, λ idt is the corresponding weight; is the background texture loss, λ background is the corresponding weight; is the background style loss, λ style is the corresponding weight.

[0098] The adversarial loss function is used to evaluate the authenticity of the generated samples. Generally speaking, the adversarial loss function for the generator G and the discriminator D is as follows:

[0099]

[0100] Among them, P data (x) and P data(y) is the data distribution of the two domains, G(·) and D(·) represent the generator G and the discriminator D. Based on the above formula, the adversarial generation loss of generator 1 and discriminator 1 and the adversarial generation loss of generator 2 and discriminator 2 can be calculated, thereby calculating the adversarial loss function of the entire model

[0101] Because of the reversibility of transformations in the CycleGAN network, the cycle consistency loss function is:

[0102]

[0103] Among them, G(·) represents generator one, F(·) represents generator two, and ||·||1 represents the L1 norm. The cycle consistency loss function is calculated from this This loss function ensures that the source domain sample is used as the input image, and after the cyclic transformation of the generator module where the two generators are located, the original sample can be reconstructed (that is, the defect image is reconstructed and the defect-free image is reconstructed).

[0104] In generator 1, a mask is used to divide the training image into defect areas and non-defect areas based on the defect image; therefore, the defect image generation model also introduces identity loss, and the corresponding loss function is as follows:

[0105]

[0106] Among them, M is the mask information of the defect area in the image, G(·) represents generator 1, and the loss is used to optimize the division of the image into defect area and non-defect area.

[0107] Considering that the texture in the fabric image needs to be intact, the background preservation loss is introduced. The background preservation loss combines the two-level constraints of background texture loss and background style loss. In the background preservation network, the image is first extracted with the VGG19 feature extraction network to extract feature maps of different levels, which can be expressed as follows: l (I)=Networks l (I),l∈{relu11,relu21,relu31,relu41},

[0108] Among them, φ l represents the feature map of the first layer of the VGG19 feature extraction network, relu11, relu21, relu31, relu41 represent the four layers of the VGG19 feature extraction network, and Networks represents the VGG19 feature extraction network. The loss function corresponding to the background texture loss of the non-defect area is:

[0109]

[0110] Among them, M is the mask information of the defect area in the image, λ l is the weight coefficient of the l-th layer feature, I gen and I real They are the images generated by the generation modules (the first generation module and the second generation module) and the real image respectively.

[0111] The corresponding background style loss function is:

[0112]

[0113] in, is the Gram matrix, and the corresponding calculation formula is:

[0114]

[0115] Among them, C l / H l / W l is the number of channels / height / width of the feature map, is the normalization factor.

[0116] Step S3, based on the trained defect image generation model, a number of real defect-free images are input, and the defect image generation model outputs corresponding generated defect images to expand the data set.

[0117] The real defect-free image is passed through the first generator module to generate defect images with different scales and categories.

[0118] The following is an explanation of the process of a defect image generation algorithm based on deep learning by combining specific experiments.

[0119] like Figure 9-14 As shown in the figure, the description experiment of the semantic descriptor in similar scenes is verified. In order to verify the performance of the proposed FDIE-GAN model, an objective and subjective evaluation scheme is designed. In the objective evaluation stage, four representative benchmark models are selected for comparison: Cycle-GAN, DFM-GAN, Dual-GAN and SDGAN. And six indicators including SSIM (Structural similarity index measure), PSNR (Peak signal-to-noise ratio), MSE (Mean squared error), MAE (Mean absolute error), LPIPS (Learned perceptual image patch similarity) and FID (Frechet inception distance) are used to evaluate the images generated by each model as a whole. Fig. 9 ,10 The results show that the proposed method outperforms the comparison method in terms of overall performance. It is worth noting that there are significant differences in the difficulty of processing different defect types. Although Metal defects generally achieve good processing results in terms of indicators, the actual performance is poor. This reflects that the Metal defect is too small and the background is similar, so the network model often ignores the defect. In addition, although SDGAN performs well in individual indicators such as PSNR, its high FID value shows that this method still has room for improvement in terms of the diversity and authenticity of generated images.

[0120] In summary, this method can obtain higher structural similarity (SSIM) and peak signal-to-noise ratio (PSNR) while maintaining lower pixel-level errors (MSE, MAE), which shows that this method can maintain the overall structure of the image well while maintaining the accuracy of details. In contrast, other comparative methods such as DFM-GAN and SDGAN may perform well in certain indicators, but it is difficult to maintain excellent performance in all evaluation dimensions. These findings not only verify the effectiveness of this method, but also provide an important reference for subsequent research on the optimization of processing strategies for different defect types and the balanced improvement of model performance. The results in Tables 11 and 12 show that the texture of the image generated by the FDIE-GAN model is highly maintained in the non-defect area. The experimental results show that this method is superior to the comparative methods in overall performance.

[0121] like Fig.13 As shown in , in another set of experiments, in order to verify the effectiveness of defect images generated by different models for downstream tasks, a dataset expansion experiment was designed. Fig.11 As shown in the figure, for each generation model, 1000 generated defect images are combined with 200 real defect images to form an extended dataset Datasets1, and 1000 purely generated fabric defect image datasets Datasets2. These extended datasets are respectively input into the recognition networks based on ResNet50 and DenseNet121 for training to obtain the recognition network weights. In order to evaluate the quality of the generated images and their contribution to downstream tasks, the weights are trained using the pre-trained weights of ResNet50 and DenseNet121 for testing.

[0122] The results show that on the D1 test set using only generated images, the classification network trained by FDIE-GAN exhibits excellent generalization ability, with ResNet50 and DenseNet121 reaching 83.18% and 80.37% accuracy respectively, which is significantly better than other comparison methods. And it has obvious advantages over other migration models in the recognition task of pure generated images. On the D2 test set containing real images, the performance of FDIE-GAN is further improved, with the accuracy of ResNet50 and DenseNet121 reaching 92.52% and 86.91% respectively, indicating that the images generated by FDIE-GAN can effectively improve the performance of the recognition network.

[0123] Embodiment 2.

[0124] Corresponding to the first embodiment of the present invention, the second embodiment of the present invention provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the following steps are implemented according to the method of the first embodiment.

[0125] Step S1, constructing a defect image generation model based on the CycleGAN network.

[0126] Step S2: training and optimizing the defect image generation model through adversarial game.

[0127] Step S3, based on the trained defect image generation model, a number of real defect-free images are input, and the defect image generation model outputs corresponding generated defect images to expand the data set.

[0128] The above storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), optical disk and other media that can store program codes.

[0129] The specific limitations on the steps implemented after the program in the computer-readable storage medium is executed can be found in Example 1, and will not be described in detail here.

[0130] Embodiment three.

[0131] Corresponding to the first embodiment of the present invention, the third embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and when the processor executes the program, the following steps are implemented according to the method of the first embodiment.

[0132] Step S1, constructing a defect image generation model based on the CycleGAN network.

[0133] Step S2: training and optimizing the defect image generation model through adversarial game.

[0134] Step S3, based on the trained defect image generation model, a number of real defect-free images are input, and the defect image generation model outputs corresponding generated defect images to expand the data set.

[0135] The specific limitations on the above steps of implementing the computer device can be found in Example 1, and will not be described in detail here.

[0136] It should be noted that each box in the block diagram and / or flow chart in the accompanying drawings of the specification of the present invention, as well as the combination of boxes in the block diagram and / or flow chart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and machine instructions.

[0137] The present invention is described above by way of example in conjunction with the accompanying drawings. It is obvious that the specific implementation of the present invention is not limited to the above-mentioned method. As long as various non-substantial improvements are made using the inventive concept and technical solution of the present invention, or the inventive concept and technical solution are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.

Claims

1. A defect image generation algorithm based on deep learning, characterized by: The following steps are involved: Step S1, constructing a defect image generation model; Step S2, training and optimizing the defect image generation model; Step S3: Based on the trained defect image generation model, a number of real defect-free images are input, and the defect image generation model outputs corresponding generated defect images to expand the data set; Among them, the defect image generation model includes two generator modules and two discriminators. The two generator modules are respectively a first generator module and a second generator module, corresponding to generator one and generator two respectively; the two discriminators are respectively discriminator one and discriminator two, the first generator module is used to convert the input real defect-free image into a generated defect image, and to convert the input generated defect-free image into a reconstructed defect image; the second generator module is used to convert the input real defect image into a generated defect-free image, and to convert the input generated defect image into a reconstructed defect-free image; discriminator one is used to judge whether the generated defect image is real, and discriminator two is used to judge whether the generated defect-free image is real; the first generator modules all include a mask module, generator one and a background preservation network in sequence, and the second generator modules all include generator two and a background preservation network in sequence.

2. The defect image generation algorithm based on deep learning according to claim 1, characterized in that: The mask module is used to use a mask to divide the input image into a defect area and a non-defect area. The generator includes a multi-scale feature extraction and bidirectional LSTM fusion module. The multi-scale feature extraction and bidirectional LSTM fusion module is used to capture defect features of different scales, and use bidirectional LSTM technology to perform time series modeling to adaptively fuse defect features of different scales.

3. The defect image generation algorithm based on deep learning according to claim 2, characterized in that: The extracted features of different scales are connected to form the result of multi-scale feature extraction F out , the corresponding formula is expressed as: F out =Concat[F1,F2,F3], Among them, F i represents the feature map of the i-th scale branch, i = 1, 2, 3; in the feature fusion stage, the feature map of each scale branch is first reshaped, and the corresponding operation process is expressed as the following formula: in, represents the result of the i-th scale branch obtained by the view operation, x i represents the result of the i-th scale branch obtained by the permute operation, B / C / H / W are the batch number / channel number / height / width of the feature map, and the result x obtained by each scale branch after processing i Together they form multi-scale defect features; Afterwards, bidirectional LSTM is used to perform time series modeling on multi-scale defect features, and bidirectional LSTM is used to determine the adaptive fusion weights. The defect image generation model dynamically adjusts the weights of each scale according to the importance of each defect feature extracted.

4. The defect image generation algorithm based on deep learning according to claim 3, characterized in that: In the fusion process of bidirectional LSTM, the hidden state is updated as two propagation processes, and the features in two directions are concatenated to obtain bidirectional feature representation. Then the attention mechanism is introduced to calculate the fusion weight, and the softmax function is used to convert the original score into the weight distribution w: Among them, Pool(·) represents the temporal pooling operation, and W is the output of two hidden states. and The fusion weight of each layer is calculated, and b is the corresponding bias term; then the fusion weight coefficient w of each layer is calculated i , the calculation formula is: Among them, z i represents the original score of the i-th layer obtained by temporal pooling, z j represents the j-th layer original score obtained by temporal pooling; the output feature obtained by multi-scale feature fusion is expressed as: Among them, F' out represents the output feature, w i Represents the fusion weight coefficient of each layer, W i is the fusion weight of the corresponding layer, b i is the bias term of the corresponding layer.

5. The robot algorithm for dynamic occlusion scenarios according to claim 4, characterized in that: LSTM calculates the hidden state of a single LSTM at the current moment through the collaborative work of multiple gating units. The calculation formula is: in, The feature map F representing the three scale branches i The concat operation, P(V(F i )) represents the feature map F i Execute the View operation and the Permute operation in sequence, σ[W f (·)+b f ] represents the forget gate, W f and b f Represent the corresponding weight and bias terms respectively, tanh[W c (·)+b c ] represents the unit status update, W c and b c Represent the corresponding weight and bias terms, C i represents the cell state, ⊙ represents element-wise multiplication, σ[W o (·)+b o ] indicates output gate processing, W o and b o They represent the corresponding weights and bias terms respectively, and h(t) represents the final hidden state. The fusion process of the bidirectional LSTM includes updating the hidden state h(t) through its own calculation formula into two propagation processes, including: Forward propagation process: Back propagation process: Feature stitching in two directions: Among them, f t represents the input features at time t, and Represent the hidden state outputs of the forward LSTM and the backward LSTM respectively.

6. The defect image generation algorithm based on deep learning according to claim 1, characterized in that: The background preservation network includes a VGG19 feature extraction network and a Gram matrix. The multi-scale deep features are first extracted through the VGG19 feature extraction network, which is expressed as follows: φ l (I)=Networks l (I),l∈{clock11,clock21,clock31,clock41}, Among them, φ l Represents the feature map of the lth layer of the VGG19 feature extraction network, relu11, relu21, relu31, relu41 represent the four layers of the VGG19 feature extraction network, and Networks represents the VGG19 feature extraction network; Then, the statistical information is compared through the Gram matrix to guide the subsequent training of the generator; the defect image generation model introduces the background preservation loss, which combines the two-level constraints of background texture loss and background style loss; in the background preservation network, the loss function corresponding to the background texture loss of the non-defect area is: Among them, M is the mask information of the defect area in the image, λ l is the weight coefficient of the l-th layer feature, I gen and I real are the images generated by the generation module and the real images respectively; and the corresponding background style loss function is: in, is the Gram matrix, and the corresponding calculation formula is: Among them, C l / H l / W l is the number of channels / height / width of the feature map, is the normalization factor.

7. The deep learning-based defect image generation algorithm according to claim 6, characterized in that: The mask module uses masks based on defect images to divide training images into defect areas and non-defect areas, introducing identity loss. The corresponding loss function is as follows: Among them, M is the mask information of the defect area in the image, and G(·) represents generator one.

8. The deep learning-based defect image generation algorithm according to claim 7, characterized in that: The total loss function of the defect image generation model is: in, To combat the loss function, λ GAN is the corresponding weight; is the cycle consistency loss function, λ cyc is the corresponding weight; is the identity loss, used to optimize the mask segmentation result, λ idt is the corresponding weight; is the background texture loss, λ background is the corresponding weight; is the background style loss, λ style is the corresponding weight.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a defect image generation algorithm based on deep learning as described in any one of claims 1 to 8 are implemented.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of a defect image generation algorithm based on deep learning as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Fabric defect detection method, fabric defect detection apparatus, computer equipment and computer readable medium

    CN109187579A

  • CTPN-based cloth defect detection method

    CN115239615A

Cited By

  • Industrial defect image generation method, terminal equipment and storage medium

    CN120931643A