Defect image generation method and system fusing attention mechanism and perceptual loss
By introducing SE attention mechanism and perceived loss in the WGAN-GP model, the problems of instability and insufficient details of the GAN model when generating defective images of small-sized LCD screens are solved, and the generated image quality is significantly improved, and the detection accuracy of the defect detection model is significantly improved.
Patent Information
- Application Number
- CN202510358763.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-25
AI Technical Summary
When generating defect images of small-size LCD screens, existing GAN models have problems of instability in training, insufficient details of generated images and distortion, especially the complex morphology and local details for leaking defects are difficult to capture, and the data set imbalance leads to insufficient accuracy of the detection model.
Using the WGAN-GP model and integrating the SE attention mechanism and perceived loss, the SE attention layer and high-resolution feature reconstruction module are added to the generator, and the deep features are extracted in combination with the VGG-19 network to build perceived loss, improving the loss function to improve the image generation quality.
The generated defective images have been significantly improved in indicators such as PSNR, SSIM, RMSE and CS. The model has performed excellently in the YOLOv8 detection model, with a 10.5% increase in detection accuracy, and significantly improved the quality of generated images and the detailed restoration effect.
Smart Images

Figure CN120356028A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of screen detection, and particularly relates to a defective image generation method and system integrating an attention mechanism and perceptual loss. Background Art
[0002] With the increasingly comprehensive penetration of intelligent technologies in modern production and life, the applications of mobile devices, wearable devices, and various smart home products are becoming more and more extensive. As an essential interactive interface for them, the demand for small-sized liquid crystal displays is also increasing day by day. Since various defects will inevitably occur during the manufacturing process of liquid crystal screens, such as Figure 1 the point defects, line defects, liquid leakage, and scratches shown. Therefore, liquid crystal screens must undergo strict quality inspection before being packaged and put on the market. At present, it mainly relies on manual quality inspection, which is not only time-consuming, laborious, and costly, but also its detection speed can no longer meet the real-time detection requirements of the continuously upgraded intelligent production line, which has become the key bottleneck restricting production capacity. With the rapid development of machine vision, the target detection technology based on machine vision has become the key technology for intelligent upgrading in the field of defect detection, such as the intelligent detection methods based on YOLO and Faster R-CNN. A diverse and balanced defect data set is the necessary data basis to ensure the detection accuracy, robustness, and real-time performance of the model. The small-sized liquid crystal display defect data set collected from the production line, as shown in Figure 2 the figure, point defects and scratch defects are very common, while the number of line defect and liquid leakage samples is extremely scarce, and the data set is extremely unbalanced. The reasons for this situation are very complex, and the optimization of the data collection method cannot improve this situation. Therefore, we try to obtain an effective and balanced defective image data set through intelligent data augmentation technology to create an effective data basis for the intelligent detection of liquid crystal screens.
[0003] A large number of studies have shown that for the problem of unbalanced defect data sets, using data augmentation technology can significantly improve the performance of defect detection models. Currently, there are mainly two types of data augmentation methods: (1) Traditional data augmentation, that is, performing operations such as geometric transformation, adjusting brightness and contrast, and adding noise on the original defective image, which retains the basic features and details of the original image, but the generated images are limited by the existing images and can only overcome the problem of unbalanced sample quantity, and cannot solve the imbalance in randomness and diversity of the sample set. (2) Intelligent image generation. The core task of image generation is to construct an image generation model and tune the parameters to minimize the divergence between the probability distributions of the generated data and the real data. In the field of industrial defect image generation, the most widely used generators are mainly GANs, VAEs, or DDPMs. Through the adversarial training mechanism of the generator and the discriminator, GANs can effectively generate high-quality defect images. By combining the advantages of variational inference and deep learning, VAEs can also achieve the purpose of generating defect images by reconstructing the original data through the encoder and decoder. The research generates defect samples through DDPMs and optimizes the generated data using a feature editing process. The image generation process of the GANs model is simpler and the model training is faster; moreover, GANs allow the generation of diverse samples through operations in the significant space to ensure the diversity of the dataset.
[0004] However, Figure 2 For the data distribution diagram of various defects in the display manufacturing factory from April 18th to April 26th, 2024 for 8 days, there are still the following challenges in directly generating LCD screen defect images using the GANs model: (1) The scarcity of defect dataset samples makes it difficult to maintain the balance between the generator and the discriminator, resulting in unstable training. (2) The morphology of the liquid leakage defect area is complex and irregular in shape, resulting in insufficient details and accuracy in the images generated by GANs. (3) The original network architecture cannot fully capture the deeper features of the defects, resulting in partial distortion of the generated images.
[0005] Since GAN was introduced in 2014, it has been widely used in the field of machine vision. However, when generating defect images, a too large gap in the capabilities of the generator and the discriminator in GAN will lead to unstable training. To address this problem, researchers have made various improvements: CGAN introduces label information during the generation process, guides image generation through input conditions and labels, improves the quality of the generated samples, reduces the gap in the capabilities of the generator and the discriminator, and optimizes the problem of unstable GAN training. However, the generated images of CGAN still have blurred problems in complex scenarios. The deep convolutional generative adversarial network (DCGAN) proposed by Alec et al. significantly improves the model's feature extraction ability for images by replacing the fully connected layer in the traditional GAN with a fully convolutional structure, laying a foundation for the application of GAN in image generation. However, its effect in generating complex textures or structures is still not ideal. The WGAN constructed by Arjovsky et al. solves the problem of gradient disappearance during the GAN training process by introducing the Wasserstein distance to replace the traditional JS divergence or KL divergence. However, WGAN uses a weight clipping mechanism during training, which limits the expressive ability of the model. Gulrajani et al. proposed WGAN-GP, which further improves the training stability of the model by introducing a gradient penalty mechanism to avoid weight clipping.
[0006] Leakage defect images often have irregular shapes and complex local details. WGAN-GP cannot perfectly capture these detailed features, and the local details of the generated images lack sufficient precision. The attention mechanism has been widely applied in multiple fields in recent years and has shown remarkable effects in machine vision. Vaswani et al. first proposed the self-attention mechanism, which has demonstrated excellent performance in both natural language processing and image generation tasks; Hu et al. constructed the SE module, which can help the model focus on the key regions of the image by modeling the importance of different channels and adjusting weights, contributing to the accurate restoration of key information. Therefore, we attempt to embed the SE attention module into the generator to enable the model to adaptively adjust the attention to the defective parts, increase the weights of the defective parts, and ensure that the details of the generated images are more realistic.
[0007] Traditional pixel-level loss functions cannot fully capture the high-level semantic information of images, resulting in differences between the generated images and the target images visually and causing regional distortion. Johnson et al. proposed perceptual loss, which measures the similarity between the generated image and the target image at the perceptual level through the feature extraction layer of a pre-trained neural network, improving the distortion phenomenon of the generated images. Li et al. combined perceptual loss with the image generation model, significantly improving the visual quality of the images with watermarks added. Therefore, we also attempt to construct perceptual loss by extracting deeper features in the defect images through the deep convolutional neural network VGG-19 to improve the fidelity of the generated defect images in the overall structure. Summary of the Invention
[0008] To solve the above problems, the present invention constructs a defect image generation model SWG-VGG based on WGAN-GP to achieve the accurate generation and detail restoration of wire defect and leakage defect images. This model integrates the SE attention mechanism to increase the weights of the defective parts in the model and enhance the model's ability to represent the defective morphology; it expands the depth and width of the generator to better capture the detailed features of the defect images; it uses VGG-19 to extract deep features in the defect images to construct perceptual loss to improve the loss function, overcoming the distortion problem in image generation. The technical solutions are as follows: On the one hand, an embodiment of the present invention provides a defective image generation method that combines an attention mechanism and perceptual loss. The method includes: S1: Obtain a real image, where the real image is a defective image on the surface of a liquid crystal display screen. S2: Construct and train an SWG-VGG model based on the WGAN-GP model. The SWG-VGG model includes a generator, a discriminator, and a VGG-19 network. Among them, the generator is adjusted as follows compared to the generator of the WGAN-GP model: Two upsampling units are added before the output layer as a high-resolution feature reconstruction module, an SE attention layer is added before the high-resolution feature reconstruction module as an SE attention module, and SE attention layers are added after the input layer and each upsampling unit. Correspondingly, two convolutional blocks are added to the discriminator compared to the discriminator of the WGAN-GP model. The VGG-19 network is used to extract the high-level features of the generated image and the real image during training to calculate the differences between these features, and finally, the feature differences of different layers are weighted and summed to obtain the final perceptual loss and feedback it to the generator. Among them, the loss function of the SWG-VGG model is the weighted sum of the loss function of the WGAN-GP model and the perceptual loss function of the VGG-19 network. S3: Input the real image into the trained SWG-VGG model to obtain a generated image.
[0009] On the other hand, an embodiment of the present invention further provides a defective image generation system that combines an attention mechanism and perceptual loss, including: A real image acquisition module: used to acquire a real image, where the real image is a defective image on the surface of a liquid crystal display screen. The SWG-VGG model includes a generator, a discriminator, and a VGG-19 network. Among them, the generator is adjusted as follows compared to the generator of the WGAN-GP model: Two upsampling units are added before the output layer as a high-resolution feature reconstruction module, an SE attention layer is added before the high-resolution feature reconstruction module as an SE attention module, and SE attention layers are added after the input layer and each upsampling unit. Correspondingly, two convolutional blocks are added to the discriminator compared to the discriminator of the WGAN-GP model. The VGG-19 network is used to extract the high-level features of the generated image and the real image during training to calculate the differences between these features, and finally, the feature differences of different layers are weighted and summed to obtain the final perceptual loss and feedback it to the generator. Among them, the loss function of the SWG-VGG model is the weighted sum of the loss function of the WGAN-GP model and the perceptual loss function of the VGG-19 network. A training module: used to train the SWG-VGG model through the loss function. A generated image module: used to input the real image into the trained SWG-VGG model to obtain a generated image.
[0010] On the other hand, the embodiments of the present invention also provide an application of the generated image produced by this method in the target detection model YOLOv8. The YOLOv8 model trained using the generated data has achieved 94.6%, 93.8% and 97.6% in precision, recall and mAP(50) respectively.
[0011] This patent proposes an intelligent defect image generation model SWG-VGG based on deep learning, introduces the SE channel attention module into WGAN-GP, and introduces perceptual loss into the model through VGG-19. The generated defect images are fused with the original data set. The YOLOv8 defect detection model is used to compare the generated defect image data set with the original defect image data set. The experimental results show that the data set generated by the SWG-VGG model reaches 0.976 in mAP(50), with a 10.5% improvement; the PSNR, SSIM, RMSE and CS values of the generated images also show high image quality. This model provides a high-quality data set for display defect detection and has important engineering application value. Description of the Drawings
[0012] Figure 1 are defect samples collected by the factory; Figure 2 is the distribution map of the defect samples collected by the factory; Figure 3 is the principle block diagram of the basic structure of GAN; Figure 4 is the framework diagram of the SWG-VGG model of the present invention; Figure 5 is the structural schematic diagram of the generator of the SWG-VGG model; Figure 6 is the structural schematic diagram of the SE attention module; Figure 7 is the structural schematic diagram of the discriminator of the SWG-VGG model; Figure 8 is the flow chart of perceptual loss calculation; Figure 9 is the comparison chart of the generator loss curves of each model of different modules; Figure 10 is the comparison chart of the scores of each index of 200 defect images randomly generated by three generation models; Figure 11 is the comparison chart of the scores of defect images of each background color generated by three generation models; Figure 12 is the image comparison of the line defects and leakage defects actually generated by three generation models; Figure 13 is the schematic diagram of each index value during the training process of YOLOv8 on different data sets; Figure 14 is the comparison chart of the PR curves of each type of defect during the training process of YOLOv8 on different data sets; Figure 15 is the comparison chart of the confusion matrices of different types of defects for two data sets using YOLOv8. Detailed Embodiments
[0013] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.
[0014] I. Related theoretical basis 1.1 Generative adversarial network The generative adversarial network (GAN) was proposed by Ian Goodfellow et al. in 2014. The core idea of GAN is to enable the generator to generate more realistic data through the adversarial training between the generator and the discriminator. The basic structure of GAN is as Figure 3 shown. Specifically, GAN mainly consists of a generator and a discriminator. The generator is responsible for generating data. By receiving a random noise vector that follows a Gaussian distribution or a uniform distribution, it uses deconvolution operations to gradually map the low-dimensional vector to the target data space. The purpose of the generator is to generate fake data similar to real data to deceive the discriminator. The discriminator is a binary classifier used to determine whether the input data is real data or fake data generated by the generator. The discriminator generally uses a convolutional neural network to extract the features of the input real data and generated data, and then outputs a probability through a fully connected layer to represent the authenticity of the input data. The training process of GAN is a game process of minimization and maximization. The goal of the generator is to minimize the correctness of the discriminator so that the discriminator cannot distinguish between the generated data and the real data, while the goal of the discriminator is to maximize its own correctness and accurately judge the authenticity of the input data. The objective function of GAN is defined as: (1) where, in formula (1), D(x) is the output of the discriminator for the real data x, D(G(z)) is the output of the discriminator for the generated data G(z), is the distribution of the real data, is the distribution of the input noise. In this optimization process, the generator and the discriminator will be alternately trained. First, the discriminator is trained with the generator fixed, and then the generator is trained with the discriminator fixed. Eventually, a Nash equilibrium is reached, and the samples generated by the generator are so similar to the real samples that they cannot be accurately distinguished.
[0015] 1.2 Channel attention mechanism Channel attention is an attention mechanism in deep learning, which is used to weight different channels in the channel dimension of the feature map, enabling the model to adaptively focus on the channels that are more important for the current task. In a convolutional neural network, the feature map contains multiple channels, and each channel may represent a certain type of feature (edges, textures, etc.). However, not all of these channels are useful for the task. The goal of the channel attention mechanism is to highlight important features and ignore irrelevant features by learning to assign weights to each channel. The implementation process of the channel attention mechanism mainly consists of three steps. First, feature aggregation is performed. The global average pooling operation is used to aggregate the information of each channel in order to obtain the global information of each channel. The calculation of the information of the c-th channel can be expressed as: (2) where, in Equation (2), X is the feature map, H and W are the height and width of the feature map respectively, and c is the number of channels. Next, weight generation is carried out. The Sigmoid non-linear activation function is used to map the aggregated features to channel weights, and the attention weights of each channel are calculated , and the calculation process can be expressed as: (3) where, in Equation (3), is the Sigmoid activation function. Finally, the channels are re-weighted. The learned weights are reapplied to the corresponding channels to achieve the effect of channel attention, and the calculation process can be expressed as: (4) where, in Equation (4), is the feature map after channel weighting.
[0016] 2.3 Perceptual Loss Perceptual loss is a loss function used in image generation and image translation tasks in deep learning. It aims to improve the visual quality of images by measuring the similarity between the generated image and the target image in the high-level feature space. Compared with traditional pixel-level losses, perceptual loss pays more attention to the high-level semantic information and perceptual effects of images. Therefore, it has been widely used in tasks such as style transfer, image denoising, and super-resolution. Perceptual loss calculates the differences between these features by extracting the high-level features of the generated image and the real image through a pre-trained convolutional neural network, and finally sums up the feature differences of different layers with weights to obtain the final perceptual loss. Perceptual loss consists of a content loss that compares the differences between the generated image and the target image in the feature space of a specific layer, and a style loss that compares the similarity between the generated image and the target image in the feature distribution. Among them, the style loss is used to ensure that the generated image is consistent with the target image in style, focusing on features such as the color, texture, and structure of the image. The content loss is used to ensure that the generated image is consistent with the target image in content. It calculates the differences between the feature representations of the two images on certain feature layers to ensure that the generated image retains the main structure and layout of the target image. The calculation of the style loss is defined as: (5) where, in Equation (5), and are the input generated image and the real image respectively, is the Gram matrix of the l-th layer, is a set of selected multiple feature layers. The calculation of the content loss can be defined as: (6) where, in Equation (6), is the feature map extracted at the l-th layer. The final perceptual loss can be expressed as: (7) where, in Equation (7), and are the weight hyperparameters of the style loss and the content loss respectively.
[0017] II. Structural Overview 2.1 General Description The proposed model SWG-VGG in this patent is based on the improved generative adversarial network WGAN-GP, and has been improved for the problems of complex defect morphologies of small-sized display screens and distorted generated images. The overall framework is as Figure 4As shown, the model consists of a generator, a discriminator, and VGG-19. During the training phase, the generator first receives a 100 (preferably) or 108-dimensional randomly sampled noise vector Z from a standard normal distribution. This noise vector represents random points in the latent space. The task of the generator is to map this noise through the network into an image with the structural and texture features consistent with the defective samples. This process is achieved through deconvolution operations. The generator gradually performs upsampling operations on the input noise Z, gradually generating intermediate feature maps with higher resolutions, and finally generating a color image with a size of 256×256. The discriminator learns the distribution differences between the images generated by the generator and the real images and calculates the loss to help the generator optimize. VGG-19 extracts the high-level features of the generated images and the real images and calculates the perceptual loss in these feature maps , and feeds it back to the generator, enabling the generator to not only generate realistic images at the pixel level but also be consistent with the real defective images in terms of high-level structures. During the defective image generation phase, the trained generator receives the noise vector Z and gradually maps the noise into the target defective image through upsampling.
[0018] 2.2 Generator and Discriminator Design When the traditional WGAN architecture generates high-resolution images, problems such as low image quality and missing details usually occur. To solve these problems, improve the quality of the generated images, and enhance the ability to capture key features. This patent redesigned the structure of the generator, and the structure is as Figure 5 shown. The generator structure mainly includes five parts: (1) Input layer, i.e., noise mapping layer, responsible for receiving a high-dimensional random noise vector Z and mapping it to a high-dimensional feature space to provide initial features for generating images. Z can be 100-dimensional, 108-dimensional, etc. It is specifically composed of an upsampling unit. Compared with the existing WGAN-GP model, an SE attention layer is added after the upsampling unit. (2) Upsampling module, gradually expanding the resolution of the feature map to the target size. It is specifically composed of three upsampling units. Compared with the existing WGAN-GP model, an SE attention layer is added after the existing upsampling unit. (3) SE attention module, an additional module responsible for adjusting the channel weights of the generated feature map to enable the model to capture key features of the image more accurately. It is located between the third upsampling unit and the fourth upsampling unit. (4) High-resolution feature reconstruction module, an additional module that further upsamples the feature map and gradually reduces the number of channels to make the high-resolution features match the target output size. It is specifically composed of two upsampling units. Compared with the existing WGAN-GP model, an SE attention layer is added after the upsampling unit. (5) Output layer, mainly using a transposed convolutional layer to convert the feature map into an RGB three-channel target image number of channels and restricting the pixel values within the range of [-1, 1] through the Tanh activation function. Compared with the existing WGAN-GP model, two upsampling units are added. That is, the generator of this patent has six upsampling units and seven SE attention layers.
[0019] The generator structure is composed of several 4×4 transposed convolutional layers and SE attention layers (SELayer). In this patent, the SELayer is embedded into the generator and the structure of the generator is extended to increase the resolution of the output image from 64×64 to 256×256, greatly enhancing the detail representation ability of the defective image. Each upsampling unit (block) contains a transposed convolutional layer, a batch normalization layer, a ReLU activation function layer, and the SELayer. The role of these layers is to gradually magnify the low-resolution features and optimize the feature expression of the channels using the attention mechanism. The working process of the generator is to first upsample the input feature map using a transposed convolutional operation and extract important features. The calculation process can be expressed as: (8) Among them, in Equation (8), and are the heights of the output and input feature maps respectively, is the stride, is the padding size, is the convolutional kernel size. Then, batch normalization operation is performed on the output feature map to standardize each batch of feature maps, reduce the internal covariance shift during training, and thus accelerate convergence and stabilize training. For each feature channel, the batch normalization process is: (9) (10) Among them, in formulas (9) and (10), x1 is the feature map, is a small constant, the purpose of which is to prevent division by zero; and are learnable parameters and are the scaling coefficient and the translation coefficient respectively, , . Then, nonlinearity is introduced through the ReLU activation function to enhance the feature expression ability. Finally, through an SELayer, the performance of channel features is enhanced.
[0020] Among them, the SE attention layer is a channel attention module used to improve the performance of convolutional neural networks. In the task of generating display defect images, the SE attention layer can weight the importance of each channel, thereby enhancing the generator's attention to the defect morphology to improve the clarity of the generated defect images. The structure of the SE attention layer is as Figure 6 shown.
[0021] First, the input feature Figure X is transformed into the feature map U through a standard convolutional operator Ftr, that is, each layer of the input feature Figure X passes through a convolution with a 2D spatial kernel and finally obtains C output feature maps to form the feature map U. Then, a compression operation is performed, that is, the global spatial information is compressed into a channel descriptor. The feature map U containing global information is compressed into a 1×1×C channel descriptor (a feature vector) using global average pooling of channels. The specific operation is defined as: (11) Among them, in formula (11), is the value of channel c at the spatial position , is the global average pooling of channels, After feature compression, the obtained channel descriptor is input into the weight generation network for weight generation to learn the weight of each channel. The weight generation is divided into two steps. First, the channel descriptor passes through the first fully connected layer and uses the ReLU activation function to reduce the feature dimension to C / r of the original, where r is the scaling factor. Then, the dimension is restored to C through the second fully connected layer, and the Sigmoid activation function is used to limit the weight within the range of [0, 1]. Each channel will obtain a corresponding weight value to represent its importance in the current task. This weight generation process can be expressed as: (12) Among them, in formula (12), z is the channel descriptor vector, W1 and W2 are the weight matrices of the first fully connected layer and the second fully connected layer respectively, and are the ReLU and Sigmoid activation functions respectively, s is the channel weight vector and it is ; finally, the generated channel weight s is multiplied by the feature map U channel by channel to achieve the final output of feature weighting, and the feature weighting can be expressed as: (13) Among them, in formula (13), is the weight of the c-th channel, is the value of the feature map U in channel c, and finally the weighted feature map is output.
[0022] In WGAN-GP, the objective of the discriminator is different from that of the traditional GAN. The discriminator of WGAN-GP uses the Wasserstein distance instead of the cross-entropy loss. WGAN-GP optimizes the adversarial between the generator and the discriminator by minimizing the Wasserstein distance. To ensure the Lipschitz continuity of the discriminator, WGAN-GP also introduces gradient penalty. The discriminator structure is as Figure 7 shown. The input is an RGB image with a size of 3×256×256, including the fake image generated by the generator and the images in the original dataset. To correspond to the generator, the discriminator structure contains a total of 5 convolutional blocks, each convolutional block consists of a convolutional layer, a batch normalization layer and a LeakyReLU activation function, gradually downsampling the image features, and finally outputting a real number score value D(x) through a 4×4 convolutional layer. The loss calculation part of the discriminator is composed of the weighted sum of the Wasserstein loss and the gradient penalty. The calculation process of the Wasserstein loss can be expressed as: (14) Among them, in formula (14), D(x) is the real number score value output by the discriminator to represent the probability that the sample x comes from the real data; E is the expectation; and are the real image samples and the generated image samples respectively, and are the distributions of the generated data and the real data respectively. The objective of the discriminator loss function is to maximize the real data score , while minimizing the score of the generated samples. To ensure the 1-Lipschitz continuity of the discriminator, WGAN-GP introduces a gradient penalty term. This penalty term enforces the Lipschitz continuity condition on the gradients of the discriminator over the interpolated samples between real and generated images. The calculation process of the gradient penalty term can be expressed as: (15) (16) where, in equations (9) and (10), is the real image sample and the generated image sample is the linear interpolation between them, is a random number sampled from a uniform distribution, , is 's data distribution. So the total loss of the discriminator is the weighted sum of the Wasserstein loss and the gradient penalty, which can be expressed as: (17) 2.3 Loss Function Improvement The loss function of the original WGAN-GP model cannot fully capture the high-level semantic information of images. Therefore, this model improves the problem of image distortion generated by the generator by introducing perceptual loss. The perceptual loss calculates the difference between the high-level features of the generated image and the real image by using VGG-19 to extract these features, and finally sums up the feature differences of different layers with weights to obtain the final perceptual loss. The calculation flow chart of the perceptual loss is as Figure 8 shown. First, the input image X passes through a generation network (ImageTransformNet in the figure) to output an image . The VGG network extracts the feature maps of the input image, and then calculates the mean squared errors between the style image, the content image and the generated image respectively to obtain the loss functions , . Finally, the weights of the generation network are adjusted according to the loss values to output the image . This model extracts image features by using the five convolutional layers of VGG-19 (specifically the first layer of five blocks), and calculates the style losses at different levels by Conv1_1, Conv1_12_1, Conv3_1, Conv4_1 and Conv5_1 respectively, calculates the content loss by using Conv3_1, and synthesizes them into the perceptual loss . The calculation process can be expressed as: (18) (19) (20) Among them, in formulas (18), (19) and (20), L is a set of feature layers, , , are the number of channels, height and width of the feature map respectively, is the generated image, is the style image, is the content image, is to calculate the Gram matrix for the target image, represents the feature map of the image at the third layer, l represents five layers of feature layers, i is the channel index, j is the height index, and k is the width index. The total loss function of this patent combines the WGAN-GP loss and the perceptual loss, and is defined as follows: (21) Among them, in formula (21), is the weight coefficient of the perceptual loss, which is used to adjust the weight of the perceptual loss in the overall loss. Through experimental analysis, the model loss convergence process has the best effect when this value is set to 0.001.
[0023] Correspondingly, the embodiment of the present invention also provides a defective image generation system that integrates an attention mechanism and a perceptual loss, including: Real image acquisition module: used to acquire a real image, and the real image is a defective image on the surface of a liquid crystal display screen.
[0024] The SWG-VGG model includes a generator, a discriminator, and a VGG-19 network. Among them, the generator is adjusted as follows compared to the generator of the WGAN-GP model: two upsampling units are added before the output layer as a high-resolution feature reconstruction module, an SE attention layer is added before the high-resolution feature reconstruction module as an SE attention module, and an SE attention layer is added between the third upsampling unit and the fourth upsampling unit. SE attention layers are added after the input layer and each upsampling unit; then the generator includes an input layer, an upsampling module, an SE attention module, a high-resolution feature reconstruction module, and an output layer arranged in sequence. The input layer includes an upsampling unit and is used to obtain a high-dimensional random noise vector Z and map it to a high-dimensional feature space to provide initial features for generating images. The upsampling module includes three upsampling units arranged in sequence. The SE attention module includes an SE attention layer. The high-resolution feature reconstruction module includes two upsampling units arranged in sequence and is used to upsample the feature map and gradually reduce the number of channels so that the high-resolution features match the target output size. The output layer is used to convert the feature map into the number of channels of the target image with RGB three channels through a transposed convolutional layer and limit the pixel values within the range of [-1, 1] through the Tanh activation function.
[0025] Among them, the upsampling unit includes a transposed convolutional layer, a batch normalization layer, a ReLU activation function layer, and an SE attention layer arranged in sequence. Among them, the transposed convolutional layer is specifically a 4×4 transposed convolutional layer. The SE attention layers of the input layer, the upsampling module, and the high-resolution feature reconstruction module are used to weight the importance of each channel to enhance the generator's attention to the defect morphology and improve the clarity of the generated images. The SE attention layer of the SE attention module is used to first transform the input features Figure X into a feature map U through a standard convolutional operator Ftr, then perform a compression operation to compress the global spatial information into a channel descriptor, then input the obtained channel descriptor into a weight generation network for weight generation to learn the weight s of each channel, and finally multiply the generated channel weight s with the feature map U channel by channel to achieve feature weighting. Among them, the discriminator includes five convolutional blocks arranged in sequence and a 4×4 convolutional layer. The convolutional block includes a convolutional layer, a batch normalization layer, and a LeakyReLU activation function arranged in sequence, and is used to gradually downsample the generated images generated by the generator and the real images in the original dataset; the 4×4 convolutional layer is used to output a real number score value D(x). The VGG-19 network is used to extract the high-level features of the generated images and the real images during training to calculate the differences between these features, and finally weight and sum the feature differences of different layers to obtain the final perceptual loss and feedback it to the generator.
[0026] Training module: used to train the SWG-VGG model through a loss function.
[0027] Generated Image Module: It is used to input real images into the trained SWG-VGG model to obtain generated images.
[0028] III. Effect Verification 3.1 Experimental Environment and Dataset To verify the effectiveness of the method of this patent, three groups of experiments were designed for verification. All the experiments of this patent were implemented under the Pytorch framework, with the operating system being Windows 11, the CPU being 12th Gen Intel(R) Core(TM) i5-12400F 2.50 GHz, the GPU being NVIDIA GeForce RTX 3080 Ti, and the running memory being 12G. The liquid crystal display surface defect image data used in this patent was collected on-site in the factory. To ensure the authenticity and diversity of the data, images of liquid crystal displays in different batches were collected in the actual production environment. Table 1 shows the distribution of defect sample data. The defect detection process requires the display screen to be lit first. Since the presentation forms and states of defects are different under different background colors, it is required that there are no defects when the display screen is in five background colors: white, blue, green, red, and color in sequence. In the line defect, black indicates the situation where the display screen cannot be lit normally and bright lines appear on the screen. Since the color of the leakage defect itself is black and it cannot be displayed when the display screen is not lit or in the black screen state, there are no leakage defects with a black background in the sample data. In this patent, each type of defect was collected under different background colors, and 30 different defect images of each defect were collected under different background colors, thus forming a defect image dataset with multiple backgrounds and multiple perspectives. A high-resolution industrial camera was used during the collection process. The dataset contains defect images from multiple production lines and different process conditions to ensure the representativeness of the data.
[0029] Table 1 Distribution of Sample Data of Line Defects and Leakage Defects
[0030] 3.2 Evaluation Metrics In this patent, to comprehensively evaluate the quality of the generated images, multiple commonly used image quality evaluation metrics were selected, including Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), Root Mean Square Error (RMSE), and Correlation Coefficient (CS).
[0031] 3.3 Experimental Results and Analysis 3.3.1 Ablation Analysis To verify the effectiveness of the model proposed in this patent, multiple groups of ablation experiments were conducted. By gradually adding the SE module and perceptual loss, the contribution of each module to the quality of the generated images was evaluated. The experiments were based on a dataset of surface defect images of small-sized liquid crystal screens, and the image resolution was uniformly 256×256. PSNR, SSIM, MI, and CS were used as image quality evaluation metrics. The following four comparison models were designed in the experiment: Model 1 is a traditional WGAN-GP (benchmark model); Model 2 adds the SE module on the basis of the benchmark model (adding an SE attention layer before the high-resolution feature reconstruction module as the SE attention module, and adding SE attention layers after the input layer and each upsampling unit); Model 3 introduces perceptual loss on the basis of the benchmark model; Model 4 adopts the complete model of this patent.
[0032] Under the same experimental environment, using the same computer and the same original dataset, four groups of defect image generation experiments were conducted using these four generation models respectively. The average index scores of the generated liquid leakage defects and line defect images in each group of experiments were statistically shown in Table 2, and the loss curves of the generators of each model were as Figure 9 shown. To better demonstrate the influence of each module on the model, this patent separately statistically compared the index scores of each model at different training stages. Table 3 shows the score comparison of liquid leakage defects at different training stages, and Table 4 shows the score comparison of line defects at different training stages; Figure 10 、 11 are the comparisons of the generated liquid leakage defects and line defect images when the model is trained to different stages respectively.
[0033] Table 2 Score comparison of the four models in PSNR, SSIM, RMSE, and CS
[0034] From the above experimental data, it can be seen that: (1) Compared with Model 1, the index scores of Model 2 are all better than those of Model 1. Among them, the PSNR, SSIM, and CS values are increased by 0.57, 0.0104, and 0.014 respectively, and the RMSE is reduced by 1.1001; Tables 3 and 4 show that the index scores of the defect images generated by Model 2 after 1900 epochs are all better than those of Model 1. This shows that the introduction of the SE module indeed improves the overall quality of the generated images to a certain extent, especially improving the details of the generated images at the perceptual level such as brightness, contrast, and structural features. In addition, by Figure 9It can be seen that the loss curve of the Model 2 generator converges significantly faster than that of Model 1, and the loss value is slightly lower than that of Model 1. Thus, it can be seen that the SE module indeed helps us achieve the goal of better capturing the inter-channel dependencies of images and improving the local detail quality of the generated images.
[0035] (2) Compared with Model 1, the PSNR, SSIM, and CS of Model 3 are increased by 0.39, 0.0212, and 0.0097 respectively, and the RMSE value is decreased by 0.7345. The overall improvement amplitude is slightly lower than that of Model 2; Tables 3 and 4 show that in most training stages, the various indicators of the defective images generated by Model 3 are superior to those of Model 1; Figure 9 It shows that the loss convergence speed of the Model 3 generator is roughly the same as that of Model 1, but the loss value is significantly lower than that of Model 1. This indicates that introducing perceptual loss can improve the overall quality of the generated images, and at the levels of the overall quality and perceptual quality of the images, the loss of the generator with perceptual loss introduced is significantly lower than that of Model 1 and Model 2.
[0036] (3) Compared with Model 1, the PSNR, SSIM, and CS of Model 4 are increased by 1.6, 0.0417, and 0.0442 respectively, and the RMSE value is decreased by 1.7238; Tables 3 and 4 show that among the four evaluation indicators, Model 4 achieves the best results among the four models in most training stages. And, the loss convergence speed of the generator is slightly faster than that of Model 1, and the loss value is the same as that of Model 3. This indicates that by introducing the SE module and perceptual loss simultaneously, the model has made significant breakthroughs in both the overall similarity and perceptual level of the reconstructed image quality. Figure 10 And Figure 11 It can also be clearly seen that the images of leakage defects and line defects generated by Model 4 are superior to those of the other three models in terms of both quality and details. Therefore, introducing the SE module or perceptual loss alone can improve the image quality, while introducing both simultaneously greatly enhances the performance of the model. Thus, it can be seen that the combination of the SWG-VGG architecture design and perceptual loss proposed in this patent can greatly improve the quality of the defective image generation of small-size liquid crystal displays.
[0037] Table 3 Score table of the leakage defect images generated by the four models at different training stages
[0038] Table 4 Score table of the line defect images generated by the four models at different training stages
[0039] 3.3.2 Comparative analysis To comprehensively evaluate the performance of the SWG-VGG model proposed in this patent, comparative experiments were conducted with the classical WGAN-GP and DCGAN models. The performance of different models in generating images of two typical defects (line defects and liquid leakage defects) was tested. The experimental dataset still used the dataset shown in Table 1. Under the same experimental environment, using the same computer and the same original dataset, three groups of defect image generation experiments were carried out using these three generative models respectively. The PSNR, SSIM, RMSE, and CS of the generated images in each group of experiments and the values of each index are shown in Table 5. And, the average values of the images generated by the three models at different training stages are shown in Table 6. The score values of each index for 200 randomly generated line defect and liquid leakage defect images by the three models and the score comparison of each index under various backgrounds are as Figure 12 , 13 shown, and the image comparison of the line defects and liquid leakage defects actually generated by the three models is as Figure 14 shown.
[0040] Table 5 Score Table of Defect Images Generated by WGAN-GP, DCGAN, and SWG-VGG Models in Each Index
[0041] Table 5 and Figure 12 show that: The PSNR, SSIM, RMSE, and CS scores of the line defect and liquid leakage defect images generated by the SWG-VGG model proposed in this patent are all better than those of the other two models. The scores of the liquid leakage defect images are 23.7, 0.83, 19.13, and 0.84 respectively; the scores of the line defect images are 23.41, 0.78, 22.92, and 0.81 respectively, and the average scores are 23.56, 0.80, 21.02, and 0.83 respectively. Compared with the defect images generated by the WGAN-GP model, the PSNR, SSIM, and CS values have increased by 1.6, 0.04, and 0.05 respectively; the RMSE value has decreased by 1.73. From Figure 12 the scatter plot, it can also be seen that the scores of the liquid leakage defect and line defect images generated by the SWG-VGG model in the four indexes are significantly better than those of the WGAN-GP model. Compared with the DCGAN model, the PSNR, SSIM, and CS values of the defect images generated by the model proposed in this patent have increased by 1.13, 0.03, and 0.04 respectively; the RMSE value has decreased by 0.81. Figure 12 It can be seen from the scatter plot that in the RMSE value, the score of the model proposed in this patent for the liquid leakage defect image is significantly better than that of the DCGAN model, and the score distribution of other indexes is also slightly higher than that of the defect images generated by the DCGAN model.
[0042] Table 6 Score Table of Each Index of WGAN-GP, DCGAN, and SWG-VGG Models at Different Training Stages
[0043] Figure 13 With Figure 14 It shows that for the liquid leakage defect, the PSNR of SWG-VGG is the highest under all background colors. The PSNR of WGAN-GP and DCGAN is relatively low, especially under white and red backgrounds, and WGAN-GP performs the worst. The SSIM shows a similar trend among the three models. The structural similarity of SWG-VGG is the best and is always higher than the other two models. The SSIM of DCGAN is better than that of WGAN-GP, especially under colored and white backgrounds. The RMSE value of SWG-VGG is still the lowest, proving that it has the smallest error when generating images. In contrast, the RMSE value of WGAN-GP is always the highest, indicating that the error of the images generated by it is the largest. The CS index performs the best in SWG-VGG, further proving that the images generated by it have the highest correlation with the reference images. For the line defect, SWG-VGG performs the best in terms of PSNR. However, different from the light leakage defect, the influence of the background color is relatively large, especially under white and red backgrounds, and the performance decreases slightly. The SSIM and CS are also relatively high in SWG-VGG, especially under black and blue backgrounds, showing its stable performance. The performance of WGAN-GP is the worst, especially in terms of RMSE and CS, verifying its deficiency in generating line defect images. Figure 14 It shows that the defect images generated by SWG-VGG highly restore the original images in terms of color, texture, and structure, and accurately and clearly capture the defect targets in all examples, except that the targets under certain backgrounds are slightly lighter in color; there are image distortion problems in the colored and white background line defect images generated by the WGAN-GP model, and there are fewer distortion problems in the defect images generated by the DCGAN model, but the quality of the generated defect images is relatively poor.
[0044] 3.3.3. Application Analysis To verify the effectiveness of the defective liquid crystal display images generated by the SWG-VGG model proposed in this patent in practical applications, this patent uses the current excellent target detection model YOLOv8 for testing and analysis. YOLOv8 is a single-stage target detection algorithm that can achieve a good balance between detection speed and accuracy and is widely used in the field of defect detection. In the experiment, the training set uses the defective images generated by the SWG-VGG model, and the validation set uses the original defective images. The input image size of YOLOv8 is uniformly adjusted to 416 × 416 to adapt to its network structure. During the training process, the hyperparameters used by the model are set as follows: the number of training epochs is 50, the batch size is 16, and the optimizer is Adam. During the training process, the loss function is used to calculate the error between the prediction result and the true annotation, and the model parameters are continuously optimized through backpropagation. Precision, recall, and mean average precision mAP(50) are used to evaluate the performance of the model on the validation set. The training results are as Figure 14 shown in Table 7.
[0045] Table 7 mAP(50) values of various defects in the original dataset and the dataset after adding the generated images
[0046] From Figure 15It can be clearly seen that the YOLOv8 model trained with generated data achieved 94.6%, 93.8%, and 97.6% in precision, recall, and mAP(50) respectively. In contrast, the model trained only with original images performed significantly worse on the same test set, with precision, recall, and mAP(50) being only 89.9%, 74.5%, and 87.1%. The YOLOv8 model trained with generated data had improvements of 4.7%, 19.3%, and 10.5% respectively. This difference indicates that the introduction of the generated dataset, especially its advantages in the diversity and quantity of defect samples, significantly improved the performance of YOLOv8. The number of defect samples in the original dataset was small and the sample types were single, which led to overfitting in the model training process, especially the over-reliance on specific types of defect samples. The overfitting phenomenon caused a significant decrease in recall and detection accuracy when the model faced new samples. From the average detection accuracy mAP(50) values of various types of defects in Table 7, it can be seen that the average mAP(50) value of all types using the original dataset was only 0.871. Among them, the mAP(50) for liquid leakage defects was 0.941, for point defects was 0.962, for scratch defects was 0.972, while for line defects it was only 0.609, showing obvious deficiencies. This indicates that the various changes in the shape and color of line defects further amplified the disadvantage of insufficient diversity in the original dataset, and the decrease in recall was particularly obvious. Compared with the detection results of the original dataset, after adding the generated dataset, the overall performance of the model was significantly improved, with mAP(50) reaching 0.976 and the detection accuracy increasing by 10.5%. The detection performance of each category was significantly improved. The mAP(50) for liquid leakage and scratch defects reached 0.995, for point defects was 0.923, and the mAP of line defects increased to 0.990, approaching the optimal state. It can be seen that the generated data played a key role in improving the model's detection ability for categories with weak performance in the original dataset, such as line defects. This improvement may be attributed to the fact that the generated data effectively expanded the diversity and balance of the original data and enhanced the model's learning ability for the features of each category.
[0047] To further verify the influence of the dataset quantity on the stability of the defect detection model, this patent designed a series of experiments. The generated defect images were used to enhance the dataset to 2000, 4000, 6000, and 8000 images respectively for model training and testing. By comparing the performance of the model under different dataset scales, it aimed to explore the specific influence of the dataset size on the model training effect, precision, and stability. In the experiment, we mainly focused on the precision of the model for detecting four types of defects and the stability during the training process, especially whether the model was prone to overfitting or unstable training under the condition of a small dataset. Table 8 is the data distribution table in the dataset, and Table 9 is the detection precision statistical table of various types of defects using different datasets.
[0048] Table 8 Distribution of various defects in datasets with different numbers of data
[0049] Table 9 Statistical table of the accuracy and average accuracy of each defect in different datasets
[0050] Table 9 shows that the average accuracy of dataset 2 is 97.8%, which is 0.2% higher than that of dataset 1, with a relatively small increase. Among them, the accuracy of line defects is lower than that of dataset 1, and the accuracy of point defects is significantly higher than that of dataset 1, indicating that although the increase in the amount of data has a relatively small increase in the overall accuracy, it is beneficial to the detection of small-size defects such as point defects; the average accuracy of dataset 3 is 97.5%, which is roughly the same as that of dataset 1, and the gap in the detection accuracy of various defects has narrowed, and the accuracy of point defects is higher than that of dataset 1, indicating that the increase in the amount of data increases the stability of the detection model; the average accuracy of dataset 4 reaches the highest, and the accuracy of its point defects reaches 95.4, which is significantly higher than the point defect detection accuracy of the first three datasets, and the overall improvement is relatively stable, indicating that the increase in the scale of the dataset helps to improve the detection accuracy of the model, especially when dealing with complex and difficult-to-identify defect types, the expansion of the amount of data is particularly important.
[0051] IV. Conclusion Aiming at the problems that the defect sample set of small-sized liquid crystal display screens has scarce data and the data set is extremely imbalanced, which seriously affects the accuracy of intelligent detection models, this patent proposes a defect image generation model SWG-VGG based on WGAN-GP, which realizes the accurate generation and detail restoration of line defect and leakage defect images, and creates an effective and balanced defect sample data set for the intelligent defect detection of liquid crystal screens. In this model, we introduce the Wasserstein distance and gradient penalty mechanism, which effectively improves the stability of model training; integrates the SE attention mechanism to increase the weight of the model in the defect part and enhance the model's ability to represent the defect morphology; expands the depth and width of the generator to better capture the details of the defect image in terms of texture, color, and shape; uses VGG-19 to extract deeper features in the defect image to construct the perceptual loss, improves the loss function, and overcomes the distortion problem in image generation. The results of ablation comparison experiments show that each improved design of this model has improved the quality of the generated images. Moreover, our improved overall model has significantly improved the four evaluation indicators PSNR, SSIM, MI, and CS of the generated images while ensuring more stable training and faster convergence. The comparison experiments with DCGAN and WGAN-GP show that SWG-VGG has an absolute advantage in the quality of the generated images in any background. The application implementation shows that the defect detection model trained with the defect sample set generated by SWG-VGG has increased the detection accuracy from 87.1% to 97.6%, especially a crucial improvement in the detection accuracy of line defects. It can be seen that the SWG-VGG intelligent image generation model can indeed accurately generate and detail-restore the line defect and liquid leakage defect images of small-sized liquid crystal screens with high quality, can create an effective data set for the intelligent defect detection of liquid crystal screens, and is a liquid crystal screen defect data enhancement method with general engineering application value.
Claims
1. A defective image generation method that integrates an attention mechanism and perceptual loss, characterized in that, The method includes: S1: Obtain a real image, where the real image is a surface defect image of a liquid crystal display screen; S2: Based on the WGAN-GP model, construct and train the SWG-VGG model. The SWG-VGG model includes a generator, a discriminator, and a VGG-19 network. Among them, the generator is adjusted as follows compared with the generator of the WGAN-GP model: Two upsampling units are added before the output layer as a high-resolution feature reconstruction module, an SE attention layer is added before the high-resolution feature reconstruction module as an SE attention module, and SE attention layers are added after the input layer and each upsampling unit. Correspondingly, two convolutional blocks are added to the discriminator compared with the discriminator of the WGAN-GP model. The VGG-19 network is used to extract the high-level features of the generated image and the real image during training to calculate the difference between these features, and finally the feature differences of different layers are weighted and summed to obtain the final perceptual loss and feedback it to the generator; Among them, the loss function of the SWG-VGG model is the weighted sum of the loss function of the WGAN-GP model and the perceptual loss function of the VGG-19 network; S3: Input the real image into the trained SWG-VGG model to obtain a generated image.
2. The method according to claim 1, wherein The generator includes an input layer, an upsampling module, an SE attention module, a high-resolution feature reconstruction module, and an output layer arranged in sequence; The input layer includes an upsampling unit and is used to obtain a high-dimensional random noise vector Z and map it to a high-dimensional feature space to provide initial features for the generated image; The upsampling module includes three upsampling units arranged in sequence; The SE attention module includes an SE attention layer; The high-resolution feature reconstruction module includes two upsampling units arranged in sequence and is used to match the high-resolution features with the target output size by upsampling the feature map and gradually reducing the number of channels; The output layer is used to convert the feature map into the number of channels of the target image with RGB three channels through a transposed convolutional layer and limit the pixel values within the range of [-1, 1] through a Tanh activation function; Among them, the upsampling unit includes a transposed convolutional layer, a batch normalization layer, a ReLU activation function layer, and an SE attention layer arranged in sequence; Among them, the SE attention layers of the input layer, the upsampling module, and the high-resolution feature reconstruction module are used to weight the importance of each channel to enhance the generator's attention to the defect morphology and improve the clarity of the generated image. The SE attention layer of the SE attention module is used to first transform the input feature map X into a feature map U through a standard convolutional operator Ftr, then perform a compression operation to compress the global spatial information into a channel descriptor, and then input the obtained channel descriptor into the weight generation network for weight generation to learn the weight s of each channel. Finally, the generated channel weight s is multiplied by the feature map U channel by channel to achieve feature weighting.
3. The method according to claim 2, wherein The deconvolution layer is a 4×4 deconvolution layer, and the working process of the deconvolution layer is as follows: Through the deconvolution operation, the input feature map is upsampled and important features are extracted. The calculation process can be expressed as: ; Among them, and are the heights of the output and input feature maps respectively, is the stride, is the padding size, is the convolution kernel size; For each feature channel, the batch normalization process is as follows: ; ; Among them, x1 is the feature map, is a constant, and are learnable parameters and are the scaling coefficient and the translation coefficient respectively, , .
4. The method according to claim 2, wherein In the SE attention module, the compression process is as follows: The feature map U containing global information is compressed into a 1×1×C channel descriptor through global average pooling of channels. The calculation process is: ; Among them, is the value of channel c at the spatial position ; is the global average pooling of the used channels, ; The weight generation is divided into two steps. First, the channel descriptor passes through the first fully connected layer and uses the ReLU activation function to reduce the feature dimension to C / r of the original, where r is the scaling factor and C represents the dimension. Then, it passes through the second fully connected layer to restore the dimension to C, and uses the Sigmoid activation function to limit the weight within the range of [0, 1]. Each channel will obtain a corresponding weight value to represent its importance in the current task. The weight generation process is: ; where z is the channel descriptor vector, W1 and W2 are the weight matrices of the first fully-connected layer and the second fully-connected layer respectively, and are the ReLU and Sigmoid activation functions respectively, s is the channel weight vector and it is ; finally, the generated channel weight s is multiplied by the feature map U channel by channel to achieve the final output of feature weighting. The process of feature weighting is as follows: ; Among them, is the weight of the c-th channel, is the value of the feature map U at channel c.
5. The method according to claim 2, wherein The discriminator includes five sequentially arranged convolutional blocks and a 4×4 convolutional layer; the convolutional block includes a sequentially arranged convolutional layer, a batch normalization layer, and a LeakyReLU activation function, which are used to gradually downsample the image features of the generated image generated by the generator and the real image in the original dataset; The 4×4 convolutional layer is used to output a real number score value D(x).
6. The method according to claim 5, characterized in that, The loss function of the SWG-VGG model is as follows: ; Among them, is the loss function of the WGAN-GP model; the is the perceptual loss function obtained by using VGG-19 to extract the high-level features of the generated image and the real image, calculating the differences between these features, and finally performing weighted summation on the feature differences of different layers; is the weight coefficient of the perceptual loss.
7. The method according to claim 6, characterized in that, The said The calculation process is as follows: ; Among them, is the Wasserstein loss, is the gradient penalty; Among them, ; Among them, D(x) is the real score value output by the discriminator to represent the probability that the sample x comes from real data; E is the expectation; and are the real image sample and the generated image sample respectively, and are the distributions of the generated data and the real data respectively; Among them, ; ; Among them, is the real image sample and the generated image sample is the linear interpolation between them, is a random number sampled from a uniform distribution, , is the data distribution.
8. The method according to claim 7, wherein The generation process of the perceptual loss is: The input image X passes through a generation network to output an image , and the VGG network extracts the feature maps of the input image through five convolutional layers, and then calculates the mean squared errors between the style image, the content image and the generated image to obtain the loss function , , and finally adjusts the generated image according to the loss value The five convolutional layers are Conv1_1, Conv2_1, Conv3_1, Conv4_1 and Conv5_1, and the style losses at different levels are calculated through Conv1_1, Conv2_1, Conv3_1, Conv4_1 and Conv5_1 , and the content loss is calculated through Conv3_1 ; ; ; ; Among them, L is the set of feature layers, , , are the number of channels, height, and width of the feature map respectively, is the generated image, is the style image, is the content image, is to calculate the Gram matrix for the target image, is the feature map of the image at the third layer, l is the five-layer feature layer, i is the channel index, j is the height index, and k is the width index.
9. The method according to claim 6, wherein The real images include images where the display screen cannot be normally lit and bright lines appear on the screen, images with line defects under five background colors of white, blue, green, red, and color, and images with liquid leakage defects under five background colors of white, blue, green, red, and color; The real images are from multiple production lines and different process conditions; the resolution of the generated images is 256×256; the generator receives a 100-dimensional randomly sampled noise vector Z from the standard normal distribution and finally generates a color image with a size of 256×256; is 0.
001.
10. A defective image generation system that integrates an attention mechanism and perceptual loss, characterized in that It includes: Real image acquisition module: used to acquire real images, and the real images are surface defect images of liquid crystal display screens; The SWG-VGG model includes a generator, a discriminator, and a VGG-19 network; among them, the generator of the SWG-VGG model is adjusted as follows compared with the generator of the WGAN-GP model: Two upsampling units are added before the output layer as a high-resolution feature reconstruction module, an SE attention layer is added before the high-resolution feature reconstruction module as an SE attention module, and an SE attention layer is added after the input layer and each upsampling unit; Correspondingly, the discriminator of the SWG-VGG model correspondingly adds two convolutional blocks; The VGG-19 network is used to extract the high-level features of the generated image and the real image during training to calculate the difference between these features, and finally the feature differences of different layers are weighted and summed to obtain the final perceptual loss and feedback it to the generator; among them, the loss function of the SWG-VGG model is the weighted sum of the loss function of the WGAN-GP model and the perceptual loss function of the VGG-19 network Training module: used to train the SWG-VGG model through the loss function; Generated image module: used to input the real image into the trained SWG-VGG model to obtain the generated image.
Citation Information
Patent Citations
Multi-band image synchronous fusion and enhancement method based on improved WGA-GP
CN111696066A
Underwater image enhancement method and system based on conditional generative adversarial network
CN115565056A
Method for training image generation model and computer device
US20210192275A1
Contrast-agent-free medical diagnostic imaging
US20220208355A1
Cited By
Method for correcting difference between imaging wafer optical imaging simulation and actual measurement image
CN120850943A
Industrial defect image generation method, terminal equipment and storage medium
CN120931643A