An industrial scratch generation adversarial network system for solving single semantic diversity mapping

By splicing in the generator, constants and semantic graphs can be learned, and feature pyramid module is introduced into the discriminator, which solves the problem of unnatural generation of repeatable textures and brightness under large areas of the same color semantic image, and the generated scratched images are significantly improved in clarity and naturalness.

CN115205237BActive Publication Date: 2025-07-22CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210771641.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-07-22
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

The existing generative adversarial networks have problems of image blurring, repetitive textures and unnatural brightness when generating scratched images, especially in large areas of semantic images of the same color, which are difficult to generate scratched images with rich texture details.

Method used

An industrial scratch generation adversarial network system that solves single semantic diversity mapping is adopted. By splicing learnable constants and semantic maps as inputs in the generator, and introducing feature pyramid modules in the multi-scale discriminator, combining the joint training of Gaussian noise and multi-scale discriminator, the clarity of the image and natural lighting effect are improved.

Benefits of technology

The generated scratch image has improved sharpness, reduced repeatable textures, more natural over-lighting, and significantly better generation effects than existing models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GDA0003834913910000041
    Figure GDA0003834913910000041
  • Figure HDA0003724381820000011
    Figure HDA0003724381820000011
  • Figure HDA0003724381820000012
    Figure HDA0003724381820000012
Patent Text Reader

Abstract

The present invention discloses an industrial scratch generation adversarial network system for solving single semantic diversity mapping. The network includes a generator for splicing a learnable constant and a semantic map to form input data and then outputting a forged scratch image; and a multi-scale discriminator for judging the authenticity of an image, wherein a feature pyramid module is introduced at the second downsampling and the third downsampling respectively. By using the system of the present invention, scratch images with rich texture details, natural image overexposure, and high clarity can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically, a realistic scratch image is generated by an industrial scratch generative adversarial network that adopts single semantic diversity mapping. Background Art

[0002] In industrial production, problems such as cracks and scratches exist on the surfaces of produced products (such as metal surfaces). Detecting these scratches is extremely challenging whether it is for manual inspection or machine vision inspection. Manual detection is extremely costly and inefficient, so currently the mainstream is to use machine vision to detect scratches. However, when using machine vision to detect, a large number of scratch samples are often required to train a scratch detection model. However, due to (1) different types of scratches on different products cannot be trained together, (2) the number of scratches on some products is relatively small, (3) some products involve confidentiality and too many scratch samples cannot be made public, etc., there are not enough samples to train the scratch detection model.

[0003] To solve the problem of a small number of samples, in addition to using traditional data augmentation methods (translation, rotation, scaling, etc.), generative adversarial networks (GANs) can also be used to generate forged samples to achieve the purpose of expanding samples.

[0004] Currently, the mainstream generative adversarial network (GAN) models for image translation are Pix2Pix, Pix2PixHD, and SPADE.

[0005] Pix2Pix: The initial image translation task was achieved by using a generator with a Unet structure combined with a PatchGAN discriminator. Due to the limitations of its network structure, it lacks the ability to generate diverse images. Coupled with the limitations of the loss function, the generated images are (1) blurred, (2) have repetitive textures, and (3) the brightness cannot be controlled.

[0006] Pix2PixHD: High-definition image translation was completed through a coarse-to-fine generator structure combined with a multi-scale discriminator, and high-definition images can be generated. Although the Pix2PixHD network has strong capabilities and can generate high-definition images, it still cannot solve the problems of (1) repetitive textures in the generated images and (2) overly unnatural brightness in the generated images.

[0007] SPADE: The problem of semantic loss during normalization was solved through spatial adaptive normalization, so that diverse images can also be generated from a semantic image of a single color. Through the SPADE module, the single semantic diversity mapping problem can be alleviated. However, for a semantic image with a large area of the same color, (1) repetitive textures still cannot be effectively eliminated, and (2) the generated images have overly unnatural brightness.

[0008] In summary, the images generated by existing generative adversarial networks have the following defects: (1) the generated scratch images are blurred; (2) the generated scratch images have repetitive textures; (3) the brightness distribution of the generated scratch images is unnatural. Summary of the Invention

[0009] What is mainly solved is how to generate scratch images with rich texture details under a semantic image with a large area of the same color. The present invention uses a generative adversarial network to generate more realistic scratch image samples to expand the training samples, so as to help the scratch detection model to train.

[0010] Therefore, the technical solution adopted by the present invention is an industrial scratch generative adversarial network system for solving single semantic diversity mapping, including a generator and a multi-scale discriminator;

[0011] The generator is used to process the input data formed by splicing a learnable constant and a semantic map, and then output a forged scratch image;

[0012] The multi-scale discriminator is used to judge the authenticity of an image, and a feature pyramid module is introduced in the second downsampling and the third downsampling respectively.

[0013] Further, the splicing of the learnable constant and the semantic map specifically includes that the learnable constant consists of a constant parameter with the same dimension size as the semantic map and a bias parameter with the channel dimension size, and their initial values are both 1, which are updated by backpropagation. The dimension of the semantic map is N×C×H×W, and the channel dimension is C; when inputting, the bias parameter is added to the constant parameter in the channel direction and then spliced with the semantic map in the channel direction to finally form a combined input with a size of N×2C×H×W; where N represents the semantic image batch size, C represents the channel size, H represents the height of the image, and W represents the width of the image.

[0014] Further, the generator includes an input layer, a downsampling layer, a residual block, an upsampling layer and an output layer; among them, the input layer performs convolution operation and then instance normalization and ReLU activation. Then there are 4 times of downsampling. After downsampling, it passes through 8 residual blocks. Each residual block has two convolutions. After the first convolution, instance normalization and RuLU activation are performed, and only instance normalization is performed after the second convolution. Then there are 4 times of upsampling. The upsampling uses transposed convolution. After transposed convolution, instance normalization and RuLU activation are performed. Finally, after the output layer convolution, Tanh activation is performed, and finally a forged scratch image is output.

[0015] Further, in the upsampling of the generator, a Gaussian noise is added to its feature map every time after a transposed convolution. The Gaussian noise is multiplied by a learnable weight parameter and then added to the feature map after the transposed convolution to obtain a new feature map.

[0016] Further, in the above solution, the multi-scale discriminator includes 5 layers of downsampling. Feature pyramid modules are introduced in the second layer of downsampling and the third layer of downsampling respectively. In the feature pyramid module, the number of channels of the feature map is halved by the first convolution, and the second convolution outputs a Patch with 1 output channel, and the Patch is used for true / false judgment.

[0017] The present invention has the following beneficial technical effects:

[0018] (1) Aiming at the situation of repetitive textures in the generated scratch images, the present invention concatenates a learnable constant and the semantic image in the channel dimension and uses them together as the input of the generator, thus solving the problem that large-area single-color semantics will generate repetitive textures.

[0019] (2) Aiming at the problem of unnatural over-brightness in the generated scratch images, the present invention adds Gaussian noise after each deconvolution in the upsampling layer of the generator to make the over-brightness of the generated images more natural, and at the same time improves the clarity of the images.

[0020] (3) In order to improve the discriminative guidance ability of the discriminator for image details, the present invention introduces the idea of joint training and feature pyramid in the discriminator, thereby further improving the generation effect of image details. Description of the Drawings

[0021] Figure 1 is a true scratch sample image and the corresponding semantic label;

[0022] Figure 2 is a training flow chart;

[0023] Figure 3 is the generator architecture;

[0024] Figure 4 is the upsampling structure of the generator;

[0025] Figure 5 is the discriminator architecture;

[0026] Figure 6 is a result comparison diagram of the model of the present invention and three other models;

[0027] Figure 7 is the effect diagram of different treatments of the present invention. Detailed Embodiments

[0028] The following combines the drawings to detail the specific solution of the present invention.

[0029] Production of Sample Semantic Labels

[0030] As Figure 1As shown in the figure, in the present invention, the scratch samples can be divided into two types: bright and dark according to the different lighting intensities. For bright scratches, the pixel values of the background are encoded as 0, and the pixel values of the scratches are encoded as 1. For dark scratches, the pixel values of the background are encoded as 2, and the pixel values of the scratches are encoded as 3, thus obtaining the corresponding semantic label map. The sizes of the scratch map and the semantic label map are both 512*512.

[0031] Network architecture

[0032] 1. Generator network architecture

[0033] In the main architecture of the generator, as Figure 3 shown, the generator architecture of Pix2PixHD is mainly adopted. Among them, the convolution kernel of the input layer is 64, the size is 7×7, the stride is 1. After convolution operation, instance normalization is performed, and the ReLU function is used as the activation function. Then, there are 4 downsamplings. Each time the downsampling convolution kernel size is 3×3, the stride is 2, and the number of channels is doubled. After downsampling, it passes through 8 residual blocks. Each residual block has two convolutions. The convolution kernel size is 3×3, the stride is 1, and the number of channels is 1024. After the first convolution, instance normalization and RuLU activation are performed, and only instance normalization is performed after the second convolution. Next is four upsamplings. Among them, the convolution kernel size of the transposed convolution layer is 3×3, the stride is 2, and the number of channels is halved. After transposed convolution, instance normalization and RuLU activation are performed. In the convolution operation of the final output layer, the convolution kernel is 3, the size is 7×7, the stride is 1, and after convolution, Tanh activation is performed, and finally the forged scratch image is output.

[0034] In addition, as Figure 3 shown, the present invention splices the learnable constant and the semantic map in the channel direction and uses them as the generator input together. The learnable constant consists of a constant parameter with the same dimension size as the semantic map (i.e., the dimension is N×C×H×W) and a bias parameter with the channel dimension size (i.e., the dimension is C). Their initial values are both 1 and are updated by backpropagation. When inputting, the bias parameter is added to the constant parameter in the channel direction and then spliced with the semantic map in the channel direction to finally form a combined input with a size of N×2C×H×W. Where N represents the semantic image batch size, C represents the channel size, H represents the height of the image, and W represents the width of the image.

[0035] By splicing the learnable constant on the semantic map, the semantic information of the same color on the semantic map also has certain differences, thus solving the problem of repetitive textures in the generated scratch images. The generation effect is shown in Figure 7 the third column.

[0036] In the decoding stage, a Gaussian noise is added to its feature map every time after transposed convolution. Figure 4As shown, the Gaussian noise is multiplied by a learnable weight parameter and then added to the deconvolved feature map to obtain a new feature map. The width and height of the Gaussian noise are the same as those of the feature map, and the number of channels is 1, that is, the dimension is 1×1×H×W. The size of the learnable weight is 1×C×1×1. Through this process, the present invention makes the overexposed light of the finally generated scratch image more natural and the texture clearer. The generation effect is shown in Figure 7 the fourth column.

[0037] 2. Multi-scale discriminator network architecture

[0038] In the main architecture of the multi-scale discriminator, the present invention still uses the multi-scale discriminator architecture of Pix2PixHD. There are a total of 5 layers of downsampling. The convolution kernel of the input layer is 64, the size is 4×4, the stride is 2, the padding is 2, and then it is activated by LeakyReLU with a slope of 0.2. Then, two convolutions are performed with a convolution kernel size of 4×4, a stride of 2, a padding of 2, and the number of channels doubled. After each of the two convolutions, instance normalization and LeakyReLU activation are performed. In the 4th downsampling layer, the convolution kernel size is 4×4, the stride is 1, the padding is 2, and the number of channels is doubled. The convolution kernel of the last downsampling layer is 1, the size is 4×4, and the padding is 2. Finally, a Patch for judging true or false is output.

[0039] The difference is that in order to improve the ability of the multi-scale discriminator to guide details, the present invention introduces feature pyramid modules in the second downsampling layer and the third downsampling layer respectively, so that the gradient can be transmitted back to the bottom layer faster, so that the discriminator can better focus on the details of the image. The generation effect is shown in Figure 7 the fifth column. The reason for sharing the feature pyramid module by the two discriminators is that the present invention finds through experiments that if the two discriminators use their own feature pyramid modules, it will cause instability in training and easily cause the situation where the discriminator is stronger than the generator. As Figure 5 , in the feature pyramid module (FPN), the first convolution halves the number of channels of the feature map, and the second convolution outputs a Patch with 1 channel, which is used to judge true or false. The convolution kernels of both convolutions are 3×3, the stride is 1, and the padding is 1.

[0040] 3. Loss function

[0041] The present invention also uses the adversarial loss, feature matching loss, and perceptual loss of Pix2PixHD for the loss function without any changes. The loss function formula is as follows

[0042]

[0043] Among them, G is the generator, and D k is the kth discriminator, and LGAN For the adversarial loss, L FM For the feature matching loss, L VGG For the perceptual loss, α is the weight coefficient of the feature matching loss, and β is the weight coefficient of the perceptual loss.

[0044] The only difference is that in the adversarial loss, in addition to the Patch output by the discriminator, the present invention adds the Patch output by the feature pyramid module, that is, in L GAN (G, D k )

[0045] D(x) = D l (x) + D fpn (x)

[0046] where x is the input of the discriminator, D l is the output of the last layer of the discriminator, and D fpn is the output of the feature pyramid module in the discriminator. D is the output of the final discriminator.

[0047] 4. Training details

[0048] As Figure 2 shown, the combined label is input into the generator to obtain the forged scratch image, and then the forged scratch image or the real scratch image together with the semantic label is input into the discriminator to make a true / false judgment to train the network. After the training is completed, only the generator is used to generate the scratch image.

[0049] Most of the training parameters of Pix2PixHD are adopted in the present invention. The number of multi-scale discriminators is 2, the optimizer uses Adam, the exponential decay rate of the first moment estimation is 0.5, and the exponential decay rate of the second moment estimation is 0.999. The adversarial loss adopts LSGAN, the weight coefficient of the feature matching loss is 10, and the weight coefficient of the perceptual loss is 10.

[0050] The difference is that in the experiments used, the present invention does not perform any scaling and cropping operations on the input image, trains for a total of 400 rounds, the learning rate is 0.0002 in the first 200 rounds, and linearly decays to 0 in the last 200 rounds. The Batchsize is set to 1.

[0051] 5. Quantitative evaluation

[0052] The present invention uses FID and LPIPS as the quantitative evaluation criteria. A lower FID means that the generated samples and the real samples have a higher correlation and higher image quality. Similarly, a lower LPIPS also means that they are more similar to the real samples.

[0053] It can be seen from the following table that the model of the present invention has a lower FID, which indicates that the scratch images generated by the present invention are closer to the distribution of the real scratch images.

[0054] Comparative experiment LPIPS FID Pix2Pix 0.530162 198.8238784 Pix2PixHD 0.522455 174.9594929 SPADE 0.494397 243.3492998 ours 0.512416 155.7194544

[0055] As Figure 6 shown, the model proposed by the present invention can generate scratch images that are clearer, without repetitive textures, and have natural light transitions compared to the other three models.

Claims

1. An industrial scratch generation adversarial network system for solving single semantic diversity mapping, characterized in that: It includes a generator and a multi-scale discriminator; The generator is used to take the input data formed by splicing a learnable constant and a semantic map, and then output a forged scratch image; The splicing of the learnable constant and the semantic map specifically includes that the learnable constant consists of a constant parameter with the same dimension size as the semantic map and a bias parameter with the channel dimension size, and their initial values are both 1, which are updated by backpropagation. The dimension of the semantic map is N×C×H×W, and the channel dimension is C; when inputting, the bias parameter is added to the constant parameter in the channel direction and then spliced with the semantic map in the channel direction to finally form a combined input with a size of N×2C×H×W; where N represents the semantic image batch size, C represents the channel size, H represents the height of the image, and W represents the width of the image; In the upsampling of the generator, a Gaussian noise is added to its feature map every time after a transposed convolution. The Gaussian noise is multiplied by a learnable weight parameter and then added to the feature map after the transposed convolution to obtain a new feature map; The multi-scale discriminator is used to judge the authenticity of an image, and a feature pyramid module is introduced in the second downsampling and the third downsampling respectively; The multi-scale discriminator includes 5 layers of downsampling, and a feature pyramid module is introduced in the second downsampling and the third downsampling respectively. In the feature pyramid module, the first convolution halves the number of channels of the feature map, and the second convolution outputs a Patch with the number of output channels being 1, and the Patch is used to make a authenticity judgment; The Patch output by the feature pyramid module is reflected in the adversarial loss D(x) = D l (x) + D fpn (x) where x is the discriminator input, D l is the output of the last layer of the discriminator, D fpn is the output of the feature pyramid module in the discriminator, and D is the output of the final discriminator.

2. The industrial scratch generation adversarial network system for solving single semantic diversity mapping according to claim 1, characterized in that: The generator includes an input layer, a downsampling layer, a residual block, an upsampling layer and an output layer; among them, the input layer performs a convolution operation, then instance normalization and ReLU activation, and then 4 times of downsampling. After downsampling, it passes through 8 residual blocks. Each residual block has two convolutions. After the first convolution, instance normalization and RuLU activation are performed, and only instance normalization is performed after the second convolution. Then, 4 times of upsampling are performed. The upsampling uses transposed convolution. After the transposed convolution, instance normalization and RuLU activation are performed. Finally, after the output layer convolution, Tanh activation is performed, and finally a forged scratch image is output.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on deep adversarial learning

    CN113313180A

  • Image defogging method based on generative adversarial network fused with feature pyramid

    WO2021248938A1