A Semantic Image Synthesis Method Based on Location-Aware Adversarial Generative Network

By introducing position-aware adversarial generation network and conditional group normalized blocks in semantic image synthesis technology, the image blur and artifact problems in the prior art are solved, more realistic image generation is achieved, and the dependence on batch size is reduced, which improves the stability of model training.

CN115496821BActive Publication Date: 2025-06-27DALIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211135103.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-19
Publication Date
2025-06-27
Estimated Expiration
2042-09-19

AI Technical Summary

Technical Problem

The existing semantic image synthesis technology has blur and artifact problems when generating images, and ignores position perception information, resulting in the generated images being not realistic enough; at the same time, the batch normalization method is constrained by batch size, which affects the model training effect.

Method used

A semantic image generation method based on position-aware adversarial generation network is proposed. By constructing a LA-GAN network model, the position-aware condition group normalization block LACGN is adopted, and the group normalization is combined with the modulation parameters in the semantic segmentation mask is carried out to deepen the fusion of semantic segmentation mask and image generation, and reduce the dependence on batch size.

Benefits of technology

The generated images are clearer and more realistic, reducing artifacts and spatial distortions. Model training no longer depends on batch size, improving training stability and generation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496821B_ABST
    Figure CN115496821B_ABST
Patent Text Reader

Abstract

The present invention proposes a semantic image synthesis method based on a position-aware adversarial generation network, belonging to the technical field of semantic image generation network model algorithms. The present invention inputs a 256-dimensional noise vector z sampled from a normal distribution and a semantic segmentation mask m into the constructed LA-GAN network. After semantic fusion and upsampling through multiple LACGNResBlks, and finally after being processed by the tanh function, the generated realistic image I is obtained. The LA-GAN network model of the present invention deepens the fusion with the semantic segmentation mask during the image generation process, ensuring the authenticity of image generation; and combines group normalization to reduce the dependence of model training on the batch size.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of semantic image generation network model algorithms, and particularly relates to a semantic image generation method based on a generative adversarial network, providing important theoretical research and technical support for computer vision tasks such as image generation. Background Art

[0002] Taking pictures is the main way to obtain real images of objects in daily life. The obtained images are a true reflection of the objective world. However, this technology cannot create non-existent objects out of thin air. Due to the progress of deep learning, image synthesis technology has made great progress in recent years. Now we can use artificial synthesis technology to create realistic virtual images. It not only saves time and costs, but also greatly enriches the image content. Therefore, image synthesis has become an important technology in computer vision.

[0003] Now image synthesis technology has been widely applied in fields such as data augmentation, artistic creation, and criminal profiling. Image synthesis technology can be divided into random image synthesis and conditional image synthesis; the former synthesizes uncertain images through random variables, and the latter method guides image synthesis through external conditions. The conditions for guiding image synthesis can have various forms, including text descriptions, human poses, class labels, and semantic segmentation masks, etc.

[0004] Among them, the image synthesis technology based on semantic segmentation masks is called semantic image synthesis. Due to its controllability advantages, it has attracted wide attention. Although a large amount of research has been done in this regard by current methods, there are still blurs and artifacts in the synthesized images. And these methods ignore the position perception information during the image synthesis process. Specifically, similar modulation information should be provided to similar regions, and more modulation information should be provided to key regions. In addition, the batch normalization methods of most methods are restricted by the batch size; the smaller the batch, the worse the batch normalization effect; the larger the batch, the higher the hardware requirements. Summary of the Invention

[0005] To solve the above technical problems, the present invention proposes a location-aware generative adversarial network model LA-GAN for semantic image synthesis. Its core is the location-aware conditional group normalization block LACGN, which predicts location-aware information based on the currently generated image features and uses the modulation parameters learned from the semantic segmentation mask as the conditional information for group normalization to achieve conditional group normalization. The LA-GAN network model in the present invention deepens the fusion with the semantic segmentation mask during the image generation process, ensuring the authenticity of image generation; and combining group normalization reduces the dependence of model training on the batch size.

[0006] The technical solution of the present invention is as follows:

[0007] A semantic image synthesis method based on a position-aware adversarial generation network, comprising the following steps:

[0008] I. Construction of the LA-GAN network model

[0009] S1: Take the semantic segmentation mask m and the 256-dimensional noise vector z sampled from the normal distribution ∈ R 256 , i as inputs; First, project the noise vector z into the visual domain through a fully connected (FC) layer, and then deform it so that the noise vector becomes a feature map f0 of 1024×4×4, and use f0 as the starting input of the first-layer LACGN ResBlks module and pass it layer by layer; Second, for the LACGN ResBlks modules with different depths, downsample the semantic segmentation mask m at different resolutions and use it as the input of each LACGN ResBlks module together with the feature map.

[0010] S2: Integrate the position-aware prediction module LAPM, the guidance sampling module GSM, and the group normalization module GN to form a conditional group normalization module LACGN Block with position awareness; The specific implementation process of the LACGN Block includes:

[0011] S2.1: Accept the feature map f i transmitted and input by the i-th layer LACGN ResBlks module in step S1, calculate the correlation between each pixel in the position-aware prediction module LAPM, specifically including performing average pooling and max pooling operations on it in the spatial dimension to generate two new feature maps; Then connect the two new feature maps, perform a convolution operation, and apply the s(·) activation function to calculate the spatial awareness mapping p i . This process is represented by the following formula:

[0012] p i = s(Conv(Concat(AvePool(f i ), MaxPool(f i ))))

[0013] where, f i represents the input feature map; AvePool() and MaxPool() respectively represent average pooling and max pooling in the channel dimension; Concat() represents channel connection; Conv() represents convolution; s() represents activation by the Sigmoid function.

[0014] S2.2: Input the semantic segmentation mask m accepted in step S1 into the guidance sampling module GSM to calculate the normalized activation modulation parameters γ i and βi It is expressed as:

[0015] γ i = S γ (m)

[0016] β i = S β (m)

[0017] S2.3: Input the feature map f i into the group normalization module GN for group normalization to obtain Multiply the spatial perception mapping map p i separately with the modulation parameters γ i and β i to obtain two new modulation parameters as conditional information; then perform the following operation with to obtain a new feature map f i+1 .

[0018]

[0019] S3: Combine two LACGN Blocks with two Relu activation function layers, two convolutional layers, one skip connection, and one upsampling layer to form LACGN ResBlks; Connect six LACGN ResBlks in series and then add one convolutional layer and one tanh activation layer to finally form the LA-GAN network.

[0020] II. Training of the LA-GAN network model

[0021] S4: The noise vector z and the semantic segmentation mask m are semantically fused in the LA-GAN network to obtain the generated image I. Input the generated image I and the real images in the dataset into the discriminator for training, and optimize and train the LA-GAN network model according to the GAN loss, feature matching loss, and perceptual loss, and save the optimal model mode_best.

[0022] S5: Load the model model_best obtained in step S4 and input noise and the semantic segmentation mask to generate a realistic image.

[0023] Advantages of the present invention: The method of the present invention can generate clearer and more realistic images compared with the existing methods, and at the same time eliminates the dependence on the batch size during training. This is very meaningful for the application of image generation algorithms in practice. Description of the drawings

[0024] Figure 1 It is the framework diagram of the LA-GAN network model of the present invention.

[0025] Figure 2 This is the LACGN Block module diagram used in the present invention.

[0026] Figure 3 This is a comparison diagram of the images generated by LA-GAN and the input semantic map input, the original image Groud True, the SPADE model, and the SC-GAN model on the Cityscapes dataset.

[0027] Figure 4 This is a comparison diagram of the images generated by LA-GAN and the input semantic map input, the original image Groud True, the SPADE model, and the SC-GAN model on the AED20k dataset.

[0028] Figure 5 This is a comparison diagram of the images generated by LA-GAN and the input semantic map input, the original image Groud True, the SPADE model, and the SC-GAN model on the AED20k-outdoor dataset.

[0029] Figure 6 This is a line chart of the FID metrics of the images generated by LA-GAN using BN and GN to test at different batch sizes. Detailed implementation manners

[0030] To make the problems solved by the model of the present invention, the adopted methods, and the achieved effects clearer, the following further detailed description of the present invention is made in combination with the drawings and experiments.

[0031] As Figure 1 This is the framework diagram of the LA-GAN network model constructed by the present invention; the backbone part of this network model consists of 6 LACGN ResBlks with nearest neighbor upsampling, which are used to fuse the semantic segmentation mask and improve the resolution. Using the semantic segmentation mask m and the 256-dimensional noise vector z sampled from the normal distribution ∈ R 256 as the input, and output a realistic image I with a specific resolution. Each LACGN ResBlks consists of two LACGN Blocks, two convolutional layers, two Relu activation function layers, a skip connection, and an upsampling layer. The upsampling layer doubles the width and height of the image feature map through bilinear interpolation operation. Since each LACGN ResBlks works at different scales, the present invention downsamples the semantic segmentation mask to match the spatial resolution. During the data processing, the noise vector z is projected into the visual domain through a fully connected (FC) layer, and then it is deformed to obtain a feature map f0 with a size of 1024×4×4; the first LACGN ResBlks module uses the feature map f0 as the input, and each subsequent LACGN ResBlks module uses the feature map output by the previous LACGN ResBlks module As the input, output a feature map further fused with spatial information Among them, C i , w i 、h i are the number of channels, width, and height of the image feature map generated by the i-th LACGN ResBlks block. After 6 times of upsampling by LACGN ResBlks, the resolution of the image feature map changes from 4×4 to 256×256. After being processed by multiple LACGN ResBlks, the final feature map is obtained, and after being processed by the tanh function, the synthetic realistic image I is obtained.

[0032] Figure 2 This is the LACGN Block module diagram of the present invention. For the LACGN Block, its specific implementation process is as follows:

[0033] (1) Perform spatial perception prediction on the input feature map

[0034] The structure of the proposed location awareness prediction module LAPM is as shown in the dotted box in the upper left corner. Specifically, for the input feature map f Figure 2 , first perform average pooling and max pooling operations on it in the spatial dimension to generate two new feature maps; then connect the two new feature maps, perform convolution operations, and apply the Sigmoid activation function to calculate the spatial perception map p i , which can be represented by the following formula: i p

[0035] p i = s(Conv(Concat(AvePool(f i ), MaxPool(f i ))))

[0036] Among them, f i represents the input feature map; AvePool() and MaxPool() respectively represent average pooling and max pooling in the channel dimension; Concat() represents channel connection; Conv() represents convolution; s() represents the Sigmoid activation function.

[0037] (2) Conduct guided sampling on the input semantic segmentation mask

[0038] The structure of the guided sampling module GSM is as shown in the solid box in the upper right corner. Specifically, input the input semantic segmentation mask m into the modulation network to generate the dense modulation parameter tensors γ Figure 2 and β i and β i , denoted as:

[0039] γ i = S γ (m)

[0040] β i = S β (m)

[0041] (3) Location perception condition group normalization

[0042] Normalize the input feature map f i to obtain Then, fuse it with the modulation parameters γ i and β i as well as the location perception information p i to obtain a new feature map f according to the following formula i+1 .

[0043]

[0044] The beneficial effects of the present invention can be further illustrated by the following experiments.

[0045] The experimental environment is Windows 10 system, the programming language is Python, and the hardware configuration is Intel(R) Core(TM) i7-8700 with a main frequency of 3.20GHz CPU, 16.00GB of memory, and 1 NVIDIA V100 graphics card. The datasets used are Citysacpes, ADE20k, and ADE20k-outdoor datasets.

[0046] The specific implementation steps are as follows

[0047] Step 1: Use a 256-dimensional noise vector z sampled from a normal distribution, z ∈ R 256 as the input

[0048] Step 2: At the same time, downsample the semantic segmentation mask m in the dataset as the network deepens

[0049] Step 3: Send the input into the LA-GAN network model for processing to obtain the corresponding synthesized image

[0050] Step 4: Perform semantic segmentation on the obtained synthesized image to obtain a new semantic segmentation map, and calculate the mean intersection over union mIoU and pixel accuracy Acc with the original semantic segmentation map. At the same time, calculate FID by comparing the synthesized image and the original image

[0051] According to the above steps, on the Cityscapes, ADE20k, and ADE20k-outdoor datasets, the LA-GAN network model described in the present invention is compared with multiple models such as the CRN model, SIMS model, and Pix2pixHD model. The results are shown in Table 1. It can be seen from Table 1 that the method proposed in the present invention is significantly better than other methods in terms of mIoU, Acc, and FID.

[0052] Table 1 Comparison of semantic image synthesis results of different models

[0053]

[0054] Meanwhile, Figures 3 - 5 The effect diagram of the synthesized image of the present invention is shown. It can be observed that the image synthesized by the present invention is clearer and more realistic, has fewer artifacts, and also reduces the situation of spatial distortion and deformation.

[0055] From Figure 6 it can be found that the FID value of BN changes significantly with the change of the batch size, while GN remains within a stable range. This shows that the present invention eliminates the dependence on the batch size and makes its training more stable.

[0056] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A semantic image synthesis method based on a location-aware adversarial generation network, characterized in that The method includes the following steps: (1) Construction of the LA-GAN network model S1: Take the semantic segmentation mask m and the 256-dimensional noise vector z sampled from a normal distribution, z ∈ R 256 as inputs. First, project the noise vector z into the visual domain through a fully connected layer, and then deform it so that the noise vector becomes a feature map f0 of 1024×4×4, and use f0 as the starting input of the first-layer LACGN ResBlks module and pass it layer by layer. Second, for LACGN ResBlks modules with different depths, downsample the semantic segmentation mask m at different resolutions and use it as the input of each LACGN ResBlks module together with the feature map; S2: Integrate the Location-Aware Prediction Module (LAPM), the Guidance Sampling Module (GSM), and the Group Normalization Module (GN) to form a Location-Aware Conditional Group Normalization Block (LACGN Block). Among them, the Location-Aware Prediction Module (LAPM) is used to perform spatial awareness prediction on the input feature map to obtain a spatial awareness mapping diagram p i ; the Guidance Sampling Module (GSM) is used to perform guidance sampling on the input semantic segmentation mask to obtain modulation parameters γ i and β i ; the Group Normalization Module (GN) is used to perform group normalization on the input feature map f i to obtain Multiply the spatial awareness mapping diagram p i separately with the modulation parameters γ i and β i to obtain two new modulation parameters as conditional information; then perform an operation with to obtain a new feature map f i+1 ; S3: Combine two LACGN Blocks with two Relu activation function layers, two convolutional layers, one skip connection, and one upsampling layer to form LACGN ResBlks; Connect 6 LACGN ResBlks in series and then add one convolutional layer and one tanh activation layer to finally form the LA-GAN network; (2) Training of the LA-GAN network model S4: Perform semantic fusion of the noise vector z and the semantic segmentation mask m in the LA-GAN network to obtain the generated image I; Input the generated image I and the real images in the dataset into the discriminator for training, and optimize and train the LA-GAN network model according to the GAN loss, feature matching loss, and perceptual loss, and save the optimal model mode_best; S5: Load the model model_best obtained in step S4, and input noise and the semantic segmentation mask to generate a realistic image.

2. The method according to claim 1, characterized in that, In the said step S2, the specific implementation process of the LACGN Block includes: S2.1: Receive the feature map f input by the i-th layer of LACGN ResBlks module in step S1 i , calculate the correlation between each pixel in the location-aware prediction module LAPM, specifically including performing average pooling and max pooling operations on it in the spatial dimension to generate two new feature maps; then connect the two new feature maps, perform a convolution operation, and apply the s(·) activation function to calculate the spatial awareness mapping diagram p i ; This process is represented by the following formula: p i = s(Conv(Concat(AvePool(f i ),MaxPool(f i )))) Among them, f i represents the input feature map; AvePool() and MaxPool() respectively represent average pooling and max pooling in the channel dimension; Concat() represents channel concatenation; Conv() represents convolution; s() represents activation by the Sigmoid function; S2.2: Input the semantic segmentation mask m received in step S1 into the guidance sampling module GSM to calculate the normalized activation modulation parameters γ i and β i ; which is expressed as: γ i = S γ (m) β i = S β (m) S2.3: Input the feature map f i into the group normalization module GN for group normalization to obtain Multiply the spatial perception mapping map p i with the modulation parameters γ i and β i respectively to obtain two new modulation parameters as conditional information; then perform the following operation with to obtain the new feature map f i+1 :