Cross-view image translation method based on generative adversarial network

By building an image translation network, using the generative adversarial network and cascade residual refinement module, the problem of insufficient generation of diversity and details in cross-view image translation is solved, and high-quality and diverse ground-view image generation is achieved.

CN114092587BActive Publication Date: 2025-05-09FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111292146.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-05-09
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

The prior art cannot effectively generate diverse ground-view images in cross-view image translation tasks, and the generation results are prone to problems of blur and artifacts.

Method used

Using a cross-view image translation method based on a generative adversarial network, an image translation network is constructed including a cross-view image generator, a multi-scale discriminator, an encoder and a residual estimation network. Through a cascading residual refinement module and a multi-scale discriminator, diversified images are generated and refined.

Benefits of technology

It realizes the generation of diverse and detailed ground perspective images from a single aerial image, suppresses blur and artifacts of the generated images, and optimizes the quality of the generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114092587B_ABST
    Figure CN114092587B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-view image translation method based on a generative adversarial network. The method of the present invention comprises the following steps: 1) constructing an image translation network, the image translation network comprising a cross-view image generator and a residual-based cascade refinement module, the cross-view image generator uses U-net as a backbone network for synthesizing a rough ground panoramic image; the cascade refinement module is used to synthesize a refined and refined ground panoramic image; 2) simultaneously training the image translation network in two modes; the first mode is to use the real ground image as the encoder input to generate a Gaussian-distributed latent code-merged aerial view image and a panoramic semantic image input into the image translation network for training; the second mode is to use randomly sampled latent code-merged aerial view image and panoramic semantic image input; 3) testing phase. The method of the present invention is used to generate a variety of ground view images from a single aerial image, with more realistic details.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence and computer vision technology, and specifically relates to a cross-view image translation method based on a generative adversarial network. Background Art

[0002] In recent years, artificial intelligence and deep learning technologies have been widely used in the field of computer vision, and generative adversarial networks have been widely used in image generation tasks. Generally speaking, techniques based on generative adversarial networks mostly focus on tasks such as natural image generation, image super-resolution, image style transfer, and facial image conversion. Although these works have achieved impressive results, they have still failed to extend the image translation task to cross-view scenarios.

[0003] Thanks to the rapid development of generative adversarial networks, the image translation task based on generative adversarial networks has been widely studied. Cross-view image translation is the conversion of scene images from two different perspectives, aerial and ground, at the same location. It involves translation at both the semantic and appearance levels. This task has received increasing attention in recent years.

[0004] Although existing work has achieved some results in cross-view image translation tasks, there are still many problems. First, it can only generate a small number of ground images with a fixed perspective, and the generated ground-view images cannot retain the rich semantic information of aerial-view images. Secondly, existing work ignores the generation diversity of image translation and turns image translation into a deterministic one-to-one mapping task. However, when aerial-view images are converted to ground-view images, they may present different appearance styles due to factors such as occlusion, noise, and rotation. Therefore, generating diverse ground images from aerial-view images is closer to the real situation. Finally, the generation results of existing work are prone to artifacts and blurring because the description of the appearance of objects in ground-view images is more detailed than that in aerial-view images. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention aims to provide a cross-view image translation method based on a generative adversarial network. The method of the present invention is used to generate a variety of ground-view images from a single aerial image.

[0006] A cross-view image translation method based on a generative adversarial network comprises the following steps:

[0007] 1) Building an image translation network

[0008] The image translation network consists of a cross-view image generator G coarse , multi-scale discriminator D s , encoder E and residual estimation network; cross-view image generator G coarseThe U-net network is selected as the backbone network to receive the aerial perspective image I a , panoramic semantic image S pano As input, synthesize a rough ground panoramic image I′ pano ; The residual estimation network consists of a residual map generator and a residual map discriminator D r The residual map generator consists of multiple residual refinement modules R i Cascaded sequentially, the residual refinement module R i It consists of a multi-layer convolutional neural network to predict the residual map of the input image and pass the cross-view image generator G coarse Sum the residual image and further synthesize the refined ground panoramic image After the residual refinement module, the residual estimation network generates the final refined ground panoramic image I″ pano , residual graph discriminator D r Used to judge each residual refinement module R i Generate the authenticity of the residual map and feed the result back to the residual refinement module R i ; Multi-scale discriminator D s Determine the real ground image by using image patches of different scales I pano And the final refined ground panorama image I″ pano The authenticity of the image is verified and the result is fed back to the cross-view image generator G coarse and residual estimation network; the encoder E consists of several layers of residual-based convolutional neural networks, which encodes the input image into a latent code z close to a Gaussian distribution rec , to achieve diversified generation of networks;

[0009] 2) Training phase

[0010] The image translation network is trained in two modes at the same time; in the first training mode, the real ground image I pan The latent code z generated as the input of encoder E conforms to the Gaussian distribution e , Aerial perspective image I a , panoramic semantic image S pano The image translation network is trained by inputting it into the image translation network; in the second training mode, the randomly sampled latent code z is used r Merge aerial view images I pano , panoramic semantic image S pano Input image translation network for training, encoder E is connected to the end of image translation network, from the generated image I″ pano Generate reconstruction latent code z rec , using the loss function constraint E to map the encodings of the real ground image and the generated image to the same Gaussian distribution;

[0011] 3) Testing phase

[0012] The aerial perspective image to be tested, the panoramic semantic image and different latent codes randomly sampled from the Gaussian distribution are input into the trained image translation network to obtain the final refined bottom panorama and obtain diverse generation results.

[0013] In the present invention, in step 1), in the residual estimation network, the panoramic semantic image and the aerial perspective image are combined with the refined output of the previous residual refinement module as the input of the subsequent refinement; the generation of the panoramic image refined in the i-th stage is defined as:

[0014]

[0015] in and represents the refined image output of the (i-1)th and i-th residual refinement modules, R i represents the i-th level residual refinement module.

[0016] In the present invention, in step 2),

[0017] ① In the first training mode, the real ground image I pano And the final generated ground panoramic image I″ pano The reconstruction loss between The design is as follows:

[0018]

[0019] in represents the expectation operation, I a Represents an aerial perspective image;

[0020] Using KL divergence, through the following loss function L KL To optimize the encoder:

[0021]

[0022] in E() represents the encoder, represents Gaussian distribution;

[0023] ②In the second training mode, the latent code z is reconstructed rec Due to another reconstruction loss, it is forced to approach the random latent code z r , whose expression is:

[0024]

[0025] ③Use multi-scale discriminator D s The intermediate residual map predicted by the discriminating residual estimation network is used to generate the rough panoramic image I′ generated by the cross-view image generator.pano The adversarial loss formula is expressed as:

[0026]

[0027] Among them, D s represents the multi-scale discriminator;

[0028] After multiple levels of refinement, the network generates an intermediate image Including the final refined Figure I″ pano , the aerial view image and the generated intermediate image pair The adversarial loss formula is as follows:

[0029]

[0030] The total multi-level adversarial loss obtained from the above two equations is:

[0031]

[0032] where λ refine The weight factor to balance the contribution of coarse loss and fine loss, N is the total number of stages of cascade refinement; ④ The estimated residual map is used to fill the gap between the generated intermediate image and the true ground image. Therefore, given the generated intermediate image The true residual plot is defined as Then the residual reconstruction loss is:

[0033]

[0034] Where D r is the discriminator of the residual graph, For the residual map generated by the i-th level residual refinement module, the total residual reconstruction loss is the sum of all intermediate residual reconstruction losses:

[0035]

[0036] ⑤ Overall loss

[0037] The weighted sum of the loss function is taken as the overall loss of the entire image translation network, denoted as:

[0038]

[0039] where λ is 1 , 2 , 3 , 4 It is a hyperparameter that controls the proportion of each loss in the overall loss function. Compared with the prior art, the beneficial effects of the present invention are:

[0040] The present invention makes full use of the spatial information in aerial images, takes aerial images, covers the panoramic semantic map of all objects in the aerial images, introduces random noise as input, and encodes the noise into latent code as an additional input for diversified image generation. Based on the prediction of ground perspective images, the present invention proposes an image refinement strategy based on cascade residuals. Compared with the existing methods, it can suppress the blur and artifacts of the generated images, and optimize the generated images in a progressive manner to make them have richer details. The present invention compares the generation effect of the existing technology on the commonly used datasets CVUSA and Dayton for cross-perspective image translation tasks, and evaluates the generation effect with 5 evaluation indicators. The experimental results show the superiority of the present method in generation effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 Generate adversarial network architecture diagram for residual-based progressive refinement.

[0042] Figure 2 This is the structural diagram of the residual refinement module.

[0043] Figure 3 Generate intermediate results for the step-by-step refinement strategy.

[0044] Figure 4 Comparison of the generation results of the present invention and existing work on the CVUSA dataset. DETAILED DESCRIPTION

[0045] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings and embodiments.

[0046] The present invention proposes a residual-based progressive refinement generative adversarial network model for generating diverse ground-view images from a single aerial image. The network structure proposed in the present invention is as follows: Figure 1 The training process consists of two parts, the upper part represents image-level reconstruction, and the lower part represents latent-level reconstruction. Image-level reconstruction is similar to an autoencoder and can be seen as the reconstruction of a panoramic image. e The latent level reconstruction is performed by encoding the generated panorama into the latent code z r , forcing the generator to use randomly sampled latent codes z r In the test phase, only the cross-view image generator and the cascade refinement block are activated, that is, the method does not require ground images in the test phase. The first generation module uses G coarse Represents, generating a rough image I′ pano In the subsequent network, the i-th level residual refinement module is denoted as R i , where the refined image output is defined as The final output image is I″ panoThe following is a detailed description of each structure in the network.

[0047] 1. Diverse cross-view panoramic image generation

[0048] (1) Generate ground panoramic images from aerial perspective images

[0049] The present invention proposes to gradually refine and synthesize a ground panoramic image containing richer information from an aerial view image. First, U-net is selected as a cross-view image generator to synthesize a rough ground panoramic image. In addition to the aerial view image, the present invention also uses a panoramic semantic image as an additional input to alleviate the semantic ambiguity problem in the view conversion process. Specifically, Figure 1 As shown in the figure, the aerial view image is first resized and combined with the semantic image with a resolution of 256×1024 and the latent code z e Merge, then input into the cross-view image generator, the output is recorded as I p ' ano =G coarse (I a , S pano , z e ), where I′ pano Represents the rough ground panoramic image, G coarse represents the cross-view image generator, I a Represents the aerial perspective image, S pano Represents the semantic image, z e Represents latent code.

[0050] (2) Diverse panoramic image generation.

[0051] In order to achieve diverse panoramic ground image synthesis, the present invention uses an encoder to directly translate the real ground image into a latent code, and then merges it with other inputs of the network. pano =G CRP (I a , S pano , z e ), where G CRP represents the entire image translation network, including the cross-view image generator and the cascade refinement module. Figure 1 As shown in Figure 2, during the training phase, the entire network will be trained in two modes simultaneously. In the first mode, the encoder takes the real ground image as input to generate a latent code z that conforms to the Gaussian distribution. e . The real ground image and the final generated ground panoramic image I″ pano The reconstruction loss L between rec The design is as follows:

[0052]

[0053] in Represents the expectation operation.

[0054] Encourage potential code e The reason for being close to the Gaussian distribution is that, since the real ground image is unknown during the test phase, the latent code can be directly sampled from the standard Gaussian distribution. The present invention uses KL divergence to optimize the encoder through the following loss function:

[0055]

[0056] in E() represents the encoder, Represents a Gaussian distribution.

[0057] In the second mode, randomly sampled latent codes z are used r As input. The encoder E is cascaded at the end of the network, from the generated image I″ pano Generate latent code z rec . Reconstruct the latent code z rec Due to another reconstruction loss, it is forced to approach the random latent code z r , whose expression is:

[0058]

[0059] This can further force the network to r Instead of ignoring it, more random information is encoded in the network, so that the network performs one-to-many mapping and ensures the diversity of generation.

[0060] 2. Residual Estimation Network

[0061] (1) Residual Estimation Network

[0062] Cross-view image generator G coarse It is possible to generate a ground panoramic image with a high semantic level but a fuzzy appearance, and the generation needs to be further refined to improve the generation quality. The present invention regards the refinement process of the generated image as a generative adversarial learning process, in which the panoramic image synthesis based on the residual map is used as the generation process, and in the recognition process, it is identified whether its quality reaches the standard of the real ground image. Therefore, the present invention designs a residual estimation network to predict the residual map, and further fuses the residual map with the rough panoramic image by summing.

[0063] (2) Cascaded progressive refinement

[0064] The network based on residual estimation can synthesize ground panoramic images with less blur and artifacts in most cases. Experiments have found that the network tends to produce a lot of blur in specific semantic areas in the early stages of the network. And there are differences in the quality of the images estimated by the residual estimation network. In order to further improve the network generation capability, the present invention solves the cross-view image translation task in a progressive refinement manner, and cascades multiple residual refinement modules in sequence to make full use of the residual estimation capability. As the number of residual refinement modules increases, better quality panoramas can be synthesized and smaller residuals can be predicted. Figure 3 As shown in Figure 1, the residual map gradually focuses on small details and tiny artifacts of the image. The gradual refinement is achieved by cascading multiple residual refinement modules. In order to ensure that the refined image always has the same semantic structure as the previous coarse image, the original semantic map and the aerial view image are combined with the refined output of the previous residual refinement module as the input of the subsequent refinement. The generation of the panoramic image at the i-th stage of refinement is defined as:

[0065]

[0066] in and represents the refined image output of the (i-1)th and i-th residual refinement modules, R i represents the i-th level residual refinement module.

[0067] 3. Loss Function

[0068] Since the model proposed in the present invention generates a new refined panoramic image and a residual map after the residual refinement module in each stage, the present invention adopts a multi-stage adversarial learning objective for the coarse image and the refined image, and uses a discriminator to identify all the intermediate residual maps.

[0069] 1. Multi-stage adversarial loss

[0070] Multi-scale discriminators are more helpful in synthesizing realistic images than single-scale discriminators. Therefore, the adversarial loss formula of the rough panoramic image generated by the cross-view image generator is expressed as:

[0071]

[0072] Among them, D s represents the multi-scale discriminator.

[0073] After multiple levels of refinement, the network generates intermediate refined images Including the final refined Figure I″ pano The aerial view image and the intermediate generated image pair The adversarial loss formula is as follows:

[0074]

[0075] The total multi-level adversarial loss obtained from the above two equations is:

[0076]

[0077] where λ refine is the weight factor to balance the contribution of coarse loss and fine loss, and N is the total number of stages of cascade refinement.

[0078] 2. Residual reconstruction loss

[0079] The estimated residual map is expected to fill the gap between the intermediate generated image and the true ground truth image. The present invention defines the true residual graph as Then the residual reconstruction loss is:

[0080]

[0081] Where D r is the discriminator of the residual graph, The residual map generated for the i-th level residual refinement module. The total residual reconstruction loss is the sum of all intermediate residual reconstruction losses:

[0082]

[0083] 3. Overall loss

[0084] Finally, the weighted sum of the above loss functions is taken as the overall loss of the entire network, recorded as:

[0085]

[0086] where λ is 1 , 2 , 3 , 4 is a hyperparameter that controls the proportion of each loss in the overall loss function. The training process is based on the min-max game theory. The generator network, including the cross-view image generator and the residual estimation network, is optimized to generate realistic ground-truth panoramic images to deceive the discriminator. The discriminator distinguishes the real ground-truth images from the generated images as much as possible. The encoder maps the encodings of the real ground-truth images and the generated images to nearly the same Gaussian distribution.

[0087] Example 1

[0088] (1) Selection of training data

[0089] The model proposed in this paper is trained and tested on two public datasets, Dayton and CVUSA. The Dayton dataset contains a total of 76,048 sets of open-ground perspective image pairs, of which 55,000 sets are used as training sets and the remaining 21,048 sets are used as test sets. The CVUSA dataset uses 35,531 sets of open-ground perspective image pairs as training sets and 8,884 sets as test sets.

[0090] (2) Network structure

[0091] The cross-view image generator uses U-net as the backbone network. The encoder and decoder of U-net contain 8 downsampling layers and 8 upsampling layers respectively, where the first convolutional layer has a depth of 64 channels and the bottleneck layer of the generator has a depth of 512 channels. When there is a similar spatial correspondence between the input and output, the architecture shows a strong ability to generate images with rich texture and context information. In the residual-based refinement block, a simple convolutional neural network is used as the image refinement generator. The network uses three generative adversarial network discriminators of different scales, which can judge whether overlapping image blocks of sizes 70×70, 140×140 and 280×280 are real ground truth images or fake images. The encoder consists of several layers of residual-based convolutional neural networks. The residual estimation network consists of four convolutional layers with a convolution kernel size of 3×3.

[0092] (3) Hyperparameter setting

[0093] The four hyperparameters of the loss function are set to λ 1 =100,λ 2 =1,λ 3 =1,λ 4 =1e-3.

[0094] Adam is used as the training optimizer, where the momentum term is set to β 1 =0.5,β 1 = 0.999, the learning rate is set to 0.0002, and the dimension of the latent code is set to 8.

[0095] (4) Training details

[0096] The latent code is first expanded to the same shape as the input image and then added to each intermediate layer of the generator and refinement block. The present invention adopts a three-level residual refinement module, trains the first-level residual refinement module, and uses its parameters to initialize the second-level residual refinement module. Then the parameters of the residual refinement module are updated, and the parameters of the third-level residual refinement module are initialized. When training a new residual refinement module, all previous residual refinement modules are fixed.

[0097] (5) Experimental results

[0098] Tables 1 and 2 are quantitative experimental results comparisons of the present invention and other existing works on the cross-view image translation task on the two public datasets of Dayton and CVUSA. The evaluation indicators are as follows:

[0099] 1. FID (Frechet Inception Distance): The Wasserstein-2 distance between the generated image and the real ground image. A lower FID means that the generated image is closer to the real ground image.

[0100] 2.IS (Inception Score): A widely used method to measure the quality of generated images. It can reveal the correlation between image quality and diversity. A higher IS means that the generated images are more realistic.

[0101] 3.PSNR: Peak signal-to-noise ratio between the generated image and the ground truth image. A higher PSNR generally means better generation quality.

[0102] 4.SSIM: Calculates the similarity between the generated image and the ground truth image based on brightness, contrast, and structure. A higher SSIM usually means a high similarity between the two images.

[0103] 5. SD: Sharpness loss between generated images and ground truth.

[0104] Table 1 Comparison of experimental results of the present invention and the prior art in performing cross-view image translation tasks on the CVUSA dataset

[0105] method FID IS PSNR SSIM SD X-seq 44.16 2.7028 19.3512 0.4306 17.6350 SelectionGAN 37.68 2.8071 20.9938 0.4598 18.2075 BicycleGAN 40.02 2.8445 20.8106 0.4770 18.1401 The present invention 35.02 2.8907 21.2476 0.4879 18.5000

[0106] Table 2 Comparison of experimental results of the present invention and the prior art in performing cross-view image translation tasks on the Dayton dataset

[0107] Method FID IS PSNR SSIM SD X-seq 45.12 2.5101 20.0058 0.5487 18.7254 SelectionGAN 42.23 2.6210 22.2191 0.5884 19.9068 BicycleGAN 43.16 2.5627 21.7512 0.5304 18.1124 The present invention 40.32 2.6443 22.5510 0.5626 19.6617

[0108] Figure 4 The qualitative experimental results of the present invention and other existing works on the cross-view image translation task on the CVUSA dataset are compared. The experimental results prove the superiority of the present invention over other existing methods in the cross-view image translation task.

Claims

1. A cross-view image translation method based on generative adversarial network, characterized in that: The following steps are involved: 1) Building an image translation network The image translation network consists of a cross-view image generator G coarse , multi-scale discriminator D s , encoder E and residual estimation network; Cross-view image generator G coarse The U-net network is selected as the backbone network to receive the aerial perspective image I a , panoramic semantic image S pano As input, synthesize a rough ground panoramic image I p ′ ano ; The residual estimation network consists of a residual map generator and a residual map discriminator D r The residual map generator consists of multiple residual refinement modules R i Cascaded sequentially, the residual refinement module R i It consists of a multi-layer convolutional neural network to predict the residual map of the input image and pass the cross-view image generator G coarse Sum the residual image and further synthesize the refined ground panoramic image After the residual refinement module, the residual estimation network generates the final refined ground panoramic image I p ″ ano , residual graph discriminator D r Used to judge each residual refinement module R i Generate the authenticity of the residual map and feed the result back to the residual refinement module R i ; Multi-scale discriminator D s Determine the real ground image by using image patches of different scales I pano And the final refined ground panorama image I p ″ ano The authenticity of the image is verified and the result is fed back to the cross-view image generator G coarse and residual estimation network; The encoder E consists of several layers of residual-based convolutional neural networks, which encodes the input image into a latent code z close to a Gaussian distribution. rec , to achieve diversified generation of networks; 2) Training phase The image translation network is trained in two modes simultaneously; In the first training mode, the real ground image I pano The latent code z generated as the input of encoder E conforms to the Gaussian distribution e , Aerial perspective image I a , panoramic semantic image S pano The image translation network is trained by inputting it into the image translation network; in the second training mode, the randomly sampled latent code z is used r Merge aerial view images I pano , panoramic semantic image S pano Input image translation network for training, encoder E is connected to the end of the image translation network, from the generated image I p ″ ano Generate reconstruction latent code z rec , using the loss function to constrain the encoder E to map the encodings of the real ground image and the generated image to the same Gaussian distribution; 3) Testing phase The aerial perspective image to be tested, the panoramic semantic image and different latent codes randomly sampled from the Gaussian distribution are input into the trained image translation network to obtain the final refined ground panorama and obtain diversified generation results; Wherein: in step 2), ① In the first training mode, the real ground image I pano And the final generated ground panoramic image I p ″ ano The reconstruction loss between The design is as follows: in represents the expectation operation, I a Represents an aerial perspective image; Using KL divergence, through the following loss function L KL To optimize the encoder: in E() represents the encoder, represents Gaussian distribution; ②In the second training mode, the latent code z is reconstructed rec Due to another reconstruction loss, it is forced to approach the random latent code z r , whose expression is: ③Use multi-scale discriminator D s The intermediate residual map predicted by the discriminating residual estimation network is used to generate the rough panoramic image I generated by the cross-view image generator. p ′ ano The adversarial loss formula is expressed as: Among them, D s represents the multi-scale discriminator; After multiple levels of refinement, the network generates an intermediate image Including the final refined Figure I″ pano , aerial view image and generated intermediate image pair The adversarial loss formula is as follows: The total multi-level adversarial loss obtained from the above two equations is: where λ refine is the weight factor to balance the contribution of coarse loss and fine loss, N is the total number of stages of cascade refinement; ④The estimated residual map is used to fill the gap between the generated intermediate image and the true ground image. Therefore, given the generated intermediate image The true residual plot is defined as Then the residual reconstruction loss is: Where D r is the discriminator of the residual graph, For the residual map generated by the i-th level residual refinement module, the total residual reconstruction loss is the sum of all intermediate residual reconstruction losses: ⑤ Overall loss The weighted sum of the loss function is taken as the overall loss of the entire image translation network, denoted as: Among them, λ1, λ2, λ3, and λ4 are hyperparameters that control the proportion of each loss in the overall loss function.

2. The cross-view image translation method according to claim 1, characterized in that: In the residual estimation network, the panoramic semantic image and the aerial view image are combined with the refined output of the previous residual refinement module as the input for subsequent refinement; The generation of the i-th stage refined panoramic image is defined as: in and represents the refined image output of the (i-1)th and i-th blocks, R i represents the i-th level residual refinement module.

Citation Information

Patent Citations

  • Cross-view-angle image optimization method and device, computer equipment and readable storage medium

    CN112861747A

  • Digital Image Layout Training using Wireframe Rendering within a Generative Adversarial Network (GAN) System

    US20200151508A1