Multi-modal Cross-view Image Generation Method Based on Residual-based Cascade Progressive Optimization

Through the residual-based cascaded progressive optimization method, the variational autoencoder and adversarial generation network are used, and the multi-level residual optimization network is combined with the multi-level residual optimization network, the problem of single generation mode and insufficient quality in cross-view image generation is solved, and multi-modal image generation and quality improvement are achieved.

CN116051360BActive Publication Date: 2025-07-08FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111261792.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-28
Publication Date
2025-07-08
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

The existing cross-view image generation algorithm has poor effect when generating large-view image span images, the generation mode is single, making it difficult to simulate variable outdoor scene styles, such as weather and light changes, and the degree of image quality optimization is limited.

Method used

Using a cascading progressive optimization method based on residuals, the images are mapped to Gaussian distribution hidden encoding through a variational autoencoder, combined with a U-shaped adversarial generation network and a multi-level residual optimization network, the residual convolution neural network and reconstruction loss function are used to improve image quality, and the overall loss function is constructed for training.

Benefits of technology

Multimodal cross-view image generation is realized, which can simulate target viewing images under different lighting and weather conditions, significantly improve image generation quality, reduce distortion, and explain the quality improvement process through residual graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116051360B_ABST
    Figure CN116051360B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-modal cross-view image generation method based on residual-based cascaded progressive optimization for performing view transformation on source view images, including the following steps: Step 1, obtaining the real target view image of the source view image, and constructing a variational autoencoder to extract the first hidden encoding of the real target view image; Step 2, using a generative adversarial network to generate a rough target view image; Step 3, constructing a multi-level cascaded residual optimization network to optimize the rough target view image to obtain a fine target view image; Step 4, extracting the second hidden encoding of the fine target view image through the variational autoencoder and calculating the reconstruction loss with the first hidden encoding; Step 5, constructing an overall loss function; Step 6, after training the generative adversarial network, for the source view image to be subjected to view transformation, the generative adversarial network randomly samples the second hidden encoding to generate a multi-modal rough target view image, and performs image quality optimization through the multi-level cascaded residual optimization network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer image generation, and particularly relates to a multi-modal cross-view image generation method based on residual-based cascaded progressive optimization. Background Art

[0002] Cross-view image generation is a task of predicting the image result of the current scene observed from another view. As an important algorithm in computer vision, it has a wide range of application spaces in many fields such as drone detection and terrain estimation. With the progress of technologies such as drones and remote sensing satellites, some paired image datasets with large view spans in outdoor scenes have emerged. How to achieve the task of predicting one view from another through algorithm design has become the main problem at present. In recent years, the emergence and technological progress of generative adversarial networks have made it possible for machines to generate images. Therefore, how to use generative adversarial networks to achieve cross-view image generation has received more and more attention.

[0003] In the cross-view image generation task, due to problems such as occlusion and different field of view ranges between different views, it is difficult for even humans to speculate on what new objects may appear in another view. The literature (T. Zhou, S. Tulsiani, W. Sun, J. Malik, and A. A. Efros, “View synthesis by appearance flow,” in ECCV, 2016, pp. 286–301.) adopted a method combining optical flow and adversarial training to speculate on the view after a small-angle transformation of a simple scene or a single object. However, for cross-view image generation algorithms with large view spans (such as from the perspective of a remote control satellite to the ground perspective), there are still problems of poor generation effect and single generation mode. The literature (Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al., “Spatial transformer networks,” in NIPS, 2015, pp. 2017–2025.) proposed a method based on learnable affine transformation to achieve affine transformation of views. However, when the view changes greatly, it is difficult for this method to fit the depth-of-field change in the view, and it is even more difficult to generate new objects or new regions that were previously occluded.

[0004] The literature (K. Regmi and A. Borji, “Cross-view image synthesis using conditional gans,” in CVPR, 2018, pp. 3501–3510.) proposed two cross-view generation algorithms for remote sensing-ground perspectives, mainly by cascading or paralleling a semantic estimation network to constrain the semantic distribution of the generated images. However, there is still a large gap between the generated effect and the real distribution in terms of semantic distribution, which leads to a reduction in the overall quality of the generated images and a single generation mode.

[0005] The literature (Tang, D. Xu, N. Sebe, Y. Wang, J. J. Corso, and Y. Yan, “Multi-channel attention selection GAN with cascaded semantic guidance for cross-view image translation,” in CVPR, 2019, pp. 2417–2426.) proposed a semantic-guided cross-view image generation model, which improves the quality of the generated images by introducing a semantic segmentation map as a guiding condition and adopting a coarse-to-fine generation strategy with multi-channel attention selection. However, this method still does not consider the problem of single generation mode, is difficult to simulate the styles of variable outdoor scenes (such as weather, lighting and other variable factors), and the optimization degree of its image quality is limited. Summary of the Invention

[0006] The present invention is made to solve the above problems, and aims to provide a multi-modal cross-view image generation method based on residual cascaded progressive optimization.

[0007] The present invention provides a multi-modal cross-view image generation method based on residual cascaded progressive optimization, which is used to perform perspective transformation on source perspective images to obtain multi-modal target perspective images, and has the following characteristics, including the following steps:

[0008] Step 1, obtain the real target perspective image of the source perspective image, construct a variational autoencoder based on KL-divergence constraint, and map the real target perspective image to a low-dimensional vector through the variational autoencoder to obtain the first latent code that conforms to the Gaussian distribution;

[0009] Step 2, use an adversarial generation network based on the U-shaped network to generate a rough target perspective image according to the source perspective image, the target perspective semantic segmentation map, and the first latent code;

[0010] Step 3: Construct multiple residual optimization networks, and cascade the multiple residual optimization networks to progressively optimize the rough target perspective image to obtain a fine target perspective image;

[0011] Step 4: Construct a variational autoencoder based on the reconstruction loss to extract the second hidden encoding from the fine target perspective image, calculate the reconstruction loss between the second hidden encoding and the first hidden encoding, and store the reconstructed second hidden encoding in the hidden encoding space;

[0012] Step 5: Construct an overall loss function, including an adversarial loss function and a reconstruction loss function for the multi-level cascaded residual optimization network, and a KL-divergence constraint and a reconstruction loss function for the variational autoencoder;

[0013] Step 6: Train the adversarial generation network. After training is completed, for a source perspective image that needs to be perspective-converted, the adversarial generation network randomly samples a second hidden encoding with a Gaussian distribution from the hidden encoding space, generates a multi-modal rough target perspective image through the second hidden encoding, the source perspective image, and the target perspective semantic segmentation map, and then obtains a multi-modal fine target perspective image after progressive optimization of the image quality by the multi-level cascaded residual optimization network.

[0014] In the multi-modal cross-perspective image generation method based on residual-based cascaded progressive optimization provided by the present invention, it can also have the following feature: Among them, in Step 1, the variational autoencoder is composed of a residual convolutional neural network, and the input real target perspective image is downsampled multiple times to an M-dimensional vector, and the KL-divergence is calculated with a randomly sampled M-dimensional Gaussian distribution vector. The calculation formula is as follows:

[0015]

[0016]

[0017] In Formulas (1) and (2), E() is the variational autoencoder, N(0, 1) is the standard Gaussian distribution, and p(z) and q(z) are the standard Gaussian distribution and the hidden encoding probability distribution fitted by the network, respectively.

[0018] In the multi-modal cross-perspective image generation method based on residual-based cascaded progressive optimization provided by the present invention, it can also have the following feature: Among them, in Step 2, the input layer of the adversarial generation network has six channels, and the target perspective semantic segmentation map and the source perspective image are unified in scale through bilinear interpolation.

[0019] In the multi-modal cross-view image generation method based on residual-based cascaded progressive optimization provided by the present invention, the following features may also be included: Among them, in step 3, each residual optimization network includes a residual estimation network composed of a four-layer convolutional neural network and a U-shaped image optimization network. Each level of the residual optimization network estimates the residual map of the input image through the residual estimation network, then performs weighted summation on the input image and the residual map, and then optimizes the image through the U-shaped image optimization network. The optimized image is used as the input image of the next level of the residual optimization network. After being optimized by multiple levels of the residual optimization network, a fine target view image is obtained. The calculation formula of each level of the residual optimization network is as follows:

[0020]

[0021] In formula (3), R i is the residual optimization network of the i-th level, I a is the input rough target view image, S pano is the target view semantic segmentation map, I res is the residual map estimated by the residual estimation network of this level, and are the images optimized by the residual optimization network of the previous level and the images optimized by the residual optimization network of this level, respectively.

[0022] In the multi-modal cross-view image generation method based on residual-based cascaded progressive optimization provided by the present invention, the following features may also be included: Among them, in steps 1 and 4, the variational autoencoder parameters are shared.

[0023] In the multi-modal cross-view image generation method based on residual-based cascaded progressive optimization provided by the present invention, the following features may also be included: Among them, in step 5, in the multi-level cascaded residual optimization network, the adversarial loss function and the reconstruction loss function are used as the objective functions for generating images and residual maps to perform image-level constraints on all generated images,

[0024] In the optimization of the variational autoencoder, the reconstruction loss function and the KL-divergence constraint are used to construct the objective function of the latent code,

[0025] The formula of the overall loss function is as follows:

[0026]

[0027]

[0028]

[0029]

[0030]

[0031]

[0032] Equation (4) is the reconstruction loss function of the residual optimization network,

[0033] Equation (5) is the reconstruction loss function of the variational autoencoder,

[0034] Equation (6) is the adversarial loss function of the rough target perspective image,

[0035] Equation (7) is the adversarial loss function of the optimized images at all levels in the multi-level cascaded residual optimization network,

[0036] Equation (8) is the adversarial loss function of the residual map in the multi-level cascaded residual optimization network,

[0037] In Equations (4)-(8), z r is the reconstructed second latent code in Step 4, D s and D r are discriminators for images and residual maps respectively, and λ i is the weight coefficient of different loss terms.

[0038] In the multi-modal cross-perspective image generation method based on residual cascaded progressive optimization provided by the present invention, it may further have the following feature: in Step 6, when training the adversarial generation network, the parameters in the generator and discriminator in the adversarial generation network are alternately optimized through the backpropagation algorithm.

[0039] Functions and Effects of the Invention

[0040] According to the multi-modal cross-perspective image generation method based on residual cascaded progressive optimization involved in the present invention, by jointly using a variational autoencoder and an adversarial generation network to generate multi-modal target perspective images across perspectives, it is possible to introduce multi-modal generation effects by randomly sampling latent codes from a Gaussian distribution while realizing cross-perspective image generation, thereby simulating target perspective images under different lighting and weather conditions; and the present invention also optimizes the generated rough target perspective images through a multi-level cascaded residual optimization network, which can effectively and progressively improve the image generation effect, reduce the distortion in the generated images, and can more effectively explain the process of image quality improvement through visualizing the residual map. Brief Description of the Drawings

[0041] Figure 1 is the system composition diagram of the multi-modal cross-perspective image generation method based on residual cascaded progressive optimization in the embodiment of the present invention;

[0042] Figure 2It is a flowchart of the multi-modal cross-view image generation method based on residual-based cascaded progressive optimization in the embodiments of the present invention;

[0043] Figure 3 It is a schematic diagram of the processing process of the multi-modal cross-view image generation method based on residual-based cascaded progressive optimization in the embodiments of the present invention. Detailed implementation manners

[0044] In order to make the technical means and effects achieved by the present invention easy to understand, the present invention will be specifically described below in conjunction with embodiments and drawings.

[0045] <Embodiment>

[0046] Figure 1 It is a system composition diagram of the multi-modal cross-view image generation method based on residual-based cascaded progressive optimization in the embodiments of the present invention.

[0047] As Figure 1 shown, in this embodiment, in the system 100 adopted by the multi-modal cross-view image generation method based on residual-based cascaded progressive optimization, it includes media data 101, a computing device 110, and a display device 191. The media data 101 is a source-view image, which can be extracted from remote sensing satellites, unmanned aerial vehicles, etc.

[0048] The computing device 110 is a computing device for processing the media data 101, mainly including a computer processor 120 and a memory 130. The processor 120 is a hardware processor for the computing device 110, such as a central processing unit CPU or a graphics processing unit (Graphical Process Unit). The memory 130 is a non-volatile storage device for storing computer code for the computing process of the processor 120. At the same time, the memory 130 also stores various intermediate data and parameters. The memory 130 includes a cross-view image data set 135, machine-related data, and executable code 140. The executable code 140 includes one or more software modules for performing the calculations of the computer processor 120. As Figure 1 shown, the executable code 140 includes a variational autoencoder 141, a generative adversarial network 143, and a residual-based cascaded image optimization module 147.

[0049] The variational autoencoder 141 is used to extract random information from the target-view image, that is, to map the target-view image to a latent code of a Gaussian distribution.

[0050] The generative adversarial network 143 is used to generate a multi-modal target-view image from the input source-view image, target-view semantic segmentation map, and latent code, that is, to generate a coarse-grained target-view image.

[0051] The residual-based cascaded image optimization module 147 is used to perform residual estimation on the coarse-grained target perspective image and further improve the image quality, that is, progressive image quality optimization.

[0052] The display device 191 is a device suitable for playing media data 101 and displaying the predicted results output by the computing device 101, and can be a computer, a television, or a mobile device.

[0053] Figure 2 It is a flowchart of the multi-modal cross-perspective image generation method based on residual-based cascaded progressive optimization in the embodiments of the present invention. Figure 3 It is a schematic diagram of the processing process of the multi-modal cross-perspective image generation method based on residual-based cascaded progressive optimization in the embodiments of the present invention.

[0054] As Figure 2 and Figure 3 As shown in

[0055] Step 1: Obtain the real target perspective image of the source perspective image, construct a variational autoencoder based on KL-divergence constraint, and map the real target perspective image to a low-dimensional vector through the variational autoencoder to obtain the first latent code that conforms to the Gaussian distribution.

[0056] In Step 1, the variational autoencoder is composed of a residual convolutional neural network, and the input real target perspective image is downsampled multiple times to an M-dimensional vector, and the KL-divergence is calculated with a randomly sampled M-dimensional Gaussian distribution vector. The calculation formula is as follows:

[0057]

[0058]

[0059] In Formula (1) and Formula (2), E() is the variational autoencoder, N(0, 1) is the standard Gaussian distribution, and p(z) and q(z) are the standard Gaussian distribution and the latent code probability distribution fitted by the network, respectively.

[0060] In this embodiment, the backbone model of the variational autoencoder that maps the image to a low-dimensional vector is constructed based on the residual convolutional neural network. Specifically, it is composed of four residual convolutional neural networks, and a max-pooling layer is used between each residual convolutional neural network to reduce its resolution.

[0061] Step 2: Use an adversarial generative network based on the U-shaped network to generate a rough target perspective image according to the source perspective image, the target perspective semantic segmentation map, and the first latent code.

[0062] In step 2, the input layer of the adversarial generative network has six channels, and the target perspective semantic segmentation map and the source perspective image are unified in scale through bilinear interpolation.

[0063] In this embodiment, by simultaneously inputting the source perspective image, the target perspective semantic segmentation map, and the latent code, the source perspective image and the target perspective semantic segmentation map are unified in scale, and are concatenated in the channel dimension to obtain a 6-dimensional input, which is input to the generator. In addition, the latent code is subjected to a scale transformation to obtain a tensor of the same scale as the image, and is concatenated with the shallow features of the generator in the channel dimension, so as to embed the randomness of the latent code in the generation process.

[0064] Step 3: Construct multiple residual optimization networks, and cascade the multiple residual optimization networks to progressively optimize the rough target perspective image to obtain a fine target perspective image.

[0065] In step 3, each residual optimization network includes a residual estimation network composed of a four-layer convolutional neural network and a U-shaped image optimization network.

[0066] Each stage of the residual optimization network uses the residual estimation network to estimate the residual map of the input image, and then performs weighted summation on the input image and the residual map, and then performs image optimization through the U-shaped image optimization network to constrain the image pixel values to fall within a reasonable range. The optimized image is used as the input image of the next stage of the residual optimization network to achieve progressive optimization. After being optimized by multiple stages of residual optimization networks, a fine target perspective image is obtained. The calculation formula of each stage of the residual optimization network is as follows:

[0067]

[0068] In formula (3), R i is the i-th stage of the residual optimization network, I a is the input rough target perspective image, S pano is the target perspective semantic segmentation map, I res is the residual map estimated by the residual estimation network of this stage, and are the images optimized by the previous stage of the residual optimization network and the images optimized by the current stage of the residual optimization network respectively.

[0069] In this embodiment, the subsequent stages of the residual optimization network are all initialized with the parameters of the previous stage of the residual optimization network. For each additional stage of the residual optimization network, the parameters of the previously trained network are fixed, and only the last stage of the residual optimization network is trained.

[0070] Step 4: Construct a variational autoencoder based on the reconstruction loss to extract the second latent code from the fine target perspective image, calculate the reconstruction loss between the second latent code and the first latent code, and store the reconstructed second latent code in the latent code space.

[0071] The variational autoencoder parameters in Step 1 and Step 4 are shared.

[0072] In this embodiment, by calculating the reconstruction loss between the output second latent code and the first latent code at the input end in Step 1, the generated image can encode sufficient random information.

[0073] Step 5: Construct an overall loss function, including an adversarial loss function and a reconstruction loss function for the multi-stage cascaded residual optimization network, and a KL-divergence constraint and a reconstruction loss function for the variational autoencoder.

[0074] In Step 5, in the multi-stage cascaded residual optimization network, use the adversarial loss function and the reconstruction loss function as the objective functions for the generated image and the residual map to perform image-level constraints on all generated images.

[0075] In the optimization of the variational autoencoder, use the reconstruction loss function and the KL-divergence constraint to construct the objective function of the latent code.

[0076] The formula of the overall loss function is as follows:

[0077]

[0078]

[0079]

[0080]

[0081]

[0082]

[0083] Formula (4) is the reconstruction loss function of the residual optimization network.

[0084] Formula (5) is the reconstruction loss function of the variational autoencoder.

[0085] Formula (6) is the adversarial loss function of the rough target perspective image.

[0086] Formula (7) is the adversarial loss function of each level of the optimized image in the multi-stage cascaded residual optimization network.

[0087] Formula (8) is the adversarial loss function of the residual map in the multi-stage cascaded residual optimization network.

[0088] In Formulas (4) - (8), z r is the second latent code reconstructed in Step 4, D s and D r are discriminators for images and residual images respectively, and λ i is the weight coefficient of different loss terms.

[0089] Step 6: Train the adversarial generation network. After training is completed, for a source-view image that needs to be perspective-transformed, the adversarial generation network randomly samples a second latent code of a Gaussian distribution from the latent code space, generates a multi-modal rough target-view image through the second latent code, the source-view image, and the target-view semantic segmentation map, and then obtains a multi-modal fine target-view image after progressive optimization of the image quality through a multi-stage cascaded residual optimization network.

[0090] In Step 6, when training the adversarial generation network, the parameters in the generator and discriminator of the adversarial generation network are alternately optimized through the backpropagation algorithm.

[0091] In this embodiment, the ADAM optimizer is used to train the adversarial generation network. The initial learning rate lr = 0.0002, and it decays by 0.05 every 10 epochs. The network is trained for about 50 epochs until convergence. We use the method of alternately training the generator and discriminator, that is, for each batch of data, first fix the parameters of the generator and update the parameters of the discriminator, and then fix the parameters of the discriminator and update the parameters of the generator.

[0092] Specifically, the training data in the CVUSA dataset and the Dayton dataset are used for training, and the testing is carried out in the test dataset. The training and testing data indices are the same as those in the literature (Tang, D., Xu, N., Sebe, N., Wang, Y., Corso, J. J., & Yan, Y. “Multi-channel attention selection GAN with cascaded semantic guidance for cross-view image translation,” in CVPR, 2019, pp. 2417–2426.). The generated images are evaluated using metrics such as FID, IS, PSNR, SSIM, and SD. In the CVUSA dataset, the above metrics reach 35.02, 2.8907, 21.2476, 0.4879, and 18.5000 respectively. In the Dayton dataset, the above metrics reach 40.32, 2.6443, 22.5510, 0.5626, and 19.6617 respectively.

[0093] Functions and Effects of the Embodiment

[0094] A multi-modal cross-view image generation method based on residual-based cascaded progressive optimization according to this embodiment generates multi-modal target-view images across views by jointly using a variational autoencoder and a generative adversarial network. It can, while achieving cross-view image generation, introduce a multi-modal generation effect by randomly sampling latent codes from a Gaussian distribution, thereby simulating target-view images under different lighting and weather conditions. Moreover, this embodiment also optimizes the generated rough target-view images through a multi-level cascaded residual optimization network, which can effectively and progressively improve the image generation effect, reduce the distortion in the generated images, and can more effectively explain the process of image quality improvement through visualizing the residual maps.

[0095] The above embodiments are preferred cases of the present invention and are not used to limit the protection scope of the present invention.

Claims

1. A multi-modal cross-view image generation method based on residual-based cascaded progressive optimization, which is used to perform perspective conversion on source perspective images to obtain multi-modal target perspective images, characterized in that, It includes the following steps: Step 1: Obtain the real target perspective image of the source perspective image, construct a variational autoencoder based on KL-divergence constraint, and map the real target perspective image to a low-dimensional vector through the variational autoencoder to obtain a first hidden encoding that conforms to the Gaussian distribution; Step 2: Use an adversarial generation network based on a U-shaped network to generate a rough target perspective image according to the source perspective image, the target perspective semantic segmentation map, and the first hidden encoding; Step 3: Construct multiple residual optimization networks, and cascade the multiple residual optimization networks to progressively optimize the rough target perspective image to obtain a fine target perspective image; Step 4: Construct the variational autoencoder based on the reconstruction loss to extract a second hidden encoding from the fine target perspective image, calculate the reconstruction loss between the second hidden encoding and the first hidden encoding, and store the reconstructed second hidden encoding in the hidden encoding space; Step 5: Construct an overall loss function, including an adversarial loss function and a reconstruction loss function for the multi-stage cascaded residual optimization network, and a KL-divergence constraint and a reconstruction loss function for the variational autoencoder; Step 6: Train the adversarial generation network. After training, for a source perspective image that needs to be perspective-converted, the adversarial generation network randomly samples the second hidden encoding of the Gaussian distribution from the hidden encoding space, and generates a multi-modal rough target perspective image through the second hidden encoding, the source perspective image, and the target perspective semantic segmentation map, and then obtains a multi-modal fine target perspective image after progressive optimization of the image quality by the multi-stage cascaded residual optimization network.

2. The multi-modal cross-perspective image generation method based on cascaded progressive optimization of residuals according to claim 1, wherein: Among them, In the step 1, the variational autoencoder is composed of a residual convolutional neural network, and the input real target perspective image is downsampled multiple times to an M-dimensional vector, and the KL-divergence is calculated with a randomly sampled M-dimensional Gaussian distribution vector. The calculation formula is as follows: In formulas (1) and (2), E() is the variational autoencoder, N(0, 1) is the standard Gaussian distribution, and p(z) and q(z) are the standard Gaussian distribution and the hidden encoding probability distribution fitted by the network, respectively.

3. The multi-modal cross-perspective image generation method based on cascaded progressive optimization of residuals according to claim 1, wherein: Among them, In the step 2, the input layer of the adversarial generation network is six channels, and the target perspective semantic segmentation map and the source perspective image are unified in scale through bilinear interpolation.

4. The multi-modal cross-perspective image generation method based on cascaded progressive optimization of residuals according to claim 1, wherein: Among them, In the step 3, each residual optimization network includes a residual estimation network composed of a four-layer convolutional neural network and a U-shaped image optimization network. Each level of the residual optimization network estimates the residual of the input image through the residual estimation network to obtain a residual map, then performs weighted summation on the input image and the residual map, and optimizes the image through the U-shaped image optimization network. The optimized image is used as the input image of the next level of the residual optimization network. After being optimized by multiple levels of the residual optimization network, the fine target perspective image is obtained. The calculation formula of each level of the residual optimization network is as follows: In formula (3), R i is the residual optimization network of the i-th level, I a is the input rough target perspective image, S pano is the target perspective semantic segmentation map, I res is the residual map estimated by the residual estimation network of this level, and are the image optimized by the residual optimization network of the previous level and the image optimized by the residual optimization network of this level, respectively.

5. The multi-modal cross-perspective image generation method based on residual-based cascaded progressive optimization according to claim 1, wherein: Among them, The variational autoencoder parameters in step 1 and step 4 are shared.

6. The multi-modal cross-perspective image generation method based on residual-based cascaded progressive optimization according to claim 1, wherein: Among them, In step 5, in the multi-level cascaded residual optimization network, the adversarial loss function and the reconstruction loss function are used as the objective functions for generating images and residual maps to perform image-level constraints on all generated images. In the optimization of the variational autoencoder, the reconstruction loss function and the KL-divergence constraint are used to construct the objective function of the latent code. The formula of the overall loss function is as follows: Formula (4) is the reconstruction loss function of the residual optimization network. Formula (5) is the reconstruction loss function of the variational autoencoder. Formula (6) is the adversarial loss function of the rough target perspective image. Formula (7) is the adversarial loss function of the optimized images at each level in the multi-level cascaded residual optimization network. Formula (8) is the adversarial loss function of the residual map in the multi-level cascaded residual optimization network. In Formulas (4) - (8), z r is the reconstructed second hidden code in Step 4, D s and D r are discriminators for the image and the residual map respectively, and λ i is the weight coefficient of different loss terms.

7. The multi-modal cross-perspective image generation method based on residual-based cascaded progressive optimization according to claim 1, wherein: Among them, In step 6, when training the adversarial generation network, the parameters in the generator and discriminator of the adversarial generation network are alternately optimized through the backpropagation algorithm.

Citation Information

Patent Citations

  • Multi-sequence MRI image segmentation method based on residual network

    CN111739051A

  • Multi-modal image translation using neural networks

    US20190279075A1