A network structure and method for image scene relighting based on a GAN network
The GAN-based network structure for image relighting addresses challenges in general scene relighting by enhancing feature extraction and attention to scene details, resulting in more realistic image relighting outcomes.
Patent Information
- Application Number
- CN202211308741.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-10-25
AI Technical Summary
The existing image re-lighting methods based on general scenes based on image-to-image conversion are not as progressing as those in other categories. The main reason is that there are many difficulties in achieving scene re-lighting with realism, such as removing and re-casting shadows, processing of texture details, etc.
The image scene re-illumination network structure based on GAN network is adopted, including scene reconstruction network, shadow estimation network and re-rendering network. By generating anti-aggression principles and feature encoding and decoding structures, the attention ability of lighting feature extraction and scene information changes is enhanced, and the attention mechanism and up-down sampling blocks are used to improve feature details learning.
It significantly reduces the difficulty of redrawing shadows, and the generated heavy-light images are more realistic. It shows competitiveness in PNSR, SSIM, LPIPS and MPS indicators in quantitative analysis, especially in SSIM and MPS.
Smart Images

Figure CN115578497B_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the field of deep learning, and particularly relates to a network structure and method for image scene relighting based on a GAN network. Background Art
[0002] Relighting [1] Refers to the process of synthesizing an original scene under new lighting conditions. This technology was first used in the film special effects industry to achieve high-fidelity light and shadow details of characters and objects in the scene, and later was further applied to the fields of video games and virtual reality. Image Relighting [2] Is a subfield of relighting and is also a long-term research problem in the current fields of computer vision and graphics. It refers to the simulation and reconstruction of the scene appearance under new lighting conditions. In some literature
[21] It is also referred to as Image LightSource Transfer.
[0003] In the 1990s, traditional relighting work mainly used laser scanning [5] or user-assisted reconstruction algorithms [6][7] to estimate lighting information such as the geometric information, reflectivity, or ambient light of the scene. In 2000, Debevec et al. [8] pioneeringly used a light stage to obtain the reflection field of a human face. They directly obtained the reflection field by densely sampling the reflection map generated after the incident angle reaches the object surface, and synthesized the image of human face relighting through the mapping of the ambient light map and the reflection field map. This method achieved great success in the field of human face relighting. Subsequently, the team successively launched Light Stage 2 [9] 、LightStage 3
[10] 、Light Stage 5
[11] and other image acquisition devices have been applied in many classic film special effects syntheses. In traditional classic image relighting tasks [4][5][8] scene three-dimensional reconstruction is a very difficult and cumbersome step. The main reasons are: it is very difficult to obtain light transmission information such as the geometric model, reflection characteristics, and lighting environment of the real world, and often expensive special equipment is required for acquisition, such as a light stage [8] ; the reconstruction of scenes with complex geometric structures and optical effects will consume a large amount of computing resources and have low operating efficiency. Therefore, traditional relighting methods [8][9]
[10]
[11] could only be applied to high-investment industries such as film special effects in the early stage and could not be applied to fields such as virtual reality that require real-time lighting.
[0004] In recent years, deep learning
[12] has achieved good results in various computer vision and computer graphics tasks, such as neural rendering
[13] , neural inverse rendering
[14] , scene reconstruction
[15] etc. Introducing deep learning technology into the field of image relighting has also become a current research trend, so the research on deep learning-based relighting has received increasing attention. In this field, some relighting research
[18]
[19] hopes to achieve the relighting effect under multiple views by relying on special information such as the geometric priors of human faces or buildings. However, its disadvantage is that it heavily relies on geometric priors and is only suitable for special applications such as human faces, lacking generality for general scenes; The image relighting technology based on image-to-image
[20]
[21]
[22] then hopes to establish an end-to-end rendering process and minimize the dependence on redundant scene information as much as possible, and achieve the relighting effect by extracting feature information from existing image-form data, so as to simplify the relighting process.
[0005] However, no matter which type of method, image relighting research requires a large number of high-definition and realistic image datasets. Especially in the field of image relighting for general scenes, due to the very complex scenes and lighting effects in the real world, it is often difficult to produce high-quality high-definition datasets based on the real world. In order to promote related research, in recent years, existing research has tried to introduce synthetic virtual scene datasets, such as VIDIT
[16] , which has been widely used in relighting competitions held by top conferences such as ECCV, strongly promoting the development of scene relighting.
[0006] Most of the existing general-scene image relighting methods based on image-to-image conversion use the relighting dataset VIDIT under virtual scenes
[16] , the training set of this dataset contains 300 scenes from the Unreal Engine, and each scene is collected to obtain 40 pictures under 5 different color temperatures and 8 different lighting directions, for a total of 12,000 pictures. The method of the present invention is also trained on the VIDIT dataset
[16] and comparative experiments and ablation experiments are designed.
[0007] Most of this type of research uses structures based on encoders and decoders. For example, Wang et al
[20] proposed a deep neural network DRN for the image relighting task. This network consists of three sub-networks, which respectively complete three tasks: scene reconstruction, shadow prior estimation, and re-rendering. Among them, both the scene reconstruction sub-network and the shadow estimation sub-network adopt a structure similar to the U-Net network
[17] . On the basis of the work of Wang et al, in order to effectively solve the above problems, another work
[21] A new feature self - calibration block is added as a basic block for the feature encoder and decoder in the scene reconstruction and shadow estimation tasks.
[0008] The attention mechanism has long been widely used in various image enhancement tasks in recent years
[22] In these tasks, the attention mechanism assigns weights to the feature maps to amplify and emphasize the features in important regions. The relit images contain parts with large contrasts such as shadows and highlights. Therefore, the attention mechanism helps to guide the deep neural network to learn to focus on positions such as shadows and highlights in the image, enabling the network to better remove and re - project the shadows of the original image and enhance the realism of the predicted relit image.
[0009] However, no matter which method, the progress in the field of image - based relighting for general scenes from image - to - image is not as good as that of other types of relighting and is still basically in its infancy [3] The main reason is that there are many difficulties in achieving realistic scene relighting, such as removing and re - projecting shadows, and dealing with texture details.
[0010] References:
[0011] [1] Debevec P. Virtual cinematography: Relighting through computation[J]. Computer, 2006, 39(8): 57 - 65.
[0012] [2] Debevec P. Image - based lighting[M] / / ACM SIGGRAPH 2006 Courses. 2006: 4 - es.
[0013] [3] Einabadi F, Guillemaut J Y, Hilton A. Deep neural models for illumination estimation and relighting: A survey[C] / / Computer Graphics Forum. 2021, 40(60: 315 - 331.
[0014] [4] Dobashi Y, Kaneda K, Nakatani H, et al. A quick rendering method using basis functions for interactive lighting design[C] / / Computer Graphics Forum. Edinburgh, UK: Blackwell Science Ltd, 1995, 14(3): 229 - 240.
[0015] [5] Marschner S R, Greenberg D P. Inverse lighting for photography[C] / / Color and Imaging Conference. Society for Imaging Science and Technology, 1997, 1997(1): 262 - 265.
[0016] [6] Loscos C, Frasson M C, Drettakis G, et al. Interactive virtual relighting and remodeling of real scenes[C] / / Eurographics Workshop on Rendering Techniques. Springer, Vienna, 1999: 329 - 340.
[0017] [7] Yu Y, Debevec P, Malik J, et al. Inverse global illumination: Recovering reflectance models of real scenes from photographs[C] / / Proceedings of the 26th annual conference on Computer graphics and interactive techniques. 1999: 215 - 224.
[0018] [8]Debevec P, Hawkins T, Tchou C, et al. Acquiring the reflectance field of a human face[C] / / Proceedings of the 27th annual conference on Computer graphics and interactive techniques. 2000:145-156.
[0019] [9]Hawkins T, Cohen J, Debevec P. A photometric approach to digitizing cultural artifacts[C] / / Proceedings of the 2001 conference on Virtual reality, archeology, and cultural heritage. 2001:333-342.
[0020]
[10] Debevec P, Wenger A, Tchou C, et al. A lighting reproduction approach to live-action compositing[J]. ACM Transactions on Graphics(TOG), 2002, 21(3):547-556.
[0021]
[11] WENGER A., GARDNER A., TCHOU C., UNGER J., HAWKINS T., DEBEVEC P.: Performance relighting and reflectance transformation with time-multiplexed illumination. 756–764.
[0022]
[12] LeCun Y, Bengio Y, Hinton G. Deep learning[J]. nature, 2015, 521(7553):436-444.
[0023]
[13] Tewari A, Fried O, Thies J, et al. State of the art on neural rendering[C] / / Computer Graphics Forum. 2020, 39(2): 701-727.
[0024]
[14] Yu Y, Smith W A P. Inverse renderernet: Learning single image inverse rendering[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 3155-3164.
[0025]
[15] Srinivasan P P, Mildenhall B, Tancik M, et al. Lighthouse: Predicting lighting volumes for spatially - coherent illumination[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020: 8080-8089.
[0026]
[16] Helou M E, Zhou R, Barthas J, et al. VIDIT: Virtual image dataset for illumination transfer[J]. arXiv preprint arXiv:2005.05460, 2020.
[0027]
[17] Ronneberger O, Fischer P, Brox T. U - net: Convolutional networks for biomedical image segmentation[C] / / International Conference on Medical image computing and computer - assisted intervention. Springer, Cham, 2015: 234-241.
[0028]
[18] Philip J, Gharbi M, Zhou T, et al. Multi-view relighting using ageometry-aware network[J]. ACM Trans. Graph., 2019, 38(4): 78:1-78:14.
[0029]
[19] Guo K, Lincoln P, Davidson P, et al. The relightables: Volumetricperformance capture of humans with realistic relighting[J]. ACM Transactionson Graphics(ToG), 2019, 38(6): 1-19.
[0030]
[20] Wang L W, Siu W C, Liu Z S, et al. Deep relighting networks for imagelight source manipulation[C] / / European Conference on ComputerVision. Springer, Cham, 2020: 550-567.
[0031]
[21] Wang Y, Lu T, Zhang Y, et al. Multi-scale self-calibrated network forimage light source transfer[C] / / Proceedings of the IEEE / CVF Conference onComputer Vision and Pattern Recognition. 2021: 252-259.
[0032]
[22] Yang H H, Chen W T, Kuo S Y. S3Net: A single stream structure fordepth guided image relighting[C] / / Proceedings of the IEEE / CVF Conference onComputer Vision and Pattern Recognition. 2021: 276-283. Summary of the Invention
[0033] The present invention improves the deficiencies existing in the current image relighting algorithms for general scenarios based on image-to-image, and provides a network structure and method for image scene relighting based on a GAN network. The present invention adopts a new feature encoding and decoding structure, strengthens the ability to extract illumination features and the ability to pay attention to changes in scene information in a deep neural network, focuses on the detailed learning of image features, and enables the network to learn more rich and effective features.
[0034] The purpose of the present invention is achieved through the following technical solutions:
[0035] A network structure for image scene relighting based on a GAN network includes a scene reconstruction network, a shadow estimation network, and a re-rendering network arranged in sequence; the scene reconstruction network and the shadow estimation network adopt the principle of generative adversarial, and both are composed of a generator and a discriminator, and the generator adopts an encoder-decoder structure;
[0036] The scene reconstruction network is composed of a 7*7 convolutional layer, four upsampling blocks, a residual block, four downsampling blocks, and a 3*3 convolutional layer arranged in sequence. The feature information of the four upsampling blocks is fused together by skip connections, and the feature information output by the 7*7 convolutional layer and the feature information output by the 3*3 convolutional layer are fused together by skip connections;
[0037] The shadow estimation network is composed of a 7*7 convolutional layer, four upsampling blocks, a residual block, four downsampling blocks, and a 3*3 convolutional layer arranged in sequence;
[0038] The re-rendering network is composed of a convolutional module, an average pooling layer, two fully connected layers, an activation layer, a 3*3 convolutional layer, and a 7*7 convolutional layer arranged in sequence; the convolutional module is composed of several convolutional layers with different convolutional kernel sizes.
[0039] The present invention also provides a generation method for image scene relighting based on a GAN network, including:
[0040] Input an image under a given illumination condition into the scene reconstruction network, remove the original illumination effect of the input image through the scene reconstruction network, and extract the inherent scene information from the input image, and output a scene reconstruction image containing the inherent scene information;
[0041] Input an image under a given illumination condition into the shadow estimation network, and output a shadow estimation image containing the target illumination information;
[0042] Input the splicing result of the scene reconstruction image and the shadow estimation image into the re-rendering network, and output a relit image under the target illumination condition;
[0043] During the process of training the scene reconstruction network and the shadow estimation network, the discriminator continuously fine-tunes the generator, making the scene reconstruction image and the shadow estimation image continuously approach the corresponding target images.
[0044] Further, the re-rendering network first inputs the input image into a total of 12 convolutional layers including a 3×3 convolutional layer, a 5×5 convolutional layer, a 7×7 convolutional layer, …, and a 25×25 convolutional layer, then concatenates the feature information output by the 12 convolutional layers, and next undergoes an average pooling operation, two fully connected layers, a sigmoid function, a 1×1 convolutional layer, and a 7×7 convolutional layer.
[0045] Further, the scene reconstruction network, the shadow estimation network, and the re-rendering network are trained separately. First, the scene reconstruction network is trained through a loss function and paired input images and shadowless target images; second, the shadow estimation network is trained using paired input images and target images; finally, the re-rendering network is trained using the loss function.
[0046] During the process of training the scene reconstruction network and the shadow estimation network, the discriminator continuously fine-tunes the generator, making the scene reconstruction image and the shadow estimation image continuously approach the corresponding target images.
[0047] Further, during training, the size of all images is adjusted from 1024×1024 to 512×512, and the mini-batch size is set to 6. The network parameters are adjusted using the Adam optimizer, the momentum is set to 0.5, the learning rate is set to 0.0001, and each network is trained for 20 epochs.
[0048] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method for generating image scene relighting based on the GAN network.
[0049] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method for generating image scene relighting based on the GAN network.
[0050] Compared with the prior art, the beneficial effects brought by the technical solution of the present invention are:
[0051] (1) In order to improve the extraction ability of scene features and shadow features and reduce the difficulty of image relighting, the present invention designs an image relighting network structure based on the Retinex theory. This network structure includes a scene reconstruction network, a shadow estimation network, and a re-rendering network. The method of the present invention uses the scene reconstruction network to extract scene information, and uses the shadow estimation network to transfer the lighting effect from the input image X to the target image Y, which significantly reduces the difficulty of redrawing shadows. Finally, a re-rendering network is used to combine the scene information and the lighting effect.
[0052] (2) The training process of the GAN network follows the adversarial learning strategy, and the latent distribution within the target image is learned and used during the training process. At the end of the training, a dynamic balance will be achieved, where the images generated by the generator have a similar latent perceptual structure to the target image. In order to generate a relighted image closer to the real image and make the relighted image more realistic, both the scene reconstruction network and the shadow estimation network of the present invention are based on the GAN network, and the training also adopts the adversarial learning method.
[0053] (3) In order to make full use of the spatial information of the features to better reconstruct the scene and shadows and retain as much texture detail as possible, the present invention designs a novel upsampling and downsampling block. This upsampling and downsampling block can make full use of the feature information and is used in the scene reconstruction network and the shadow estimation network.
[0054] (4) The present invention verifies the effectiveness of the method of the present invention through experiments on the VIDIT dataset. Under quantitative analysis, the evaluation metrics adopt the widely recognized PNSR, SSIM, LPIPS, and MPS in the field. Compared with the existing methods Retinex-Net, DRN, MCN, and S3Net in this field (image relighting in the general scenario based on image-to-image conversion), the method of the present invention shows competitive and significant effects in all other metrics except that it is slightly lower than S3Net in PNSR, which proves the effectiveness of the method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a schematic diagram of the overall network structure of the method of the present invention.
[0056] Figure 2a and Figure 2b They are respectively schematic diagrams of the upsampling and downsampling block structures used in the scene reconstruction network.
[0057] Figure 3a and Figure 3b They are respectively schematic diagrams of the upsampling and downsampling block structures used in the shadow estimation network.
[0058] Table 1 shows the quantitative comparison results between the method of the present invention and other methods DETAILED DESCRIPTION OF THE INVENTION
[0059] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0060] This embodiment provides a network structure for an image relighting task in a general scenario based on image-to-image. The network structure includes a scene reconstruction network, a shadow estimation network, and a re-rendering network. As Figure 1 shown, from left to right are the scene reconstruction network, the shadow estimation network, and the re-rendering network. Among them, the scene reconstruction network and the shadow estimation network adopt the principle of generative adversarial, which are divided into a generator and a discriminator. The generator part adopts an encoder-decoder structure. During the training process of the scene reconstruction network and the shadow estimation network, the discriminator continuously fine-tunes the generator, so that the scene reconstruction image output by the generator of the scene reconstruction network and the shadow estimation image output by the generator of the shadow estimation network continuously approach the corresponding target images.
[0061] The network structure of the scene reconstruction network consists of a 7*7 convolutional layer, four upsampling blocks, a residual block, four downsampling blocks, and a 3*3 convolutional layer. The scene reconstruction network also adopts skip connections to fuse the feature information of the first, second, third, and fourth upsampling blocks together, and fuse the feature information output by the 7*7 convolutional layer and the feature information output by the 3*3 convolutional layer together.
[0062] The network structure of the shadow estimation network consists of a 7*7 convolutional layer, four upsampling blocks, a residual block, four downsampling blocks, and a 3*3 convolutional layer. The network of the shadow estimation network is similar to the scene reconstruction network, and the difference is that the shadow estimation network removes the skip connections. Removing the skip connections helps the shadow estimation network to focus on the global light effect.
[0063] The re-rendering network inputs the input image into a total of 12 convolutional layers, namely a 3*3 convolutional layer, a 5*5 convolutional layer, a 7*7 convolutional layer,..., a 25*25 convolutional layer, and then splices the feature information output by the 12 convolutional layers together. Next, through average pooling operations, two fully connected layers, a sigmoid function, a 1*1 convolutional layer, and a 7*7 convolutional layer.
[0064] Specifically, the method for generating relighting based on the above network structure includes:
[0065] (1) The input of the scene reconstruction network is an image under a given lighting condition, and the output is a scene reconstruction image containing the inherent information of the scene. The purpose of the scene reconstruction network is to remove the original lighting effect of the input image and extract the inherent information of the scene from the input image.
[0066] (2) The input of the shadow estimation network is an image under a given lighting condition, and the output is a shadow estimation image containing the target lighting information. The purpose of the shadow estimation network is to obtain the lighting effect information under the target lighting condition.
[0067] (3) The input of the re-rendering network is the result of splicing the scene reconstruction image output by the scene reconstruction network and the shadow estimation image output by the shadow estimation network, and the output is the re-illuminated image under the target lighting condition. The purpose of the re-rendering network is to obtain the re-illuminated image under the target lighting condition.
[0068] In the entire rendering process, the design of the downsampling block and the upsampling block is extremely important, and they determine the ability to extract important feature information. In this embodiment, the upsampling and downsampling blocks of the scene reconstruction network use the processing flow as shown in Figure 2a and Figure 2b , and the upsampling and downsampling blocks of the shadow estimation network use the processing flow as shown in Figure 3a and Figure 3b . Taking the downsampling block of the scene reconstruction network as an example, the downsampling block of the scene reconstruction network first inputs the input feature into the first 3×3 convolutional layer with a stride of 2, maps the input feature to the small-scale space, and obtains a small-scale feature F small ; then inputs the small-scale feature F small into the 4×4 deconvolutional layer with a stride of 2 and maps it back to the input scale space to obtain the feature F normal . At the same time, another branch generates the calibration weight λ1 by passing the input feature through a 1×1 convolutional layer and the LeakyReLu function. The weight calibration λ1 and F normal perform a multiplication operation and obtain the calibrated feature with the size of the input scale space is remapped to the small-scale space through the second 3×3 convolutional layer with a stride of 2 to obtain the feature The feature F small obtains the calibration weight λ2 through the branch containing a dense residual block. The feature and the calibration weight λ2 go through a summation operation to obtain the output feature F out with half the size of the input feature size. According to the above process, the formula of the downsampling block can be expressed as:
[0069]
[0070] where and represent 2 different 3×3 convolutional layers, DeCon4 represents a 4×4 deconvolutional layer, λ1 and λ2 are the calibration weights respectively, and λ1 and λ2 in the upsampling and downsampling blocks of the scene reconstruction network can be defined as follows
[0071] λ1 = LeakyReLu(Con1(F))
[0072]
[0073] Among them, Con1 represents a 1×1 convolutional layer, and DR_Block represents the newly designed dense residual block.
[0074] Similar to the above process, the formula of the downsampling block of the scene reconstruction network can be expressed as:
[0075]
[0076] Specifically, this embodiment relates to the image relighting task under the general image-to-image scenario, specifically, rendering the scene under any illumination condition into the scene under a certain unified illumination condition, abbreviated as the AnyToOne task.
[0077] Assume the input image is An image representing the illumination condition Φ. The relighting task is to render the image under the given target illumination condition Ψ. According to the Retinex theory, the input image can be described as:
[0078] X = L Φ (S)
[0079] where S represents the inherent scene information of the image under different illumination conditions, and L Φ (·) is a defined illumination function, which is responsible for providing the global illumination and shadow effects under the illumination condition Φ. The image relighting task can be further described as:
[0080] Y = L Ψ (L Φ -1 (X))
[0081] The overall relighting task can be divided into the following two-step operations:
[0082] (1) Scene reconstruction operation: L Φ -1 (·). This step is to recover the scene structure information S from the input image X, and the goal is to remove the original illumination effect Φ. The key to this step is to remove shadows.
[0083] (2) Relighting operation: L Ψ (·). This step is to relight the scene structure information S with the target light source to obtain the output image Y. The goal is to add the new illumination effect Ψ. The key to this step is to add shadows.
[0084] Since single-image relighting has no additional geometric information input, the relighting operation is more difficult. The method of the present invention does not directly find a certain relighting operation LΨ (·), but rather seeks a transfer operation L that migrates the lighting effect from the input image X to the target image Y Φ→Ψ (X), which significantly reduces the difficulty of redrawing shadows. Finally, a re-rendering process R(·) is used to combine the scene information and the lighting effect. The entire process can be expressed as:
[0085] Y′ = R(L Φ -1 (X), L Φ→Ψ (X))
[0086] Since the scene information S extracted from the input image is itself difficult to define, only the manually observed image can be used as the label for model training. In this embodiment, the exposure fusion image recommended by DRN
[20] is used as the real image of the scene reconstruction network. Such an exposure fusion image may contain some redundant information that does not belong to the scene information S. Therefore, strictly speaking, the scene reconstruction network in this embodiment belongs to semi-supervised learning rather than fully supervised learning.
[0087] Both the scene reconstruction network and the shadow estimation network are based on the GAN network, so their training also adopts the adversarial learning method. During the training of the scene reconstruction network, its discriminator is used to distinguish the scene reconstruction image output by the scene reconstruction network generator and the corresponding exposure fusion image, and its loss function is defined as:
[0088]
[0089] where Y no-shadow represents the exposure fusion image, X represents the input image of the scene reconstruction network, G represents the generator of the scene reconstruction network, and D represents the discriminator of the scene reconstruction network.
[0090] The shadow estimation network uses a shadow region discriminator to distinguish the shadow estimation image output by the shadow estimation network generator and the corresponding real image, and its loss function is defined as:
[0091]
[0092] where X represents the input image of the shadow estimation network, Y represents the target image of the shadow estimation network, G represents the generator of the shadow estimation network, D represents the discriminator of the shadow estimation network, and D′ represents the shadow region discriminator of the shadow estimation network.
[0093] The loss function of the re-rendering network is defined as:
[0094]
[0095] Among them, λ is set to 0.01 in all trainings of this embodiment, Y represents the target image of the re-rendering network, Y′ represents the predicted image of the re-rendering network, and feat(·) represents the feature map obtained after passing through the VGG network.
[0096] The network training of this embodiment uses the VIDIT dataset, which contains 390 different virtual scenes, including 300 scenes for training, 45 scenes for validation, and 45 scenes for testing. Each scene is rendered with 8 light directions and 5 color temperatures, resulting in 40 images with a resolution of 1024×1024. The VIDIT dataset was initially used in the relighting challenge, and the test set is not publicly available, so the 45 validation scenes and 45 test scenes in it cannot be used in this invention. Therefore, this invention uses 280 of the 300 training scenes for training and 20 scenes for testing.
[0097] Limited by the GPU memory and computing power, the three sub-networks (scene reconstruction network, shadow estimation network, and re-rendering network) in this embodiment are trained separately. First, the scene reconstruction network is trained using paired input images and shadow-free target images through a designed loss function. Then, the shadow estimation network is trained using paired input images and target images. Finally, with the scene reconstruction network and shadow estimation network fixed and their last convolutional layer and discriminator removed, the re-renderer network is trained using the designed loss function. During training, the size of all images is adjusted from 1024×1024 to 512×512, and the mini-batch size is set to 6. The network parameters are adjusted using the Adam optimizer with a momentum of 0.5 and a learning rate of 0.0001, and each network is trained for 20 epochs. All experiments are carried out using the deep learning training framework PyTorch on a machine equipped with 3 NVIDIA GTX3090Ti GPUs.
[0098] In quantitative analysis, the evaluation metrics are the widely recognized PNSR, SSIM, LPIPS, and MPS in the field. Compared with the existing methods in this field, Retinex-Net, DRN, MCN, and S3Net, the method of this invention shows significant effects in all metrics except that the PNSR is slightly lower than that of S3Net. For details, see Table 1.
[0099] Table 1. Quantitative comparison results of the method of this invention and other methods
[0100]
[0101] Preferably, an embodiment of the present application further provides a specific implementation manner of an electronic device capable of implementing all steps in the above-mentioned image scene relighting generation method based on a GAN network. The electronic device specifically includes the following:
[0102] A processor, a memory, a communications interface, and a bus;
[0103] Among them, the processor, the memory, and the communications interface complete mutual communication through the bus; the communications interface is used to implement information transmission between related devices such as server-side devices, metering devices, and client-side devices.
[0104] The processor is used to call the computer program in the memory. When the processor executes the computer program, all steps in the above-mentioned image scene relighting generation method based on a GAN network are implemented.
[0105] An embodiment of the present application further provides a computer-readable storage medium capable of implementing all steps in the above-mentioned image scene relighting generation method based on a GAN network. A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, all steps in the above-mentioned image scene relighting generation method are implemented.
[0106] Each embodiment in this specification is described in a progressive manner. For the same or similar parts between the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the hardware + program type embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0107] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0108] Although the present application provides method operation steps such as those in the embodiments or flowcharts, more or fewer operation steps may be included based on routine or non-creative labor. The order of steps listed in the embodiments is only one way among many execution orders of the steps and does not represent the only execution order. When the actual device or client product is executed, it may be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing).
[0109] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0110] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.
[0111] These computer program instructions may also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.
[0112] The present invention is not limited to the above-described embodiments. The above description of the specific embodiments is intended to describe and illustrate the technical solutions of the present invention. The above specific embodiments are merely illustrative and not restrictive. Without departing from the spirit of the present invention and the scope protected by the claims, those of ordinary skill in the art can make many specific transformations in various forms under the inspiration of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A generation method for image scene relighting based on a GAN network, based on the network structure of image scene relighting, characterized in that, The network structure includes a scene reconstruction network, a shadow estimation network, and a re-rendering network arranged in sequence; the scene reconstruction network and the shadow estimation network adopt the generative adversarial principle and are both composed of a generator and a discriminator, and the generator adopts an encoder-decoder structure; the scene reconstruction network consists of a 7*7 convolutional layer, four upsampling blocks, a residual block, four downsampling blocks, and a 3*3 convolutional layer arranged in sequence. The feature information of the four upsampling blocks is fused together by skip connection, and the feature information output by the 7*7 convolutional layer and the feature information output by the 3*3 convolutional layer are fused together by skip connection; The shadow estimation network consists of a 7*7 convolutional layer, four upsampling blocks, a residual block, four downsampling blocks, and a 3*3 convolutional layer arranged in sequence; The re-rendering network consists of a convolutional module, an average pooling layer, two fully connected layers, an activation layer, a 3*3 convolutional layer, and a 7*7 convolutional layer arranged in sequence; the convolutional module is composed of several convolutional layers with different convolutional kernel sizes; The generation method includes the following steps: Input an image under a given lighting condition into the scene reconstruction network. The scene reconstruction network removes the original lighting effect of the input image and extracts the scene inherent information from the input image, and outputs a scene reconstruction image containing the scene inherent information; Input an image under a given lighting condition into the shadow estimation network, and output a shadow estimation image containing the target lighting information; Input the splicing result of the scene reconstruction image and the shadow estimation image into the re-rendering network, and output a re-lighted image under the target lighting condition.
2. The generating method for image scene relighting based on a GAN network according to claim 1, wherein, The re-rendering network first inputs the input image into a total of 12 convolutional layers, namely the 3*3 convolutional layer, 5*5 convolutional layer, 7*7 convolutional layer,..., 25*25 convolutional layer. Then, the feature information output by the 12 convolutional layers is spliced together. Next, it undergoes an average pooling operation, two fully connected layers, a sigmoid function, a 1*1 convolutional layer, and a 7*7 convolutional layer.
3. The generation method of image scene relighting based on GAN network according to claim 1, characterized in that Train the scene reconstruction network, the shadow estimation network, and the re-rendering network respectively. First, train the scene reconstruction network through a loss function and paired input images and shadowless target images; second, train the shadow estimation network using paired input images and target images; finally, train the re-renderer network using the loss function; During the process of training the scene reconstruction network and the shadow estimation network, the discriminator continuously fine-tunes the generator to make the scene reconstruction image and the shadow estimation image continuously approach the corresponding target images.
4. The generation method of image scene relighting based on the GAN network according to claim 3, characterized in that, During training, the size of all images is adjusted from 1024×1024 to 512×512, and the mini-batch size is set to 6. The network parameters are adjusted using the Adam optimizer, the momentum is set to 0.5, the learning rate is set to 0.0001, and each network is trained for 20 epochs.
5. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the generation method for image scene re-lighting based on the GAN network described in claims 1-4.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the generation method for image scene re-lighting based on the GAN network described in claims 1-4.
Citation Information
Patent Citations
Image super-resolution method and device based on optical-field collection device
CN108074218A
Three-dimensional (3D) reconstructions of dynamic scenes using reconfigurable hybrid imaging system
CN111344746A