A method and device for full-body portrait video relighting based on convolutional neural network
Through the full-body portrait video reillumination method based on convolutional neural network, the problem of inconsistent results of the reillumination video under dynamic lighting conditions is solved, and high-quality full-body portrait reillumination is achieved, enhancing the sense of reality.
Patent Information
- Application Number
- CN202210612418.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-05-31
AI Technical Summary
The prior art is difficult to achieve consistent lighting video results under dynamic lighting conditions, especially in full-body portrait lighting, where artifacts and masks are present, which lacks realism.
Using a full-body portrait video reillumination method based on a convolutional neural network, through a pre-trained image processing model, multiple image frames of the to-process video image and the target illumination scene are rendered into rendered image frames in the target illumination scene, and time consistency processing is performed to finally synthesize the reillumination video image.
The consistent re-illumination video effect under dynamic lighting conditions is achieved, improving the reality and quality of full-body portrait re-illumination.
Smart Images

Figure CN115100337B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to a method and device for relighting a full-body portrait video based on a convolutional neural network. Background Art
[0002] With the rise of digital photography technology, people have an increasing demand for digital image processing. Manual photo editing requires a lot of work and has high requirements for the professionalism of users, which has prompted the industry to pay attention to technologies such as automatic image enhancement. Due to the sensitivity of the human eye to light, relighting has gradually become one of the most important and cutting-edge technologies. Realistic relighting provides immersive visual effects for augmented reality, virtual reality, and digital special effects, and has a wide range of applications in the multimedia technology industry. Especially in stage performance scenes, complex and changeable lighting settings are often required to assist the performance effects. How to achieve consistent relighting video results under challenging dynamic lighting conditions remains a difficult problem to be solved.
[0003] There are many existing methods in the field of video and image relighting:
[0004] Traditional relighting methods require dense sampling of illumination through multi-cameras, and then quantify the illumination and remap it to the image to be illuminated through methods such as temporal differentiation and integration. This method not only requires powerful hardware support, but also requires considerable computational effort. There is much room for improvement in terms of algorithm efficiency and generated image quality. At present, deep learning-based methods have realized end-to-end relighting operations, predicting normal maps and albedo images through neural networks, and finally performing synthetic rendering of relighted images. Google proposed a method to simulate complex light transmission by building an implicit reflection model. However, the prediction of the albedo image of human clothing is not accurate enough, and there is no clear temporal consistency modeling, which will lead to certain temporal instability. Tsinghua University proposed to embed facial albedo, geometric structure, specular reflection and shadow by explicitly modeling multiple reflection channels, but did not take into account complex light transmission effects such as global illumination and subsurface scattering. Moreover, this method only experiments with facial relighting, and full-body relighting is still a difficult problem that needs to be solved; ShanghaiTech University proposed a video portrait relighting solution, which achieves real-time face relighting through adversarial training, but due to the lack of facial geometry information, this method leads to certain artifacts and false face phenomena in the relighting results, and lacks realism.
[0005] In summary, existing relighting methods all have limitations and cannot achieve consistent relighting video results under challenging dynamic lighting conditions.
[0006] Application Contents
[0007] The present application aims to solve one of the technical problems in the related art at least to some extent.
[0008] To this end, the purpose of this application is to achieve a consistent relighting video effect under dynamic lighting conditions, and a full-body portrait video relighting method based on convolutional neural network is proposed.
[0009] Another object of the present application is to propose a full-body portrait video re-lighting device based on a convolutional neural network.
[0010] According to a first aspect of an embodiment of the present application, a full-body portrait video relighting method based on a convolutional neural network is proposed, comprising the following steps:
[0011] Acquire a video image to be processed, wherein the video to be processed includes a full-body portrait video image;
[0012] Inputting a plurality of image frames and a target lighting scene of the video image to be processed into a pre-trained image processing model to obtain a rendered image frame sequence, wherein the image processing model is used to render the image frames and the target lighting scene into the rendered image frames under the target lighting scene, and perform temporal consistency processing on the rendered image frames;
[0013] The sequence of rendered image frames is synthesized into a re-illuminated video image.
[0014] In some possible embodiments, the pre-trained image processing model includes a first convolutional neural network and a second convolutional neural network, and the inputting of the plurality of image frames of the video image to be processed and the target lighting scene into the pre-trained image processing model to obtain a rendered image frame sequence includes:
[0015] Inputting each of the image frames into the first convolutional neural network frame by frame for de-illumination processing, so as to obtain a portrait albedo image and a normal map under a standard lighting scene;
[0016] Inputting the portrait albedo images and normal maps of a plurality of adjacent image frames and the target lighting environment map into a plurality of second convolutional neural networks at the same time, to obtain a plurality of synthesized re-illuminated frames, wherein the plurality of second convolutional neural networks encode temporal consistency through an inter-frame attention mechanism;
[0017] The background image in the image frame and the multiple re-illuminated frames are synthesized to generate the rendered image frame sequence.
[0018] In some possible embodiments, before inputting each of the image frames into the first convolutional neural network frame by frame for de-illumination processing, the method further includes:
[0019] Acquire training data, wherein the training data includes video image data under different lighting scenarios;
[0020] Inputting the video image data under two different lighting scenes into a pre-constructed first initial convolutional neural network respectively, and obtaining albedo images and normal maps corresponding to the video image data under the two different lighting scenes respectively;
[0021] Calculating the Euclidean space distances of the albedo image and the normal map corresponding to the video image data under the two different lighting scenes, and the feature space distance of the convolutional layer feature map of the first initial convolutional neural network;
[0022] The Euclidean space distance between the albedo image and the normal map and the feature space distance are used together as a loss function to train the network, thereby obtaining the first convolutional neural network after training.
[0023] In some possible embodiments, before simultaneously inputting the portrait albedo images and normal maps of a plurality of adjacent image frames and the target lighting environment map into a plurality of second convolutional neural networks to obtain a plurality of synthesized re-illuminated frames, the method further includes:
[0024] Performing temporal consistency encoding on a plurality of pre-constructed second initial convolutional neural networks according to the inter-frame attention mechanism;
[0025] Inputting the portrait albedo images and normal maps of the plurality of adjacent image frames corresponding to any one of the two lighting scenes into the plurality of second initial convolutional neural networks simultaneously and outputting the plurality of re-illuminated frames;
[0026] A plurality of the second initial convolutional neural networks are trained based on the plurality of re-illuminated frames and the target illumination environment maps of the plurality of adjacent image frames to obtain a plurality of the trained second convolutional neural networks.
[0027] According to a second aspect of an embodiment of the present application, a full-body portrait video relighting device based on a convolutional neural network is proposed, comprising:
[0028] A first acquisition module is used to acquire a video image to be processed, wherein the video to be processed includes a full-body portrait video image;
[0029] A rendering module, used for inputting a plurality of image frames of the video image to be processed and a target lighting scene into a pre-trained image processing model to obtain a sequence of rendered image frames, wherein the image processing model is used for rendering the image frames and the target lighting scene into the rendered image frames under the target lighting scene, and performing temporal consistency processing on the rendered image frames;
[0030] A synthesis module is used to synthesize the rendered image frame sequence into a re-illuminated video image.
[0031] In some possible embodiments, the rendering module includes:
[0032] A first input unit, configured to input each of the image frames into the first convolutional neural network frame by frame for de-illumination processing, so as to obtain a portrait albedo image and a normal map under a standard lighting scene;
[0033] A second input unit is used to input the portrait albedo images and normal maps of a plurality of adjacent image frames and the target lighting environment map into a plurality of the second convolutional neural networks at the same time, so as to obtain a plurality of synthesized re-illuminated frames, wherein the plurality of second convolutional neural networks encode temporal consistency through an inter-frame attention mechanism;
[0034] A synthesis unit is used to synthesize the background image in the image frame and the multiple re-illumination frames to generate the rendered image frame sequence.
[0035] In some possible embodiments, the device further includes:
[0036] A second acquisition module, used to acquire training data, wherein the training data includes video image data under different lighting scenes;
[0037] A first input module is used to input the video image data under two different lighting scenes into a pre-constructed first initial convolutional neural network, respectively, to obtain an albedo image and a normal map corresponding to the video image data under the two different lighting scenes;
[0038] A calculation module, used to calculate the Euclidean space distance of the albedo image and the normal map corresponding to the video image data under the two different lighting scenes, and the feature space distance of the convolutional layer feature map of the first initial convolutional neural network;
[0039] The first training module is used to use the Euclidean space distance between the albedo image and the normal map and the feature space distance as a loss function to train the network and obtain the first convolutional neural network after training.
[0040] In some possible embodiments, the device further includes:
[0041] An encoding module, used for performing temporal consistency encoding on a plurality of pre-constructed second initial convolutional neural networks according to the inter-frame attention mechanism;
[0042] A second input module is used to input the portrait albedo images and normal maps of the plurality of adjacent image frames corresponding to any one of the two lighting scenes into the plurality of second initial convolutional neural networks at the same time, and output the plurality of re-illuminated frames;
[0043] The second training module is used to train the plurality of second initial convolutional neural networks based on the plurality of re-illumination frames and the target lighting environment maps of the plurality of adjacent image frames to obtain the plurality of trained second convolutional neural networks.
[0044] According to a third aspect of an embodiment of the present application, there is provided an electronic device, including:
[0045] processor;
[0046] a memory for storing instructions executable by the processor;
[0047] The processor is configured to execute the instructions to implement the full-body portrait video relighting method based on a convolutional neural network as described in any one of the first aspects.
[0048] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the full-body portrait video relighting method based on a convolutional neural network as described in any one of the first aspects.
[0049] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the full-body portrait video relighting method based on a convolutional neural network as described in any one of the first aspects.
[0050] Beneficial effects of this application:
[0051] According to a method for relighting a full-body portrait video based on a convolutional neural network in an embodiment of the present application, a video image to be processed is obtained, wherein the video to be processed includes a full-body portrait video image, and multiple image frames and a target lighting scene of the video image to be processed are input into a pre-trained image processing model to obtain a rendered image frame sequence, wherein the image processing model is used to render the image frames and the target lighting scene into rendered image frames under the target lighting scene, and perform temporal consistency processing on the rendered image frames, and synthesize the rendered image frame sequence into a relighted video image. The present application can achieve relighting of full-body portraits and improve the relighting effect.
[0052] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0054] Figure 1 is a flow chart of a full-body portrait video relighting method based on a convolutional neural network according to an embodiment of the present application;
[0055] Figure 2 Schematic diagram of the training process of the first convolutional neural network according to an embodiment of the present application;
[0056] Figure 3 Schematic diagram of the training process of the second convolutional neural network according to an embodiment of the present application;
[0057] Figure 4 is a structural schematic diagram of a full-body portrait video re-lighting device based on a convolutional neural network according to an embodiment of the present application;
[0058] Figure 5 The block diagram is a block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0059] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0060] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0061] Explanation of relevant terms:
[0062] (1) Image (video) generation: Image (video) generation refers to generating a target image (video) based on an input vector. The input vector can be random noise or a conditional vector specified by the user, which is used to guide the machine learning algorithm to output an image (video) that meets the requirements. Specific application scenarios include handwritten digit generation, two-dimensional face synthesis, style transfer, image restoration, image enhancement, etc.
[0063] (2) Image (video) enhancement: In the process of acquiring images (videos), the image (video) quality is degraded due to various factors, such as low brightness, strong noise, poor color, missing details, etc. Image (video) enhancement refers to improving the visual effect of images (videos) through a series of technologies, converting them into a form that is more suitable for computer or human analysis and application. Typical methods include histogram equalization, Retinex algorithm, defogging algorithm, gamma correction, etc.
[0064] (3) Image (video) relighting: Image (video) relighting tasks belong to image (video) enhancement tasks. Given an input image and ambient light, the algorithm regularly changes the brightness of the input image pixel values according to the lighting characteristics, and performs detail processing such as highlights, shadows, and surface scattering, and outputs a new rendered image under the corresponding lighting effect.
[0065] (4) Realistic rendering: Realistic rendering refers to the process of generating realistic graphics images of a three-dimensional scene in a computer, so that the human eye cannot distinguish between the generated image and the image of the same scene taken by a camera. There are three main standards for realism: photorealism, physical correctness, and high performance. The main realistic rendering algorithms include scanline algorithm, ray tracing algorithm, and radiosity algorithm.
[0066] (5) Temporal consistency modeling: Temporal consistency between video frames is one of the basic characteristics of video information. Currently, many video processing algorithms are usually edited frame by frame without fully considering the continuity between frames. When the generated video is played, subtle unnatural oscillations can often be observed in some areas of the screen, which is called temporal inconsistency. Considering that there is strong continuity information between the upper and lower frames, an algorithm can be designed for this continuity, which is called temporal consistency modeling.
[0067] The following describes a full-body portrait video relighting method and device based on a convolutional neural network proposed in an embodiment of the present application with reference to the accompanying drawings. First, the full-body portrait video relighting method based on a convolutional neural network proposed in an embodiment of the present application will be described with reference to the accompanying drawings.
[0068] Figure 1 This is a flowchart of a full-body portrait video relighting method based on a convolutional neural network according to an embodiment of the present application.
[0069] like Figure 1 As shown, the full-body portrait video relighting method based on convolutional neural network includes the following steps:
[0070] Step S110, obtaining a video image to be processed.
[0071] The video to be processed includes a full-body portrait video image, and the video to be processed may be a video image that needs to be re-illuminated.
[0072] In the embodiment of the present application, the video images to be processed may be full-body portrait video images, and the meaning of being processed may be re-illumination of these full-body portrait video images.
[0073] It should be noted that the video image to be processed may be in the form of a segment of video, or may be a video frame image decomposed from a segment of video.
[0074] Step S120 , inputting a plurality of image frames of the video image to be processed and the target lighting scene into a pre-trained image processing model to obtain a rendered image frame sequence.
[0075] Among them, the image processing model is used to render the image frame and the target lighting scene into a rendered image frame under the target lighting scene, and perform time consistency processing on the rendered image frame. The target lighting scene can be a re-lighting scene that needs to be achieved for processing the video image to be processed.
[0076] In the embodiment of the present application, after obtaining the video to be processed, multiple image frames of the video image to be processed and the target lighting scene can be input into a pre-trained image processing model to obtain a rendered image frame sequence. In other words, the image processing model is pre-trained, and a re-illumination process based on time consistency can be performed on multiple image frames of the video image to be processed according to the target lighting scene to obtain a rendered image frame sequence after the re-illumination process corresponding to the multiple image frames of the video image to be processed.
[0077] Step S130: synthesize the rendered image frame sequence into a re-illuminated video image.
[0078] In an embodiment of the present application, after a plurality of image frames of a video image to be processed and a target lighting scene are input into a pre-trained image processing model to obtain a rendered image frame sequence, the rendered image frame sequence can be synthesized into a re-illuminated video image. That is, after obtaining the rendered image frame sequence, the image re-illuminating process has been completed, and the image frame sequence after the re-illuminating process can be synthesized into a re-illuminated video image in sequence order.
[0079] In some possible embodiments, the pre-trained image processing model includes a first convolutional neural network and a second convolutional neural network, and a plurality of image frames of a video image to be processed and a target lighting scene are input into the pre-trained image processing model to obtain a rendered image frame sequence, including:
[0080] Input each image frame in the image frame into the first convolutional neural network for de-illumination processing, and obtain the portrait albedo image and normal map under the standard lighting scene;
[0081] Inputting portrait albedo images and normal maps of multiple adjacent image frames and target lighting environment maps into multiple second convolutional neural networks simultaneously to obtain multiple synthesized re-illuminated frames, wherein the multiple second convolutional neural networks encode temporal consistency through an inter-frame attention mechanism;
[0082] The background image in the image frame and multiple relighting frames are synthesized to generate a rendered image frame sequence.
[0083] Among them, the first convolutional neural network can be an image de-illumination processing network, which is used to process the input image to obtain the portrait albedo image and normal map under the standard lighting scene; the second convolutional neural network can be an image re-illumination processing network, which is used to re-illumination the input portrait albedo image and normal map under the standard lighting scene according to the target lighting environment in the form of multiple adjacent image frames to obtain multiple synthesized re-illumination frames.
[0084] In an embodiment of the present application, after obtaining the video image to be processed, each image frame in the image frame can be input into the first convolutional neural network frame by frame for de-illumination processing to obtain the portrait albedo image and normal map under the standard lighting scene, and then the portrait albedo image and normal map of multiple adjacent image frames and the target lighting environment map can be simultaneously input into multiple second convolutional neural networks to obtain multiple synthesized re-illumination frames, wherein the multiple second convolutional neural networks encode temporal consistency through the inter-frame attention mechanism, and then the background image in the image frame and the multiple re-illumination frames can be synthesized to generate a rendered image frame sequence. In other words, the image processing model includes a first convolutional neural network and a second convolutional neural network, and the first convolutional neural network performs de-illumination processing on each image frame frame by frame to obtain the portrait albedo image and normal map under the standard lighting scene, and the second convolutional neural network can then perform temporal consistency-based re-illumination on the portrait albedo image and normal map under the standard lighting scene according to the target lighting environment to obtain a re-illumination frame.
[0085] It should be noted that the standard lighting environment may be a lighting environment of white light with average brightness.
[0086] In some possible embodiments, before inputting each of the image frames into the first convolutional neural network frame by frame for de-illumination processing, the method further includes:
[0087] Acquire training data, where the training data includes video image data under different lighting scenarios;
[0088] The video image data under two different lighting scenes are respectively input into a pre-constructed first initial convolutional neural network to obtain albedo images and normal maps corresponding to the video image data under the two different lighting scenes;
[0089] Calculating the Euclidean space distances of the albedo image and the normal map corresponding to the video image data under two different lighting scenes, and the feature space distance of the convolutional layer feature map of the first initial convolutional neural network;
[0090] The Euclidean space distance and feature space distance of the albedo image and the normal map are used together as the loss function to train the network, and the first convolutional neural network after training is obtained.
[0091] In an embodiment of the present application, before inputting each image frame in the image frame into the first convolutional neural network for de-illumination processing frame by frame, the first convolutional neural network may be trained:
[0092] To obtain training data, we can build a lighting stage and adjust the brightness and color of the lights to simulate the complex and changeable ambient lighting in stage performances, including white light and colored light. We can use a monocular RGB-D (Red Green Blue-Depth) camera to shoot full-body portrait static videos of different volunteers in standard lighting environments. The video length can be set as needed, for example, 3 to 5 seconds, to generate color albedo images and depth volume data without light interference, and map the depth information to a normal map. We can then randomly adjust the lighting settings to obtain full-body portrait static videos in non-standard lighting environments.
[0093] The full-body portrait static video under standard lighting environment is used as supervision, and the full-body portrait static video under non-standard lighting environment is used to train the first convolutional neural network. For the full-body portrait static video under non-standard lighting environment of the same volunteer, two groups of full-body portrait static videos with different lighting environments are randomly decomposed into image sequences, and the two groups of image sequences with different lighting environments can be represented by frame A and frame B. Two groups of first initial convolutional neural networks with the same structure can be constructed, which can be represented by network A and network B. Frame A is input into network A, and frame B is input into network B, and the albedo image and normal map corresponding to frame A and frame B are obtained.
[0094] The Euclidean space distance of the albedo image and the normal map corresponding to frame A and frame B, as well as the feature space distance of the convolutional layer feature map corresponding to network A and network B are calculated. The Euclidean space distance and the feature space distance are used as loss functions, and the first initial convolutional neural network is iterated to optimize the parameters to obtain the first convolutional neural network.
[0095] It should be noted that the first convolutional neural network, as a de-illumination network, can remove the impact of random and complex lighting on the color and shadow of the portrait. The encoder of the first convolutional neural network consists of a series of downsampling convolutional layers, which encodes the input image frame into a 1*1 feature vector representation. The decoder of the first convolutional neural network upsamples the feature vector through a series of transposed convolutional layers, restores the feature map layer by layer to the size of the input image frame, and finally outputs the albedo image and the normal map. Among them, the encoder and the decoder are skipped.
[0096] In some possible embodiments, before simultaneously inputting the portrait albedo images and normal maps of a plurality of adjacent image frames and the target lighting environment map into a plurality of second convolutional neural networks to obtain a plurality of synthesized re-illuminated frames, the method further includes:
[0097] Temporal consistency encoding is performed on multiple pre-built second initial convolutional neural networks according to the inter-frame attention mechanism;
[0098] Inputting portrait albedo images and normal maps of a plurality of adjacent image frames corresponding to any one of the two lighting scenes into a plurality of second initial convolutional neural networks at the same time, and outputting a plurality of re-illuminated frames;
[0099] A plurality of second initial convolutional neural networks are trained based on a plurality of re-illuminated frames and a target illumination environment map of a plurality of adjacent image frames to obtain a plurality of trained second convolutional neural networks.
[0100] In the embodiment of the present application, after obtaining the first convolutional neural network, the second convolutional neural network can also be trained based on the first convolutional neural network:
[0101] Construct multiple second initial convolutional neural networks, such as Figure 3 As shown, according to the inter-frame attention mechanism, multiple pre-constructed second initial convolutional neural networks are temporally consistent encoded, and multiple second initial convolutional neural networks can be represented by network 1, network 2, ... network k, where k is a positive integer greater than 1. For the same volunteer, the portrait albedo images and normal maps of multiple adjacent image frames corresponding to any of the two lighting scenes after the first convolutional neural network is trained are simultaneously input into multiple second initial convolutional neural networks. Correspondingly, multiple image frames can also be represented by frame 1, frame 2 ... frame k to obtain re-illuminated frames corresponding to k image frames, and the target lighting environment maps of k adjacent image frames are used as supervision of k networks respectively. The parameters of the second initial convolutional neural network are iteratively optimized according to the re-illuminated frames corresponding to the output k image frames to obtain k second convolutional neural networks.
[0102] It should be noted that the second convolutional neural network, as a synthesis network, outputs the re-illuminated video frame under the target lighting environment, and performs collaborative processing on colors, shadows and highlights to synthesize a highly realistic re-illuminated effect. The encoder of the second convolutional neural network consists of a series of down-sampling convolutional layers, which encode the input albedo image and normal map into a 1*1 feature vector representation, which is called a bottleneck layer. The bottleneck layers of the above k second initial convolutional neural networks are encoded into a 1*128 vector, and the vectors of the k second initial convolutional neural networks are encoded by a temporal attention mechanism to capture the illumination dependency between different frames and perform temporal consistency modeling. The decoder of the second convolutional neural network upsamples the feature vector through a series of transposed convolutional layers, restores the features layer by layer to the size of the input image frame, and finally outputs a synthesized re-illuminated image. Among them, a skip connection is performed between the encoder and the decoder.
[0103] Through the above steps, a video image to be processed is obtained, wherein the video to be processed includes a full-body portrait video image, and multiple image frames and a target lighting scene of the video image to be processed are input into a pre-trained image processing model to obtain a rendered image frame sequence, wherein the image processing model is used to render the image frames and the target lighting scene into rendered image frames under the target lighting scene, and perform temporal consistency processing on the rendered image frames, and synthesize the rendered image frame sequence into a re-illuminated video image. The present application can achieve re-illumination of full-body portraits and improve the re-illumination effect.
[0104] In order to implement the above embodiment, Figure 4 As shown, this embodiment also provides a full-body portrait video re-lighting device 400 based on a convolutional neural network. The device 400 includes: a first acquisition module 410, a rendering module 420, and a synthesis module 430.
[0105] A first acquisition module 410 is used to acquire a video image to be processed, wherein the video to be processed includes a full-body portrait video image;
[0106] A rendering module 420 is used to input multiple image frames of the video image to be processed and the target lighting scene into a pre-trained image processing model to obtain a sequence of rendered image frames, wherein the image processing model is used to render the image frames and the target lighting scene into rendered image frames under the target lighting scene, and perform temporal consistency processing on the rendered image frames;
[0107] The synthesis module 430 is used to synthesize the rendered image frame sequence into a re-illuminated video image.
[0108] In some possible embodiments, the rendering module 420 includes:
[0109] A first input unit is used to input each image frame in the image frames into the first convolutional neural network for de-illumination processing, so as to obtain a portrait albedo image and a normal map under a standard lighting scene;
[0110] A second input unit is used to input the portrait albedo images and normal maps of a plurality of adjacent image frames and the target lighting environment map into a plurality of second convolutional neural networks at the same time to obtain a plurality of synthesized re-illuminated frames, wherein the plurality of second convolutional neural networks encode temporal consistency through an inter-frame attention mechanism;
[0111] The synthesis unit is used to synthesize the background image and multiple re-illumination frames in the image frame to generate a rendered image frame sequence.
[0112] In some possible embodiments, the full-body portrait video re-lighting device 400 based on a convolutional neural network further includes:
[0113] A second acquisition module is used to acquire training data, where the training data includes video image data under different lighting scenes;
[0114] A first input module is used to input the video image data under two different lighting scenes into a pre-constructed first initial convolutional neural network, respectively, to obtain albedo images and normal maps corresponding to the video image data under the two different lighting scenes;
[0115] A calculation module, used to calculate the Euclidean space distance of the albedo image and the normal map corresponding to the video image data under two different lighting scenes, and the feature space distance of the convolution layer feature map of the first initial convolutional neural network;
[0116] The first training module is used to use the Euclidean space distance and the feature space distance of the albedo image and the normal map as loss functions to train the network, thereby obtaining a first convolutional neural network after training.
[0117] In some possible embodiments, the full-body portrait video relighting 400 based on a convolutional neural network further includes:
[0118] An encoding module, used for temporally consistent encoding of a plurality of pre-built second initial convolutional neural networks according to an inter-frame attention mechanism;
[0119] A second input module is used to input the portrait albedo images and normal maps of a plurality of adjacent image frames corresponding to any one of the two lighting scenes into a plurality of second initial convolutional neural networks at the same time, and output a plurality of re-illuminated frames;
[0120] The second training module is used to train multiple second initial convolutional neural networks based on multiple re-illumination frames and target lighting environment maps of multiple adjacent image frames to obtain multiple trained second convolutional neural networks.
[0121] According to the full-body portrait video relighting device based on a convolutional neural network in the embodiment of the present application, by obtaining a video image to be processed, wherein the video to be processed includes a full-body portrait video image, multiple image frames of the video image to be processed and a target lighting scene are input into a pre-trained image processing model to obtain a rendered image frame sequence, wherein the image processing model is used to render the image frame and the target lighting scene into a rendered image frame under the target lighting scene, and perform temporal consistency processing on the rendered image frame, and synthesize the rendered image frame sequence into a relighted video image. The present application can achieve relighting of full-body portraits and improve the relighting effect.
[0122] It should be noted that the aforementioned explanation of the embodiment of the full-body portrait video relighting method based on a convolutional neural network is also applicable to the full-body portrait video relighting device based on a convolutional neural network in this embodiment, and will not be repeated here.
[0123] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a computer-readable storage medium, and a computer program product.
[0124] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device 500 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0125] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0126] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0127] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as a full-body portrait video relighting method based on a convolutional neural network. For example, in some embodiments, a full-body portrait video relighting method based on a convolutional neural network may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the full-body portrait video relighting method based on a convolutional neural network described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform a full-body portrait video relighting method based on a convolutional neural network in any other appropriate manner (e.g., by means of firmware).
[0128] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0129] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0130] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0131] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0132] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0133] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0134] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0135] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0136] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A full-body portrait video relighting method based on convolutional neural network, It is characterized in that include: Acquire a video image to be processed, wherein the video to be processed includes a full-body portrait video image; Inputting a plurality of image frames and a target lighting scene of the video image to be processed into a pre-trained image processing model to obtain a rendered image frame sequence, wherein the image processing model is used to render the image frames and the target lighting scene into the rendered image frames under the target lighting scene, and perform temporal consistency processing on the rendered image frames; synthesizing the rendered image frame sequence into a re-illuminated video image; The pre-trained image processing model includes a first convolutional neural network and a second convolutional neural network, and the multiple image frames of the video image to be processed and the target lighting scene are input into the pre-trained image processing model to obtain a rendered image frame sequence, including: Inputting each of the image frames into the first convolutional neural network frame by frame for de-illumination processing, so as to obtain a portrait albedo image and a normal map under a standard lighting scene; Inputting the portrait albedo images and normal maps of a plurality of adjacent image frames and the target lighting environment map into a plurality of second convolutional neural networks at the same time, to obtain a plurality of synthesized re-illuminated frames, wherein the plurality of second convolutional neural networks encode temporal consistency through an inter-frame attention mechanism; The background image in the image frame and the multiple re-illuminated frames are synthesized to generate the rendered image frame sequence.
2. The method according to claim 1, It is characterized in that Before inputting each of the image frames into the first convolutional neural network frame by frame for de-illumination processing, the method further includes: Acquire training data, wherein the training data includes video image data under different lighting scenarios; Inputting the video image data under two different lighting scenes into a pre-constructed first initial convolutional neural network respectively, and obtaining albedo images and normal maps corresponding to the video image data under the two different lighting scenes respectively; Calculating the Euclidean space distances of the albedo image and the normal map corresponding to the video image data under the two different lighting scenes, and the feature space distance of the convolutional layer feature map of the first initial convolutional neural network; The Euclidean space distance between the albedo image and the normal map and the feature space distance are used together as a loss function to train the network, thereby obtaining the first convolutional neural network after training.
3. The method according to claim 2, It is characterized in that Before simultaneously inputting the portrait albedo images and normal maps of the adjacent plurality of image frames and the target lighting environment map into the plurality of second convolutional neural networks to obtain the synthesized plurality of re-illuminated frames, the method further includes: Performing temporal consistency encoding on a plurality of pre-constructed second initial convolutional neural networks according to the inter-frame attention mechanism; Inputting the portrait albedo images and normal maps of the plurality of adjacent image frames corresponding to any one of the two lighting scenes into the plurality of second initial convolutional neural networks simultaneously and outputting the plurality of re-illuminated frames; A plurality of the second initial convolutional neural networks are trained based on the plurality of re-illuminated frames and the target illumination environment maps of the plurality of adjacent image frames to obtain a plurality of the trained second convolutional neural networks.
4. A full-body portrait video relighting device based on convolutional neural network, It is characterized in that include: A first acquisition module is used to acquire a video image to be processed, wherein the video to be processed includes a full-body portrait video image; A rendering module, used for inputting a plurality of image frames of the video image to be processed and a target lighting scene into a pre-trained image processing model to obtain a sequence of rendered image frames, wherein the image processing model is used for rendering the image frames and the target lighting scene into the rendered image frames under the target lighting scene, and performing temporal consistency processing on the rendered image frames; A synthesis module, used for synthesizing the rendered image frame sequence into a re-illuminated video image; The pre-trained image processing model includes a first convolutional neural network and a second convolutional neural network, and the rendering module includes: A first input unit, configured to input each of the image frames into the first convolutional neural network frame by frame for de-illumination processing, so as to obtain a portrait albedo image and a normal map under a standard lighting scene; A second input unit is used to input the portrait albedo images and normal maps of a plurality of adjacent image frames and the target lighting environment map into a plurality of the second convolutional neural networks at the same time, so as to obtain a plurality of synthesized re-illuminated frames, wherein the plurality of second convolutional neural networks encode temporal consistency through an inter-frame attention mechanism; A synthesis unit is used to synthesize the background image in the image frame and the multiple re-illumination frames to generate the rendered image frame sequence.
5. The device according to claim 4, It is characterized in that The device further comprises: A second acquisition module, used to acquire training data, wherein the training data includes video image data under different lighting scenes; A first input module is used to input the video image data under two different lighting scenes into a pre-constructed first initial convolutional neural network, respectively, to obtain an albedo image and a normal map corresponding to the video image data under the two different lighting scenes; A calculation module, used to calculate the Euclidean space distance of the albedo image and the normal map corresponding to the video image data under the two different lighting scenes, and the feature space distance of the convolutional layer feature map of the first initial convolutional neural network; The first training module is used to use the Euclidean space distance between the albedo image and the normal map and the feature space distance as a loss function to train the network and obtain the first convolutional neural network after training.
6. The method according to claim 5, It is characterized in that The device further comprises: An encoding module, used for performing temporal consistency encoding on a plurality of pre-constructed second initial convolutional neural networks according to the inter-frame attention mechanism; A second input module is used to input the portrait albedo images and normal maps of the plurality of adjacent image frames corresponding to any one of the two lighting scenes into the plurality of second initial convolutional neural networks at the same time, and output the plurality of re-illuminated frames; The second training module is used to train the plurality of second initial convolutional neural networks based on the plurality of re-illumination frames and the target lighting environment maps of the plurality of adjacent image frames to obtain the plurality of trained second convolutional neural networks.
7. An electronic device, It is characterized in that include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the full-body portrait video relighting method based on a convolutional neural network as described in any one of claims 1 to 3.
8. A computer-readable storage medium, It is characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the full-body portrait video relighting method based on a convolutional neural network as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Human body scene image eigen-decomposition and relighting method and device
CN113240622A
Method and device for generating relighting image and electronic equipment
CN113554739A