Multi-modal portrait video editing method, electronic device, and storage medium

By constructing a three-dimensional dynamic Gaussian feature field and a convolutional neural renderer, combined with an expression similarity loss function and a facial perception model, the problems of low quality and inter-frame inconsistency in two-dimensional portrait editing are solved, achieving efficient and high-quality multimodal portrait video editing.

WO2026000851A1PCT designated stage Publication Date: 2026-01-02UNIV OF SCI & TECH OF CHINA

Patent Information

Application Number
PCT/CN2024/138411
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2024-12-11
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing 2D portrait editing methods produce portraits of low quality and inconsistent frame timing. Pre-trained image diffusion models cannot generate satisfactory video results and have long computation times.

Method used

By embedding Gaussian points in a 3D Gaussian field into a 2D texture space of a parameterized human geometric representation, a 3D dynamic Gaussian feature field is constructed. A convolutional neural renderer is used for 3D reconstruction and optimization. By combining the expression similarity loss function and the reconstruction loss function, video frames and 3D portraits are edited alternately and iteratively. A face-aware portrait editing model is used to enhance facial region editing.

Benefits of technology

It enables high-quality, multimodal portrait video editing, ensuring the three-dimensional and temporal consistency of portraits in the video, improving the quality and efficiency of editing results, and avoiding the problem of facial expression degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024138411_02012026_PF_FP_ABST
    Figure CN2024138411_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a multi-modal portrait video editing method, which can be applied to the technical field of video editing. The method comprises: given a portrait video, preprocessing same to obtain camera parameters, a human identity coefficient, a human expression coefficient, a human pose coefficient, a human semantic segmentation map, and a two-dimensional portrait mask; on the basis of a neural Gaussian texture mechanism, embedding a learnable three-dimensional Gaussian feature into a parameterized human geometric surface, using a neural renderer to convert a three-dimensional Gaussian splatting feature map into an image, and optimizing the reconstruction of a three-dimensional portrait on the basis of RGB and segmentation map information of the video; and using an iterative dataset update technique to distill knowledge of a multi-modal two-dimensional image generation model into three-dimensional portrait editing, and using expression similarity guidance and a face-aware portrait editing model to improve editing quality. The method elevates a two-dimensional editing task to three-dimensional space, and ensures good three-dimensional consistency and temporal consistency. By means of the knowledge of the multi-modal generation model, high-quality portrait video editing functions can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

A multi-modal portrait video editing method, electronic device and storage medium TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to the technical field of portrait video editing, and more particularly to a multi-modal portrait video editing method, an electronic device and a storage medium. BACKGROUND

[0002] Portrait video editing has a wide range of applications in the fields of movies, arts and augmented reality / virtual reality. How to ensure the structural similarity and temporal consistency of the edited video while achieving high-quality, multi-modal editing effects has always been an important problem in the field of image processing technology.

[0003] Early two-dimensional portrait editing methods are based on generative adversarial networks (GANs). However, two-dimensional portrait editing methods based on GANs are limited by the representation ability of GAN models, and the generated portraits often have low quality. Other two-dimensional portrait editing methods based on denoising diffusion models (Denoising Diffusion Probabilistic Models) have better performance than two-dimensional portrait editing methods based on GANs, but the pictures generated by the above two-dimensional portrait editing methods based on denoising diffusion models often have temporal inconsistency problems between video frames when applied to video editing.

[0004] To solve the problems of low quality and temporal inconsistency between frames of the generated portraits in the above two-dimensional portrait editing methods, the skilled person in the art uses a pre-trained image diffusion model for training-free video editing. However, due to the lack of 3D priori and face or body priori, the method of using a pre-trained image diffusion model for training-free video editing may not be able to generate satisfactory video results in terms of quality and temporal consistency. At the same time, due to the complex denoising process of the denoising diffusion model, this method takes several minutes of computing time to generate only 1 second of video segment. SUMMARY

[0005] In view of the above problems, the present application provides a multi-modal portrait video editing method, an electronic device and a storage medium.

[0006] According to a first aspect of the present application, there is provided a multi-modal portrait video editing method, comprising:

[0007] The frame images of the portrait video are subjected to a dimensionality increasing operation by embedding Gaussian points in a three-dimensional Gaussian field into a two-dimensional texture space of a parameterized human body geometry representation, to obtain a three-dimensional dynamic Gaussian feature field;

[0008] The three-dimensional Gaussian dynamic feature field is subjected to three-dimensional Gaussian splash to obtain three-dimensional Gaussian rendering features, and the three-dimensional Gaussian rendering features are converted by using a convolutional neural renderer to obtain a neural rendering RGB image;

[0009] The target person in the portrait video is reconstructed in three dimensions by using the three-dimensional dynamic Gaussian feature field, and the three-dimensional portrait reconstruction process is optimized by using the neural rendering RGB image and the original RGB image of the portrait video through the loss function in the reconstruction stage.

[0010] The editing information of the neural rendering RGB image is distilled to the optimized three-dimensional portrait reconstruction result to complete the updating of the three-dimensional portrait, and the edited neural rendering RGB image is used to replace the original frame image of the portrait video to complete the editing of the video frame dataset.

[0011] The editing process of the video frame dataset and the updating process of the three-dimensional portrait are alternately iterated, and the edited neural rendering RGB image and the original frame image of the portrait video are mapped to the same expression space.

[0012] The convolutional neural renderer is optimized by using the expression similarity loss function and the loss function in the reconstruction stage, and the dynamic three-dimensional Gaussian feature field is optimized again until the editing of all video frame images in the portrait video is completed.

[0013] According to the embodiment of the present application, the above-mentioned dimension increasing operation on the frame image of the portrait video by embedding the Gaussian points in the three-dimensional Gaussian field into the two-dimensional texture space of the parameterized human body geometry representation includes:

[0014] An initial parameterized human body geometry representation is constructed by using the portrait identity coefficient, the portrait posture coefficient and the expression coefficient of the target person obtained from each frame image of the portrait video.

[0015] The difference between the initial parameterized human body geometry representation and the real human body of the target person is corrected by using the learnable vertex offset to obtain a corrected parameterized human body geometry representation.

[0016] Based on the neural Gaussian texture mechanism, the Gaussian points in the three-dimensional Gaussian field mapped in the two-dimensional texture space are uniformly embedded on the surface of the corrected parameterized human body geometry representation to obtain a three-dimensional dynamic Gaussian feature field.

[0017] According to the embodiment of the present application, the above-mentioned construction of the initial parameterized human body geometry representation by using the portrait identity coefficient, the portrait posture coefficient and the expression coefficient of the target person obtained from each frame image of the portrait video includes:

[0018] The image identity coefficient, the image posture coefficient and the expression coefficient of the target person in each frame image in the portrait video are obtained by tracking the portrait in each frame image in the portrait video.

[0019] Based on the image identity coefficient, the image posture coefficient and the expression coefficient of the target person, an initial parameterized human body geometry representation is constructed by using a predefined skin function and skin weight, a predefined joint matrix and human body geometry in a T-shaped posture.

[0020] According to the embodiment of the present application, the three-dimensional Gaussian rendering feature is obtained by performing three-dimensional Gaussian splashing on the three-dimensional dynamic Gaussian feature field, and the neural rendering RGB image is obtained by converting the three-dimensional Gaussian rendering feature by using the convolutional neural renderer.

[0021] The three-dimensional Gaussian rendering feature is obtained by performing three-dimensional Gaussian splashing on the three-dimensional dynamic Gaussian feature field by using the camera intrinsic parameter and the camera pose of each frame image in the portrait video and the attribute information of the three-dimensional dynamic Gaussian feature field.

[0022] The convolutional neural renderer including an input layer, a convolutional layer, a nonlinear layer, a pooling layer, a fully connected layer and a loss layer is constructed, and the neural rendering RGB image is obtained by converting the three-dimensional Gaussian rendering feature by using the convolutional neural renderer.

[0023] According to the embodiment of the present application, the attribute information of the three-dimensional dynamic Gaussian feature field includes spatial position information of Gaussian points, a quaternion representing a three-dimensional orientation of the Gaussian field, a scaling vector, an opacity and a learnable feature vector.

[0024] According to the embodiment of the present application, the three-dimensional reconstruction of the target person in the portrait video is performed by using the three-dimensional dynamic Gaussian feature field, and the three-dimensional portrait reconstruction process is optimized by using the neural rendering RGB image and the original RGB image of the portrait video through the loss function in the reconstruction stage, which includes:

[0025] The two-dimensional portrait mask of each frame image in the portrait video is obtained by using a predefined image processing model to pre-process the portrait video.

[0026] The reconstruction loss function, the mask loss function and the stability loss function are used to process the neural rendering RGB image, the original RGB image and the two-dimensional portrait mask, so as to obtain the loss value in the reconstruction stage.

[0027] The three-dimensional portrait reconstruction process is supervised and trained based on the loss value in the reconstruction stage, so as to obtain the optimized three-dimensional portrait reconstruction result.

[0028] According to the embodiment of the present application, the editing information of the neural rendering RGB image is distilled to the optimized three-dimensional portrait reconstruction result to complete the update of the three-dimensional portrait, which includes:

[0029] The portrait video is preprocessed by using a predefined image processing model to obtain a human body semantic segmentation map of each frame image in the portrait video, wherein the human body semantic segmentation map includes a head region and a torso part of a target person;

[0030] Based on the multi-modal editing instruction input by the user, a two-dimensional multi-modal generation module of the pre-trained face-aware portrait editing model is used to edit the head region and the torso part in the neural rendering RGB image respectively;

[0031] The edited head region and the edited torso part are synthesized by using the human body semantic segmentation map of each frame image, and the synthesized image is inserted into the video frame data set to further update the video frame data set;

[0032] The optimization of the three-dimensional dynamic high-strength characteristic field is completed, and the editing information is distilled to the optimized three-dimensional portrait reconstruction result by the face-aware portrait editing model to further update the three-dimensional portrait.

[0033] The updating operation of the three-dimensional portrait and the updating operation of the video frame data set are alternately iterated to continuously optimize the three-dimensional dynamic high-strength characteristic field.

[0034] According to the embodiment of the present application, the two-dimensional multi-modal generation module includes a two-dimensional multi-modal generation module based on text editing, a two-dimensional multi-modal generation module based on image editing, and a two-dimensional multi-modal image generation model of portrait relighting.

[0035] The second aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the multi-modal portrait video editing method.

[0036] The third aspect of the present application also provides a computer readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions are executed by a processor to implement the steps of the multi-modal portrait video editing method.

[0037] The multi-modal portrait video editing method provided by the application upgrades the portrait in the video, uses dynamic three-dimensional Gaussian feature field to represent the person in the video by means of parameterized human geometry, and effectively ensures the three-dimensional consistency and time consistency of the portrait in the video; unlike the original three-dimensional Gaussian splash represented by spherical harmonic coefficients, the feature is represented by self-adaptive learning, so that the model capacity is higher; the convolutional neural renderer of the application can effectively combine the three-dimensional Gaussian splash feature map in depth, further improve the expression ability of the portrait representation, so that the edited result is more in line with the editing instruction, and at the same time, higher quality is shown, especially good at fitting non-three-dimensional visual style effect; the face-aware portrait editing model of the application can make the editing process pay more attention to the face area, and significantly enhance the editing quality of the face area; at the same time, the expression similarity guide term of the application can ensure the correctness of the expression of the three-dimensional portrait, and avoid the expression degradation problem. BRIEF DESCRIPTION OF DRAWINGS

[0038] The above and other objects, features and advantages of the present application will become more apparent from the following description of embodiments of the present application taken in conjunction with the accompanying drawings, in which:

[0039] Fig. 1 is a flowchart of a multi-modal portrait video editing method according to an embodiment of the application;

[0040] Fig. 2 is a pipeline diagram of a portrait video editing method based on three-dimensional dynamic Gaussian feature field and multi-modal generative prior according to an embodiment of the application;

[0041] Fig. 3 is a visual result comparison diagram based on text editing according to an embodiment of the application;

[0042] Fig. 4 is a visual result comparison diagram based on picture editing according to an embodiment of the application;

[0043] Fig. 5 is a visual result comparison diagram based on relighting according to an embodiment of the application;

[0044] Fig. 6 schematically shows a block diagram of an electronic device suitable for implementing the multi-modal portrait video editing method according to an embodiment of the application. DETAILED DESCRIPTION

[0045] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary, and are not intended to limit the scope of the present application. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the present application. However, it will be apparent to one skilled in the art that one or more embodiments can be practiced without these specific details. In other instances, well-known structures and techniques have been omitted in order to avoid obscuring the concepts of the present application.

[0046] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the terms "comprises", "comprising", "includes", "including" and the like are specifically intended to be open-ended and to mean that other features, steps, operations, and / or components can be added.

[0047] All terms used herein including technical and scientific terms have the meanings commonly understood by one of ordinary skill in the art unless otherwise defined. It should be noted that the terms used herein are defined as consistent with the context where used, and should not be construed as ideal or overly formal unless expressly so defined.

[0048] In the case of using expressions similar to "at least one of A, B, and C, etc.", it should generally be interpreted to include at least one of each of the items, unless otherwise defined, for example, "a system having at least one of A, B, and C" should be interpreted to include systems having at least one of A, B, or C individually, systems having at least one of A and B, systems having at least one of A and C, systems having at least one of B and C, and / or systems having at least one of A, B, and C, etc.

[0049] In the technical solutions of the present application, the user information (including but not limited to user personal information, user image information, user equipment information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved are information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user to choose authorization or refusal.

[0050] Portrait video editing has a wide range of application scenarios and is the focus of research by those skilled in the art. Early two-dimensional portrait editing technical solutions mainly use generative adversarial networks (GAN) to edit portraits or generate portrait animations, and specify the style required for editing by giving a style label or a reference picture. A contrastive language-image pre-training (CLIP) model builds a bridge between text and images, and under the guidance of CLIP similarity, a GAN model can generate pictures that meet the text description. However, such work is limited by the representation ability of the GAN model, and the quality of the generated portraits is often not high. Recently, denoising diffusion probabilistic models (DDPM) have shown better performance than GANs in terms of generation ability. Based on this paradigm, a series of high-quality stylized portrait image generation models, fine-tuning methods, and adapters have emerged. However, directly using these image generation models for video editing often makes it difficult to maintain temporal consistency between frames.

[0051] To further improve the continuity of the edited video, some works choose to use pre-trained image diffusion models for training-free video editing. They constrain the motion of the edited video by using dense correspondences, or enhance the perception of the structural information of the input video through cross-frame attention or feature injection. Some works also attempt to connect frames in the time dimension and train a temporal attention module to ensure temporal or multi-view consistency. However, due to the lack of 3D priors and face / body priors, they may not be able to generate satisfactory video results in terms of quality and temporal consistency. At the same time, due to the complex denoising process of denoising diffusion models, these methods require several minutes of computation time to generate only 1 second of video footage.

[0052] To solve the technical problems existing in the prior art, the present application provides a high-efficiency and high-quality multi-modal portrait video editing method, which raises the portrait video editing problem to three dimensions, uses dynamic three-dimensional Gaussian field to ensure three-dimensional consistency and temporal consistency of the video, distills the knowledge of existing two-dimensional multi-modal generation modules to three-dimensional portraits to achieve multi-modal and high-quality portrait editing, and uses three-dimensional Gaussian splashing technology to render the three-dimensional Gaussian portrait to obtain the final video editing result. With the help of three-dimensional Gaussian field construction and the efficiency of three-dimensional Gaussian splashing, the technical solution provided by the present application can realize efficient portrait video editing and high-frame-rate portrait video rendering.

[0053] FIG. 1 is a flowchart of a multi-modal portrait video editing method according to an embodiment of the present application.

[0054] As shown in FIG. 1, the multi-modal portrait video editing method includes operations S110-S160.

[0055] In operation S110, a frame image of a portrait video is subjected to a dimensionality operation by embedding a Gaussian point in a three-dimensional Gaussian field into a two-dimensional texture space of a parameterized human body geometry, to obtain a three-dimensional dynamic Gaussian feature field.

[0056] The dimensionality of a portrait in a portrait video (or video, the same below) is increased, and a dynamic three-dimensional Gaussian feature field is used to represent a person in the video by means of a parameterized human body geometry. The present application proposes a mechanism of neural Gaussian texture, which embeds a Gaussian point into a parameterized human body surface The Gaussian attributes include a Gaussian point position X0, a quaternion Q representing a three-dimensional orientation of the Gaussian, a scaling vector S, an opacity O, and a learnable feature F. A uniform neural Gaussian field φ is present in the UV space (i.e., U-VEEZ, a two-dimensional texture coordinate with vertex component information of a polygon and a subdivision surface mesh) of the parameterized human body geometry, and based on the neural Gaussian field φ of the human body geometry and the UV space, a three-dimensional Gaussian feature field is obtained.

[0057] In operation S120, a three-dimensional Gaussian splashing is performed on the three-dimensional dynamic Gaussian feature field to obtain a three-dimensional Gaussian rendering feature, and a convolutional neural renderer is used to convert the three-dimensional Gaussian rendering feature to obtain a neural rendering RGB image.

[0058] Given a camera intrinsic K, a camera pose P, and the above three-dimensional Gaussian feature field, a feature map I F is rendered. Then a convolutional neural network U is used as a neural renderer to convert it to an RGB image.

[0059] In operation S130, a three-dimensional reconstruction is performed on a target person in the portrait video using the three-dimensional dynamic Gaussian feature field, and a neural rendering RGB image and an original RGB image of the portrait video are used to optimize the three-dimensional portrait reconstruction process through a loss function in the reconstruction stage.

[0060] A three-dimensional reconstruction is performed on a dynamic portrait in a video. Given a portrait video, a camera intrinsic K and a portrait identity coefficient β are obtained by tracking the portrait, and a camera pose P, a portrait pose θ, an expression coefficient ψ, a human body semantic segmentation map, and a two-dimensional portrait mask of each frame are also obtained. Given the human body parameters β, θ, ψ of each frame, the camera parameters K and P, a three-dimensional portrait RGB picture is rendered, and the three-dimensional portrait is optimized by the RGB and segmentation map information of the video.

[0061] In operation S140, the editing information of the neural rendering RGB image is distilled to the optimized three-dimensional portrait reconstruction result to complete the update of the three-dimensional portrait, and the edited neural rendering RGB image is used to replace the original frame image of the portrait video to complete the editing of the video frame dataset.

[0062] In operation S150, the editing process of the video frame dataset and the updating process of the three-dimensional portrait are alternately iterated, and the edited neural rendering RGB image and the original frame image of the portrait video are mapped to the same expression space.

[0063] Using the iterative dataset updating technology, the dataset video frame editing and three-dimensional portrait updating are alternately performed, and the editing of the two-dimensional multi-modal generation module is distilled to three dimensions. Specifically, the three-dimensional portrait is optimized every few steps, and the two-dimensional multi-modal generation module is used to edit a frame in the dataset, and then the edited frame is replaced by the original frame. With the progressive updating of the three-dimensional portrait and the dataset, the knowledge of the two-dimensional multi-modal generation module can be distilled to the three-dimensional portrait.

[0064] In operation S160, the convolutional neural renderer is optimized using the expression similarity loss function and the loss function in the reconstruction stage, and the dynamic three-dimensional high-frequency feature field is optimized again until the editing of all video frame images in the portrait video is completed.

[0065] The above operation S150 and operation S160 also involve the editing of the head (or face, the same below) region and the torso part of the target person in the portrait video. Specifically, when editing an upper body image, if the face occupies a small proportion, the editing may not be fine enough to handle the face structure. The present application further proposes a face-aware portrait editing model, which does not need to be trained on the two-dimensional multi-modal generation module, and can significantly enhance the editing quality of the face region. First, the face region is cropped, then the face part and the portrait part are edited by the image editing model, and the edited two parts are synthesized into the final edited image through the head-torso semantic segmentation map.

[0066] In order to further avoid the expression degradation problem caused by the accumulation of editing errors in the iterative dataset updating process, the present application uses a pre-trained expression recognition network to map the rendered frame and the original frame into the same expression space and calculate the expression similarity.

[0067] The multi-modal portrait video editing method provided by the application upgrades the portrait in the video, uses a dynamic three-dimensional Gaussian feature field to represent the person in the video by means of a parameterized human body geometry, and effectively ensures the three-dimensional consistency and time consistency of the portrait in the video; unlike the original three-dimensional Gaussian splash that uses spherical harmonic coefficients to represent color, the application uses self-adaptive learning features to represent, so that the model capacity is higher; the convolutional neural renderer of the application can effectively perform deep combination on the three-dimensional Gaussian splash feature map, further improve the expression ability of the portrait representation, make the edited result more consistent with the editing instruction, and at the same time show higher quality, especially good at fitting non-three-dimensional visual style effects; the face-aware portrait editing model of the application can make the editing process pay more attention to the face area, and significantly enhance the editing quality of the face area; at the same time, the expression similarity guide term of the application can ensure the correctness of the expression of the three-dimensional portrait and avoid the problem of expression degradation.

[0068] The multi-modal portrait video editing method provided by the application will be further described in detail below through specific embodiments and FIG. 2.

[0069] FIG. 2 is a pipeline diagram of the portrait video editing method based on a three-dimensional dynamic Gaussian feature field and a multi-modal generative prior according to an embodiment of the application.

[0070] As shown in FIG. 2, the multi-modal portrait editing provided by the application is roughly divided into three-dimensional portrait reconstruction, video frame dataset updating, and three-dimensional portrait editing stages.

[0071] According to the embodiment of the application, the upgrading operation of the frame image of the portrait video by embedding the Gaussian point in the three-dimensional Gaussian field into the two-dimensional texture space of the parameterized human body geometry representation includes: constructing an initial parameterized human body geometry representation by using the portrait identity coefficient, the portrait pose coefficient and the expression coefficient of the target person obtained from each frame image of the portrait video; correcting the difference between the initial parameterized human body geometry representation and the real human body of the target person by using a learnable vertex offset to obtain a corrected parameterized human body geometry representation; based on a neural Gaussian texture mechanism, the Gaussian points in the two-dimensional texture space are uniformly embedded on the surface of the corrected parameterized human body geometry representation, and a three-dimensional dynamic Gaussian feature field is obtained.

[0072] According to the embodiment of the present application, the above-mentioned constructing the initial parameterized human body geometry representation by using the portrait identity coefficient, the portrait pose coefficient and the expression coefficient of the target person obtained from each frame image of the portrait video includes: obtaining the portrait identity coefficient, the portrait pose coefficient and the expression coefficient of the target person in each frame image by performing portrait tracking on each frame image in the portrait video; and constructing the initial parameterized human body geometry representation by using the predefined skinning function and skinning weight, the predefined joint matrix and the human body geometry in the T-type pose based on the portrait identity coefficient, the portrait pose coefficient and the expression coefficient of the target person.

[0073] According to the embodiment of the present application, the above-mentioned performing three-dimensional Gaussian splatting on the three-dimensional dynamic Gaussian feature field to obtain three-dimensional Gaussian rendering features and converting the three-dimensional Gaussian rendering features by using the convolutional neural renderer to obtain the neural rendering RGB image includes: performing three-dimensional Gaussian splatting on the three-dimensional dynamic Gaussian feature field by using the camera intrinsic parameters and the camera pose of each frame image in the portrait video and the attribute information of the three-dimensional dynamic Gaussian feature field to obtain three-dimensional Gaussian rendering features; and constructing the convolutional neural renderer including an input layer, a convolutional layer, a nonlinear layer, a pooling layer, a fully connected layer and a loss layer, and converting the three-dimensional Gaussian rendering features by using the convolutional neural renderer to obtain the neural rendering RGB image.

[0074] According to the embodiment of the present application, the attribute information of the three-dimensional dynamic Gaussian feature field includes the spatial position information of the Gaussian point, the quaternion representing the three-dimensional orientation of the Gaussian field, the scaling vector, the opacity and the learnable feature vector.

[0075] The above-mentioned acquisition and correction process of the parameterized human body geometry representation, the principle of the neural Gaussian texture mechanism and the acquisition process of the neural rendering RGB image will be further described below by specific embodiments and in conjunction with FIG. 2,

[0076] The present application improves the portrait video editing problem to three dimensions, ensures the three-dimensional consistency and temporal consistency of the portrait video by using the three-dimensional dynamic Gaussian feature field, and constructs the three-dimensional dynamic Gaussian feature field based on the parameterized human body geometry representation, as shown in the three-dimensional Gaussian feature in FIG. 2. The three-dimensional portrait representation is obtained by using the following operations.

[0077] The parameterized human body geometry is determined by the portrait identity coefficient β, the portrait pose θ and the expression coefficient ψ, and is represented as As shown in formula (1):

[0078] where W is the skinning function, is the skinning weight, and J(β) is the joint. The human body geometry T(β, θ, ψ) in the T-type pose is obtained by formula (2):

[0079] The present application uses a learnable vertex offset to further correct the difference between the parameterized human geometry and the real human body, and obtains the corrected human geometry As formula (3):

[0080] The present application proposes a mechanism of neural Gaussian texture, which embeds Gaussian points on the surface of the parameterized human body The Gaussian attributes include: Gaussian point position X0, quaternion Q representing Gaussian three-dimensional orientation, stretch vector S, opacity O, and a learnable feature F, as shown in Figure 2. The present application exists a uniform neural Gaussian field φ in the UV space of the parameterized human geometry, and obtains a three-dimensional Gaussian feature field based on the human geometry And the neural Gaussian field φ of the UV space, as shown in formula (4):

[0081] Among them, Is the UV mapping.

[0082] Given the camera intrinsic K, the camera pose P, and the above three-dimensional Gaussian feature field, the feature map I is rendered F As shown in formula (5):

[0083] Then use a convolutional neural network as a neural renderer Convert it into an RGB image, as shown in formula (6):

[0084] According to the embodiment of the present application, the above-mentioned three-dimensional reconstruction of the target person in the portrait video by using the three-dimensional dynamic Gaussian feature field, and the optimization of the three-dimensional portrait reconstruction process by using the neural rendered RGB image and the original RGB image of the portrait video through the loss function in the reconstruction stage include: using a pre-defined image processing model to pre-process the portrait video to obtain a two-dimensional portrait mask of each frame image in the portrait video; using the reconstruction loss function, the mask loss function and the stability loss function to process the neural rendered RGB image, the original RGB image and the two-dimensional portrait mask, to obtain the loss value in the reconstruction stage; based on the loss value in the reconstruction stage, the three-dimensional portrait reconstruction process is supervised and trained to obtain the optimized three-dimensional portrait reconstruction result.

[0085] According to the embodiment of the present application, the above-mentioned updating of the three-dimensional portrait by distilling the editing information of the neural rendering RGB image to the optimized three-dimensional portrait reconstruction result comprises: pre-processing the portrait video by using a pre-defined image processing model to obtain a human body semantic segmentation map of each frame image in the portrait video, wherein the human body semantic segmentation map comprises a head region and a torso part of the target person; based on a multi-modal editing instruction input by a user, a two-dimensional multi-modal generation module of a pre-trained face-aware portrait editing model is used to edit the head region and the torso part in the neural rendering RGB image respectively; the edited head region and the edited torso part are synthesized by using the human body semantic segmentation map of each frame image, and the synthesized image is inserted into a video frame data set to update the video frame data set; the three-dimensional portrait reconstruction result is optimized by optimizing a three-dimensional dynamic high-strength feature field, and the editing information is distilled to the optimized three-dimensional portrait reconstruction result by the face-aware portrait editing model to update the three-dimensional portrait; the updating operation of the three-dimensional portrait and the updating operation of the video frame data set are alternately iterated to continuously optimize the three-dimensional dynamic high-strength feature field.

[0086] According to the embodiment of the present application, the two-dimensional multi-modal generation module comprises a two-dimensional multi-modal generation module based on text editing, a two-dimensional multi-modal generation module based on image editing, and a two-dimensional multi-modal image generation model of portrait relighting.

[0087] The above-mentioned reconstruction, optimization and updating process of the three-dimensional portrait and the editing (or updating) process of the video frame data set provided by the present application will be further described in detail below by specific embodiments and in conjunction with the accompanying drawings 2.

[0088] Given a portrait video, first, the three-dimensional portrait of the original video is reconstructed, and then the editing knowledge of the two-dimensional multi-modal generation module is distilled to the three-dimensional.

[0089] In the data preprocessing stage, given a portrait video, the camera intrinsic parameter K and the portrait identity coefficient β are obtained by tracking the portrait, and the camera pose P, the portrait pose θ, the expression coefficient ψ, the human body semantic segmentation map and the two-dimensional portrait mask of each frame are also obtained.

[0090] In the reconstruction stage, the present application uses a reconstruction loss function, a mask loss function and a stability loss function.

[0091] Wherein, the reconstruction loss function (I is a rendered RGB image, I src is a frame image of the original video) is as shown in formula (7): L recon (I, I src )=||I-I src ||1 (7).

[0092] where A is the alpha value of the rendered image, A is the ground truth alpha value of the image, and L is the mask loss function. The mask loss function is shown in equation (8) as follows: src L (A, A) = ||A-A||1 (8). mask src src

[0093] where I is the input image, I is the ground truth image, and L is the stability loss function. The stability loss function is shown in equation (9) as follows: F F stable F src F src

[0094] In addition, in order to improve the details, the present application uses a VGG network (Visual Geometry Group Network) as a backbone to calculate the perception loss function.

[0095] After the reconstruction is completed, the present application uses an iterative dataset update technique to alternately perform dataset video frame editing and three-dimensional portrait updating, and distills the editing of the two-dimensional multi-modal generation module to three dimensions. Specifically, the present application uses the two-dimensional multi-modal generation module to edit a frame in the dataset every few optimization steps, and then replaces the original frame with the edited frame. With the progressive updating of the three-dimensional portrait and the dataset, the knowledge of the two-dimensional multi-modal generation module can finally be distilled to the three-dimensional portrait.

[0096] In addition to the loss function used during reconstruction, the present application additionally uses an expression similarity loss function to further avoid the expression degradation problem caused by the accumulation of editing errors in the iterative dataset update process. The present application uses a pre-trained expression recognition network to map the rendered frame and the original frame into the same expression space and calculate the expression similarity, as shown in equation (10):

[0097] where ε is the network for expression encoding of the picture, I is the rendered video frame, and I is the original video frame. exp src

[0098] ​​​​​​​​​​​​When editing an upper body image, if the proportion of the face is small, the editing may not be fine enough to handle the facial structure. The present application further proposes a face-aware portrait editing model (which includes the above-mentioned two-dimensional multi-modal generation module), which can significantly enhance the editing quality of the face region without training the two-dimensional multi-modal generation module. As shown in FIG. 2, first, the face region is cropped, and then the face part and the portrait part are both edited by the face-aware portrait editing model, and the edited two parts are synthesized into the final edited image through the head-torso semantic segmentation map.

[0099] In order to support multi-modal portrait editing requirements in different application scenarios, the present application uses pre-trained image generation models based on text editing, image editing, and portrait relighting. The above method is also applicable to other structure-preserving image generation models.

[0100] The present application provides the above-mentioned portrait video editing method based on three-dimensional dynamic Gaussian field and multi-modal generative prior. After a given portrait video is obtained, the camera parameters, human identity, expression, pose coefficients, human semantic segmentation map, and two-dimensional portrait mask are pre-processed; based on the neural Gaussian texture mechanism, the learnable three-dimensional Gaussian feature is embedded on the parameterized human geometry surface, the neural renderer is used to convert the three-dimensional Gaussian splatter feature map into an image, and the three-dimensional portrait is reconstructed by optimizing the RGB and segmentation map information of the video; the iterative dataset update technology is used to distill the knowledge of the multi-modal two-dimensional image generation model into three-dimensional portrait editing, and the expression similarity guide and face-aware portrait editing model are used in this process to further improve the editing quality. This method improves the two-dimensional editing task to three-dimensional, and uses three-dimensional human prior and three-dimensional Gaussian splatter technology to make the edited video have good three-dimensional consistency and temporal consistency. With the knowledge of the multi-modal generation model, high-quality portrait video editing functions such as text-based editing, image-based editing, and portrait relighting can be realized.

[0101] The advantages and effectiveness of the above-mentioned method provided by the present application are verified by comparative experiments and combined with FIGS. 3-5. FIG. 3 is a visual result comparison diagram of text-based editing according to an embodiment of the present application, FIG. 4 is a visual result comparison diagram of image-based editing according to an embodiment of the present application, and FIG. 5 is a visual result comparison diagram of relighting-based editing according to an embodiment of the present application.

[0102] The methods to be compared are: (1) TokenFlow: Consistent Diffusion Features for Consistent Video Editing (ICLR 2024); (2) Rerender A Video (RAV): Zero-Shot Text-Guided Video-to-Video Translation (SIGGRAPH Asia 2023 Conference Proceedings); (3) CoDeF: Content Deformation Fields for Temporally Consistent Video Processing (CVPR 2024 Highlight); and (4) AnyV2V: A Plug-and-Play Framework For Any Video-to-Video Editing Tasks (arXiv preprint).

[0103] It should be particularly noted that the image data involved in FIGS. 3-5 is only used for illustrative purposes and will not be used for commercial purposes. The images of public figures involved in FIGS. 3-5 come from research institutions for scientific research purposes, and the image data of the above-mentioned public figures provided by them has obtained the authorization of the relevant public figures and is only used for scientific research purposes. It is an open data set for the public. The image data of non-public figures in FIGS. 3-5 has obtained the authorization of the relevant parties, and the processing of their portraits is only used to illustrate the advantages and effectiveness of the method provided by the present application. In addition, all the above-mentioned image data has taken effective security measures, and the process of processing it also strictly follows the requirements of relevant laws and regulations.

[0104] In terms of editing video quality comparison, as shown in FIGS. 3-5, the multi-modal portrait video editing method provided by the present application is superior to other methods in terms of the degree of fit with editing instructions, structure preservation, temporal consistency, and overall video quality.

[0105] In terms of video editing efficiency comparison method, by analyzing the number of frames processed per minute by each method, it is further verified that the editing efficiency of the method provided by the present application is superior to other methods, as shown in Table 1. Different methods use different inference modes. In order to compare fairly, the present application calculates the time cost including reconstruction and editing, and averages on the number of frames processed. It can be seen that the method of the present application is superior to the previous video editing methods in terms of efficiency.

[0106] Table 1 Comparison of editing efficiency of each method

[0107] FIG. 6 schematically illustrates a block diagram of an electronic device suitable for implementing the multi-modal portrait video editing method according to an embodiment of the present application.

[0108] As shown in FIG. 6, the electronic device 600 according to an embodiment of the present application includes a processor 601 which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 can include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a related chipset, and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), and the like. The processor 601 can also include an on-board memory for cache use. The processor 601 can include a single processing unit or multiple processing units for executing different actions of the method processes according to embodiments of the present application.

[0109] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method processes according to embodiments of the present application by executing the programs in the ROM 602 and / or the RAM 603. Note that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method processes according to embodiments of the present application by executing the programs stored in the one or more memories.

[0110] According to an embodiment of the present application, the electronic device 600 can further include an input / output (I / O) interface 605 which is also connected to the bus 604. The electronic device 600 can further include one or more of the following components connected to the input / output (I / O) interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as necessary. A removable medium 611 such as a magnetic disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 610 as necessary, so that a computer program read out therefrom is installed in the storage section 608 as necessary.

[0111] The application further provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments, or can exist independently without being assembled into the device / apparatus / system. The computer readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the application.

[0112] According to the embodiments of the application, the computer readable storage medium can be a non-volatile computer readable storage medium, which can include, but is not limited to, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination thereof. In this application, a computer readable storage medium can be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. For example, in the embodiments of the application, a computer readable storage medium can include one or more of the above-described ROM 602 and / or RAM 603 and / or one or more memories other than the ROM 602 and the RAM 603.

[0113] The flowcharts and block diagrams in the drawings illustrate the possible architectures, functionality, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the reverse order, depending on the functionality involved. It will also be noted that each block in the block diagrams or flowcharts, and combinations of blocks in the block diagrams or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0114] Those skilled in the art can understand that the features described in various embodiments of the application can be combined and / or integrated in various combinations, even if such combinations are not explicitly described in the application. In particular, the features described in various embodiments of the application can be combined and / or integrated in various combinations without departing from the spirit and teachings of the application. All such combinations fall within the scope of the application.

[0115] The embodiments of the application have been described. However, these embodiments are merely for illustration and are not intended to limit the scope of the application. Although each embodiment is described above separately, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Various alternatives and modifications to the embodiments described herein will be apparent to those skilled in the art in view of the foregoing without departing from the scope of the application.

Claims

1. A multimodal portrait video editing method, characterized in that, The method includes: By embedding Gaussian points in a three-dimensional Gaussian field into a two-dimensional texture space representing parameterized human geometry, the frame images of the portrait video are subjected to dimensionality upscaling to obtain a three-dimensional dynamic Gaussian feature field. The three-dimensional dynamic Gaussian feature field is subjected to three-dimensional Gaussian splashing to obtain three-dimensional Gaussian rendering features, and the three-dimensional Gaussian rendering features are transformed using a convolutional neural renderer to obtain a neurally rendered RGB image. The target person in the portrait video is reconstructed in three dimensions using the three-dimensional dynamic Gaussian feature field, and the three-dimensional portrait reconstruction process is optimized by using the RGB image of the neural rendering and the original RGB image of the portrait video through the loss function of the reconstruction stage. The editing information of the neurally rendered RGB image is distilled onto the optimized 3D human portrait reconstruction result to complete the 3D human portrait update, and the original frame image of the human portrait video is replaced with the edited neurally rendered RGB image to complete the editing of the video frame dataset. The video frame dataset editing process and the 3D human portrait update process are performed alternately and iteratively, and the edited neurally rendered RGB image is mapped to the same expression space as the original frame image of the human portrait video; The convolutional neural renderer is optimized using the facial expression similarity loss function and the loss function of the reconstruction stage, and the dynamic three-dimensional Gaussian feature field is optimized again until the editing of all video frame images in the portrait video is completed.

2. The method according to claim 1, characterized in that, By embedding Gaussian points in a three-dimensional Gaussian field into a two-dimensional texture space representing parameterized human geometry, the frame images of the portrait video are subjected to dimensionality upscaling to obtain a three-dimensional dynamic Gaussian feature field, including: Using the facial identity coefficient, facial pose coefficient, and facial expression coefficient of the target person obtained from each frame of the facial video, an initial parameterized human geometric representation is constructed. The difference between the initial parametric human geometry representation and the real human body of the target person is corrected by using learnable vertex offsets, resulting in a corrected parametric human geometry representation. Based on the neural Gaussian texture mechanism, the three-dimensional Gaussian field is mapped onto Gaussian points in the two-dimensional texture space and uniformly embedded into the surface of the corrected parameterized human geometry representation to obtain the three-dimensional dynamic Gaussian feature field.

3. The method according to claim 2, characterized in that, Constructing an initial parametric human geometric representation using the target person's facial identity coefficient, facial pose coefficient, and facial expression coefficient obtained from each frame of the facial video includes: By performing portrait tracking on each frame of the portrait video, the portrait identity coefficient, portrait pose coefficient, and expression coefficient of the target person in each frame are obtained; Based on the target person's facial identity coefficient, facial pose coefficient, and facial expression coefficient, the initial parameterized human geometric representation is constructed using predefined skinning functions and skinning weights, predefined joint matrix, and human geometry under T-pose.

4. The method according to claim 1, characterized in that, The three-dimensional dynamic Gaussian feature field is subjected to three-dimensional Gaussian splashing to obtain three-dimensional Gaussian rendering features. These features are then transformed using a convolutional neural renderer to obtain a neurally rendered RGB image, including: Using the camera intrinsic parameters and camera pose of each frame in the portrait video, as well as the attribute information of the three-dimensional dynamic Gaussian feature field, the three-dimensional dynamic Gaussian feature field is subjected to three-dimensional Gaussian splashing to obtain three-dimensional Gaussian rendering features. A convolutional neural renderer is constructed, comprising an input layer, a convolutional layer, a nonlinear layer, a pooling layer, a fully connected layer, and a loss layer. The convolutional neural renderer is then used to transform the 3D Gaussian rendering features to obtain a neurally rendered RGB image.

5. The method according to claim 4, characterized in that, The attribute information of the three-dimensional dynamic Gaussian feature field includes the spatial location information of the Gaussian points, the quaternion representing the three-dimensional orientation of the Gaussian field, the scaling vector, the opacity, and the learnable feature vector.

6. The method according to claim 1, characterized in that, The method involves using the three-dimensional dynamic Gaussian feature field to perform three-dimensional reconstruction of the target person in the portrait video, and optimizing the three-dimensional portrait reconstruction process using the neurally rendered RGB image and the original RGB image of the portrait video through a loss function in the reconstruction stage, including: The portrait video is preprocessed using a predefined image processing model to obtain a two-dimensional portrait mask for each frame of the portrait video; The neurally rendered RGB image, the original RGB image, and the two-dimensional human face mask are processed using the reconstruction loss function, the masking loss function, and the stabilization loss function to obtain the loss value in the reconstruction stage. The 3D human portrait reconstruction process is supervised and trained based on the loss value of the reconstruction stage to obtain the optimized 3D human portrait reconstruction result.

7. The method according to claim 1, characterized in that, Distilling the editing information from the neurally rendered RGB image onto the optimized 3D portrait reconstruction result to update the 3D portrait includes: The portrait video is preprocessed using a predefined image processing model to obtain a human semantic segmentation map for each frame of the portrait video, wherein the human semantic segmentation map includes the head region and torso of the target person. Based on user-input multimodal editing commands, the two-dimensional multimodal generation module of the pre-trained facial perception portrait editing model is used to edit the head region and the torso in the neurally rendered RGB image respectively. Using the human semantic segmentation map of each frame image, the edited head region and the edited torso are synthesized, and the synthesized image is inserted into the video frame dataset to complete the update of the video frame dataset; The 3D portrait reconstruction result is optimized by optimizing the 3D dynamic Gaussian feature field, and the editing information is distilled onto the optimized 3D portrait reconstruction result through the face perception portrait editing model to complete the 3D portrait update. The three-dimensional dynamic Gaussian feature field is continuously optimized by iteratively updating the three-dimensional human image and the video frame dataset.

8. The method according to claim 7, characterized in that, The two-dimensional multimodal generation module includes a two-dimensional multimodal generation module based on text editing, a two-dimensional multimodal generation module based on image editing, and a two-dimensional multimodal image generation model for portrait relighting.

9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, It stores a computer program or instructions thereon, characterized in that, when the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • High-fidelity three-dimensional face reconstruction and generation method based on implicit neural function

    CN116071494A

  • Digital human synthesis method and device, digital human synthesis model training method and device, equipment and storage medium

    CN117315211A

  • Three-dimensional face generating and driving method based on Gaussian splashing method

    CN118071898A

  • High-fidelity dynamic character rendering method, system and device based on compact Gaussian splashing, chip and medium

    CN118172470A

  • Three-dimensional human body generation method based on text prompt

    CN118229860A

Cited By

  • Multi-modal enhanced rendering method based on 3D Gaussian

    CN121837483A

  • Video editing method and device based on two-dimensional Gaussian function, equipment and medium

    CN121937608A

  • Video feature extraction method based on double-branch collaboration and perception consistency representation

    CN121937944A