Video Processing Method, System and Related Devices Based on Character Avatar Model
Through the video processing method based on the character avatar model, the problem of expression-driven in the existing technology is solved, and efficient video display and face replacement effects are achieved, improving the user experience.
Patent Information
- Application Number
- CN202211130893.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-09-16
AI Technical Summary
In the prior art, by processing video frame by frame, it is impossible to drive the second user's face to make corresponding expressions with the expression of the first user, resulting in poor video display and face replacement effects.
Using a video processing method based on the character avatar model, a video of the driver object, permission verification information, and a character avatar model and reference image of the driven object are obtained, and the driven object is generated, so that the driven object can execute the same expression and posture as the driver object in the driver video.
It realizes the driving object's face to make corresponding expressions with the expression of the driving object, which improves the effect of video display and face replacement and improves the user experience.
Smart Images

Figure CN115643349B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and in particular to a video processing method, system and related devices based on a role avatar model. Background Art
[0002] With the development of science and technology, especially the development of video processing technology, users' requirements for video processing have gradually increased. For example, users want to achieve face replacement based on video processing, such as using the expression of a first user to drive the face of a second user to make corresponding expressions in the video.
[0003] In the prior art, usually the video is processed frame by frame. It is required that the first user and the second user record a video respectively. For each frame image in the video, the face regions in the images of the first user and the second user are intercepted and replaced. The problem with the prior art is that the face regions in the images of the first user and the second user are intercepted and replaced, and although the expression in the face region of the corresponding image of the second user after replacement is the expression of the first user, in fact, the corresponding facial features are still the facial features of the first user, and the purpose of using the expression of the first user to drive the face of the second user to make corresponding expressions is not achieved.
[0004] The problem with the prior art is that the video processing solution of only intercepting and replacing the face regions in each frame image of the videos of two users cannot achieve the purpose of using the expression of the first user to drive the face of the second user to make corresponding expressions, which is not conducive to improving the video display effect and is also not conducive to improving the video face replacement effect.
[0005] Therefore, the prior art still needs to be improved and developed. Summary of the Invention
[0006] The main purpose of the present invention is to provide a video processing method, system and related devices based on a role avatar model, aiming to solve the problem that the video processing solution of only intercepting and replacing the face regions in each frame image of the videos of two users in the prior art is not conducive to improving the video display effect.
[0007] To achieve the above object, in the first aspect of the present invention, a video processing method based on a role avatar model is provided, wherein the video processing method based on a role avatar model includes:
[0008] Obtain a driving video of a driving object, permission verification information of the driving object, and a driven object corresponding to the driving object, wherein the driving video is obtained by photographing the expression and posture of the driving object;
[0009] When the permission verification information of the above driving object meets the permission verification conditions of the above driven object, obtain the role avatar model and reference image corresponding to the above driven object;
[0010] Obtain multiple frames of facial geometry rendering images corresponding to the above driving object according to the above driving video, wherein the above facial geometry rendering images are used to reflect the expressions and postures corresponding to the above driving object;
[0011] Obtain the time encoding corresponding to each of the above facial geometry rendering images, and generate a driven video through the above role avatar model according to the above reference image, each of the above facial geometry rendering images, and the time encoding corresponding to each of the above facial geometry rendering images, wherein the driven object in the above driven video performs the same expressions and postures as the driving object in the above driving video.
[0012] Optionally, the above reference image is used to provide the image texture details corresponding to the above driven object for the above role avatar model, and the image texture details of the above driven video are the same as those of the above reference image.
[0013] Optionally, the above obtaining multiple frames of facial geometry rendering images corresponding to the above driving object according to the above driving video includes:
[0014] Split the above driving video to obtain multiple frames of driving images;
[0015] Extract and obtain the three-dimensional facial parameters corresponding to each of the above driving images respectively;
[0016] Obtain the three-dimensional facial meshes corresponding to each of the above driving images respectively according to the three-dimensional facial parameters corresponding to each of the above driving images;
[0017] Render the three-dimensional facial meshes corresponding to each of the above driving images to obtain the facial geometry rendering images corresponding to each of the above driving images, wherein the above facial geometry rendering images are grayscale images.
[0018] Optionally, before the above obtaining the three-dimensional facial meshes corresponding to each of the above driving images respectively according to the three-dimensional facial parameters corresponding to each of the above driving images, the above method further includes:
[0019] Align the three-dimensional facial parameters based on the facial spatial position corresponding to the above driven object in the above role avatar model to update the three-dimensional facial parameters.
[0020] Optionally, the above three-dimensional facial parameters include individual coefficients, expression coefficients, and posture coefficients.
[0021] Optionally, obtaining the time encoding corresponding to each of the above facial geometry rendering images, and generating a driven video through the above avatar model according to the above reference image, each of the above facial geometry rendering images, and the time encoding corresponding to each of the above facial geometry rendering images, includes:
[0022] Obtaining the time encoding corresponding to each of the above facial geometry rendering images according to a preset time encoding calculation formula;
[0023] Sequentially inputting each group of data to be processed into the above avatar model to obtain the driven images corresponding to each group of the data to be processed, wherein one group of the data to be processed consists of the above reference image, one of the above facial geometry rendering images, and the time encoding corresponding to this facial geometry rendering image, and the same expressions and postures as those in the corresponding facial geometry rendering image are performed by the driven object in the above driven image;
[0024] Sequentially connecting each of the above driven images according to the corresponding time encoding and generating the above driven video.
[0025] Optionally, the above time encoding is used to input time information into the above avatar model, and the above time encoding calculation formula is: TPE t =(sin(2 0 πt), cos(2 0 πt), …, sin(2 N-1 πt), cos(2 N-1 πt)), where TPE t represents the time encoding corresponding to the facial geometry rendering image numbered t, N is a preset constant, and the spatial dimensions of the above time encoding, the facial geometry rendering image corresponding to the above time encoding, and the above reference image are the same.
[0026] Optionally, the above avatar model is pre-trained according to the following steps:
[0027] Inputting the reference image, the training facial geometry rendering image, and the training time encoding corresponding to this training facial geometry rendering image in the training data into a deep neural network generator, and generating a training driven image for the above reference image and the above training facial geometry rendering image through the above deep neural network generator, wherein the above training data includes multiple groups of training image groups, and each group of training image groups includes a reference image corresponding to the above driven object, a training facial geometry rendering image corresponding to the above driving object, the training time encoding corresponding to this training facial geometry rendering image, and the training driven image corresponding to this training facial geometry rendering image;
[0028] Based on the above-mentioned training driven image and the above-mentioned training driving image, adjust the model parameters of the above-mentioned deep neural network generator, and continue to execute the above-mentioned step of inputting the reference image, the training facial geometry rendering image in the training data, and the training time encoding corresponding to the training facial geometry rendering image into the above-mentioned avatar model until the preset training conditions are met to obtain the above-mentioned avatar model;
[0029] Among them, the above-mentioned preset training conditions include the convergence of the reconstruction loss between the above-mentioned training driven image and the above-mentioned training driving image.
[0030] Optionally, the above-mentioned training data is obtained by processing the collected image data through a preset data augmentation method, and the above-mentioned preset data augmentation method includes spatial random cropping.
[0031] Optionally, the above-mentioned reconstruction loss is obtained by calculating through a multi-joint image reconstruction loss function, and the above-mentioned multi-joint image reconstruction loss function is used to jointly calculate at least two losses among the L1 reconstruction loss, the perceptual loss, and the GAN discriminator loss.
[0032] The second aspect of the present invention provides a video processing system based on an avatar model. Among them, the above-mentioned video processing system based on an avatar model includes:
[0033] A driving information acquisition module, configured to acquire the driving video of the driving object, the permission verification information of the driving object, and the driven object corresponding to the driving object, where the above-mentioned driving video is obtained by shooting the expression and posture of the driving object;
[0034] A permission verification module, configured to obtain the avatar model and the reference image corresponding to the driven object when the permission verification information of the driving object meets the permission verification conditions of the driven object;
[0035] A driving video processing module, configured to obtain multiple frames of facial geometry rendering images corresponding to the driving object according to the above-mentioned driving video, where the above-mentioned facial geometry rendering images are used to reflect the expression and posture corresponding to the driving object;
[0036] A driven video generation module, configured to obtain the time encoding corresponding to each of the above-mentioned facial geometry rendering images, and generate a driven video through the above-mentioned avatar model according to the above-mentioned reference image, each of the above-mentioned facial geometry rendering images, and the time encoding corresponding to each of the above-mentioned facial geometry rendering images, where the expression and posture executed by the driven object in the above-mentioned driven video are the same as those of the driving object in the above-mentioned driving video.
[0037] In a third aspect of the present invention, an intelligent terminal is provided. The intelligent terminal includes a memory, a processor, and a video processing program based on a role avatar model stored on the memory and executable on the processor. When the video processing program based on the role avatar model is executed by the processor, the steps of any one of the above-mentioned video processing methods based on the role avatar model are implemented.
[0038] In a fourth aspect of the present invention, a computer-readable storage medium is provided. A video processing program based on a role avatar model is stored on the computer-readable storage medium. When the video processing program based on the role avatar model is executed by a processor, the steps of any one of the above-mentioned video processing methods based on the role avatar model are implemented.
[0039] As can be seen from the above, in the solution of the present invention, a driving video of a driving object, permission verification information of the driving object, and a driven object corresponding to the driving object are obtained, wherein the driving video is obtained by photographing the expressions and postures of the driving object; when the permission verification information of the driving object meets the permission verification conditions of the driven object, a role avatar model and a reference image corresponding to the driven object are obtained; multiple frames of facial geometry rendering images corresponding to the driving object are obtained according to the driving video, wherein the facial geometry rendering images are used to reflect the expressions and postures of the driving object; time codes corresponding to the facial geometry rendering images are obtained, and according to the reference image, each facial geometry rendering image, and the time codes corresponding to the facial geometry rendering images, a driven video is generated through the role avatar model, wherein the driven object in the driven video performs the same expressions and postures as the driving object in the driving video.
[0040] Compared with the prior art, in the solution of the present invention, it is not just to intercept and replace the facial regions in the images of different objects, but a role avatar model for the driven object is preset in advance. After obtaining the driving video corresponding to the driving object and passing the permission verification, the trained role avatar model and the reference image of the corresponding driven object are obtained. Then, facial geometry rendering images for reflecting the expressions and postures of the driving object are obtained according to the driving video, and combined with the time codes, the facial geometry rendering images and the reference image are fused through the role avatar model to obtain the driven video.
[0041] The driven video is not obtained by simply replacing the image of the face region, but by fusing the expressions and postures of the driving object with the actual texture of the driven object, so as to enable the driven object to execute the same expressions and postures as the driving object in the driving video. Since the face geometry rendering image only reflects the expressions and postures and does not reflect the actual texture of the face of the driving object, and the actual texture is only provided by the reference image of the driven object, the actual texture corresponding to the driving object will not be incorrectly retained when the character avatar model performs image information fusion, that is, the image texture details of the object shown in the final driven video are the same as those of the driven object. In this way, using the video of the driving object to drive the character avatar model is beneficial to improving the video display effect of the character avatar model. It can realize driving the face of the driven object to make corresponding expressions with the expressions of the driving object, which is beneficial to improving the effect of video face replacement and the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0043] Figure 1 It is a schematic flowchart of a video processing method based on a character avatar model provided by an embodiment of the present invention;
[0044] Figure 2 It is a specific flowchart of generating a driven video based on a character avatar model of user A provided by an embodiment of the present invention;
[0045] Figure 3 It is a schematic diagram of the component modules of a video processing system based on a character avatar model provided by an embodiment of the present invention;
[0046] Figure 4 It is a schematic block diagram of the internal structure principle of an intelligent terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0048] It should be understood that, as used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0049] It should also be understood that the terminology used in the specification of the present invention is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0050] It should be further understood that the term "and / or" as used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0051] As used in this specification and the appended claims, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to classifying". Similarly, the phrase "if determined" or "if classified into [the described condition or event]" may be construed, depending on the context, as meaning "once determined", "in response to determining", "once classified into [the described condition or event]", or "in response to classifying into [the described condition or event]".
[0052] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0053] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention, but the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the spirit of the present invention, so the present invention is not limited by the specific embodiments disclosed below.
[0054] With the development of science and technology, especially the development of video processing technology, users' requirements for video processing are gradually increasing. For example, users want to achieve face replacement based on video processing, such as achieving driving the face of a second user to make corresponding expressions with the expressions of a first user in a video, so as to achieve an entertainment effect.
[0055] In the prior art, video frames are usually processed one by one. It is required that the first user and the second user respectively record a video. For each frame image in the video, the facial regions in the images of the first user and the second user are intercepted and replaced. The problem with the prior art is that, although the expression in the facial region of the image corresponding to the second user after replacement is the expression of the first user, in fact, the corresponding facial features are still those of the first user, and the purpose of driving the face of the second user to make corresponding expressions with the expression of the first user is not achieved.
[0056] The problem with the prior art is that the video processing solution of only intercepting and replacing the facial regions in each frame image of the videos of two users is not conducive to improving the effect of video facial replacement, and it is impossible to drive the face of the second user to make corresponding expressions with the expression of the first user.
[0057] In one application scenario, based on 3D reconstruction, animation, and CG rendering technologies, multiple cameras are used to scan and 3D reconstruct a static person, and then it is bound to the key points of the driving model to complete effects such as action driving. Finally, the 2D image is rendered and reproduced through rendering technologies such as relighting and PBR. However, this solution requires a large number of additional devices (such as a multi-camera array) to collect the details of the human figure in advance for high-precision restoration of geometric shapes and textures, and manual 3D modeling adjustment is also required, resulting in a large production cost.
[0058] In another application scenario, a portrait video can be used. By using deep learning and other methods, some explicit attributes such as pose and camera position, or implicit feature expressions can be learned and predicted. These attributes and feature expressions can be adjusted and manipulated to restore the portrait image. For example, one solution (such as FOMM) is to establish the key point correspondence between the target portrait and the driving portrait based on unsupervised learning and convert it into a generated dense optical flow field, then use the optical flow field to map the face image, and finally generate an image through a generation network. However, this solution is based on 2D pixel mapping, and the face does not have 3D consistency, and the dense optical flow field easily makes the background and the face move together. In another solution (such as DVP), a large number of explicit attributes including face correspondence maps and implementation maps are estimated first, and combined with image-to-image conversion technology to restore the real portrait. However, this solution uses a large number of explicit attributes such as 3D face correspondence maps, but the generated results will have artifacts and blurs, and it is not smooth when processing videos. In another solution (such as NerFace), a neural radiance field is used as a renderer to generate high-definition portraits, which can improve the 3D consistency at large angles, but its rendering results will still lose details and the rendering efficiency is low.
[0059] To solve at least one of the above-mentioned multiple problems, the solution of the present invention takes into account both cost and effect, and provides a solution with higher efficiency, better effect and lower training duration required for the model, so that the results of video processing and rendering are clearer and more realistic, achieving driving the face of the driven object to make corresponding expressions based on the expressions of the driving object, generating the corresponding driven video, and making the generated driven video approach the effect of a real video.
[0060] Specifically, in the solution of the present invention, a driving video of a driving object, the permission verification information of the driving object, and the driven object corresponding to the driving object are obtained, wherein the driving video is obtained by shooting the expressions and postures of the driving object; when the permission verification information of the driving object meets the permission verification conditions of the driven object, the role avatar model and reference image corresponding to the driven object are obtained; according to the driving video, multiple frames of facial geometry rendering images corresponding to the driving object are obtained, wherein the facial geometry rendering images are used to reflect the expressions and postures of the driving object; time encodings corresponding to each of the facial geometry rendering images are obtained, and according to the reference image, each of the facial geometry rendering images, and the time encodings corresponding to each of the facial geometry rendering images, a driven video is generated through the role avatar model, wherein in the driven video, the driven object executes the same expressions and postures as the driving object in the driving video.
[0061] Compared with the prior art, in the solution of the present invention, it does not simply intercept and replace the facial regions in the images of different objects, but a role avatar model for the driven object is preset in advance. After obtaining the driving video corresponding to the driving object and passing the permission verification, the trained role avatar model and reference image corresponding to the driven object are obtained. Then, according to the driving video, facial geometry rendering images for reflecting the expressions and postures of the driving object are obtained, and combined with the time encoding, the role avatar model is used to fuse the facial geometry rendering images and the reference image, so as to obtain the driven video.
[0062] The driven video is not obtained by simply replacing the image of the face area, but by fusing the expression and posture of the driving object with the actual texture of the driven object, so as to enable the driven object to perform the same expressions and postures as the driving object in the driving video. Since the face geometry rendering image only reflects the expression and posture and does not reflect the actual texture of the face of the driving object, and the actual texture is only provided by the reference image of the driven object, the actual texture corresponding to the driving object will not be wrongly retained when the character avatar model performs image information fusion. That is, the image texture details of the object shown in the final driven video are the same as those of the driven object. In this way, using the video of the driving object to drive the character avatar model is beneficial to improving the video display effect of the character avatar model. It can realize driving the face of the driven object to make corresponding expressions with the expression of the driving object, which is beneficial to improving the effect of video face replacement and the user experience.
[0063] Exemplary method
[0064] As Figure 1 shown, an embodiment of the present invention provides a video processing method based on a character avatar model. Specifically, the above method includes the following steps:
[0065] Step S100, obtain the driving video of the driving object, the permission verification information of the above driving object, and the driven object corresponding to the above driving object, where the above driving video is obtained by shooting the expression and posture of the driving object.
[0066] Among them, the above driving object is an object that needs to retain the corresponding expression and posture but not the corresponding facial details during the video processing (such as user B). In this embodiment, a character avatar model is pre-trained for the driven object (such as user A) (such as a digital portrait model of the driven object speaking in a specific scenario, or a digital portrait model obtained by training based on the video of the driven object speaking in a specific scenario). The driven object is an object that needs to retain the corresponding facial details during the video processing (that is, user A). Therefore, the video processing process in this embodiment is equivalent to driving the character avatar model corresponding to user A with the expression and posture provided by user B, so that the character avatar model corresponding to user A generates a corresponding driven video. In this driven video, user A's image makes the same expressions and postures as in user B's driving video, so as to achieve the effect of driving user A's image with user B through video processing. The above posture represents the head posture of the corresponding object, and the expression represents the facial expression of the corresponding object. The above driving video can be obtained by shooting the driving object with devices such as cameras and mobile phones, and specifically is a speaking video or a video with lip movements obtained by shooting, so as to perform modeling and matching in combination with real lip movements.
[0067] In one embodiment, corresponding avatar models can be pre-trained for multiple other users, and the driving object determines the avatar model to be selected and used by specifying the corresponding driven object.
[0068] It should be noted that the above driving object and driven object can be animals, animated images, virtual characters or real people, and the driving object and the driven object can be the same or different; in this embodiment, a real person is taken as an example for illustration, but it is not a specific limitation.
[0069] Furthermore, in this embodiment, the images of the head regions of the driving object and the driven object are processed according to the above video processing method. Correspondingly, the above avatar model is also a model for processing head region avatars, but based on this solution, the above avatar model can also be used to process the entire human figure in the video, including the head region and limbs.
[0070] Step S200, when the permission verification information of the above driving object meets the permission verification conditions of the above driven object, obtain the avatar model and reference image corresponding to the above driven object.
[0071] Among them, the above permission verification information is used to verify whether the driving object has the permission to use the avatar model and / or reference image corresponding to the driven object. Specifically, in order to protect the privacy and security of the driven object and avoid the situation where any user can use the avatar model of the driven object to generate a video with the driven object, in this embodiment, permission verification conditions are pre-set for the avatar model of the driven object. Only when the permission verification information of the driving object meets the permission verification conditions of the driven object can the avatar model and reference image corresponding to the driven object be obtained. It should be noted that there are various ways to set the permission verification conditions and the corresponding permission verification information, such as password matching, authorization through a permission table, etc., which are not specifically limited here.
[0072] In this embodiment, the above reference image is used to provide the image texture details corresponding to the above driven object for the above avatar model, and the image texture details of the above driven video are the same as those of the above reference image. Specifically, the driving video provides expressions and postures, and the corresponding driven video is generated by combining the image texture details of the face in the reference image. The above reference image is obtained by photographing the driven object, and the image texture details corresponding to the background area in the scene where the driven object is located can be photographed. Thus, the corresponding background in the driven video is also the same as the reference image. It should be noted that the image texture details can include details other than expressions and postures, such as facial features (e.g., facial features, glasses, etc.), and the image texture of the background area, which are not specifically limited here.
[0073] Step S300, obtain multiple frames of facial geometry rendering images corresponding to the driving object according to the above driving video, where the facial geometry rendering images are used to reflect the expressions and postures of the driving object.
[0074] In this embodiment, multiple consecutive frames of facial geometry rendering images corresponding to the driving object are sequentially obtained according to the above driving video, where each frame of facial geometry rendering image is used to reflect the expressions and postures of the driving object, but the image texture details of each frame in the driving video are not retained.
[0075] Specifically, in this embodiment, obtaining multiple frames of facial geometry rendering images corresponding to the driving object according to the above driving video includes: splitting the above driving video to obtain multiple frames of driving images; respectively extracting three-dimensional facial parameters corresponding to each of the above driving images; respectively obtaining three-dimensional facial meshes corresponding to each of the above driving images according to the three-dimensional facial parameters corresponding to each of the above driving images; and rendering the three-dimensional facial meshes corresponding to each of the above driving images to obtain facial geometry rendering images corresponding to each of the above driving images, where the facial geometry rendering images are grayscale images.
[0076] Among them, the above three-dimensional facial parameters include individual coefficients, expression coefficients, and pose coefficients. Further, in order to improve the accuracy in the video processing process and improve the authenticity of the finally obtained driven video and the vividness of the expressions therein, in this embodiment, before respectively obtaining the three-dimensional facial meshes corresponding to each of the above driving images according to the three-dimensional facial parameters corresponding to each of the above driving images, the method further includes: aligning the three-dimensional facial parameters based on the facial spatial position corresponding to the driven object in the above character avatar model to update the three-dimensional facial parameters. That is, the facial rendering images are obtained according to the updated three-dimensional facial parameters.
[0077] Specifically, the above driving video is sequentially split into image frames to obtain multiple frames of driving images, and then for each frame of driving image, 3D face parameter estimation is performed to obtain the corresponding three-dimensional facial parameters, including individual coefficients, expression coefficients, and pose coefficients. In one application scenario, the corresponding three-dimensional facial parameters are extracted by a pre-trained parameter extraction model, which is trained to output the corresponding three-dimensional facial parameters according to the input face image, and the parameter extraction model can be a pre-trained neural network model. The above expression coefficients are used to reflect the expression characteristics of the driving object, such as grinning, pouting, etc.; the above pose coefficients are used to reflect the head pose of the driving object, such as turning the head left and right, nodding up and down, shaking the head, etc.; and the above individual coefficients are used to reflect the facial characteristics of the driving object, such as face shape. Different users have different face shapes, so the individual coefficients of different driving objects are also different. Combining the above three types of three-dimensional facial parameters can make the expressions in the generated driven video more accurate.
[0078] Further, the individual coefficient, expression coefficient, and pose coefficient are transformed into the facial space of the avatar model of the driven object (i.e., face pose correction is performed), thereby improving the generation effect of the driven video. During the process of three-dimensional facial parameter alignment (or transformation), the head poses of both the driving object and the driven object are aligned, that is, the spatial size and spatial position of the head need to be roughly aligned. In one application scenario, the three-dimensional facial parameters corresponding to the driven object are stored in the avatar model corresponding to the driven object, which can be used to reflect the facial spatial position corresponding to the driving object. During the alignment of the three-dimensional facial parameters of the driving object, the goal is to align the mean and variance of each coefficient in the transformed three-dimensional facial parameters of the driving object with those of each coefficient in the three-dimensional facial parameters of the driven object.
[0079] After obtaining the three-dimensional facial parameters (i.e., individual coefficient, expression coefficient, and pose coefficient) corresponding to each driving image, the three-dimensional facial meshes corresponding to each driving image are calculated through a pre-set 3D face model (such as BFM, FLAME, 3D face model, represented by the function f). The above three-dimensional facial meshes represent the facial geometric information in the driving image. Further, a pre-set renderer (such as Pytorch3D, the renderer is represented by Render) is used to render each three-dimensional facial mesh to obtain the facial geometric rendering images corresponding to each of the above driving images. In this embodiment, the above facial geometric rendering images are grayscale images with 1 channel, and a driving video corresponds to multiple driving images, so there are also multiple corresponding facial geometric rendering images (i.e., a set of facial geometric rendering images can be obtained).
[0080] In step S400, the time encoding corresponding to each of the above facial geometric rendering images is obtained. According to the above reference image, each of the above facial geometric rendering images, and the time encoding corresponding to each of the above facial geometric rendering images, the driven video is generated through the above avatar model, where the driven object in the above driven video performs the same expressions and poses as the driving object in the driving video.
[0081] In this embodiment, after obtaining the above-mentioned facial geometric rendering image, time encoding is added, and then in combination with the reference image, the prediction and generation of the driven object with the current pose and expression are completed through the avatar model, and then the corresponding driven video is generated. Specifically, the above-mentioned reference image is an RGB image with 3 channels, and the reference image can be any image containing the face of the driven object. In this embodiment, the avatar model used is the same as the reference image used when training the avatar model, and it can be any frame in the training video containing the driven object used when training the above-mentioned avatar model. The above-mentioned reference image is used to provide the texture of the person and the background, so that the avatar model (i.e., a trained neural network generator, such as a UNet model) can restore more details.
[0082] It should be noted that for an avatar model, after the reference image is selected, the same reference image is globally fixed and no longer changed.
[0083] Furthermore, obtaining the time encoding corresponding to each of the above-mentioned facial geometric rendering images, and generating a driven video through the above-mentioned avatar model according to the above-mentioned reference image, each of the above-mentioned facial geometric rendering images, and the time encoding corresponding to each of the above-mentioned facial geometric rendering images, includes: obtaining the time encoding corresponding to each of the above-mentioned facial geometric rendering images according to a preset time encoding calculation formula; sequentially inputting each group of data to be processed into the above-mentioned avatar model to obtain the driven images corresponding to each group of the above-mentioned data to be processed, where a group of the above-mentioned data to be processed is composed of the above-mentioned reference image, one of the above-mentioned facial geometric rendering images, and the time encoding corresponding to this facial geometric rendering image, and the same expression and pose as those in the corresponding facial geometric rendering image are executed by the driven object in the above-mentioned driven image; sequentially connecting each of the above-mentioned driven images according to the corresponding time encoding to generate the above-mentioned driven video.
[0084] Among them, the above-mentioned time encoding is used to input time information to the above-mentioned avatar model to improve the temporal stability when generating the driven image. In this embodiment, the above-mentioned time encoding is a time information encoding with 2N channels, and its spatial dimension is the same as that of the driven image and the facial geometric rendering image, and the values in each channel are the same.
[0085] Specifically, the above-mentioned preset time encoding calculation formula is shown as formula (1) below:
[0086] TPE t =(sin(2 0 πt),cos(2 0 πt),…,sin(2 N-1 πt),cos(2 N-1 πt)) (1)
[0087] Among them, TPE t represents the time encoding corresponding to the facial geometry rendering image numbered t. The facial geometry rendering image numbered t corresponds to the driving image numbered t. In this embodiment, for the driving video, multiple driving images are obtained by splitting according to the time position corresponding to each frame, and the time position corresponding to each driving image is used as the number (i.e., label) of the driving image. Therefore, in this embodiment, the value of t can start from 0 (t = 0, 1, 2, 3, 4...). N is a preset value, which can be set and adjusted according to the actual situation (for example, set to 3). The number of channels of the time encoding is 2N because there are two sets of encodings, sin and cos.
[0088] After obtaining the time encoding, each group of data to be processed can be formed in sequence. A group of data to be processed includes a reference image, a frame of facial geometry rendering image, and a time encoding. Each group of data to be processed is input into the above-mentioned avatar model in sequence, and each frame of the driven image can be obtained in sequence, and finally the driven video is combined. The human body in the above-mentioned driven video is the driven object, and the expressions and postures made by the driven object in the driving video are performed by the human body.
[0089] Specifically, the above-mentioned avatar model is pre-trained according to the following steps:
[0090] Input the reference image, the training facial geometry rendering image, and the training time encoding corresponding to the training facial geometry rendering image in the training data into the deep neural network generator. Through the above-mentioned deep neural network generator, a training driven image for the above-mentioned reference image and the above-mentioned training facial geometry rendering image is generated. Among them, the above-mentioned training data includes multiple groups of training image groups, and each group of training image groups includes a reference image corresponding to the above-mentioned driven object, a training facial geometry rendering image corresponding to the above-mentioned driving object, the training time encoding corresponding to the training facial geometry rendering image, and the training driving image corresponding to the training facial geometry rendering image;
[0091] According to the above-mentioned training driven image and the above-mentioned training driving image, adjust the model parameters of the above-mentioned avatar model, and continue to execute the above step of inputting the reference image, the training facial geometry rendering image, and the training time encoding corresponding to the training facial geometry rendering image in the training data into the above-mentioned avatar model until the preset training conditions are met, so as to obtain a trained deep neural network generator, and use the above-mentioned trained deep neural network generator as the above-mentioned avatar model;
[0092] Among them, the above-mentioned preset training conditions include the convergence of the reconstruction loss between the above-mentioned training driven image and the above-mentioned training driving image.
[0093] Specifically, the above training data is obtained by processing the collected image data through a preset data augmentation method, and the preset data augmentation method includes spatial random cropping. Among them, the above collected image data is the image data directly collected during training, and the training data is the data obtained after performing data augmentation operations on the collected image data. The above reconstruction loss is calculated through a multi-joint image reconstruction loss function, and the multi-joint image reconstruction loss function is used to jointly calculate at least two of the L1 reconstruction loss, perceptual loss, and GAN discriminator loss
[0094] In this embodiment, a specific application scenario is also used to specifically illustrate the training process of the above deep neural network generator and the process of video processing based on the avatar model. It should be noted that during the training and use of the model, the data used is corresponding or the same. For example, the training face geometry rendering image and the geometry rendering image are corresponding, and their name differences are used to distinguish the data used during the training process and the data used during the video processing process using the model. Their acquisition methods or processing methods can be referred to each other
[0095] Specifically, in this embodiment, the trained avatar model corresponds to the driven object (i.e., user A). Therefore, the training face geometry rendering image used during the training process corresponds to the driven object. The training face geometry rendering image is obtained through the training 3D face mesh during the training process, and the training 3D face mesh can be obtained through the training 3D face parameters, and the training 3D face parameters can be obtained through the training driving image, and the training driving image can be obtained through the captured training driving video of user A
[0096] Specifically, first, user A is photographed to obtain a video of user A speaking (i.e., the training driving video), and then it is split into multiple frames of training driving images in sequence. In this embodiment, during the training and use of the avatar model, the method of dividing the video into frames is the same. Therefore, during the training process, the time position corresponding to each frame is also marked as t (t = 0, 1, 2, 3, 4...). The obtained training driving image is denoted as I t , and then for each frame of training driving image I t 3D face parameter estimation is performed to obtain the training 3D face parameters, including the individual coefficient β t , expression coefficient and pose coefficient θ t . Then, based on a preset 3D face model (which can be denoted as f), the corresponding training 3D face mesh is calculated, and then based on a preset renderer Render, a 1-channel training face geometry rendering image M t ∈R H×W×1, where H and W respectively represent the height and width of the training facial geometry rendering image (or the training driving image), and the rendering process is shown in the following formula (2):
[0097]
[0098] Among them, Render represents the processing process of the renderer, and the function f represents the processing process of the 3D face model.
[0099] It should be noted that in the training stage of the avatar model, the purpose is to train a deep neural network generator to recover the original training driving image I with a face from the training facial geometry rendering image M t In addition, a reference image t and a time encoding TPE are introduced. t The reference image is an RGB image with 3 channels, that is Specifically, for the same avatar model, during the training of the deep neural network generator and the video processing based on the avatar model, the way of dividing image frames is the same, the reference image used is the same, and the setting method of the time encoding is also the same. Therefore, the time encoding during the training process can be specifically set with reference to the above formula (1), which will not be elaborated here. In this embodiment, the spatial dimension of the time encoding TPE t is the same as that of the training facial geometry rendering image M t and the reference image , and TPE t ∈R H×W×2N .
[0100] In this embodiment, the function of the above avatar model (denoted as g) is to generate an image with a face (i.e., the training driven image) I' t when given M t , TPE and t ∈R H×W×3 , as shown in the following formula (3):
[0101]
[0102] Among them, I' t represents the training driven image corresponding to the t-th frame training driving image, and g represents the processing process of the avatar model. The avatar model (i.e., the neural network generator) used in this embodiment is a convolutional neural network (such as UNet) from image to image (both the input and output are images and the spatial dimensions remain unchanged). During the training process, the neural network generator g is trained to make the predicted training driven image I' tThe corresponding training driving image I in the training video can be reconstructed t , specifically by minimizing the training driven image I′ t and the training driving image I t The model parameters in the avatar model g are iteratively optimized and updated with the goal of minimizing the reconstruction loss between them until the reconstruction loss converges to the minimum. It should be noted that the above preset training conditions may also include that the number of iterations reaches the iteration number threshold.
[0103] After obtaining the trained avatar model corresponding to the above driven object (i.e., user A), when the driving object (i.e., user B) wants to drive the driven object to perform corresponding actions and expressions, a driving video is collected for the driving object, and then video processing is performed to generate the corresponding driven video. Figure 2 is a specific process schematic diagram for generating a driven video based on the avatar model of user A provided by an embodiment of the present invention. As Figure 2 shown, in this embodiment, for each frame of the driving image of user B, 3D face parameters (including individual coefficients, expression coefficients, and pose coefficients) are extracted, and then they are converted into the space corresponding to user A. The goal of the conversion is to align the mean and variance of the converted coefficients, as shown in the following formula (4):
[0104]
[0105] Among them, and respectively represent the individual coefficient, expression coefficient, and pose coefficient corresponding to the space of user A obtained after conversion. T represents the alignment conversion process, and respectively represent the individual coefficient, expression coefficient, and pose coefficient corresponding to user B before conversion.
[0106] Using and render to obtain the facial rendering geometric image corresponding to user B Combined with time encoding and reference images, through the avatar model, the prediction of the driven image corresponding to user A with the current pose and expression is completed to obtain the corresponding driven image. The processing process can refer to the above formula (3) and its specific steps, which will not be elaborated here.
[0107] Thus, in this embodiment, a low-cost and highly realistic character avatar model is provided. For the character avatar model of user A, user A only needs to use a daily shooting device (such as a mobile phone) to shoot a corresponding training driving video (for example, a speaking video with a duration of about 2 minutes) in a scene (scene one), which can be used as the material for training the character avatar model. Moreover, based on a training driving video, multiple groups of training image sets can be obtained by using data augmentation and extension techniques. In a specific application scenario, the above training driving video is used to train a digital avatar model on a training platform for about 4 hours, and a digital avatar model that can support any head movements and expressions of user A in the same scene (scene one) can be obtained. In the subsequent usage process, user B can record a driving video in any scene (scene two) to drive the above digital avatar model to generate a driven video of user A speaking in scene one, and the postures and expressions of user A in the driven video are the same as those of user B in the driving video. At the same time, in this embodiment, time encoding is combined to ensure that the generated driven video is real and natural, and has high fluency and stability in the time domain, and can achieve a generation result similar to that of a real shooting video.
[0108] Specifically, during the training process of the character avatar model, data augmentation methods including spatial random cropping can be used, and the neural network generator is trained by optimizing the multi-joint image reconstruction loss function (such as L1 reconstruction loss, perceptual loss, GAN discriminator loss). The training is carried out on an NVIDIA A100-SXM4-40GB GPU, with a Batch Size of 20 and an input-output image resolution of 512*512. Among them, data augmentation means adding spatial random cropping during the training process to enhance data diversity. For example, the video frames shot by user A are randomly cropped spatially and then input into the neural network generator, and the parameters of the neural network generator are adjusted by calculating with the multi-joint image reconstruction loss function to obtain a neural network generator that meets the training expectations. Among them, calculating the loss means calculating the loss between the image generated by the neural network generator (i.e., the character avatar model) for the driven object and its corresponding original driving image.
[0109] It should be noted that the video processing method based on the character avatar model provided in this embodiment has high efficiency both in the model training and the process of rendering a new video. In an application scenario, a model with stable convergence is trained under the same conditions (the same duration of training materials and an image resolution of 512*512). The DVP scheme requires an average of 42 hours of training and 0.2 seconds per frame on average when rendering the video; the NerFace scheme requires an average of 55 hours of training and 6 seconds per frame on average when rendering the video; while in this embodiment, the average training time is 4 hours and 0.03 seconds per frame on average when rendering the video. It can be seen that the scheme of this embodiment is beneficial to improving the training and processing efficiency.
[0110] Furthermore, for the quality of the generated driven video, the Structural Similarity Index (SSIM) and the Peak Signal-to-Noise Ratio (PSNR) can be used as reconstruction metrics for evaluation. The higher these two metrics are, the smaller the time gap between the generated driven video and the real driving video, and the higher the quality of the generated driven video. The structural similarity corresponding to the driven video generated based on the solution of this embodiment is greater than 95%, and the peak signal-to-noise ratio is greater than 26.85, indicating that the quality of the generated driven video is relatively high.
[0111] As can be seen from the above, in the video processing method based on the avatar model provided by the embodiment of the present invention, not only the face regions in the images of different objects are intercepted and replaced, but an avatar model for the driven object is preset. After obtaining the driving video corresponding to the driving object and passing the permission verification, the avatar model and the reference image of the corresponding driven object are obtained. Then, according to the driving video, a facial geometry rendering image for reflecting the expression and posture of the driving object is obtained. Combining with the time encoding, the facial geometry rendering image and the reference image are fused through the avatar model of the driven object, thereby obtaining the driven video.
[0112] The driven video is not obtained by simply replacing the image of the face region, but by fusing the expression and posture of the driving object with the actual texture of the driven object, so as to enable the driven object to execute the same expression and posture as the driving object in the driving video. Since the facial geometry rendering image only reflects the expression and posture and does not reflect the actual texture of the face of the driving object, and the actual texture is only provided by the reference image of the driven object, the actual texture corresponding to the driving object will not be incorrectly retained when the avatar model performs image information fusion. That is, the image texture details of the object shown in the final driven video are the same as those of the driven object. In this way, using the video of the driving object to drive the avatar model is beneficial to improving the video display effect of the avatar model. It can realize driving the face of the driven object to make corresponding expressions with the expression of the driving object, which is beneficial to improving the effect of video face replacement and the user experience.
[0113] Exemplary Device
[0114] As Figure 3 shown in [reference], corresponding to the above video processing method based on the avatar model, the embodiment of the present invention further provides a video processing system based on the avatar model. The above video processing system based on the avatar model includes:
[0115] A driving information acquisition module 510 is configured to acquire a driving video of a driving object, authentication information of the driving object, and a driven object corresponding to the driving object, where the driving video is obtained by photographing the expression and posture of the driving object;
[0116] An authentication module 520 is configured to obtain a role avatar model and a reference image corresponding to the driven object when the authentication information of the driving object meets the authentication conditions of the driven object;
[0117] A driving video processing module 530 is configured to obtain multiple frames of facial geometry rendering images corresponding to the driving object according to the driving video, where the facial geometry rendering images are used to reflect the expression and posture of the driving object;
[0118] A driven video generation module 540 is configured to obtain time codes corresponding to the facial geometry rendering images, and generate a driven video through the role avatar model according to the reference image, the facial geometry rendering images, and the time codes corresponding to the facial geometry rendering images, where the driven object in the driven video performs the same expression and posture as the driving object in the driving video.
[0119] It should be noted that the specific structures and implementation manners of the above video processing system based on the role avatar model and its various modules or units can refer to the corresponding descriptions in the above method embodiments, and will not be elaborated here.
[0120] It should be noted that the division manner of the various modules of the above video processing system based on the role avatar model is not unique, and will not be specifically limited here.
[0121] Based on the above embodiments, the present invention further provides an intelligent terminal, and its principle block diagram can be as Figure 4 shown. The intelligent terminal includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a video processing program based on the role avatar model. The internal memory provides an environment for the operation of the operating system and the video processing program based on the role avatar model in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. When the video processing program based on the role avatar model is executed by the processor, it implements the steps of any one of the above video processing methods based on the role avatar model. The display screen of the intelligent terminal can be a liquid crystal display screen or an electronic ink display screen.
[0122] Those skilled in the art can understand, Figure 4The principle block diagram shown only shows the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the intelligent terminal to which the solution of the present invention is applied. The specific intelligent terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0123] In one embodiment, an intelligent terminal is provided. The intelligent terminal includes a memory, a processor, and a video processing program based on a role avatar model stored on the memory and executable on the processor. When the video processing program based on the role avatar model is executed by the processor, the steps of any one of the video processing methods based on the role avatar model provided by the embodiments of the present invention are implemented.
[0124] The embodiments of the present invention also provide a computer-readable storage medium. A video processing program based on a role avatar model is stored on the computer-readable storage medium. When the video processing program based on the role avatar model is executed by a processor, the steps of any one of the video processing methods based on the role avatar model provided by the embodiments of the present invention are implemented.
[0125] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0126] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the above device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working process of the units and modules in the above device can refer to the corresponding process in the foregoing method embodiments and will not be described herein again.
[0127] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0128] Those of ordinary skill in the art will appreciate that the units and algorithm steps of the various examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0129] In the embodiments provided by the present invention, it should be understood that the disclosed system / terminal device and method can be implemented in other ways. For example, the system / terminal device embodiments described above are merely illustrative. For example, the above-mentioned division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0130] If the above-mentioned integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present invention, it can also be completed by a computer program instructing the relevant hardware. The above-mentioned computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the above-mentioned computer program includes computer program code, and the above-mentioned computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The above-mentioned computer-readable medium can include: any entity or device capable of carrying the above-mentioned computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the above-mentioned computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.
[0131] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not deviate from the spirit and scope of the technical solutions of the present invention in essence, and should all be included in the protection scope of the present invention.
Claims
1. A video processing method based on an avatar model, characterized in that, The video processing method based on the role avatar model includes: Obtaining a driving video of a driving object, permission verification information of the driving object, and a driven object corresponding to the driving object, wherein the driving video is obtained by photographing the expression and posture of the driving object; When the permission verification information of the driving object meets the permission verification condition of the driven object, obtaining a role avatar model and a reference image corresponding to the driven object; Obtaining multiple frames of facial geometry rendering images corresponding to the driving object according to the driving video, wherein the facial geometry rendering images are used to reflect the expression and posture corresponding to the driving object; Obtaining time codes corresponding to the facial geometry rendering images according to a preset time coding calculation formula; Successively form each group of data to be processed, and successively input each group of data to be processed into the avatar model to successively obtain the driven images corresponding to each group of the data to be processed. Wherein, a group of the data to be processed consists of the reference image, one frame of the facial geometry rendering image, and a time code corresponding to the facial geometry rendering image. The time code is used to input time information into the avatar model. The time code, the facial geometry rendering image corresponding to the time code, and the reference image have the same spatial dimensions. In the driven image, the driven object performs the same expression and posture as those in the corresponding facial geometry rendering image. The time code calculation formula is: , wherein, represents the time code corresponding to the facial geometry rendering image numbered t, and N is a preset constant; Sequentially connecting the driven images according to the corresponding time codes to generate a driven video, wherein in the driven video, the driven object performs the same expressions and postures as the driving object in the driving video.
2. The video processing method based on an avatar model according to claim 1, characterized in that, The reference image is used to provide image texture details corresponding to the driven object for the role avatar model, and the image texture details of the driven video are the same as those of the reference image.
3. The video processing method based on an avatar model according to claim 1, characterized in that, The obtaining multiple frames of facial geometry rendering images corresponding to the driving object according to the driving video includes: Splitting the driving video to obtain multiple frames of driving images; Respectively extracting three-dimensional facial parameters corresponding to the driving images; Respectively obtaining three-dimensional facial meshes corresponding to the driving images according to the three-dimensional facial parameters corresponding to the driving images; Rendering the three-dimensional facial meshes corresponding to the driving images to obtain facial geometry rendering images corresponding to the driving images, wherein the facial geometry rendering images are grayscale images.
4. The video processing method based on an avatar model according to claim 3, characterized in that, Before respectively obtaining the three-dimensional facial meshes corresponding to the driving images according to the three-dimensional facial parameters corresponding to the driving images, the method further includes: Aligning the three-dimensional facial parameters based on the facial spatial position corresponding to the driven object in the role avatar model to update the three-dimensional facial parameters.
5. The video processing method based on an avatar model according to claim 3, characterized in that, The three-dimensional facial parameters include a body coefficient, an expression coefficient, and a posture coefficient.
6. The video processing method based on an avatar model according to any one of claims 1-5, characterized in that, The role avatar model is pre-trained according to the following steps: Inputting a reference image, a training facial geometry rendering image, and a training time code corresponding to the training facial geometry rendering image in the training data into a deep neural network generator, and generating a training driven image for the reference image and the training facial geometry rendering image through the deep neural network generator, wherein the training data includes multiple groups of training image groups, and each group of training image groups includes a reference image corresponding to the driven object, a training facial geometry rendering image corresponding to the driving object, a training time code corresponding to the training facial geometry rendering image, and a training driving image corresponding to the training facial geometry rendering image; Adjust the model parameters of the deep neural network generator according to the training driven image and the training driving image, and continue to execute the step of inputting the reference image, the training face geometry rendering image, and the training time encoding corresponding to the training face geometry rendering image in the training data into the deep neural network generator until a preset training condition is met, so as to obtain the avatar model; Among them, the preset training condition includes the convergence of the reconstruction loss between the training driven image and the training driving image.
7. The video processing method based on an avatar model according to claim 6, characterized in that, The training data is obtained by processing the collected image data through a preset data augmentation method, and the preset data augmentation method includes spatial random cropping.
8. The video processing method based on an avatar model according to claim 6, characterized in that, The reconstruction loss is obtained by calculating through a multi-joint image reconstruction loss function, and the multi-joint image reconstruction loss function is used to jointly calculate at least two of the L1 reconstruction loss, the perceptual loss, and the GAN discriminator loss.
9. A video processing system based on an avatar model, characterized in that, The video processing system based on the avatar model includes: A driving information acquisition module, configured to acquire the driving video of the driving object, the permission verification information of the driving object, and the driven object corresponding to the driving object, wherein the driving video is obtained by photographing the expression and posture of the driving object; A permission verification module, configured to obtain the avatar model and the reference image corresponding to the driven object when the permission verification information of the driving object meets the permission verification condition of the driven object; A driving video processing module, configured to obtain multiple frames of face geometry rendering images corresponding to the driving object according to the driving video, wherein the face geometry rendering image is used to reflect the expression and posture corresponding to the driving object; A driven video generation module, configured to obtain the time code corresponding to each of the facial geometry rendering images according to a preset time coding calculation formula; sequentially form groups of data to be processed, and sequentially input the groups of data to be processed into the character avatar model, and sequentially obtain the driven images corresponding to the groups of data to be processed, wherein a group of data to be processed is composed of the reference image, one frame of the facial geometry rendering image, and a time code corresponding to the facial geometry rendering image, the time code is used to input time information into the character avatar model, the spatial dimensions of the time code, the facial geometry rendering image corresponding to the time code, and the reference image are the same, and in the driven image, the driven object performs the same expression and posture as in the corresponding facial geometry rendering image, and the time coding calculation formula is: , where represents the time code corresponding to the facial geometry rendering image numbered t, and N is a preset constant; sequentially connect the driven images according to the corresponding time codes to generate a driven video, wherein in the driven video, the driven object performs the same expression and posture as the driving object in the driving video.
10. An intelligent terminal, characterized in that, The intelligent terminal includes a memory, a processor, and a video processing program based on the avatar model stored on the memory and executable on the processor. When the video processing program based on the avatar model is executed by the processor, it implements the steps of the video processing method based on the avatar model according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a video processing program based on the avatar model. When the video processing program based on the avatar model is executed by a processor, it implements the steps of the video processing method based on the avatar model according to any one of claims 1-8.
Citation Information
Patent Citations
Face driving and live broadcasting method and device, computer equipment and storage medium
CN113486787A