Video generation method, device, equipment and storage medium
By combining the FaceFormer model and the styleGANv2 model, the naturalness and accuracy of lip movements in virtual human videos are improved, solving the problem of unnatural lip movements in existing technologies and enhancing the interactive capabilities of digital humans.
Patent Information
- Application Number
- CN202410693165.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-05-30
AI Technical Summary
In the existing technology, the lip movements of virtual humans in virtual human videos have a low degree of correspondence with voice, the lip movements are not natural enough, and the facial movements are inaccurate, resulting in insufficient interactive capabilities of the digital human.
The FaceFormer model, based on the Transformer architecture, combines the speech file and the target facial video to generate speech-driven facial videos through a two-stage process. The first stage generates a 3D facial mesh sequence, which is then rendered and fitted to the target facial video to ensure the continuity and accuracy of lip movements. The improved styleGANv2 model is then used to generate high-fidelity speech-driven facial videos.
It improves the naturalness and accuracy of lip movements in virtual human videos, enhances the interactive capabilities of digital humans, makes them more realistic and humane, and is suitable for a variety of business scenarios.
Smart Images

Figure CN118691717B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as computer vision and deep learning, and can be applied to scenarios such as content generation based on artificial intelligence. Background Art
[0002] With the development of artificial intelligence (AI), generative AI (AIGC) content creation tools have become a crucial technology for assisting humans, significantly improving technological productivity and work efficiency. Virtual human (or digital human) technology is a key component of AIGC. It uses deep learning and other technologies to create virtual characters with human-like interactive capabilities. Summary of the Invention
[0003] The present disclosure provides a video generation method, apparatus, device, and storage medium.
[0004] According to one aspect of the present disclosure, there is provided a video generation method, comprising:
[0005] Inputting a voice file and a face video to be driven into a pre-trained first model, and having the first model output a three-dimensional face mesh sequence; wherein the three-dimensional face mesh sequence corresponds to the voice features of the voice file, and corresponds to the facial features and speaking style features of the face video to be driven;
[0006] Based on the three-dimensional face mesh sequence and the face video to be driven, a voice-driven face video matching the voice file is generated.
[0007] According to another aspect of the present disclosure, a model training method is provided, comprising:
[0008] Generate multiple face video samples using a face reconstruction model;
[0009] Using the multiple face video samples and voice samples, establish a first training set and a first test set;
[0010] The first model is trained using the first training set and the first test set so that the first model can generate a three-dimensional face mesh sequence based on the voice file and the face video to be driven; wherein the three-dimensional face mesh sequence corresponds to the voice features of the voice file, and corresponds to the face features and speaking style features of the face video to be driven.
[0011] According to another aspect of the present disclosure, a model training method is provided, comprising:
[0012] Generate multiple face video samples using a face reconstruction model;
[0013] Using the plurality of face video samples and the two-dimensional face image sequence samples, a second training set and a second test set are established;
[0014] The second model is trained using the second training set and the second test set, so that the second model can generate a voice-driven face video based on the two-dimensional face image sequence and the face video to be driven.
[0015] According to another aspect of the present disclosure, there is provided a video generating apparatus, comprising:
[0016] An input / output module, configured to input a voice file and a facial video to be driven into a pre-trained first model, and have the first model output a three-dimensional facial mesh sequence; wherein the three-dimensional facial mesh sequence corresponds to the voice features of the voice file, and corresponds to the facial features and speaking style features of the facial video to be driven;
[0017] The first generating module is used to generate a voice-driven face video that matches the voice file based on the three-dimensional face mesh sequence and the face video to be driven.
[0018] According to another aspect of the present disclosure, there is provided a model training device, comprising:
[0019] The second generation module is used to generate multiple face video samples using a face reconstruction model;
[0020] A first training set generating module, configured to establish a first training set and a first test set using the plurality of face video samples and voice samples;
[0021] The first training module is used to train the first model using the first training set and the first test set, so that the first model can generate a three-dimensional face mesh sequence based on the voice file and the face video to be driven; wherein the three-dimensional face mesh sequence corresponds to the voice features of the voice file, and corresponds to the face features and speaking style features of the face video to be driven.
[0022] According to another aspect of the present disclosure, there is provided a model training device, comprising:
[0023] The third generation module is used to generate multiple face video samples using the face reconstruction model;
[0024] A second training set generating module is used to establish a second training set and a second test set by using the plurality of face video samples and the two-dimensional face image sequence samples;
[0025] The second training module is used to train the second model using the second training set and the second test set, so that the second model can generate a voice-driven face video based on the two-dimensional face image sequence and the face video to be driven.
[0026] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0027] at least one processor; and
[0028] a memory communicatively connected to the at least one processor; wherein,
[0029] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.
[0030] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.
[0031] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.
[0032] This video generation method, proposed in this disclosure, is used to generate facial videos driven by voice files. It consists of two stages: the first stage generates a three-dimensional facial mesh sequence based on the features of the voice file and the facial video to be driven; the second stage, based on this three-dimensional facial mesh sequence, generates a voice-driven facial video that matches the voice file. The two stages have clear divisions of labor: the first stage extracts lip movement trends, while the second stage ensures the naturalness of the generated voice-driven facial video, achieving lossless facial reconstruction, thereby improving the naturalness and accuracy of the generated voice-driven facial video.
[0033] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0035] Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure;
[0036] Figure 2 is a flow chart of an implementation method of a video generation method according to an embodiment of the present disclosure;
[0037] Figure 3 is a schematic diagram of input and output content of the FaceFormer model in one embodiment of the present disclosure;
[0038] Figure 4 This is a schematic diagram of the structure and input and output contents of a first model according to an embodiment of the present disclosure;
[0039] Figure 5 This is a video generation process and effect diagram of an embodiment of the present disclosure;
[0040] Figure 6 This is a flowchart of a model training method implementation in an embodiment of the present disclosure;
[0041] Figure 7 is a flowchart of another model training method implementation in an embodiment of the present disclosure;
[0042] Figure 8 is a structural diagram of a video generating device 800 according to an embodiment of the present disclosure;
[0043] Figure 9 is a structural diagram of a video generating device 900 according to another embodiment of the present disclosure;
[0044] Figure 10 1 is a schematic structural diagram of a model training device 1000 according to an embodiment of the present disclosure;
[0045] Figure 11 is a structural diagram of a model training device 1100 according to an embodiment of the present disclosure;
[0046] Figure 12 A schematic block diagram of an example electronic device 1200 is shown, which may be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION
[0047] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0048] The “and / or” in the embodiments of the present disclosure indicates that there may be three relationships. For example, A and / or B may indicate three situations: A exists alone, A and B exist at the same time, and B exists alone. The term “at least one” herein indicates any combination of at least two of any one or more of a plurality of. For example, at least one of A, B, and C may indicate any one or more elements selected from the set consisting of A, B, and C. The terms “first” and “second” herein refer to and distinguish between multiple similar technical terms, and do not mean to limit the order or to limit the meaning to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature may be one or more, and the second feature may also be one or more.
[0049] Virtual humans (digital humans) are avatars built using deep learning and other technologies. They possess the same interactive capabilities as humans, speaking and gesturing like humans. These characteristics give digital humans enormous potential for development and application value. In fields such as digital finance, live streaming, news broadcasting, customer service, and education, they have evolved into financial services digital humans, customer service digital humans, virtual digital humans, virtual teachers, digital avatars, live streaming digital humans, and virtual news anchors. Combined with large models and multimodal technologies, digital humans possess excellent interactive capabilities. Existing technologies often lack the ability to accurately align the lip movements of virtual humans in virtual human videos with their speech. Their lip movements are unnatural, and their lip or facial movements are inaccurate.
[0050] How to make digital humans more humane, realistic, and natural is a technical problem that needs to be solved in current digital humans. To solve the above problem, the present disclosure proposes a video generation method. Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure, such as Figure 1As shown, the application scenario diagram of the embodiment of the present disclosure may include, but is not limited to, a client device 110 and a video generation device 120. The client device 110 and the video generation device 120 can communicate with each other via any type of wired or wireless network. Specifically, based on a user request, the client device 110 can input a facial video to be driven and a voice file to the video generation device 120. The facial video to be driven is a pre-arranged full-length or half-length portrait video, which serves as the driven video base; the voice file contains a voice file of a period of time, which is used to drive the character image in the video base. The video generation device 120 receives the facial video to be driven and the voice file sent by the client device 110, and generates a voice-driven facial video that matches the voice file. The character image in the voice-driven facial video is the same as the character image in the facial video to be driven. The so-called "match" means that the lip movements of the character in the voice-driven facial video correspond to the voice file it matches. For example, the voice-driven face video corresponding to the voice file includes multiple frames of images. Each short segment of voice in the voice file corresponds to one frame of image, and the pronunciation of the voice segment is consistent with the lip shape of the person in the corresponding image.
[0051] The client device 110 proposed in the embodiments of the present disclosure includes, but is not limited to, electronic devices such as mobile phones, computers, intelligent voice interaction devices, smart home appliances, in-vehicle terminals, game consoles, e-book readers, multimedia playback devices, and wearable devices. The video generation device 120 may include an electronic device or server that provides video generation capabilities for the behavioral client device 110. Furthermore, the embodiments of the present disclosure do not impose any specific restrictions on the number of client devices 110. For example, the application scenario diagram of the embodiments of the present disclosure may include one or more client devices 110.
[0052] Figure 2 1 is a flowchart of a video generation method according to an embodiment of the present disclosure, including:
[0053] S210: Inputting the voice file and the face video to be driven into a pre-trained first model, and having the first model output a three-dimensional face mesh sequence; wherein the three-dimensional face mesh sequence corresponds to the voice features of the voice file and corresponds to the facial features and speaking style features of the face video to be driven;
[0054] S220: Generate a voice-driven face video that matches the voice file based on the three-dimensional face mesh sequence and the face video to be driven.
[0055] In some examples, an audio file (audio) refers to a continuous or discontinuous segment of speech on a timeline, such as a word, phrase, or sentence with specific semantics. It can be natural speech spoken by a natural person, pre-recorded speech, or speech obtained by synthesizing text information through speech synthesis technology.
[0056] The face video to be driven can be a pre-edited full-length or half-length portrait video, serving as the base video to be driven. The voice file then drives the person in the face video to speak or sing, while the background remains unchanged. Because the base video serves as the base for the dynamic spoken word, the lip movements in the generated voice-driven face video are more accurate and coherent. The person in the face video to be driven can be a natural person or a virtual person.
[0057] The video generation method proposed in the disclosed embodiments utilizes two stages to achieve complete voice-driven face recognition. The first stage generates a 3D face mesh sequence corresponding to an audio file (audio); the second stage renders the 3D face mesh sequence into a 2D face image sequence. The rendered 2D face sequence and the face video to be driven (i.e., the base video) are then used to generate a 2D face video that matches the lip shape of the rendered face, i.e., the voice-driven face video described above. The two stages have clear divisions of labor: the first stage extracts lip movement trends, while the second stage renders the face to a real person, ensuring naturalness. The video generation method proposed in the disclosed embodiments incorporates the face video to be driven in the first 3D driving stage, incorporating facial features and speaking style features into the generated 3D face mesh sequence. Furthermore, in the second stage, facial features, speaking style features, lip movement features, and spatial features of the face frame are integrated to improve the naturalness and accuracy of the generated video.
[0058] In some embodiments, the first model used in the first stage includes a face-driven model based on a Transformer architecture, such as a FaceFormer model. The FaceFormer model is an autoregressive model capable of synthesizing a 3D face mesh sequence corresponding to a speech file. The sequence includes multiple 3D face meshes, each of which contains multiple face vertex data.
[0059] In some embodiments, a voice file is continuous voice information on a timeline. Accordingly, the generated voice-driven facial video includes multiple frames of facial images that are time-matched with the voice file. For example, the voice file includes multiple voice segments, each corresponding to a facial frame in the voice-driven facial video, and the pronunciation of the voice segment is consistent with the lip shape of the person in the corresponding image.
[0060] The video file generated using the embodiment of the present disclosure can be used to show the interaction between a virtual person and its interactive object in a human-computer interaction scenario, generate corresponding voice and video files to provide feedback on the input information of the interactive object, and play the voice and video files synchronously.
[0061] The first model used in the disclosed embodiments can be a FaceFormer model based on the Transformer architecture, an autoregressive model based on the Transformer structure. The disclosed embodiments modify the existing FaceFormer model, using not only voice files as input but also the video of the face to be driven as input. Figure 3 FIG. 1 is a schematic diagram of input and output content of the FaceFormer model in one embodiment of the present disclosure. Figure 3 As shown, the voice file and the face video to be driven are input into the FaceFormer model. The FaceFormer model first outputs a three-dimensional face mesh (3D face mesh, or 3D mesh) that matches the first video frame corresponding to the voice file, referred to as the first 3D mesh. The first 3D mesh is then input into the FaceFormer model. Based on the voice file, the face video to be driven, and the first 3D mesh, the FaceFormer model generates a 3D mesh that matches the second video frame corresponding to the voice file, referred to as the second 3D mesh. This continues in this manner, with each 3D mesh generated being based on the previously generated 3D mesh. Ultimately, the FaceFormer model outputs a series of 3D meshes, forming a 3D mesh sequence.
[0062] Unlike the existing FaceFormer model, the embodiment of the present disclosure extracts the facial features and speaking style features of the face video to be driven from the face video to be driven, and generates a corresponding three-dimensional face mesh sequence (3D Face mesh sequence, referred to as 3D mesh sequence) based on the voice file and the facial features and speaking style features of the face video to be driven. Therefore, the generated 3D mesh sequence can retain the appearance features and speaking style features of the character.
[0063] The first model proposed in the embodiment of the present disclosure may include an encoding module and a decoding module. Figure 4 This is a schematic diagram of the structure and input and output contents of a first model of an embodiment of the present disclosure. Figure 4 As shown, the output end of the encoding module is connected to the input end of the decoding module; the voice file and the face video to be driven are input into the encoding module of the first model;
[0064] The encoding module is used to extract the voice features of the voice file and extract the facial features and speaking style features of the face video to be driven;
[0065] The decoding module is used to obtain a three-dimensional face mesh sequence based on the voice feature, the face feature and the speaking style feature.
[0066] The decoding module generates each 3D mesh in the 3D mesh sequence one by one. The previous 3D mesh generated by the decoding module is returned to the decoding module as input information for generating the next 3D mesh; multiple 3D meshes generated by the decoding model constitute a 3D mesh sequence.
[0067] The disclosed embodiment performs contextual encoding of speech based on the FaceFormer model, so that the model's currently predicted 3D facial mesh (or vertices) is obtained based on the model's previous prediction results. In this way, as the length of speech accumulates, the FaceFormer model will continuously correct the previous and current prediction results, ensuring the continuity and accuracy of the lip movements of the entire model.
[0068] In some embodiments, the decoding module of the FaceFormer model may include a waveform-to-vector conversion (wav2vec) unit, an identifier extraction unit (ID_extractor), and a style extraction unit (style extractor). The wav2vec unit is used to extract speech features from a speech file; the ID extractor is used to extract facial features of the face video to be driven; and the style_extractor is used to extract speaking style features of the face video to be driven. This embodiment excludes features that are irrelevant or weakly related to speech, such as facial angle, so that the features learned by the FaceFormer model can be applied to each ID.
[0069] In one example, when extracting facial features of a face video to be driven, the encoding module (such as the IDextractor in the encoding module) is specifically used to:
[0070] Extracting facial feature parameters from each frame of the face video to be driven;
[0071] Based on the facial feature parameters of each frame image, the facial features of the face video to be driven are determined.
[0072] Here, the extracted facial feature parameters may include ID coefficients. This embodiment retains the ID coefficients and removes coefficients such as exp (expression), trans (translation coefficient), and angle (rotation angle) from each video series frame. Based on the ID coefficients of each frame, an average ID coefficient is determined to obtain ID template (ID_template) coefficients. Rendering is then performed based on the ID template (ID_template) coefficients to obtain the facial features of the face video to be driven.
[0073] Since everyone's speaking style is different, in order to perfectly clone a digital human, in addition to the person's appearance features, the speaking style should also be retained. In view of this, the embodiment of the present disclosure uses the face ID coefficient and expression (exp) coefficient of all frames of each ID video as speaking style features.
[0074] In one example, when extracting the speaking style features of a face video to be driven, the encoding module (such as the style extractor in the encoding module) is specifically used to:
[0075] Extracting facial feature parameters and expression parameters from each frame of the face video to be driven;
[0076] Based on the facial feature parameters and expression parameters of each frame image, the speaking style features of the face video to be driven are determined.
[0077] Here, the extracted facial feature parameters may include ID coefficients, and the extracted expression parameters may include exp coefficients.
[0078] In the embodiment of the present disclosure, the speech features of the speech file are obtained through the self-supervised wav2vec model (i.e., the above-mentioned wav2vec unit). In the process of training the FaceFormer model, the wav2vec model does not participate in the training, that is, the parameters of the wav2vec model are fixed, ensuring the richness and recognizability of the speech features. The FaceFormer model proposed in the embodiment of the present disclosure uses an autoregressive method to learn lip movement features from speech features (excluding coefficients such as angle and trans that are not related to lip movement). Lip movement depends not only on the current speech features, but also on the historical lip movements predicted by all previous frames. As the speech features accumulate, the lip movements are constantly self-corrected to ensure the accuracy and consistency of the lip movements. At the same time, the model has the ability to learn speaking styles, so that the driven digital person speaks in the same style as the person himself, greatly improving the recognizability of the digital person and having the ability to be a true digital clone.
[0079] The above describes the relevant technologies of the first stage. The first stage uses the first model (such as the FaceFormer model) to generate a 3D face mesh sequence based on the voice file and the face video to be driven. The first stage determines the lip movement ability of the digital human.
[0080] The second stage generates a voice-driven face video that matches the voice file based on the three-dimensional face mesh sequence generated in the first stage and the face video to be driven. This stage ultimately determines the naturalness and high fidelity of the digital human. To balance video quality and real-time performance, the disclosed embodiment uses a style-based generative adversarial network (styleGAN) model, such as the styleGANv2 model, in the second stage. The styleGANv2 model is improved to design a lightweight neural network structure for generating voice-driven face videos that match the voice file. The model has a simple structure and can save computing resources.
[0081] The three-dimensional face mesh sequence generated in the first stage is a high-dimensional feature. Extracting lip movement features from high-dimensional features will increase the burden on the neural network and may lose important information such as spatial position, which will increase the difficulty of model training and slow down the training speed. Taking this into account, in some embodiments, the output content of the first stage is not directly used as the input of the second stage model. Instead, the three-dimensional face mesh sequence output by the first stage (plus the angle and trans coefficients of the base face) is first rendered to a 2D face, and the high-dimensional features are directly mapped to low-dimensional 2D face image features; then the lower half of the face is pasted back to the face of the video base (thus turning the purpose of the model to learning to generate a real lower half of the face), and at the same time, any frame of the video base, or a specific frame of the video base, or a screened frame of the video base is used as a reference image (providing texture and face ID information for the model). This design greatly simplifies the difficulty of model training and improves the generation effect. The generated result retains the same quality as the rendered 2D face. Figure 1 The lip movements are similar, and the facial spatial position that is consistent with the video base frame is retained in detail, and the facial ID, color and texture information that are consistent with the video base frame are retained as a whole. Among them, the filtered frame of the video base can be an image that can provide the model with clearer texture and face ID information, such as an image with the head facing forward, no occlusion on the face, clear facial texture, and / or a non-exaggerated facial expression. For this reason, after the two stages, the digital human presents the characteristics of high fidelity, high naturalness and authenticity. In addition, in the above process, the three-dimensional face mesh sequence output by the first stage (plus the angle and trans coefficients of the base face) is first rendered to a 2D face. This method is simple and clear, and has super strong spatial position information, which makes it easy for the model to learn real lip movements.
[0082] Taking the above into consideration, in some embodiments, in the second stage, the embodiments of the present disclosure generate a voice-driven face video that matches the voice file based on the three-dimensional face mesh sequence and the face video to be driven, including:
[0083] Rendering the three-dimensional face mesh sequence to obtain a two-dimensional face image sequence;
[0084] Fitting the two-dimensional face image sequence to the face video to be driven to obtain a fitted image sequence;
[0085] The fitted image sequence and a frame of image in the face video to be driven are input into the second model to obtain the voice-driven face video.
[0086] When the two-dimensional face image sequence is fitted with the face video to be driven, processing can be performed frame by frame, and the lower half of each two-dimensional face image in the two-dimensional face image sequence is pasted back to the corresponding frame in the face video to be driven.
[0087] Figure 5 This is a video generation process and effect diagram of an embodiment of the present disclosure. Figure 5 As shown in Figure 2, the video generation process consists of two stages:
[0088] (1) Phase 1:
[0089] The first phase mainly includes the following steps:
[0090] (1) Input the speech file into the waveform to vector conversion (wav2vec) unit of the FaceFormer model, and the wav2vec unit extracts the speech features of the speech file; render the face video to be driven, and input the rendering result into the identification extraction unit (ID extractor) and style extraction unit (style_extractor) of the FaceFormer model. The identification extraction unit (ID_extractor) extracts the facial features of the face video to be driven, and the style extraction unit (style_extractor) extracts the speaking style features of the face video to be driven.
[0091] (2) The speech features of the speech file, the facial features of the face video to be driven, and the speaking style features are input into the decoding module of the FaceFormer model. The decoding module outputs a sequence of three-dimensional face meshes (3DFace meshes). The decoding module outputs one 3D face mesh in the sequence each time. The current output 3D face mesh is used as the input information of the decoding model and is used by the decoding module to generate the next 3D face mesh.
[0092] (2) Second stage:
[0093] The second phase mainly includes the following steps:
[0094] (1) Render the 3D Face mesh output from the first stage to obtain a two-dimensional face image sequence.
[0095] (2) The two-dimensional face image sequence is aligned with the face video to be driven to obtain an aligned image sequence; Figure 5 In the example, the image with the upper half being the face image and the lower half being the 3D face mesh rendering result is called a fitted image. Multiple fitted images constitute a fitted image sequence. During the fitting process, the 2D face sequence and the face video to be driven are processed frame by frame. For example, the 2D face sequence includes N 2D faces, and the face video to be driven includes N frames. The N 2D faces correspond one-to-one with the N frames, where N is a positive integer. During the fitting process, the lower half of the 2D face is combined with the upper half of the corresponding frame frame by frame to obtain the corresponding fitted image. Multiple fitted images constitute a fitted image sequence.
[0096] (3) The fitted image sequence and any frame image in the face video to be driven (such as Figure 5 The reference image (ref_img) shown is input into a generator (such as the styleGANv2 model), which then outputs a speech-driven facial video. The facial image in this speech-driven facial video is consistent with the facial image in the face video to be driven, and the lip shape of the character in this speech-driven facial video matches the speech file. Specifically, the duration of this speech-driven facial video is the same as that of the speech file. The speech-driven facial video includes multiple frames, each of which corresponds to a short speech segment in the speech file, and the lip shape of the face in this frame matches the corresponding short speech segment.
[0097] It can be seen from the above process that the video generation method proposed in the embodiment of the present disclosure implements a speech-to-face mesh (audio2mesh) prediction based on the FaceFormer autoregressive method in the first stage of video generation. The extracted lip movements not only depend on the current speech features, but also on the historical lip movements predicted in all previous frames; as the speech (audio) accumulates, the lip movements continuously self-correct to ensure the accuracy and consistency of the lip movements; and the FaceFormer model realizes expression (lip shape) decoupling, eliminates the influence of facial texture, color, angle and displacement, and greatly improves the accuracy of lip movements; and gives the model the ability to transfer speaking style learning. For each different ID character, the model can learn the person's speaking style based on the video base, and finally display it on the digital human clone. In the second stage of video generation, a high-performance and high-fidelity generator model structure was designed. The styleGANv2 model structure was optimized to achieve real-time inference. In the second stage, the predicted face mesh sequence from the first stage is not directly used. Instead, the face mesh sequence is first rendered into a 2D face. The 2D face and the video base frame are then used as input, thus avoiding the redundancy of spatial features in high-dimensional feature learning and reducing the difficulty of model training. A reference face (ref_img) is also injected. ref_img contains stable ID, texture, and color information, guiding the model to generate the final voice-driven face video. This ensures that the ID of the generated voice-driven face video remains unchanged, thus ensuring consistency between the generated face and the original face in the driven face video. Ultimately, the digital human has high fidelity, high naturalness, and strong identity restoration capabilities. Through the division of these two stages, the entire video generation process is simple and practical, with real-time performance and good results. It can cover all scenarios of digital humans, and the cost of building a virtual image anchor is low and simple to produce.
[0098] The video generation method proposed in the disclosed embodiments offers a simple and clear overall solution, eliminating the cumbersome and complex modules of other existing 3D technologies in the digital human field. This method can, to a certain extent, avoid problems such as poor lip movement, jittery faces, unrealistic faces, and color differences. The relevant models involved in the disclosed embodiments are universal models, applicable to various business scenarios, and capable of individual fine-tuning. Only simple fine-tuning training is required to achieve customized digital human solutions, demonstrating high value and universal applicability.
[0099] The disclosed embodiments also propose a model training method, which uses a face reconstruction model to generate face video samples, and constructs a training set and a test set of the above-mentioned first model (such as the FaceFormer model) and / or the second model (such as the styleGANv2 model) based on the face video samples, which are used to train the above-mentioned first model (such as the FaceFormer model) and / or the second model (such as the styleGANv2 model). The face reconstruction model may include a high-fidelity face swapping (hifiFace, high-fidelity face swapping method) model, or a hifiFace 3D face reconstruction model. In some examples, more than 1,000 hours of high-quality data and more than 600 IDs are collected, with rich data scenes, half of which are male and half are female, and contain diverse speaking and / or singing characteristics, and contain a variety of noisy speaking data. By building 3D training and test datasets based on the hifiFace model, the problem of missing data can be solved. Through this method, any facial video can be converted into trainable 3D face driving data. The method is simple and practical, and data missing is no longer a bottleneck restricting 3D virtual human training. It achieves the same data capabilities as 2D digital human data and solves the problems of poor model generalization and low lip movement accuracy that may be caused by missing 3D data. It has great applicability.
[0100] Figure 6 : is a flow chart of a model training method according to an embodiment of the present disclosure. The training method can be used to train the first model described above. The training method includes the following steps:
[0101] S610, generating multiple face video samples using a face reconstruction model;
[0102] S620: Create a first training set and a first test set using the multiple face video samples and voice samples;
[0103] S630. Use the first training set and the first test set to train the first model so that the first model can generate a three-dimensional face mesh sequence based on the voice file and the face video to be driven; wherein the three-dimensional face mesh sequence corresponds to the voice features of the voice file, and corresponds to the face features and speaking style features of the face video to be driven.
[0104] In some examples, the face reconstruction model includes a high fidelity face swapping method (hifiFace) model.
[0105] The present disclosure also proposes a model training method. Figure 7: is a flow chart of a model training method according to an embodiment of the present disclosure. The training method can be used to train the second model described above. The training method includes the following steps:
[0106] S710, generating multiple face video samples using a face reconstruction model;
[0107] S720: Using the plurality of face video samples and the two-dimensional face image sequence samples, establish a second training set and a second test set;
[0108] S730. Use the second training set and the second test set to train the second model, so that the second model can generate a voice-driven face video based on the two-dimensional face image sequence and the face video to be driven.
[0109] In some examples, the face reconstruction model includes a high fidelity face swapping method (hifiFace) model.
[0110] It can be seen that the facial data production method constructed in the embodiment of the present disclosure greatly enriches 3D facial audio-visual data, is simple and easy to operate, and can solve the bottleneck of training data for the first model and / or second model used to generate video files. It makes it possible to train a general model based on 3D face talking drive (3D-face-talking), which can be applied to various business scenarios.
[0111] The present disclosure also provides a video generation device. Figure 8 FIG. 8 is a structural diagram of a video generating apparatus 800 according to an embodiment of the present disclosure, comprising:
[0112] Input / output module 801 is configured to input a voice file and a facial video to be driven into a pre-trained first model, and the first model outputs a three-dimensional facial mesh sequence; wherein the three-dimensional facial mesh sequence corresponds to the voice features of the voice file and corresponds to the facial features and speaking style features of the facial video to be driven;
[0113] The first generating module 802 is configured to generate a voice-driven face video that matches the voice file based on the three-dimensional face mesh sequence and the face video to be driven.
[0114] Figure 9 is a schematic structural diagram of a video generation device 900 according to another embodiment of the present disclosure. In some implementations, the first model includes an encoding module 903 and a decoding module 904;
[0115] The encoding module 903 is used to extract the voice features of the voice file and extract the facial features and speaking style features of the face video to be driven;
[0116] The decoding module 904 is configured to obtain the three-dimensional face mesh sequence based on the voice feature, the face feature, and the speaking style feature.
[0117] In some embodiments, the encoding module 903 is configured to:
[0118] Extracting facial feature parameters from each frame image of the face video to be driven;
[0119] Based on the facial feature parameters of each frame of image, the facial features of the face video to be driven are determined.
[0120] In some embodiments, the encoding module 903 is configured to:
[0121] Extracting facial feature parameters and expression parameters from each frame image of the face video to be driven;
[0122] Based on the facial feature parameters and expression parameters of each frame image, the speaking style features of the face video to be driven are determined.
[0123] In some implementations, the first model includes a facial drive model based on a Transformer architecture.
[0124] In some implementations, the first generating module 802 is configured to:
[0125] Rendering the three-dimensional face mesh sequence to obtain a two-dimensional face image sequence;
[0126] Fitting the two-dimensional face image sequence to the face video to be driven to obtain a fitted image sequence;
[0127] The fitted image sequence and a frame of image in the face video to be driven are input into the second model to obtain the voice-driven face video.
[0128] In some implementations, the second model includes a style-based generative adversarial network model.
[0129] Figure 10 is a structural diagram of a model training device 1000 according to an embodiment of the present disclosure, as shown in FIG. Figure 10 As shown, in some embodiments, including:
[0130] The second generating module 1001 is used to generate a plurality of face video samples using a face reconstruction model;
[0131] A first training set generating module 1002 is configured to create a first training set and a first test set using the plurality of face video samples and voice samples;
[0132] The first training module 1003 is used to train the first model using the first training set and the first test set, so that the first model can generate a three-dimensional face mesh sequence based on the voice file and the face video to be driven; wherein the three-dimensional face mesh sequence corresponds to the voice features of the voice file, and corresponds to the face features and speaking style features of the face video to be driven.
[0133] In some embodiments, the face reconstruction model includes a high-fidelity face swap model.
[0134] Figure 11 1 is a schematic diagram of the structure of a model training device 1100 according to an embodiment of the present disclosure. Figure 11 As shown, in some embodiments, including:
[0135] The third generating module 1101 is used to generate multiple face video samples using a face reconstruction model;
[0136] A second training set generating module 1102 is configured to create a second training set and a second test set using the plurality of face video samples and two-dimensional face image sequence samples;
[0137] The second training module 1103 is used to train the second model using the second training set and the second test set, so that the second model can generate a voice-driven face video based on the two-dimensional face image sequence and the face video to be driven.
[0138] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0139] In the technical solution disclosed herein, the acquisition, storage and application of personal information of users involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0140] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0141] Figure 12A schematic block diagram of an example electronic device 1200 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0142] like Figure 12 As shown, the device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. Various programs and data required for the operation of the device 1200 can also be stored in the RAM 1203. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0143] Various components in device 1200 are connected to I / O interface 1205, including: an input unit 1206, such as a keyboard, a mouse, etc.; an output unit 1207, such as various types of displays, speakers, etc.; a storage unit 1208, such as a magnetic disk, an optical disk, etc.; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 allows device 1200 to exchange data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0144] The computing unit 1201 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1201 performs the various methods and processes described above, such as the video generation method or the video training method. For example, in some embodiments, the video generation method or the video training method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the video generation method described above can be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to execute the video generation method or the video training method in any other appropriate manner (eg, by means of firmware).
[0145] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0146] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0147] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0148] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0149] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0150] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0151] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0152] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A video generation method, comprising: Inputting a voice file and a face video to be driven into a pre-trained first model, and having the first model output a three-dimensional face mesh sequence; wherein the three-dimensional face mesh sequence corresponds to the voice features of the voice file, and corresponds to the facial features and speaking style features of the face video to be driven; The three-dimensional face mesh sequence is rendered to obtain a two-dimensional face image sequence; the two-dimensional face image sequence is fitted with the face video to be driven to obtain a fitted image sequence; the fitted image sequence and a frame image in the face video to be driven are input into a second model to obtain the voice-driven face video.
2. The method according to claim 1, wherein The first model includes an encoding module and a decoding module; The encoding module is used to extract the voice features of the voice file and extract the facial features and speaking style features of the face video to be driven; The decoding module is used to obtain the three-dimensional face mesh sequence based on the voice features, the face features and the speaking style features.
3. The method according to claim 2, wherein: The encoding module is used for: Extracting facial feature parameters from each frame image of the face video to be driven; Based on the facial feature parameters of each frame of image, the facial features of the face video to be driven are determined.
4. The method according to claim 2, wherein: The encoding module is used for: Extracting facial feature parameters and expression parameters from each frame image of the face video to be driven; Based on the facial feature parameters and expression parameters of each frame image, the speaking style features of the face video to be driven are determined.
5. The method according to any one of claims 1 to 4, wherein: The first model includes a facial drive model based on a Transformer architecture.
6. The method according to claim 1, wherein The second model includes a style-based generative adversarial network model.
7. A video generation device, comprising: An input / output module, configured to input a voice file and a face video to be driven into a pre-trained first model, and output a three-dimensional face mesh sequence from the first model; wherein the three-dimensional face mesh sequence corresponds to the voice features of the voice file, and corresponds to the facial features and speaking style features of the face video to be driven; The first generation module is used to render the three-dimensional face mesh sequence to obtain a two-dimensional face image sequence; fit the two-dimensional face image sequence with the face video to be driven to obtain a fitted image sequence; and input the fitted image sequence and a frame image from the face video to be driven into the second model to obtain the voice-driven face video.
8. The device according to claim 7, wherein The first model includes an encoding module and a decoding module; The encoding module is used to extract the voice features of the voice file and extract the facial features and speaking style features of the face video to be driven; The decoding module is used to obtain the three-dimensional face mesh sequence based on the voice features, the face features and the speaking style features.
9. The device according to claim 8, wherein The encoding module is used for: Extracting facial feature parameters from each frame image of the face video to be driven; Based on the facial feature parameters of each frame of image, the facial features of the face video to be driven are determined.
10. The device according to claim 8, wherein The encoding module is used for: Extracting facial feature parameters and expression parameters from each frame image of the face video to be driven; Based on the facial feature parameters and expression parameters of each frame image, the speaking style features of the face video to be driven are determined.
11. The device according to any one of claims 7 to 10, wherein: The first model includes a facial drive model based on a Transformer architecture.
12. The device according to claim 7, wherein The second model includes a style-based generative adversarial network model.
13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.
15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Video generation method and device, electronic equipment and readable storage medium
CN111294665A
Three-dimensional face synthesis method and device, equipment and storage medium
CN118037909A