Voice-driven digital human video generation method and device
By inputting the drive voice and character reference images to the pre-trained digital human video generation model, continuous video frames matching the audio are generated, and digital human video is obtained through audio and video encoding, which solves the problem of limitations in the model application scenarios and incoherence of videos in the prior art, and high-quality digital human video generation is achieved.
Patent Information
- Application Number
- CN202510213146.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-23
AI Technical Summary
When the prior art generates digital human videos that match audio, the model application scenarios are limited, making it difficult to generate videos of other digital humans, and the actions do not match the audio and the video picture is incoherent.
By obtaining the driver voice and character reference images, inputting them into a system based on a pre-trained digital human video generation model, continuous video frames are generated, and digital human video is obtained through audio and video encoding. This model is based on character reference images and audio of character videos and is trained by training labels based on continuous video frames.
It realizes the generation of high-quality digital human videos based on driver voice and character reference images, and can flexibly generate videos of different digital human images without retraining the model, which improves the generation quality of voice-driven digital human videos.
Smart Images

Figure CN120034706A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method and device for generating a voice-driven digital human video. Background Art
[0002] In the field of artificial intelligence technology, one application scenario requires generating a digital human video that matches the audio so that the movements and expressions of the digital human in the video can match the audio.
[0003] In actual applications, the appearance of digital humans varies greatly, and it is often necessary to perform model training for the required digital human appearance in order to obtain a digital human video generation model with better quality. Although this method can generate digital human videos, the model has limitations in application scenarios. The effect of generating videos of other digital humans through the model is not good, and there may be situations such as mismatches between movements and audio, and incoherent video images.
[0004] How to improve the generation quality of voice-driven digital human videos is a technical problem to be solved by this application. Summary of the invention
[0005] The purpose of the embodiments of the present application is to provide a voice-driven digital human video generation method, a model training method and a device, so as to improve the generation quality of voice-driven digital human videos.
[0006] In a first aspect, a method for generating a digital human video driven by speech is provided, comprising: Obtain driving voice and character reference images; Inputting the driving voice and the character reference image into a digital human video generation model to obtain continuous video frames, wherein the digital human video generation model is trained based on training samples constructed based on the character reference image and audio of the character video and training labels constructed based on continuous video frames of the character video; Audio and video encoding is performed on the driving voice and the continuous video frames to obtain a digital human video.
[0007] In a second aspect, a voice-driven digital human video generation device is provided, comprising: An acquisition module is used to acquire driving voice and character reference images; A generation module inputs the driving voice and the character reference image into a digital human video generation model to obtain continuous video frames, wherein the digital human video generation model is trained based on training samples constructed based on the character reference image and audio of the character video and training labels constructed based on continuous video frames of the character video; The encoding module performs audio and video encoding on the driving voice and the continuous video frames to obtain a digital human video.
[0008] According to a third aspect, an electronic device is provided. The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the method according to the first aspect are implemented.
[0009] According to a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method according to the first aspect are implemented.
[0010] In a fifth aspect, a computer program product is provided, the computer program product comprising a non-transitory computer-readable storage medium storing a computer program, the computer program being operable to cause a computer to execute part or all of the steps of the method of the first aspect.
[0011] In an embodiment of the present application, first, a driving voice and a character reference image are obtained; then, the driving voice and the character reference image are input into a digital human video generation model to obtain continuous video frames, wherein the digital human video generation model is trained based on training samples constructed based on the character reference image and audio of the character video and training labels constructed based on continuous video frames of the character video; subsequently, audio and video encoding is performed on the driving voice and continuous video frames to obtain a digital human video. Through the scheme provided in the embodiment of the present application, continuous video frames can be generated based on the driving voice and the character reference image by a pre-trained digital human video generation model. Among them, the digital human video generation model is trained by constructing training samples with the character reference image and audio of the character video, and constructing training labels with continuous video frames of the character video, and the digital human video generation model can make the continuous video frames show continuous video frames that match the character reference image and audio. Furthermore, audio and video encoding is performed on the driving voice and continuous video frames to obtain a digital human video with matching sound and picture, which effectively improves the generation quality of the voice-driven digital human video. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 This is one of the flowcharts of a method for generating a digital human video driven by a voice according to an embodiment of the present application; Figure 2 This is a second flow chart of a method for generating a digital human video driven by voice according to an embodiment of the present application; Figure 3 This is a flowchart of a method for generating a digital human video driven by voice according to an embodiment of the present application; Figure 4 This is a fourth flow chart of a method for generating a digital human video driven by a voice according to an embodiment of the present application; Figure 5a It is a structural schematic diagram of a digital human video generation model of a voice-driven digital human video generation method according to an embodiment of the present application; Figure 5b This is a flowchart of a method for generating a digital human video driven by voice according to an embodiment of the present application; Figure 6 This is a sixth flow chart of a method for generating a digital human video driven by a voice according to an embodiment of the present application; Figure 7a This is a flowchart of a method for generating a digital human video driven by a voice according to an embodiment of the present application; Figure 7b It is a flowchart of a method for generating a digital human video driven by voice in a scenario where multiple human images are inputted, according to an embodiment of the present application; Figure 8a This is a flowchart of a method for generating a digital human video driven by voice according to an embodiment of the present application; Figure 8b It is a flowchart of a method for generating a digital human video driven by a voice in a scene with multiple candidate character reference images according to an embodiment of the present application; Fig. 9 This is a ninth flow chart of a method for generating a digital human video driven by a voice according to an embodiment of the present application; Fig.10 This is a flowchart of a voice-driven digital human video generation method according to an embodiment of the present application; Fig.11 This is a flowchart of a method for generating a digital human video driven by voice according to an embodiment of the present application; Fig.12 This is a flowchart diagram twelfth of a method for generating a digital human video driven by a voice according to an embodiment of the present application; Fig.13 This is a flowchart of a method for generating a digital human video driven by voice according to an embodiment of the present application; Fig.14 It is a structural schematic diagram of a voice-driven digital human video generation device according to an embodiment of the present application. DETAILED DESCRIPTION
[0013] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application. The numbering of the drawings in this application is only used to distinguish the various steps in the scheme, and is not used to limit the execution order of the various steps. The specific execution order is subject to the description in the specification.
[0014] In the field of artificial intelligence technology, digital human application scenarios often require digital humans to make movements that match audio, especially those involving digital human mouth movements. If a model is trained for a specific digital human image, the trained model can often only be used to generate videos of that specific digital human. Once a new digital human image is replaced, the model needs to be retrained, which has poor universality and requires more resources.
[0015] In order to solve the problems existing in the related art, the embodiment of the present application provides a voice-driven digital human video generation method, which relates to the fields of artificial intelligence, digital human, voice-driven, video generation, etc. In the case of inputting driving voice and character reference image, the solution provided by the embodiment of the present application can efficiently generate digital human video and has universal applicability. By changing the character reference image, videos with different digital human images can be flexibly generated without retraining the model.
[0016] The solution provided in the embodiment of this application is as follows Figure 1 As shown, including: S11: Acquire driving voice and character reference image.
[0017] The driving voice may be an audio containing voice, which drives the facial expressions and movements of the characters. The driving voice may also contain sounds corresponding to body movements, such as clapping, snapping fingers, etc.
[0018] The above-mentioned character reference image refers to an image containing the appearance of the character, and the image may include a background and a foreground, and the foreground at least includes the character image. The character reference image can be used as a benchmark to generate facial expressions in subsequent steps.
[0019] Optionally, the character reference image includes a frontal face image of the character, which can help generate a digital human with vivid and lifelike facial expressions and movements in subsequent steps.
[0020] S12: Input the driving voice and the character reference image into a digital human video generation model to obtain continuous video frames, wherein the digital human video generation model is trained based on training samples constructed based on character reference images and audio of character videos and training labels constructed based on continuous video frames of the character videos.
[0021] In this step, the driving voice and the character reference image can be directly input into the digital human video generation model, or the driving voice and the character reference image can be pre-processed and then input into the digital human video generation model. The pre-processing can include optimization methods such as noise reduction, beautification, and style processing.
[0022] The digital human video generation model in this solution is trained based on training samples constructed based on the character reference image and audio of the character video and training labels constructed based on the continuous video frames of the character video. The character video can be decomposed into audio frames and video frames based on frame correspondence, the audio frame provides the video playback sound, and the video frame provides the video playback screen. The character reference image can be a video frame containing the character's appearance in the character video, or it can be generated based on the integration of multiple video frames containing the character's appearance.
[0023] The character reference image, audio and continuous video frames of the character video are associated with each other. The audio provides the sound of the character video, and the continuous video frames provide the picture of the video, and the sound and the picture correspond to each other. The character reference image can show the appearance of the character, and the character appearance can be used as a benchmark for the sound and the picture, so that the model can learn the matching relationship between the audio and the character's expression and action in the corresponding video frame.
[0024] Optionally, the above-mentioned digital human video generation model can be constructed based on a Stable Diffusion (SD) model framework, and the pre-trained digital human video generation model is used to generate continuous video frames corresponding to the driving voice based on the input character reference image, so that the character's facial expressions and movements in the continuous video frames correspond to the driving voice.
[0025] S13: Performing audio and video encoding on the driving voice and the continuous video frames to obtain a digital human video.
[0026] In this step, the driving voice and continuous video frames are arranged in chronological order, and a frame-by-frame correspondence is established through audio and video encoding, so that the audio frames correspond to the video frames, and a digital human video with synchronized audio and video is generated.
[0027] Through the solution provided by the embodiment of the present application, continuous video frames can be generated based on driving voice and character reference images through a pre-trained digital human video generation model. Among them, the digital human video generation model is trained by constructing training samples with character reference images and audio of character videos, and constructing training labels with continuous video frames of character videos. The digital human video generation model can make continuous video frames show continuous video frames that match the character reference images and audio. Then, audio and video encoding is performed on the driving voice and continuous video frames to obtain a digital human video with matching audio and video, effectively improving the generation quality of voice-driven digital human videos.
[0028] Based on the solution provided in the above embodiment, optionally, Figure 2 As shown, in the above step S12, the driving voice and the character reference image are input into the digital human video generation model to obtain continuous video frames, including: S21: Input the driving voice and the character reference image into the audio key point mapping model of the digital human video generation model to obtain a facial key point sequence, wherein the facial key point sequence includes multiple groups of facial key point coordinates arranged in time sequence.
[0029] S22: predicting continuous video frames corresponding to the facial key point sequence based on the character reference image through the diffusion model of the digital human video generation model.
[0030] In the solution provided in the embodiment of the present application, the digital human video generation model includes an audio key point mapping model and a diffusion model, wherein the audio key point mapping model can be expressed as audio2landmark, which can generate a corresponding facial key point sequence according to the input driving voice. The facial key point sequence is arranged based on time sequence, and the facial key point coordinates corresponding to any time point correspond to the audio frame of the driving voice at this time point. That is, the facial key point sequence output by the audio key point corresponds to the driving voice, and can express the facial action expression corresponding to the driving voice.
[0031] Based on the facial key point sequence output by the audio key point mapping model, a diffusion model is used to predict the character reference image, and continuous video frames in which the facial expressions of the character reference image change correspondingly with the facial key point sequence are obtained.
[0032] Through the scheme provided by the embodiment of the present application, the audio key point mapping model in the digital human video generation model can generate a corresponding sequence of facial key points according to the driving voice, thereby converting the audio into the corresponding facial expression action. The diffusion model in the digital human video generation model can perform prediction according to the facial key point sequence and the character reference image, and obtain continuous video frames in which the facial expression of the digital human changes correspondingly with the facial key point sequence, so that the facial expression of the digital human in the continuous video frames matches the driving voice, and realizes audio and video synchronization.
[0033] The present solution is explained below with an example. In one application scenario, the driving voice is input into the audio key point mapping model to obtain multiple facial key point images arranged in time sequence, forming a facial key point sequence. Then, the facial key point sequence and the character reference image are input into the diffusion model to obtain a sequence of digital human speech action images of the character reference image, forming a continuous video frame arranged in time sequence. Finally, the driving voice is combined with the continuous video frames through audio and video coding to obtain the digital human speech video.
[0034] Based on the solution provided in the above embodiment, optionally, Figure 3 As shown, before the above step S12, before the driving voice and the character reference image are input into the audio key point mapping model of the digital human video generation model to obtain the facial key point sequence, it also includes: S31: extracting facial key points from the character reference image using a key point detection model to obtain a character reference image containing reference coordinates of facial key points.
[0035] In this embodiment, preprocessing is performed on the person reference image, and the facial key points in the person reference image are extracted through a key point detection model, which may specifically include key points corresponding to the facial features. These facial key points can represent the facial expression state.
[0036] Subsequently, the character reference image containing the reference coordinates of the facial key points and the driving voice are input into the above-mentioned digital human video generation model. Among them, the reference coordinates of the facial key points in the character reference image can represent the facial expression characteristics of the character in the character reference image, which can help the model predict the facial key point coordinates that change with the driving voice based on the facial key point reference coordinates, thereby improving the matching degree of the facial key point sequence and the driving voice, and optimizing the quality of the facial key point sequence.
[0037] Based on the solution provided in the above embodiment, optionally, Figure 4 As shown, in the above step S21, the driving voice and the character reference image are input into the audio key point mapping model of the digital human video generation model to obtain a facial key point sequence, including: S41: performing feature extraction on the driving speech to obtain speech features.
[0038] In this step, the driving speech may be input into a pre-trained audio feature extractor to obtain speech features. Optionally, before extracting speech features from the driving speech, the original driving speech may be normalized to optimize audio quality.
[0039] S42: Input the speech features into a shallow convolutional model to obtain a sequence of predicted key point relative offsets.
[0040] In this step, the speech features are input into a shallow convolutional model (e.g., Multi-Layer Perceptron (MLP)) to obtain the relative offset of the key points. The coordinates of the facial key points corresponding to the current frame are obtained by superimposing the relative offset of the key points based on the reference coordinates of the facial key points. The coordinates of the facial key points arranged in time sequence constitute the facial key point sequence.
[0041] S43: Determine the facial key point sequence based on the person reference image containing the facial key point reference coordinates and the predicted key point relative offset sequence.
[0042] The above-mentioned person reference image is pre-input into the trained facial key point model to obtain the facial key point reference coordinates, and then obtain the person reference image containing the facial key point reference coordinates.
[0043] In this step, based on the reference coordinates of the facial key points in the character reference image, the relative offset sequence of the predicted key points is superimposed to obtain the facial key point sequence.
[0044] The solution provided in the embodiment of the present application performs prediction based on the driving voice through a shallow convolution model to obtain the relative offset of the key points of the face. When the reference coordinates of the key points of the face represent the facial expression of the digital human with the mouth closed in a calm state, this solution can effectively predict the facial expression of the digital human that matches the driving voice. Among them, the facial key points and the key point offsets can efficiently express the facial expression characteristics of the digital human, which is conducive to the efficient generation of the sequence of facial key points, and then the efficient construction of continuous video frames representing the facial expression movements of the digital human. Among them, based on the facial key points, the expression of mouth details can be effectively enhanced, especially the expression of the teeth features of the mouth can be achieved, ensuring that the mouth movement is coordinated with the overall face, and improving the authenticity of the final generated digital human video effect.
[0045] Optionally, in actual applications, a mouth feature optimization module may be applied in the process of predicting the offset of facial key points to improve the prediction accuracy of the offset of the mouth key points.
[0046] Based on the solution provided in the above embodiment, optionally, Figure 5aAs shown, the diffusion model includes an image encoder (VAE (Variational Auto-Encoders) Encoder), an image reference module, an image optimization module, a key point encoder (Landmark Encoder) and an image decoder (VAE Decoder). The diffusion model and the audio key point module together constitute a digital human video generation model, and the audio key point module includes an audio key point mapping model (audio2landmark). The model provided in the embodiment of the present application is constructed based on the stable diffusion (SD) model graph generation framework, which is used to redraw the input character reference image and generate a new image that matches the driving voice.
[0047] See also Figure 5b In the above step S22, the continuous video frames corresponding to the facial key point sequence are predicted based on the character reference image by using the diffusion model of the digital human video generation model, including: S51: Encode a person reference image including reference coordinates of facial key points into a first latent space feature through the image encoder.
[0048] In the embodiment of the present application, the latent space is also called the latent space, which can effectively realize feature expression by mapping data from high-dimensional observation data to a low-dimensional representation space. In this step, the reference coordinates of the key points of the face of the character reference image are encoded by the image encoder to obtain the first latent space features. This representation space can effectively capture the potential features and structure of the data. By sampling and operating on this latent space, tasks such as data generation and feature extraction can be achieved. The low dimensionality of the latent space helps to simplify data processing, reduce computational complexity, and effectively retain key information of the data.
[0049] S52: Perform feature extraction on the first latent space feature through the image reference module to obtain a facial appearance feature based on the latent space.
[0050] The network architecture of the image reference module can be U-Net, which is consistent with the image optimization module. In each Transformer block of the image reference module, the self-attention mechanism is used to extract the reference image features, which are then fused with the hidden state variables at the corresponding positions in the denoising U-Net of the image optimization module as the key-value input of the next layer of Transformer blocks in the image optimization module.
[0051] The role of this image reference module is to encode the reference image without introducing noise, inject reference image coding features at different levels of the image optimization module, and it is very efficient, requiring only one forward pass to be performed during the diffusion process.
[0052] S53: Encode the facial key point sequence into a second latent space feature sequence through the key point encoder.
[0053] S54: Add latent variable noise to the second latent space feature sequence.
[0054] In the embodiments of the present application, latent variable noise refers to random noise introduced in the latent space, which can be used to generate new data or enhance the robustness of the model. This noise can help the model learn the distribution of data, making the generated data more diverse and realistic.
[0055] S55: Generate a video frame feature sequence through the image optimization module based on the facial appearance features and the second latent space feature sequence after adding latent variable noise.
[0056] The key point encoder can be used to encode the key points of the face into the second latent space features, thereby obtaining the second latent space feature sequence. The second latent space feature sequence is input into the U-Net network after adding noise in the latent space to guide the image generation process so that the subsequently generated image matches the facial key point sequence corresponding to the driving speech.
[0057] The above-mentioned facial appearance features can be used to guide the video frame features generated by the image optimization module to have the facial appearance features in the person reference image, so that the face in the subsequently generated video frame is close to the face in the person reference image.
[0058] S56: Decoding the video frame feature sequence into time-series continuous video frames through the image decoder.
[0059] In the embodiment of the present application, the character reference image is first passed into the image encoder (VAE encoder) to generate latent space features, and the latent space features are input into the image optimization module (U-Net network) to perform optimization iterations to obtain latent variable features. Among them, the key point encoder encodes the facial key point image into latent space features, and the encoded features are input into the U-Net network after adding noise to the latent space to guide the image generation process. Subsequently, the latent variable features are input into the image decoder (VAE Decoder) for decoding and reconstruction into a new image.
[0060] Through the scheme provided by the embodiment of the present application, image feature extraction and image reconstruction can be achieved through various modules in the diffusion model. Under the guidance of the facial key point sequence corresponding to the driving voice, image reconstruction is performed based on the facial appearance features of the character reference image, so that the generated video frame features have the facial appearance features of the character reference image, and then the continuous video frames obtained by decoding are matched with the driving voice to form a dynamic video picture based on the face in the character reference image.
[0061] Based on the solution provided in the above embodiment, optionally, Figure 6 As shown, the image optimization module includes multiple transformation modules, and the transformation module includes a reference attention layer and a temporal attention layer.
[0062] In an embodiment of the present application, each Transformer block of the image optimization module has a reference attention layer and a temporal attention layer, so that the U-Net component can better utilize external injected information.
[0063] Among them, in the above step S55, the image optimization module generates a video frame feature sequence based on the face appearance features and the second latent space feature sequence after adding latent variable noise, including: S61: Generate a first video frame feature having at least part of the facial appearance features through the reference attention layer.
[0064] The above-mentioned reference attention layer can promote the accurate encoding of the relationship between the current frame and the person reference image, so that the generated new image can better retain the facial appearance characteristics of the person reference image.
[0065] S62: Generate a second video frame feature with an inter-frame correlation relationship through the temporal attention layer.
[0066] The above-mentioned temporal attention layer learns the temporal order of consecutive video frames by adding a self-attention mechanism in the time domain, which can effectively improve the coherence between generated video frames. The temporal attention layer captures the dependency between consecutive frames by adjusting the hidden state shape and applying a self-attention mechanism along the time axis of the frame sequence. Specifically, given the hidden state variable is , where b represents the batch size, l represents the frame sequence length, c represents the number of channels, h represents the height, and w represents the width. The hidden state variable needs to be reshaped to , after adjustment, the self-attention mechanism can be applied in the time dimension. Through this process, the temporal attention layer can discern and learn subtle motion patterns, ensuring smooth and natural transitions of synthesized frames. Based on this, consecutive video frames can show a high degree of temporal consistency, reflecting natural and smooth motion, effectively improving the visual quality and realism of the generated content.
[0067] S63: Generate a video frame feature sequence having the first video frame feature and the second video frame feature based on the second latent space feature sequence after adding latent variable noise through a denoising U-Net.
[0068] In this step, the facial appearance features of the first latent space features are fused with the hidden state variables at the corresponding positions in the denoising U-Net and used as the key-value input of the next layer of Transformer block in the image optimization module, thereby achieving encoding without introducing noise and injecting reference image encoding features at different levels of the image optimization module.
[0069] The solution provided in the embodiment of the present application improves the quality of continuous video frames from two aspects, namely, the consistency of digital human appearance and the continuity of pictures, through an image optimization module, thereby effectively improving the authenticity of digital human videos.
[0070] Based on the solution provided in the above embodiment, optionally, Figure 7a As shown, in the above step S11, obtaining the driving voice and the character reference image includes: S71: Acquire multiple images of the target person; S72: extracting mouth key points from each of the character images; S73: performing cluster analysis on the mouth key points of each of the character images to obtain at least one cluster; S74: Determine the person image at the center of at least one cluster as a candidate person reference image.
[0071] Clustering described in the embodiments of the present application is a data analysis technique used to divide objects in a data set into multiple clusters. Its goal is to make objects in the same cluster as similar as possible in a certain sense, while objects in different clusters are as different as possible.
[0072] See also Figure 7b When the input is multiple person images, the coordinates of the mouth key points are extracted for each image respectively, and the extracted coordinates are standardized to make them suitable for clustering algorithm input to improve the clustering effect.
[0073] Then, clustering analysis is performed on the coordinates of the mouth key points using a clustering algorithm to obtain multiple clusters. According to the clustering results, the face images corresponding to the mouth key points of multiple cluster centers are used as candidate character reference images and stored in the target portrait reference image library.
[0074] For application scenarios where there are multiple character images available, in order to make full use of the materials, this solution obtains several representative face templates corresponding to different mouth shapes through clustering methods, which can effectively improve the image quality of the digital human's mouth area and the similarity with the target human's mouth features.
[0075] Based on the solution provided in the above embodiment, optionally, Figure 8aAs shown, in the above step S12, the driving voice and the character reference image are input into the digital human video generation model to obtain continuous video frames, including: S81: inputting the driving voice into the audio key point mapping model of the digital human video generation model to obtain a facial key point sequence corresponding to the driving voice; S82: determining, from each of the candidate character reference images, a target character reference image having the highest similarity to the mouth key points in the facial key point sequence; S83: Inputting the facial key point sequence and the target person reference image into the diffusion model in the digital human video generation model to obtain the continuous video frames.
[0076] like Figure 8b As shown, the driving speech is input into the audio key point mapping model to obtain a sequence of facial key points arranged in time sequence. Then, the most similar candidate character reference image (i.e., the cluster center mentioned above) is found by the mouth key point coordinates in the facial key point coordinates, and this candidate character reference image is used as the target character reference image and the facial key point coordinates are input into the diffusion model together to obtain a digital human speaking image generated based on the target character reference image. Finally, the generated digital human speaking images are arranged in time sequence to construct continuous video frames. The continuous video frames are used to perform audio and video encoding with the driving speech to obtain a digital human video with synchronized audio and video.
[0077] Based on the solution provided in the above embodiment, optionally, Fig. 9 As shown, before the above step S12, before the driving voice and the character reference image are input into the digital human video generation model to obtain continuous video frames, it also includes: S91: Acquire multiple character videos, wherein any of the character videos includes audio and continuous video frames corresponding to the audio.
[0078] The solution provided in the embodiment of the present application is used to perform pre-training on a digital human video generation model. First, multiple character videos are obtained as materials for model training, wherein the characters in each character video can be different. To optimize the training effect, any character video at least includes the face of the character, so that the model learns the correlation between audio and the facial movements of the character.
[0079] S92: Determine the face reference images corresponding to the respective character videos based on the continuous video frames of the respective character videos.
[0080] In this step, a video frame can be selected from the continuous video frames as the face reference image corresponding to the character video. In practical applications, a video frame that can clearly show the facial features can be selected as the face reference image. If there are multiple video frames that can clearly show the facial features, then a video frame with no expression and closed mouth is selected from the multiple video frames that can clearly show the facial features as the face reference image.
[0081] S93: Train an audio key point mapping model based on a first training sample and a first training label, wherein the first training sample is constructed based on face reference images and audio of multiple videos of the character, and the first training label is constructed based on continuous video frames of multiple videos of the character.
[0082] In the solution provided in the embodiment of the present application, the audio key point mapping model and the diffusion model can be trained in sequence to optimize the overall quality of the model. For the audio key point mapping model, the first training sample and the corresponding first training label are used to perform training in this step. Among them, the first training sample is constructed based on a face reference image and audio, the face reference image is used to represent the facial features of the character, the audio includes the character's voice, and the first training label is used to represent the facial action picture of the character speaking. Based on the first training sample and the first training label, the audio key point mapping model can be trained to learn the correspondence between the facial expressions of the character and the character's voice.
[0083] S94: Train a diffusion model based on a second training sample and a second training label to obtain a digital human video generation model including the audio key point mapping model and the diffusion model, wherein the second training sample is constructed based on a plurality of facial reference images and audios of the human video through the pre-trained audio key point mapping model, and the second training label is constructed based on a plurality of continuous video frames of the human video.
[0084] In this step, based on the audio key point mapping model trained in the above step S93, the second training sample and the corresponding second training label are used to perform training for the diffusion model. The second training sample is constructed based on the face reference image and audio through the trained audio key point mapping model. Specifically, the face reference image and audio are input into the trained audio key point mapping model, and the second training sample is constructed based on the model output. The second training label is used to represent the facial action picture of the person speaking. Based on the second training sample and the second training label, the diffusion model can be trained to learn the correspondence between the facial expression of the person and the voice of the person on the basis of the above audio key point mapping model.
[0085] Through the solution provided in the embodiment of the present application, the audio key point mapping model and the diffusion model can be trained in sequence to improve the overall quality of the digital human video generation model.
[0086] Based on the solution provided in the above embodiment, optionally, Fig.10 As shown, before the above step S93, before training the audio key point mapping model based on the first training sample and the first training label, it also includes: S101: extracting corresponding facial key point sequences from the continuous video frames of each of the character videos using a key point detection model.
[0087] The solution provided in the embodiment of the present application constructs a first training sample and a first training label based on a character video. Among them, a key point detection model is applied to respectively detect the facial key point sequence of the continuous video frames of each character video. Among them, the facial key point sequence may include multiple facial key point coordinates arranged in time sequence. The above-mentioned key point detection model is used to perform frame-by-frame analysis and recognition on each video frame, extract the facial key points in the frame, and express the facial expression and movement in the corresponding frame with the facial key points. Based on this, the facial key point sequence corresponding to the character video can represent the continuous movement of the facial expression within the time corresponding to the video.
[0088] S102: Generate facial key point reference images corresponding to each of the character videos based on the character closed-mouth video frames in the continuous video frames, wherein the facial key point reference images include facial key point reference coordinates.
[0089] In this step, the video frames in which the person is in a closed-mouth state are identified from the continuous video frames. If there are multiple video frames in which the person is in a closed-mouth state in a person video, a video frame with no obvious expression can be selected as the above-mentioned closed-mouth video frame of the person.
[0090] In this step, facial key point detection is performed on the closed-mouth video frame of the person, for example, by performing detection using the above-mentioned key point detection model, to obtain the facial key point reference coordinates corresponding to the closed-mouth video frame of the person, and then the closed-mouth video frame of the person containing the facial key point reference coordinates is used as the facial key point reference image of the corresponding person video. The facial key point reference image can represent the facial features of the person in the corresponding person video without expression, which is conducive to the model generating other expression and action images that match the appearance of the person.
[0091] S103: constructing the first training sample based on the reference coordinates and audio of the facial key points of each of the character videos, and constructing the first training label based on the facial key point sequence of each of the character videos.
[0092] In this step, the first training sample constructed based on the reference coordinates of the facial key points and the audio can represent the closed-mouth appearance features and the voice features of the person, and the corresponding first training label can represent the facial expression movements made by the person corresponding to the audio. The first training sample and the first training label constructed in this way can help the audio key point mapping model learn the correlation between the changes in the audio features and the changes in the facial key points.
[0093] Based on the solution provided in the above embodiment, optionally, Fig.11 As shown, in the above step S93, training the audio key point mapping model based on the first training sample and the first training label includes: S111: Inputting the first training sample into the audio key point mapping model to obtain a facial key point prediction sequence, wherein the facial key point prediction sequence includes a plurality of facial key point prediction coordinates arranged in time sequence; S112: Training the audio key point mapping model with minimizing the key point difference loss as a training goal, wherein the key point difference loss includes the distance difference between the predicted coordinates of the face key points corresponding to the same video frame and the coordinates of the face key points of the first training label.
[0094] The solution provided by the embodiment of the present application can effectively train the audio key point mapping model, wherein the training optimization goal is to minimize the key point difference loss. Specifically, the training set may include character videos of multiple different speakers, each character video corresponding to its own facial key point reference image.
[0095] Optionally, the first frame of the person video can be a reference image of the facial key points of the person video, which can help the model learn the correlation between audio and person expression based on the person's expressionless appearance features.
[0096] Specifically, the key point detection model can be used to extract a sequence of facial key points from continuous video frames of a person video. The coordinates of the facial key points extracted from the closed-mouth speaking image in the first frame are used as the reference coordinates of the facial key points, and the key point coordinates corresponding to the remaining frames are used as the true value labels to construct a sequence of facial key points.
[0097] The audio of the person video is first input into the pre-trained audio feature extractor to obtain the audio features. Then it is input into the shallow convolution model to output the predicted key point relative offset. The predicted key point relative offset is superimposed with the above-mentioned face key point reference coordinates to obtain the predicted coordinates of the face key point.
[0098] Furthermore, the predicted coordinates of the facial key points are compared with the coordinates of the facial key points in the video frame, and the audio key point mapping model is trained by minimizing the differences.
[0099] Based on the solution provided in the above embodiment, optionally, the distance difference includes the difference of the Euclidean distance of the plane coordinates and / or the difference of the Euclidean distance of the graph Laplace coordinates.
[0100] The solution provided in the embodiment of the present application provides an optional method for calculating the distance difference, wherein the Euclidean distance of the plane coordinates of the key point sequence can be calculated, or the Euclidean distance of the graph Laplace coordinates of each key point can be calculated, or the above two distance differences can be superimposed through a statistical algorithm to calculate a comprehensive difference.
[0101] Among them, the Laplace coordinate calculation formula is as follows:
[0102] in is with A set of key points where key points are connected. Different facial regions have their own connected key point areas.
[0103] Based on the solution provided in the above embodiment, optionally, Fig.12 As shown, before the above step S94, that is, before training the diffusion model based on the second training sample and the second training label to obtain the digital human video generation model including the audio key point mapping model and the diffusion model, it also includes: S121: Generate facial key point sequences corresponding to the audio of each character video respectively through the pre-trained audio key point mapping model; S122: constructing the second training samples based on the facial key point reference images and facial key point sequences of each of the character videos, and constructing the second training labels based on the continuous video frames of each of the character videos.
[0104] In the solution provided by the embodiment of the present application, the audio key point mapping model parameters are fixed and the diffusion model is trained to achieve the purpose of training the overall algorithm model. The audio of the character video is respectively input into the pre-trained audio key point mapping model to obtain the facial key point sequence corresponding to the audio of each character video output by the model. Furthermore, the facial key point reference image and the facial key point sequence of the character video are constructed as the second training sample, and the continuous video frames of the character video are constructed as the second training label, so that the diffusion model learns to generate continuous video frames matching the facial key point sequence based on the facial key point reference image.
[0105] Based on the solution provided in the above embodiment, optionally, Fig.13 As shown, in the above step S94, the diffusion model is trained based on the second training sample and the second training label to obtain a digital human video generation model including the audio key point mapping model and the diffusion model, including: S131: inputting the second training sample into the diffusion model to obtain a predicted video frame based on time sequence; S132: Training the diffusion model with minimizing the video frame difference loss as a training goal, wherein the video frame difference loss includes latent space loss and / or pixel space loss corresponding to the same video frame.
[0106] This solution minimizes the pixel space loss while minimizing the latent space loss in diffusion model training, that is, the difference loss between the generated image output by the VAE decoder and the corresponding actual real image. The difference loss between the generated image and the actual real image can be measured by the mean squared error (MSE) loss and the LPIPS (Learned Perceptual Image Patch Similarity) loss.
[0107] L total =L 潜空间 +L 像素空间 L 像素空间 =L MSE +L LPIPS LPIPS loss is a metric used to measure the perceptual difference between two images. LPIPS takes into account the characteristics of the human visual system, extracts features through a deep convolutional neural network, and calculates the difference between images in the feature space.
[0108] This application is a privately provided solution for extracting features from audio, and predicting the offset of key points by training an audio key point mapping model. The offset refers to the offset between the mouth key points of the target face in the current frame and the facial key points when the target face is in a calm state and the mouth is closed. This makes it easier for the audio key point mapping model to pay more attention to the dynamic changes of the face corresponding to the audio, especially the changes in the mouth shape, rather than the differences in the facial key point structures caused by different faces, thereby improving the generalization of the trained model.
[0109] In order to solve the problems existing in the related art, the embodiment of the present application also provides a voice-driven digital human video generation device 140, such as Fig.14 As shown, including: An acquisition module 141 acquires a driving voice and a character reference image; A generation module 142 inputs the driving voice and the character reference image into a digital human video generation model to obtain continuous video frames, wherein the digital human video generation model is trained based on training samples constructed based on the character reference image and audio of the character video and training labels constructed based on continuous video frames of the character video; The encoding module 143 performs audio and video encoding on the driving voice and the continuous video frames to obtain a digital human video.
[0110] Through the device provided in the embodiment of the present application, continuous video frames can be generated based on driving voice and character reference images through a pre-trained digital human video generation model. The digital human video generation model is trained by constructing training samples with character reference images and audio of character videos, and constructing training labels with continuous video frames of character videos. The digital human video generation model can make continuous video frames show continuous video frames that match the character reference images and audio. Furthermore, audio and video encoding is performed on the driving voice and continuous video frames to obtain a digital human video with matching audio and video, effectively improving the generation quality of voice-driven digital human videos.
[0111] Among them, the above modules in the device provided by the embodiment of the present application can also implement the method steps provided by the above method embodiment. Alternatively, the device provided by the embodiment of the present application can also include other modules in addition to the above modules to implement the method steps provided by the above method embodiment. And the device provided by the embodiment of the present application can achieve the technical effects that can be achieved by the above method embodiment.
[0112] Preferably, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned voice-driven digital human video generation method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0113] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each process of the above-mentioned voice-driven digital human video generation method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0114] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. The computer program can be operated to enable a computer to execute part or all of the steps of the above-mentioned voice-driven digital human video generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0115] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0116] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0117] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0119] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0120] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0121] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0122] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0123] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0124] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A method for generating a digital human video driven by speech, characterized in that: include: Obtain driving voice and character reference images; Inputting the driving voice and the character reference image into a digital human video generation model to obtain continuous video frames, wherein the digital human video generation model is trained based on training samples constructed based on the character reference image and audio of the character video and training labels constructed based on continuous video frames of the character video; Audio and video encoding is performed on the driving voice and the continuous video frames to obtain a digital human video.
2. The method according to claim 1, characterized in that Inputting the driving voice and the character reference image into a digital human video generation model to obtain continuous video frames, including: Inputting the driving voice and the character reference image into the audio key point mapping model of the digital human video generation model to obtain a facial key point sequence, wherein the facial key point sequence includes multiple groups of facial key point coordinates arranged in time sequence; The continuous video frames corresponding to the facial key point sequence are predicted based on the character reference image through the diffusion model of the digital human video generation model.
3. The method according to claim 2, characterized in that Before inputting the driving voice and the character reference image into the audio key point mapping model of the digital human video generation model to obtain a facial key point sequence, the method further includes: The key point detection model is used to extract the key points of the face of the person reference image, so as to obtain the person reference image containing the reference coordinates of the key points of the face.
4. The method according to claim 3, characterized in that Inputting the driving voice and the character reference image into the audio key point mapping model of the digital human video generation model to obtain a facial key point sequence, including: Performing feature extraction on the driving speech to obtain speech features; Inputting the speech features into a shallow convolutional model to obtain a sequence of relative offsets of predicted key points; The facial key point sequence is determined based on a character reference image containing reference coordinates of facial key points and the predicted key point relative offset sequence.
5. The method according to claim 3, characterized in that The diffusion model includes an image encoder, an image reference module, an image optimization module, a key point encoder and an image decoder; The method of predicting the continuous video frames corresponding to the facial key point sequence based on the character reference image by using the diffusion model of the digital human video generation model includes: Encoding a person reference image including reference coordinates of key facial points into a first latent space feature by means of the image encoder; Performing feature extraction on the first latent space feature by the image reference module to obtain a face appearance feature based on the latent space; Encoding the facial key point sequence into a second latent space feature sequence by the key point encoder; Adding latent variable noise to the second latent space feature sequence; Generate a video frame feature sequence based on the face appearance features and the second latent space feature sequence after adding latent variable noise by the image optimization module; The video frame feature sequence is decoded into continuous video frames based on time sequence by the image decoder.
6. The method according to claim 5, characterized in that The image optimization module includes a plurality of transformation modules, wherein the transformation module includes a reference attention layer and a temporal attention layer; The method of generating a video frame feature sequence based on the face appearance features and the second latent space feature sequence after adding latent variable noise by the image optimization module includes: Generate a first video frame feature having at least part of the face appearance feature through the reference attention layer; Generate a second video frame feature with an inter-frame correlation relationship through the temporal attention layer; A video frame feature sequence having the first video frame feature and the second video frame feature is generated based on the second latent space feature sequence after adding latent variable noise through a denoising U-Net.
7. The method according to any one of claims 1 to 6, characterized in that Get driving voice and character reference images, including: Acquire multiple person images of a target person; Extracting mouth key points from each of the character images; Performing cluster analysis on the mouth key points of each of the character images to obtain at least one cluster; A person image at the center of at least one cluster is determined as a candidate person reference image.
8. The method according to claim 7, characterized in that Inputting the driving voice and the character reference image into a digital human video generation model to obtain continuous video frames, including: Inputting the driving voice into the audio key point mapping model of the digital human video generation model to obtain a facial key point sequence corresponding to the driving voice; Determine, from each of the candidate character reference images, a target character reference image having the highest similarity to the mouth key points in the facial key point sequence; The facial key point sequence and the target person reference image are input into the diffusion model in the digital human video generation model to obtain the continuous video frames.
9. The method according to any one of claims 1 to 6, characterized in that: Before inputting the driving voice and the character reference image into the digital human video generation model to obtain continuous video frames, the method further includes: Acquire multiple character videos, wherein any of the character videos includes audio and continuous video frames corresponding to the audio; Determine the face reference images corresponding to the respective character videos based on the continuous video frames of the respective character videos; Training an audio key point mapping model based on a first training sample and a first training label, wherein the first training sample is constructed based on a plurality of face reference images and audios of the character videos, and the first training label is constructed based on a plurality of continuous video frames of the character videos; A diffusion model is trained based on a second training sample and a second training label to obtain a digital human video generation model including the audio key point mapping model and the diffusion model, wherein the second training sample is constructed based on a plurality of facial reference images and audio of the human video through the pre-trained audio key point mapping model, and the second training label is constructed based on a plurality of continuous video frames of the human video.
10. The method according to claim 9, characterized in that Before training the audio key point mapping model based on the first training sample and the first training label, the method further includes: Extracting corresponding facial key point sequences from continuous video frames of each character video using a key point detection model; Generate facial key point reference images corresponding to each of the character videos based on the closed-mouth character video frames in the continuous video frames, wherein the facial key point reference images include facial key point reference coordinates; The first training sample is constructed based on the reference coordinates and audio of the facial key points of each of the character videos, and the first training label is constructed based on the facial key point sequence of each of the character videos.
11. The method according to claim 10, characterized in that Training an audio key point mapping model based on a first training sample and a first training label includes: Inputting the first training sample into the audio key point mapping model to obtain a facial key point prediction sequence, wherein the facial key point prediction sequence includes a plurality of facial key point prediction coordinates arranged in time sequence; The audio key point mapping model is trained with minimizing the key point difference loss as a training goal, wherein the key point difference loss includes the distance difference between the predicted coordinates of the face key points corresponding to the same video frame and the coordinates of the face key points of the first training label.
12. The method according to claim 11, characterized in that The distance difference includes a difference in Euclidean distance of plane coordinates and / or a difference in Euclidean distance of graph Laplacian coordinates.
13. The method according to claim 9, characterized in that Before training the diffusion model based on the second training sample and the second training label to obtain the digital human video generation model including the audio key point mapping model and the diffusion model, the method further includes: Generate facial key point sequences corresponding to the audio of each character video through the pre-trained audio key point mapping model; The second training samples are constructed based on the facial key point reference images and facial key point sequences of each of the character videos, and the second training labels are constructed based on the continuous video frames of each of the character videos.
14. The method according to claim 13, characterized in that The diffusion model is trained based on the second training sample and the second training label to obtain a digital human video generation model including the audio key point mapping model and the diffusion model, including: Inputting the second training sample into the diffusion model to obtain a predicted video frame based on time sequence; The diffusion model is trained with minimizing the video frame difference loss as a training objective, wherein the video frame difference loss includes a latent space loss and / or a pixel space loss corresponding to the same video frame.
15. A voice-driven digital human video generation device, characterized in that: include: An acquisition module is used to acquire driving voice and character reference images; A generation module inputs the driving voice and the character reference image into a digital human video generation model to obtain continuous video frames, wherein the digital human video generation model is trained based on training samples constructed based on the character reference image and audio of the character video and training labels constructed based on continuous video frames of the character video; The encoding module performs audio and video encoding on the driving voice and the continuous video frames to obtain a digital human video.
16. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the method according to any one of claims 1 to 14 when executed by the processor.
17. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 14 are implemented.
18. A computer program product, characterized in that The computer program product comprises a non-transitory computer-readable storage medium storing a computer program, the computer program being operable to cause a computer to perform the steps of the method according to any one of claims 1 to 14.
Citation Information
Cited By
Method and device for generating mouth shape video of digital human
CN120640101A
Digital human lip shape video generation method and device
CN120640101B
Personalized audio-driven lip shape generation method and system based on reference frame guidance
CN120726192A
Voice-driven digital human generation method and device
CN121458844A
Voice-driven digital human generation methods and devices
CN121458844B