Digital human video generation method and device, and storage medium
By introducing an audio pose sequence correspondence learning module into the digital human video generation method, the audio signal is converted into pose sequence data, and the problem of insufficient maintenance of movement and audio consistency and character image in the prior art is solved, and a natural and smooth movement and highly consistent audio synchronization effect is achieved.
Patent Information
- Application Number
- CN202510233748.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to generate natural and smooth body movements that are consistent with the driving audio, and there are shortcomings in maintaining the consistency of the character image.
By introducing an audio pose sequence correspondence learning module, the audio signal is converted into pose sequence data, and the audio guidance network is trained using pre-trained speech digital human video frames, so that the actions of the audio driven are combined with the generation process of the pose driven to ensure that the characters in the video are natural and synchronized with the audio.
It achieves a high degree of consistency between the actions and driving audio in the generated digital human video, while effectively maintaining the consistency of the character image and ensuring the natural and smooth video content.
Smart Images

Figure CN120050483A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of deep learning, and more particularly relates to a method, device, and storage medium for generating digital human videos. Background Art
[0002] Existing 2D digital human video generation methods are usually pose-driven, that is, given a person's picture and a pose sequence, an action video of the corresponding person is generated. However, in practical applications, especially in interactive scenarios, people more expect the corresponding actions of the digital human to be speech-driven, that is, given a person's picture and a speech segment, a speaking video with corresponding limb movements is generated.
[0003] The technology of speech-driven 2D digital human video generation aims to transform a portrait picture or video of a given person into a speaking video synchronized with the driving speech. It is a multi-modal generation technology that shows great application value in fields such as film production, virtual assistants, online education, and video conferencing. Appropriate limb movements, as a supplement to human language, help to enhance the credibility of virtual digital humans. However, most of the existing technologies focus on the generation of the facial or head region. Comparing these regions, especially the lip-sync part, the limb movements have a weak correlation with the driving audio, which makes it more challenging to generate limb movements that are consistent with the driving audio and natural and smooth.
[0004] A direct generation scheme is to first use the co-speech gesture generation method to convert speech into a pose sequence, and then use the pose sequence-driven video generation method to render the pose sequence into a person video. However, the two steps of this scheme are completely independent, and the generation errors in the first step will be transmitted and accumulated, resulting in serious jitter in the final generation result and being difficult to satisfy. Recently, DiffTED and S2G-MDDiffusion integrate the above two steps into a unified framework by using the Thin-Plate Spline (TPS) motion model. Although the jitter is effectively reduced, the human image cannot be well maintained, and there are obvious blurring and deformation. Summary of the Invention
[0005] In view of the above analysis, the embodiments of the present disclosure aim to provide a method, device, and storage medium for generating digital human videos, aiming to effectively maintain the consistency of the human image while ensuring the consistency between the actions in the generated digital human video and the driving audio.
[0006] In the first aspect of the present disclosure, a method for generating a digital human video is provided, including:
[0007] Receiving an audio signal and a reference person image;
[0008] Input the audio signal into the audio-pose sequence correspondence learning module, and output the pose sequence data corresponding to the audio signal; wherein, the audio-pose sequence correspondence learning module is pre-trained and generated using existing speaking digital human video frames, generating a pose skeleton image and audio features based on the existing speaking digital human video frames, inputting the pose skeleton image into a pose guide, and inputting the audio features into an audio guide. The pose guide uses a pre-trained pose guide network, and only trains the audio guide network. The objective of the network learning during the training process is to minimize the difference between the output of the audio guide and the output of the pose guide;
[0009] Input the reference person image and the pose sequence data into a generation model, and sequentially generate video frames according to the pose sequence data;
[0010] Synthesize the generated video frames in chronological order and output a digital human video.
[0011] Optionally, the network structures of the audio guide network and the pose guide network are the same; each network includes four convolutional layers; wherein, the first convolutional network is: the convolutional kernel size is 3x3, the stride size is 1x1, the padding size is 1x1, and the number of output channels is 16; the second convolutional network is: the convolutional kernel size is 4x4, the stride size is 2x2, the padding size is 1x1, and the number of output channels is 32; the third convolutional network is: the convolutional kernel size is 4x4, the stride size is 2x2, the padding size is 1x1, and the number of output channels is 64; the fourth convolutional network is: the convolutional kernel size is 4x4, the stride size is 2x2, the padding size is 1x1, and the number of output channels is 128.
[0012] Optionally, the inputting the reference person image and the pose sequence data into a generation model and sequentially generating video frames according to the pose sequence data includes:
[0013] Extract semantic features from the reference person image through a CLIP image encoder;
[0014] Encode the reference person image through a VAE encoder to generate compressed features;
[0015] Input the semantic features and the compressed features into a reference network for processing to integrate the appearance information of the image and the dynamic information of the audio;
[0016] Generate video frames from the compressed features extracted from the reference person image through a VAE decoder.
[0017] Optionally, the inputting the reference person image and the pose sequence data into a generation model and sequentially generating video frames according to the pose sequence data further includes:
[0018] Use a denoising network to optimize the details of the generated video frames;
[0019] Among them, the denoising network is based on the U-Net architecture and includes an encoder, a bottleneck layer, and a decoder; the hierarchical modules of the encoder and the decoder are symmetric in the spatial dimension and the number of channels;
[0020] Among the four downsampling modules of the encoder, the first three modules introduce a residual and cross-attention mechanism and perform downsampling; the fourth module only contains a residual convolution mechanism and does not perform downsampling.
[0021] Optionally, after receiving the audio signal, it further includes:
[0022] Map the audio signal to a feature space aligned with the video frames in the time dimension;
[0023] Expand the dimension of the latent variable with the noise signal to the same spatial size as the feature space, and align the audio signal and the noise signal;
[0024] Add the features of the aligned audio signal and the noise signal to obtain a fused feature;
[0025] Use the fused feature as the input to the denoising network.
[0026] Optionally, the audio feature input to the audio guide is the high-level audio feature after being processed by the audio encoder, and the output of the audio guide is the encoded feature of the corresponding pose sequence data.
[0027] Optionally, the audio feature input to the audio guide is the low-level audio feature of the original audio signal, and the output of the audio guide is a pose skeleton image; the output pose skeleton image is input to the pose guide.
[0028] Optionally, the correspondence between the audio signal and the pose sequence data is established by a discriminator, the generation result is evaluated by comparing the skeleton image and the real pose sequence data, and the generation network is optimized by a loss function.
[0029] In a second aspect of the present disclosure, there is provided a digital human video generation device, including a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, it implements the digital human video generation method according to any one of the above.
[0030] In a third aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the digital human video generation method according to any one of the above.
[0031] The digital human video generation method provided by the present disclosure converts an audio signal into pose sequence data by introducing an audio-pose sequence correspondence learning module, which describes the actions of a person driven by the audio. This module is generated through pre-training of existing speaking digital human video frames, and an existing speaking video is used to train an audio guidance network so that the audio guidance network learns how to generate appropriate pose sequence data according to the audio signal, enabling the audio-driven actions to be combined with the original pose-driven generation process to ensure that the actions of the person in the video are natural and synchronized with the audio. Traditional pose sequence-driven video generation methods mainly rely on pose data to generate video frames, while in this method, the audio signal guides and synchronizes the generation of the pose sequence. By converting the audio signal into pose sequence data, the model can closely link the actions of the person with the audio content (such as head movements and facial expressions during speaking) when generating the video, ensuring a high degree of consistency between the actions of the person in the video and the audio content. Through this audio-driven pose generation method, it is possible to utilize the existing pose-driven method to ensure the naturalness of the person's actions and adjust the video content according to the audio signal, so that the generated actions have temporal consistency and emotional consistency with the audio signal.
[0032] In addition, the present disclosure also provides a digital human video generation device and a computer-readable storage medium having the above technical effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0034] Figure 1 It is a flowchart of the digital human video generation method provided by some embodiments of the present disclosure;
[0035] Figure 2 It is a flowchart of the implementation process of the digital human video generation method provided by the present disclosure;
[0036] Figure 3 It is a structural block diagram of the digital human video generation device provided by the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of them. It should be noted that, without conflict, the implementation manners and features in the implementation manners of the present disclosure can be combined, separated, interchanged, and / or rearranged. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.
[0038] The terms used herein are for the purpose of describing particular embodiments and are not intended to be limiting. As used herein, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are also intended to include the plural forms. In addition, when the terms "comprise" and / or "include" and their variants are used in this specification, it indicates that there are the stated features, wholes, steps, operations, components, assemblies, and / or groups thereof, but does not exclude the existence or addition of one or more other features, wholes, steps, operations, components, assemblies, and / or groups thereof. It should also be noted that, as used herein, the terms "substantially", "about", and other similar terms are used as approximate terms rather than degree terms, so they are used to explain the inherent deviations of measured values, calculated values, and / or provided values that those of ordinary skill in the art will recognize.
[0039] Some embodiments of the present disclosure provide a method for generating a digital human video. Figure 1 The flowchart of the method for generating a digital human video provided by some embodiments of the present disclosure is shown. The method specifically includes:
[0040] S101: Receive an audio signal and a reference person image.
[0041] Receive the audio signal, which can be the speech recording of a person. Specifically, audio features can be extracted through an audio encoding model (such as Whisper, VGGish, etc.). These features may include information such as the rhythm, intonation, and emotion of the audio.
[0042] Receive the reference person image, which is used to provide the appearance features of the target person, including the face, hairstyle, clothing, body shape, etc. These visual features will ensure the consistency of the person image when generating the video.
[0043] S102: Input the audio signal into the audio-pose sequence correspondence learning module, and output the pose sequence data corresponding to the audio signal.
[0044] Among them, the audio-pose sequence correspondence learning module is pre-trained and generated using existing speaking digital human video frames. Based on the existing speaking digital human video frames, a pose skeleton image and audio features are generated. The pose skeleton image is input into a pose guide, and the audio features are input into an audio guide. The pose guide uses a pre-trained pose guidance network, and only the audio guidance network is trained. The objective of the network learning during the training process is to minimize the difference between the output of the audio guide and the output of the pose guide.
[0045] The role of the audio-pose sequence correspondence learning module is to map the audio signal into pose sequence data, that is, to convert the audio signal into corresponding actions or poses. This module learns the relationship between the audio signal and the pose sequence through pre-trained speaking digital human video frames.
[0046] The audio features are input into the audio guide, which generates audio-driven pose features. The input of the pose guide is the pose skeleton image and is processed by a pre-trained pose guidance network. This pose guide generates more refined pose information to ensure that the generated poses and actions conform to the laws of real body movements.
[0047] During the training process, by minimizing the difference between the output of the audio guide and the output of the pose guide, the generated result of the audio guide is made to be consistent with the result generated by the pose guide in terms of action performance.
[0048] S103: Input the reference person image and the pose sequence data into the generation model, and sequentially generate video frames according to the pose sequence data.
[0049] According to the input pose sequence data, the generation model gradually generates video frames. The content of these video frames is both consistent with the pose sequence and maintains the visual style of the reference person image in appearance. The generation model drives the actions of the person according to the pose sequence to ensure that the generated actions and expressions of the person are synchronized with the audio signal.
[0050] Each generated frame not only needs to conform to the pose data but also needs to be consistent with the reference image in terms of the face, clothing, etc. The generation network processes these two input conditions through conditional generation to generate video frames that meet the visual style requirements.
[0051] S104: Synthesize the generated video frames in chronological order and output a digital human video.
[0052] Arrange each generated video frame in chronological order to synthesize a complete video. These video frames should be temporally continuous to ensure natural transitions between actions. The generated complete video consists of multiple frames, showing the actions, facial expressions, and appearance of the person, and is synchronized with the audio signal to ensure that the video content is natural and smooth and conforms to the input reference image.
[0053] The generation of video frames can adopt a diffusion model, which is a generative model that includes a noise-adding diffusion process and a denoising process. The noise-adding diffusion process is that during the training stage, noise is gradually added to the data, and the signal is reconstructed from the noise through a denoising network to generate high-quality images or video frames consistent with the input conditions.
[0054] In the digital human video generation method provided by the present disclosure, by introducing an audio-pose sequence correspondence learning module, the audio signal is converted into pose sequence data, which describes the actions of the person under the audio drive. This module is generated through the pre-training of existing speaking digital human video frames, and the relationship between audio and pose is trained with the help of past speaking videos. The goal of training the audio-guided network is to minimize the output difference between the audio guide and the pose guide. Through this joint training, the audio-guided network learns how to generate appropriate pose features according to the audio signal, so that the audio-driven actions can be combined with the original pose-driven generation process to ensure that the actions of the person in the video are natural and synchronized with the audio. Traditional pose-sequence-driven video generation methods mainly rely on pose data to generate video frames, while in this method, the audio signal guides and synchronizes the generation of the pose sequence. By converting the audio signal into pose sequence data, the model can closely link the actions of the person with the audio content (such as head movements and facial expressions during speech) when generating the video, ensuring a high degree of consistency between the actions of the person in the video and the audio content. Through this audio-driven pose generation method, it is possible to utilize the existing pose-driven method to ensure the naturalness of the person's actions and adjust the video content according to the audio signal, so that the generated actions have temporal consistency and emotional consistency with the audio signal.
[0055] Compared with existing technologies such as DiffTED and S2G-MDDiffusion that adopt thin-plate spline motion models, the present disclosure uses common pose sequences as motion representations, which can be seamlessly integrated with existing pose-sequence-driven video generation systems without having to train the rendering network from scratch. In addition, the present disclosure models the relationship between audio and limb movement in the encoding space rather than the pixel space, which is more concise and efficient.
[0056] Taking the network structure of Animate Anyone as the basic framework as an example, based on the audio-visual synchronization characteristics in the speaking digital human video, a convolutional network is trained to obtain the pose-related features in the audio. First, given the speaking digital human video frame v i (which can be 3x768x768), the pose skeleton image p i (which can be 3x768x768) can be extracted through a pose encoding model such as DWPose, and the audio feature a i (with a size of 50x384) corresponding to the video frame i is obtained through an audio encoding model such as Whisper. Then, the obtained pose skeleton image and audio feature are respectively input into the pose-guided network and the audio-guided network.
[0057] The network structures of the audio-guided network and the pose-guided network are the same. The parameters in the pose-guided network remain unchanged, and only the audio-guided network is trained. As a specific embodiment, the basic framework can be Animate Anyone, its pose-guided network is a four-layer convolutional network, and the audio-guided network is also a four-layer convolutional network like it.
[0058] Each network contains four convolutional layers; among them, the first convolutional network is: the convolutional kernel size is 3x3, the stride size is 1x1, the padding size is 1x1, and the number of output channels is 16; the second convolutional network is: the convolutional kernel size is 4x4, the stride size is 2x2, the padding size is 1x1, and the number of output channels is 32; the third convolutional network is: the convolutional kernel size is 4x4, the stride size is 2x2, the padding size is 1x1, and the number of output channels is 64; the fourth convolutional network is: the convolutional kernel size is 4x4, the stride size is 2x2, the padding size is 1x1, and the number of output channels is 128.
[0059] Furthermore, in order to calculate the loss and adapt to video frames of different sizes, the present disclosure transforms the size of the input audio vector so that its spatial size is consistent with that of the video frame.
[0060] Finally, the output of the audio guide is constrained by the L1 loss to be consistent with the pre-trained pose guide so as to complete the replacement, where:
[0061]
[0062] In some specific embodiments, referring to Figure 2 , the implementation process of the digital human video generation method provided by the present disclosure specifically includes the following process:
[0063] S201: Input the audio signal.
[0064] S202: Encode the audio signal through a pre-trained audio encoder.
[0065] S203: Obtain the pose control information in the audio through an audio guide.
[0066] S204: Add the features of the audio signal and the noise signal to obtain a fused feature, and input the fused feature into the denoising network.
[0067] Specifically, use the trained audio-pose sequence correspondence learning module to map the audio information to a feature space of 128×96×96 that is aligned with the video image in the time dimension. At the same time, convert the latent variable with noise 4×96×96 into a feature of size 128×96×96 through a 3×3 convolutional layer. Finally, add the two sets of obtained features as a mechanism for guiding the audio-conditioned denoising process, and use the obtained fused feature z t (128×96×96) as the input of the denoising network Unet
[0068] S205: Input the reference person image.
[0069] S206: Obtain the semantic features of the reference person image through the CLIP image encoder.
[0070] Use the CLIP image encoder to further extract the semantic features of the reference person image. The CLIP image encoder is a multimodal model that can understand the relationship between image content and natural language.
[0071] S207: Encode the reference image through the VAE encoder.
[0072] Use the variational autoencoder (VAE) to encode the reference person image and extract the compressed latent features.
[0073] Adopt the VAE pre-trained by Stability AI, which compresses the reference frame of size 3×768×768 into a latent variable of size 4×96×96 as the input of the reference network. Use the CLIP ViT-L / 14 pre-trained by OpenAI to map the representation of the reference frame in the latent space to a vector embedding with a length of 768 and high-level semantic information.
[0074] S208: Obtain the person appearance information in the reference image through the reference network.
[0075] The reference network (ReferenceNet) uses the extracted features to generate or adjust specific image content. It can contain multiple convolutional layers and attention mechanisms to handle spatial relationships and feature integration.
[0076] Specifically, spatial attention can be adopted: focusing on specific regions of the image to better understand and generate key parts in dynamic images. Cross-attention: used to fuse different types of inputs (such as images and pose data) or features (such as images and text). Temporal attention: focusing on the temporal relationships between different frames in a sequence to ensure the coherence and smoothness of video animations over time.
[0077] S209: Obtain the fused features, semantic features, and human appearance information, and perform iterative denoising through a denoising network to obtain the initial latent variable.
[0078] The denoising network (Denoising U-Net) gradually generates high-quality images through iterative denoising steps. Passing through the denoising network multiple times, each iteration may involve adjusting the network's parameters or using different training strategies to gradually improve the output quality.
[0079] Among them, both the reference network and the denoising network adopt the same U-net architecture to ensure that the denoising network can selectively introduce reference frame features from the feature space of the same level and dimension in the reference network into different hierarchical modules. The overall architecture of U-net is symmetric, consisting of a downsampling path (encoder), a bottleneck layer, and an upsampling path (decoder). Each hierarchical module of the decoder uses the same design as the corresponding module in the encoder in terms of spatial dimension and number of channels. At the same time, through skip connections, each hierarchical module in the decoder receives the input from the corresponding module in the encoder during the forward process as part of the input features of this decoder module. Such an architecture helps to effectively extract high-level features and recover low-level features; the bottleneck layer further processes the features extracted by the encoder, enabling the network to understand global semantic information on the feature map with the smallest resolution and providing the overall semantics for the upsampling path to decode back to the original resolution. At the same time, in the modules at different levels of U-net, the reference network uses the cross-attention mechanism to introduce z ref into the feature extraction process to ensure the consistency of the obtained features at the high-level semantic level.
[0080] The denoising network first uses the feature map of the reference frame extracted by the reference network and the fused feature z t obtained above to perform spatial cross-attention to receive the spatial information of the reference frame, and then performs cross-attention with to receive semantic information, and finally uses the self-attention mechanism in the temporal dimension to maintain the consistency and continuity between frames.
[0081] Specifically, the encoder uses four downsampling modules to gradually extract the high-level features of the image and reduce the spatial resolution: the first three are cross-attention downsampling modules, which introduce residual and cross-attention mechanisms and perform downsampling; the last module is a residual module, which only contains a residual convolution mechanism and does not perform downsampling. The input dimension of the first module is 128×96×96, and the output dimension is 320×48×48. The second module takes the output of the first module as the input, and the output dimension is 640×24×24. The third module takes the output of the second module as the input, and the output dimension is 1280×12×12. The fourth module takes the output of the third module as the input, and the output dimension is 1280×12×12. The bottleneck layer contains residual convolution and cross-attention mechanisms, takes the output of the fourth module of the encoder as the input, and the output dimension is 1280×12×12. The decoder adopts a design symmetric to the encoder, that is, the first module takes the output of the bottleneck layer, only contains a residual convolution mechanism and does not perform upsampling, while the second to fourth modules introduce residual and cross-attention mechanisms and perform upsampling, and the output dimensions are 1280×12×12, 640×24×24, 320×48×48, 128×96×96 in turn.
[0082] S210: Decode the initial latent variable through the VAE decoder.
[0083] Use the VAE decoder to reconstruct the latent variable into the pixel space, generate video frames of 3×768×768, and finally obtain a sequence of video frames corresponding to the audio length to synthesize the digital human action video.
[0084] S211: Output the video frames.
[0085] It can be understood that the audio features input to the audio guide in the present disclosure are the high-level audio features processed by the audio encoder, and the output of the audio guide is the encoded features of the corresponding pose sequence data. In this embodiment, the existing pose guide module is directly replaced with an audio guide module to generate the pose video.
[0086] In addition, another specific implementation can be: the audio features input to the audio guide are the low-level audio features of the original audio signal, and the output of the audio guide is the pose skeleton image; the output pose skeleton image is input to the pose guide.
[0087] In this embodiment, the input of the audio guide is the frequency feature of the original audio signal, the output of the audio guide is the pose skeleton image, and then it is input to the pose guide in the original framework. At this time, the correspondence between the audio and the pose sequence is established by the discriminator D. The trained pose guide in the original framework is used as the discriminator without change, and only the generation network G is updated.
[0088] In addition, the present disclosure also provides a digital human video generation device, such as Figure 3 As shown in the structural block diagram of the digital human video generation device provided by the present disclosure, the device includes a memory 31 and a processor 32. The memory 31 stores a computer program, and when the computer program is executed by the processor 32, it realizes the digital human video generation method according to any one of the above.
[0089] In addition, the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the digital human video generation method according to any one of the above.
[0090] Computer-readable storage media include both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0091] Those skilled in the art should also be able to further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0092] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination of both. The software modules may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0093] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present disclosure. It should be understood that the above description is only for the specific embodiments of the present disclosure and is not used to limit the protection scope of the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for generating a digital human video, characterized in that: include: receiving an audio signal and a reference character image; The audio signal is input into an audio posture sequence correspondence learning module, and posture sequence data corresponding to the audio signal is output; wherein the audio posture sequence correspondence learning module is generated by pre-training using an existing speaking digital human video frame, and a posture skeleton image and audio features are generated based on the existing speaking digital human video frame, the posture skeleton image is input into a posture guide, and the audio features are input into an audio guide, the posture guide uses a pre-trained posture guidance network, and only trains the audio guidance network, and the goal of network learning during the training process is to minimize the difference between the output of the audio guide and the output of the posture guide; Inputting the reference character image and the posture sequence data into a generation model, and sequentially generating video frames according to the posture sequence data; The generated video frames are synthesized in time sequence and the digital human video is output.
2. The method for generating a digital human video according to claim 1, characterized in that: The audio-guided network and the posture-guided network have the same network structure; each network contains four convolutional layers; wherein, the first convolutional network has a convolution kernel size of 3x3, a step size of 1x1, a padding size of 1x1, and 16 output channels; the second convolutional network has a convolution kernel size of 4x4, a step size of 2x2, a padding size of 1x1, and 32 output channels; the third convolutional network has a convolution kernel size of 4x4, a step size of 2x2, a padding size of 1x1, and 64 output channels; the fourth convolutional network has a convolution kernel size of 4x4, a step size of 2x2, a padding size of 1x1, and 128 output channels.
3. The method for generating a digital human video according to claim 1, characterized in that: The step of inputting the reference character image and the posture sequence data into a generation model and sequentially generating video frames according to the posture sequence data comprises: Extracting semantic features from the reference person image by a CLIP image encoder; Encoding the reference character image through a VAE encoder to generate compressed features; Inputting the semantic features and the compressed features into a reference network for processing, integrating the appearance information of the image and the dynamic information of the audio; The video frames are generated by the VAE decoder using compressed features extracted from the reference person images.
4. The method for generating a digital human video according to claim 3, characterized in that: The step of inputting the reference character image and the posture sequence data into a generation model and sequentially generating video frames according to the posture sequence data further comprises: A denoising network is used to optimize the details of the generated video frames; The denoising network is based on a U-Net architecture, including an encoder, a bottleneck layer and a decoder; the hierarchical modules of the encoder and the decoder are symmetrical in terms of spatial dimension and number of channels; Among the four downsampling modules of the encoder, the first three modules introduce residual and cross-attention mechanisms and perform downsampling; the fourth module only contains a residual convolution mechanism and does not perform downsampling.
5. The method for generating digital human video according to claim 4, characterized in that: After receiving the audio signal, the method further comprises: Mapping the audio signal into a feature space aligned with the video frame in a temporal dimension; Expanding the dimension of the latent variable with the noise signal to the same spatial size as the feature space, and aligning the audio signal and the noise signal; Add the features of the aligned audio signal and the noise signal to obtain the fusion feature; The fused features are used as input of the denoising network.
6. The method for generating a digital human video according to any one of claims 1 to 5, characterized in that: The audio features input to the audio guide are high-level audio features processed by an audio encoder, and the output of the audio guide is the encoded features of the corresponding posture sequence data.
7. The method for generating a digital human video according to any one of claims 1 to 5, characterized in that: The audio features input to the audio guide are audio low-level features of the original audio signal, and the output of the audio guide is a posture skeleton image; the output posture skeleton image is input to the posture guide.
8. The method for generating digital human video according to claim 7, characterized in that: The corresponding relationship between the audio signal and the posture sequence data is established through a discriminator, the generation result is evaluated by comparing the skeleton image and the real posture sequence data, and the generation network is optimized through a loss function.
9. A digital human video generation device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method for generating a digital human video according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the method for generating a digital human video according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Digital human video generation method and device, equipment and storage medium
CN121217996A
Automatic movie and television video script extraction method based on multi-modal large model
CN121388985A