Digital human video generation method based on motion driving
By integrating multiple action-driven information and appearance information, using posture guidance network and denoising UNet to generate digital human videos, the problems of monotony, inconsistent body and unstable generation of videos in traditional methods are solved, and stable, coordinated and high-quality digital human video generation is achieved.
Patent Information
- Application Number
- CN202510006401.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-09
AI Technical Summary
The traditional digital video generation method is limited by the selection of driver information. The generated character videos are monotonous and inconsistent, lack of visual coherence, and unstable generation quality, and there is inter-frame jitter problem.
By obtaining a variety of action-driven information, such as DWPose, SMPL-CS and HaMeR, and using the gesture guidance network to fuse this information, combining the appearance and background information of the target person, input denoising UNet to generate digital human videos. The method integrates a motion module to smooth the sequence of action and inter-frame jitter.
The generated digital human videos have improved stability, enhanced coordination of body movements, improved visual coherence, effectively solved interframe jitter problems, and improved fidelity and quality of the generated videos have been improved.
Smart Images

Figure CN119967257A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for generating digital human videos based on action driving, which is applicable to the fields of image processing and computer vision. Background Art
[0002] With the continuous development of artificial intelligence (AI) technology in the field of content creation, the technology for producing digital humans has become more mature and diverse. Digital humans produced by AI have achieved an unprecedented level of realism in visual presentation. Digital humans have the potential for wide application across industries. For example, in the field of cultural tourism, they can provide tourists with detailed scenic spot guide services; in the film and television industry, they can replace real people to complete the shooting of high-risk action scenes; in the field of education, they can play a role in assisting teaching guidance; and in the financial services industry, they can also guide customers to complete business operations independently.
[0003] Based on the research of traditional methods, the process of generating digital humans can be divided into three steps. The first step is to determine the driving source and driving area. The role of the driving source is to clarify the content attributes of the output video. The driving source can be a visual clue (such as an image, video or posture sequence), a text prompt (such as a text or dialogue describing an action), and an audio signal (such as voice, music). The role of the driving area is to clearly target a specific part of the human body (such as the face) or the entire human body for generation. The second step is action planning. This step aims to plan a logical, natural and smooth human action sequence based on the input driving source. The current methods generally map the input driving source to the action sequence of the video frame through feature mapping. The third step is to replicate the action video output, which is the core step of the entire generation process. After obtaining the planned action sequence or action signal latent space, it needs to be converted into an action video of a given person.
[0004] However, traditional digital human generation methods are limited by the selection of driving information, and the generated character videos often have monotonous or uncoordinated limbs and lack visual coherence.
[0005] In recent years, with the rise of diffusion model technology, the use of Stable Diffusion models has become a mainstream technology for generating digital human videos. DreamPose (an action-based digital human algorithm) uses a pre-trained Stable Diffusion model to build an adapter that can adapt a given character to a given posture. DisCo (an action-based digital human algorithm) is inspired by ControlNet and uses a Stable Diffusion model to decouple the character posture and background in the driving video, and then uses the decoupled posture and background to drive the given character separately. Although the above two methods can output a given character's action imitation video, the generated quality has many shortcomings:
[0006] First, the generation effect is not stable. The above methods use OpenPose (a pose point extraction algorithm) and DWpose (a pose point extraction algorithm) to generate action sequences. When the pose point extraction of a frame in the action sequence is wrong, the generation quality of the corresponding frame in the generated video is likely to be poor.
[0007] Second, only using pose points as hand driving information makes it difficult to generate realistic hands due to the lack of 3D / depth information or inaccurate extraction of hand pose points.
[0008] Third, video jitter: Since the action sequences of the above methods are extracted by static image pose estimation methods (OpenPose and DWpose), inter-frame jitter (the character's pose is sometimes larger and sometimes smaller) in the action sequences is inevitable. Although they all integrate motion modules in the denoising UNet (U-Net Convolutional Network) to smooth the video, the driving influence of the action sequence is too significant to completely eliminate the video jitter. Summary of the invention
[0009] The technical problem to be solved by the present invention is: in view of the above-mentioned existing problems, a method for generating digital human videos based on action driving is provided.
[0010] The technical solution adopted by the present invention is: a method for generating digital human video based on action drive, comprising:
[0011] Acquire a person video as a driving source, and extract multiple action driving information from the video frame of the driving source, wherein the multiple action driving information includes DWPose, SMPL-CS and HaMeR;
[0012] The encoding features of each action driving information are obtained through the posture guidance network, and the encoding features are fused to obtain the fusion features;
[0013] Obtain a target person image as a driving image, and extract the target person's appearance information and background information through an appearance information extraction network;
[0014] The fused features, as well as the target person’s appearance information and background information, are input into the denoising UNet to generate a digital human video.
[0015] The posture guidance network is integrated with a motion module for smoothing the jitter of action-driven information; the denoising UNet is integrated with a motion module for smoothing the character image.
[0016] The encoding features of each action driving information are obtained through the posture guidance network respectively, and each encoding feature is fused to obtain a fused feature, including:
[0017] Each action driving information is respectively passed through a conditional encoder to obtain its own encoding features;
[0018] The coded features are fused through the adaptive gating layer to obtain the fused features;
[0019] The adaptive gating layer concatenates the encoding features to obtain concatenated features; the concatenated features are input into the bypass network to generate gating weights for each feature pixel; the concatenated features are multiplied by the gating weights, and then fused features are obtained through 1×1 convolution.
[0020] The fusion features, as well as the target person’s appearance information and background information are input into the denoising UNet to generate a digital human video, including:
[0021] The frame-by-frame fusion features of the driving source are combined with the appearance information and background information of the target person to generate a digital human image corresponding to each video frame of the driving source. The digital human image corresponding to each video frame is combined to form a digital human video.
[0022] A digital human video generation device based on action drive, comprising:
[0023] A driving information acquisition module, used to acquire a person video as a driving source, and extract a variety of action driving information from the video frame of the driving source, the multiple action driving information including DWPose, SMPL-CS and HaMeR;
[0024] A driving feature extraction module is used to obtain the encoding features of each action driving information through the posture guidance network, and fuse the encoding features to obtain the fusion features;
[0025] An appearance information extraction module is used to obtain a target person image as a driving image, and extract the target person's appearance information and background information through an appearance information extraction network;
[0026] The video generation module is used to input the fusion features, the appearance information and the background information of the target person into the denoising UNet to generate a digital human video.
[0027] A storage medium stores a computer program that can be executed by a processor, and the steps of the digital human video generation method are implemented when the computer program is executed.
[0028] A digital human video generation device comprises a memory and a processor. The memory stores a computer program executable by the processor. When the computer program is executed, the steps of the digital human video generation method are implemented.
[0029] The beneficial effects of the present invention are as follows: the present invention extracts a variety of action driving information of a driving source, including DWPose, SMPL-CS and HaMeR, so that the network not only relies on posture points as action guides, but also obtains other information of the character from SMPL (Skinned Multi-Person Linear Model), and subsequently fuses a variety of action driving information together, so that different action driving information can complement and improve each other, and finally generate a stable digital human video under the action of complete driving information.
[0030] The present invention introduces HaMeR to add three-dimensional and depth information to the hands of characters, so that the network can obtain accurate hand information and further generate realistic hands.
[0031] The present invention integrates motion modules in both the denoising UNet and the posture guidance network, which is equivalent to smoothing the character from both the input source and the output source, better solving the inter-frame jitter problem and keeping the character proportions consistent between frames. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 A flowchart of an embodiment.
[0033] Figure 2 It is a visualization diagram of various action driving information in the embodiment.
[0034] Figure 3 FIG. 4 is an architecture diagram of a posture guidance network in an embodiment. DETAILED DESCRIPTION
[0035] In order to better understand the technical solution of the present application, the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0036] It should be clear that the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0037] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0038] Example 1: Figure 1 As shown, this embodiment is a method for generating digital human video based on action driving, which specifically includes the following steps:
[0039] S100, obtaining a person video as a driving source, and extracting a variety of action driving information from the video frame of the driving source, and the driving area is automatically included in the selected action driving information.
[0040] This embodiment extracts three different types of driving information from the driving source, namely DWPose, SMPL-CS (SMPL-Colorful Surface) and HaMeR (Hand Mesh Recovery, a kind of hand SMPL). Figure 2 The extracted images and visualization of the three action-driven information are shown.
[0041] DWPose is a pose point extraction method that uses two-dimensional coordinates to mark key points of the human body, hand and facial feature points. During the training and inference process of this algorithm, DWPose is rendered as points and edges with different colors to enhance the network's ability to distinguish different parts of the human body.
[0042] SMPL-CS is a method for estimating the three-dimensional parameters of the target human body. During the training and reasoning process of this algorithm, gradient colors are used to render the human skin surface so that the rendering result can simultaneously integrate three-dimensional, depth and continuous semantic information, increasing the amount of human body input information of the network.
[0043] HaMeR is a state-of-the-art method for estimating 3D human hand gestures. Compared with DWPose, HaMeR can provide 3D and depth information to help the network understand the spatial structure of gestures. Compared with SMPL-CS, HaMeR can estimate complex gestures more accurately. Relying on the sequence information of HaMeR, the quality of hand generation has been significantly improved.
[0044] S200, respectively obtaining the encoding features of each action driving information through the posture guidance network, and fusing the encoding features to obtain the fused features.
[0045] In the posture guidance network, the three action-driven information of DWPose, SMPL-CS and HaMeR pass through three independent conditional encoders to obtain their respective encoding features; then an adaptive gating layer is used to fuse the three encoding features to obtain the fused features.
[0046] In this example, the conditional encoder consists of several convolutional layers and SiLU (Sigmoid-Weighted Linear Unit) activation layers.
[0047] In this embodiment, the adaptive gating layer first concatenates the three features, then inputs the concatenated features into a simple bypass network to generate a gating weight for each feature pixel, and finally multiplies the concatenated features by the gating weights, and obtains the final fused features through a 1×1 convolution.
[0048] In this embodiment, the fused features are passed through several convolution and attention layers, as well as motion modules, to obtain smooth fused features.
[0049] The function of the motion module is to smooth the jitter of the driving information. As an input source, the motion driving information itself has jitter characteristics, which will directly affect the character generation effect. Therefore, after the posture guidance network is integrated with the running module, some of the effects of jitter can be eliminated from the source.
[0050] S300: Obtain a target person image as a driving image, and extract the target person's appearance information and background information through an appearance information extraction network.
[0051] In this embodiment, the driving image is input into the appearance information extraction network. Through the guidance of the attention mechanism, the appearance information extraction network can accurately extract the appearance information and background information of the target person and decouple the two.
[0052] S400, input the fusion features, the appearance information and the background information of the target person into the denoising UNet to generate a digital human video.
[0053] S410, input the fusion features of each video frame of the driving source, together with the appearance information and background information of the target person, into the denoising UNet, and guide the denoising UNet to generate a digital human image by using the fusion features, the appearance information and background information of the target person, with the help of the powerful generation ability of the diffusion model.
[0054] The target person in the digital human image corresponding to each video frame of the driving source performs the action in the corresponding video frame of the driving source, while the background remains unchanged.
[0055] S420 , generating a digital human video of the target person frame by frame, and finally outputting the digital human video of the target person, wherein the target person in the digital human video of the target person performs the same actions as the person in the driving source video.
[0056] In this embodiment, the denoising UNet is integrated with a motion module to smooth the character image from the output and eliminate undesirable jitter.
[0057] Embodiment 2: This embodiment is a digital human video generation device based on action drive, including: a drive information acquisition module, a drive feature extraction module, an appearance information extraction module and a video generation module.
[0058] In this embodiment, the driving information acquisition module is used to acquire a person video as a driving source, and extract a variety of action driving information from the video frame of the driving source, and the multiple action driving information includes DWPose, SMPL-CS and HaMeR.
[0059] In this example, the driving feature extraction module is used to obtain the encoding features of each action driving information through the posture guidance network, and fuse the encoding features to obtain the fused features.
[0060] In this embodiment, the appearance information extraction module is used to obtain a target person image as a driving image, and extract the appearance information and background information of the target person through the appearance information extraction network.
[0061] In this example, the video generation module is used to input the fusion features, the appearance information and background information of the target person into the denoising UNet to generate a digital human video.
[0062] Embodiment 3: This embodiment is a storage medium on which a computer program that can be executed by a processor is stored. When the computer program is executed, the steps of the digital human video generation method in Embodiment 1 are implemented.
[0063] Embodiment 4: This embodiment is a digital human video generation device, which has a memory and a processor. The memory stores a computer program that can be executed by the processor. When the computer program is executed, the steps of the digital human video generation method in Embodiment 1 are implemented.
Claims
1. A method for generating digital human video based on action drive, characterized in that: include: Acquire a person video as a driving source, and extract multiple action driving information from the video frame of the driving source, wherein the multiple action driving information includes DWPose, SMPL-CS and HaMeR; The encoding features of each action driving information are obtained through the posture guidance network, and the encoding features are fused to obtain the fusion features; Obtain a target person image as a driving image, and extract the target person's appearance information and background information through an appearance information extraction network; The fused features, as well as the target person’s appearance information and background information, are input into the denoising UNet to generate a digital human video.
2. The method for generating digital human video based on action drive according to claim 1, characterized in that: The posture guidance network is integrated with a motion module for smoothing the jitter of action-driven information; the denoising UNet is integrated with a motion module for smoothing the character image.
3. The method for generating digital human video based on action drive according to claim 1, characterized in that: The encoding features of each action driving information are obtained through the posture guidance network respectively, and each encoding feature is fused to obtain a fused feature, including: Each action driving information is respectively passed through a conditional encoder to obtain its own encoding features; The coded features are fused through the adaptive gating layer to obtain the fused features; The adaptive gating layer concatenates the encoding features to obtain concatenated features; the concatenated features are input into the bypass network to generate gating weights for each feature pixel; the concatenated features are multiplied by the gating weights, and then fused features are obtained through 1×1 convolution.
4. The method for generating digital human video based on action drive according to claim 1, characterized in that: The fusion features, as well as the target person’s appearance information and background information are input into the denoising UNet to generate a digital human video, including: The frame-by-frame fusion features of the driving source are combined with the appearance information and background information of the target person to generate a digital human image corresponding to each video frame of the driving source. The digital human images corresponding to each video frame are combined to form a digital human video.
5. A motion-driven digital human video generation device, characterized in that: include: A driving information acquisition module, used to acquire a person video as a driving source, and extract a variety of action driving information from the video frame of the driving source, the multiple action driving information including DWPose, SMPL-CS and HaMeR; A driving feature extraction module is used to obtain the encoding features of each action driving information through the posture guidance network, and fuse the encoding features to obtain the fusion features; An appearance information extraction module is used to obtain a target person image as a driving image, and extract the target person's appearance information and background information through an appearance information extraction network; The video generation module is used to input the fusion features, the appearance information and the background information of the target person into the denoising UNet to generate a digital human video.
6. A storage medium having stored thereon a computer program executable by a processor, characterized in that: When the computer program is executed, the steps of the digital human video generation method according to any one of claims 1 to 4 are implemented.
7. A digital human video generation device, comprising a memory and a processor, wherein the memory stores a computer program executable by the processor, characterized in that: When the computer program is executed, the steps of the digital human video generation method according to any one of claims 1 to 4 are implemented.
Citation Information
Cited By
Model training method and device, information determination method and device, storage medium and computer program product
CN121305284A