2D video digital human gesture generation method

Through the 2D video digital human gesture generation method and the diffusion model and other technologies, the problem of unnatural gesture movements in the existing technology is solved, high-quality and clear gesture movement generation is achieved, and virtual character expression and user experience are improved.

CN120014090APending Publication Date: 2025-05-16GIANT MOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510114648.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to generate completely natural and smooth gestures, especially in complex or subtle gestures, resulting in limited expressiveness of virtual characters in movie production, game development, and VR/AR applications and reduced user experience.

Method used

The 2D video digital human gesture generation method is used to generate natural and realistic virtual character gesture actions by collecting data, training gesture motion signal extractor and gesture generation decoder, and using diffusion model to generate gesture signals.

Benefits of technology

It significantly reduces problems such as frame jumps, blur and jitter, generates natural and smooth gestures, improves the stability and coherence of video content, and improves visual effects and visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014090A_ABST
    Figure CN120014090A_ABST
Patent Text Reader

Abstract

The invention relates to a 2D video digital human gesture generation method. The method comprises the following steps: S1, collecting data; s2, training a gesture motion signal extractor and a gesture generation decoder by using videos and static pictures with gesture motion; and S3, extracting a motion signal of the gesture video by using a gesture motion signal extractor, and training the motion signal to generate a gesture signal generation model according to the motion signal of the gesture video and the diffusion model. According to the invention, a more natural and vivid virtual character can be created.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital human, and in particular to a method for generating gestures of digital human using 2D video. Background Art

[0002] Existing technologies often have difficulty in producing completely natural and smooth gestures, especially in complex or subtle gestures, which limits the expressiveness of virtual characters in filmmaking, game development, and VR / AR applications, and reduces the user experience.

[0003] In the existing technology, the generated gestures may lack sufficient details to simulate the subtle changes of real human hands, such as the bending angle of finger joints, the texture of the palm, etc. This unreality will affect the audience's trust and immersion in the virtual character.

[0004] Therefore, it is necessary to provide a 2D video digital human gesture generation method to create a more natural and realistic virtual character. Summary of the invention

[0005] The purpose of the present invention is to provide a 2D video digital human gesture generation method to create a more natural and realistic virtual character.

[0006] In order to solve the problems existing in the prior art, the present invention provides a 2D video digital human gesture generation method, comprising the following steps:

[0007] S1: Collect data;

[0008] S2: Use videos and static images with gesture motion to train the gesture motion signal extractor and gesture generation decoder;

[0009] S3: Use the gesture motion signal extractor to extract the motion signal of the gesture video, and train the gesture signal generation model generated by the motion signal according to the motion signal of the gesture video and the diffusion model.

[0010] Optionally, in the 2D video digital human gesture generation method, data is collected from public datasets, social media platforms, and private recordings.

[0011] Optionally, in the 2D video digital human gesture generation method, the static image is a picture showing the upper body of the person in a natural sitting posture, with both hands relaxed and without making any gestures.

[0012] Optionally, in the 2D video digital human gesture generation method, the gesture motion signal extractor is trained as follows:

[0013] An image encoder is used to encode the input static image into high-dimensional image features for neural network;

[0014] The image feature enhancer uses a multi-head self-attention mechanism to enhance the spatial and contextual relevance of high-dimensional image features and increase temporal dynamic information.

[0015] Optionally, in the 2D video digital human gesture generation method, the image encoder is designed based on a deep convolutional neural network, and the output high-dimensional image features include static visual information and posture information of the image.

[0016] Optionally, in the 2D video digital human gesture generation method, the gesture generation decoder is trained as follows:

[0017] An implicit motion decoder is used to decode the enhanced image features into gesture motion signals. The decoding process of the implicit motion decoder is based on the latent spatial distribution of the motion signal and combines the high-dimensional image features to generate each frame of the gesture motion video.

[0018] Optionally, in the 2D video digital human gesture generation method, the gesture signal generation model is trained as follows:

[0019] The input of the diffusion model includes: the time step in the diffusion process, the text description action features, the noisy motion features, the gesture motion video features corresponding to the current frame and the gesture motion video features generated by the previous time frame, the audio and the corresponding text of the audio;

[0020] The diffusion model gradually removes noise from the above input content to generate clear and coherent gesture motion videos;

[0021] Repeat the input for training to obtain the gesture signal generation model.

[0022] Compared with the prior art, the present invention has the following advantages:

[0023] The 2D video digital human gesture generation method proposed in this patent demonstrates significant technical advantages and application effects. By introducing the latent motion diffusion model, it can restore high-quality, clear gesture action sequences from noisy motion features, significantly reducing common frame jumps, blurs, and jitters in existing generation methods. The model can generate natural and smooth gestures, ensure the stability and coherence of the video content in the time dimension, and improve the visual effect and viewing experience of the generated video. The motion features generated by the previous time frame are used to ensure that the generated motion signal has good temporal coherence. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A flow chart of a method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The specific implementation of the present invention will be described in more detail below in conjunction with the schematic diagram. The advantages and features of the present invention will become clearer based on the following description. It should be noted that the drawings are all in a very simplified form and are not in exact proportions, and are only used to facilitate and clearly assist in explaining the purpose of the embodiments of the present invention.

[0026] Hereinafter, if the method described herein includes a series of steps, the order in which the steps are presented herein is not necessarily the only order in which the steps may be performed, and some of the steps described may be omitted and / or some other steps not described herein may be added to the method.

[0027] Existing technologies often have difficulty in producing completely natural and smooth gestures, especially for complex or subtle gestures. This limits the expressiveness of virtual characters in filmmaking, game development, and VR / AR applications, and reduces the user experience. In existing technologies, the generated gestures may lack sufficient details to simulate the subtle changes in real human hands, such as the bending angle of finger joints, the texture of the palm, etc. This sense of unreality will affect the audience's trust and immersion in the virtual characters.

[0028] In order to solve the problems existing in the prior art, the present invention provides a 2D video digital human gesture generation method, such as Figure 1 Said, comprising the following steps:

[0029] S1: Collect data from public datasets, social media platforms, and private recordings.

[0030] Specifically, public datasets: Use existing open source datasets, such as data provided by TANGO, EchoMimicV2, etc. These datasets contain a large amount of audio, image, and action sequences, which can be used directly for training or as benchmark data.

[0031] Social media platforms: Obtain authorization from publishers on video sharing platforms such as YouTube, Tik Tok, and Bilibili, and then collect relevant gesture videos. These platforms have a large amount of user-generated content, covering gesture performances in various scenarios, which helps to increase the diversity of data. The collected videos are segmented according to a complete gesture.

[0032] Private recording: select data personnel of different genders, regions, and age groups to record a specified gesture data set. The total number of personnel here is required to be at least 200, and the data length of each person is at least 3 hours. All collected video data needs to go through a strict annotation process to ensure that the hand key points in each video frame are accurately calibrated. The annotation process includes: hand key point annotation, gesture category annotation, and action state annotation.

[0033] S2: Use videos with gesture movements and static images to train the gesture motion signal extractor and gesture generation decoder; among them, the static images show the upper body of the person in a natural sitting position, with both hands relaxed and without any gesture movements.

[0034] S21: The training method of the gesture motion signal extractor is as follows:

[0035] An image encoder is used to encode the input static image into high-dimensional image features for the neural network. The image encoder is designed based on a deep convolutional neural network (ResNet). The output high-dimensional image features contain the static visual information and posture information of the image, providing support for subsequent modules.

[0036] The function of the image feature enhancer is to process the encoded high-dimensional image features to make them more suitable for gesture motion generation tasks. Specifically, the image feature enhancer uses a multi-head self-attention mechanism to enhance the spatial and contextual relevance of high-dimensional image features and increase the temporal dynamic information to guide the implicit motion decoder to generate reasonable gesture trajectories.

[0037] S22: The gesture generation decoder is trained as follows:

[0038] An implicit motion decoder is used to decode the enhanced image features into gesture motion signals. The decoding process of the implicit motion decoder is based on the latent spatial distribution of the motion signal and combines the high-dimensional image features to generate each frame of the gesture motion video.

[0039] The image encoder, picture feature enhancer, and implicit motion decoder work together to achieve the ability to generate gesture video signals from input static pictures.

[0040] S3: Use the gesture motion signal extractor to extract the motion signal of the gesture video, and train the gesture signal generation model based on the motion signal of the gesture video and the diffusion model. The goal is to generate high-quality, time-continuous motion features from noisy gesture motion video features.

[0041] Specifically, the training method of the gesture signal generation model is as follows:

[0042] The input of the diffusion model includes: the time step in the diffusion process, the text description action features, the noisy motion features, the gesture motion video features corresponding to the current frame and the gesture motion video features generated by the previous time frame, the audio and the corresponding text of the audio;

[0043] The diffusion model gradually removes noise from the above input content to generate clear and coherent gesture motion videos;

[0044] Repeat the input for training to obtain the gesture signal generation model.

[0045] The application scenarios of this patent include:

[0046] (1) Digital Human: Creating more natural and realistic virtual characters, suitable for film production, game development, virtual reality (VR) and augmented reality (AR) applications, etc. Through improved gesture generation technology, the movements of virtual characters can be made smoother and more natural, thereby improving the user experience.

[0047] (2) Image generation: Using deep learning algorithms to automatically generate or synthesize images containing specific gestures has broad application value in areas such as advertising design and social media content creation. For example, quickly generate images of product spokespersons showing specific gestures for brand marketing activities.

[0048] (3) Deep learning technology: as part of a training dataset to further optimize existing hand recognition models; or as a research tool to explore new neural network architectures and technologies to improve the understanding and reproduction of gestures.

[0049] (4) Education and training: Developing a learning platform based on gesture interaction, such as assisting teaching by imitating the correct pronunciation and mouth shape in language learning software, or simulating the fine movements during surgical operations in medical training programs.

[0050] (5) Remote communication: Improve the quality of non-verbal communication in online conference systems, allowing participants to express their emotional attitudes more accurately, especially in cross-cultural communication situations.

[0051] (6) Assistive technology: Provides a new way for people with disabilities to communicate, such as controlling wheelchair movement or performing other daily tasks by capturing the user's gestures.

[0052] (7) Entertainment industry: Games or performances that support real-time interactive experiences, where audiences can participate through simple gesture commands to increase their sense of immersion.

[0053] In summary, compared with the prior art, the present invention has the following advantages:

[0054] The 2D video digital human gesture generation method proposed in this patent demonstrates significant technical advantages and application effects. By introducing the latent motion diffusion model, it can restore high-quality, clear gesture action sequences from noisy motion features, significantly reducing common frame jumps, blurs, and jitters in existing generation methods. The model can generate natural and smooth gestures, ensure the stability and coherence of the video content in the time dimension, and improve the visual effect and viewing experience of the generated video. The motion features generated by the previous time frame are used to ensure that the generated motion signal has good temporal coherence.

[0055] The above is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any technician in the relevant technical field, without departing from the scope of the technical solution of the present invention, makes any form of equivalent replacement or modification to the technical solution and technical content disclosed in the present invention, which does not depart from the content of the technical solution of the present invention and still falls within the protection scope of the present invention.

Claims

1. A 2D video digital human gesture generation method, characterized in that: The following steps are involved: S1: Collect data; S2: Use videos and static images with gesture motion to train the gesture motion signal extractor and gesture generation decoder; S3: Use the gesture motion signal extractor to extract the motion signal of the gesture video, and train the gesture signal generation model generated by the motion signal according to the motion signal of the gesture video and the diffusion model.

2. The 2D video digital human gesture generation method according to claim 1, characterized in that: Collect data from public datasets, social media platforms, and private recordings.

3. The 2D video digital human gesture generation method according to claim 1, characterized in that: The static picture shows the upper body of the person in a natural sitting position, with both hands relaxed and without any gestures.

4. The 2D video digital human gesture generation method according to claim 1, characterized in that: The gesture motion signal extractor is trained as follows: An image encoder is used to encode the input static image into high-dimensional image features for neural networks; The image feature enhancer uses a multi-head self-attention mechanism to enhance the spatial and contextual relevance of high-dimensional image features and increase temporal dynamic information.

5. The 2D video digital human gesture generation method according to claim 4, characterized in that: The image encoder is designed based on a deep convolutional neural network, and the output high-dimensional image features contain the static visual information and posture information of the image.

6. The 2D video digital human gesture generation method according to claim 4, characterized in that: The gesture generation decoder is trained as follows: An implicit motion decoder is used to decode the enhanced image features into gesture motion signals. The decoding process of the implicit motion decoder is based on the latent spatial distribution of the motion signal and combines the high-dimensional image features to generate each frame of the gesture motion video.

7. The method for generating 2D video digital human gestures according to claim 6, characterized in that: The gesture signal generation model is trained as follows: The input of the diffusion model includes: the time step in the diffusion process, the text description action features, the noisy motion features, the gesture motion video features corresponding to the current frame and the gesture motion video features generated by the previous time frame, the audio and the corresponding text of the audio; The diffusion model gradually removes noise from the above input content to generate clear and coherent gesture motion videos; Repeat the input for training to obtain the gesture signal generation model.