Video Generation Model Head Shoulder Motion Coordination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video generation models for talking portrait videos often result in uncoordinated movements of the user's human body tissues in the reconstructed video, reducing the realism of the generated video.

Innovation Solution

A method for training a video generation model that extracts phonetic features, expression parameters, and head parameters from a training video, synthesizes them as a condition input, and performs network training using a single neural radiance field to generate a video generation model that estimates head pose and shoulder motion, ensuring coordinated movements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a talking portrait video is generated using related art methods, then the video can be generated with reconstructed avatar, but the human body tissues show uncoordinated movements reducing realism

Engineering Contradiction:
Improverealism of generated videoVSAvoidcoordination of body tissue movements
Core Design Contradiction:
ReliabilityVSStability of the object's composition

Solution Approach 1:

The patent segments the body into head and shoulder regions, extracting head parameters (head pose, head position) separately from the overall video. This segmentation allows independent analysis and processing of head movements, enabling the model to maintain coordinated movements between head and shoulder tissues by focusing on the head region's parameters while generating the complete talking portrait video.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If head parameters are extracted and synthesized into condition input, then head pose and shoulder motion can be estimated accurately, but the training process becomes more complex

Engineering Contradiction:
Improveaccuracy of head pose estimationVSAvoidcomplexity of training process
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary extraction of head parameters (head pose, head position) from the training video before the main training process. These pre-extracted parameters are then synthesized into the condition input for the neural radiance field model. This preliminary action simplifies the training process by providing ready-to-use head parameter data, reducing the computational complexity during model training while maintaining high measurement precision for head pose estimation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240212252A1Method and apparatus for training video generation model, storage medium, and computer device
Publication Date: 2024.06.27 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US20240212252A1 patent drawing
  • US20240212252A1 patent drawing
  • US20240212252A1 patent drawing

AI summary

This application discloses a method for training a video generation model performed by a computer device. A phonetic feature, an expression parameter, and a head parameter are extracted from a training video of a target user. Network training is performed on a neural radiance field based on the condition input, three-dimensional coordinates, and a viewing direction to obtain a video generation model. The video generation model is obtained through training based on an image reconstruction loss. By introducing the head pose information and the head position information in the training process, a consideration of a shoulder motion status can be introduced into the video generation model so that a motion between the head and the shoulder is more coordinated and stable.