Video Generation Model Head Shoulder Motion Coordination
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video generation models for talking portrait videos often result in uncoordinated movements of the user's human body tissues in the reconstructed video, reducing the realism of the generated video.
Innovation Solution
A method for training a video generation model that extracts phonetic features, expression parameters, and head parameters from a training video, synthesizes them as a condition input, and performs network training using a single neural radiance field to generate a video generation model that estimates head pose and shoulder motion, ensuring coordinated movements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a talking portrait video is generated using related art methods, then the video can be generated with reconstructed avatar, but the human body tissues show uncoordinated movements reducing realism
Solution Approach 1:
The patent segments the body into head and shoulder regions, extracting head parameters (head pose, head position) separately from the overall video. This segmentation allows independent analysis and processing of head movements, enabling the model to maintain coordinated movements between head and shoulder tissues by focusing on the head region's parameters while generating the complete talking portrait video.
2Measurement precision
If head parameters are extracted and synthesized into condition input, then head pose and shoulder motion can be estimated accurately, but the training process becomes more complex
Solution Approach 1:
The patent performs preliminary extraction of head parameters (head pose, head position) from the training video before the main training process. These pre-extracted parameters are then synthesized into the condition input for the neural radiance field model. This preliminary action simplifies the training process by providing ready-to-use head parameter data, reducing the computational complexity during model training while maintaining high measurement precision for head pose estimation.
Data Source
AI summary
This application discloses a method for training a video generation model performed by a computer device. A phonetic feature, an expression parameter, and a head parameter are extracted from a training video of a target user. Network training is performed on a neural radiance field based on the condition input, three-dimensional coordinates, and a viewing direction to obtain a video generation model. The video generation model is obtained through training based on an image reconstruction loss. By introducing the head pose information and the head position information in the training process, a consideration of a shoulder motion status can be introduced into the video generation model so that a motion between the head and the shoulder is more coordinated and stable.


