Generative Human Motion Models with Video Reconstruction Artifact Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The scarcity of large-scale motion data sets hinders the development of high-quality general-purpose motion synthesis models, as motion capture data is scarce and resource-intensive to collect.

Innovation Solution

A system utilizing motion capture data, video reconstruction data, and user feedback to train a human motion foundation model, enhanced by reinforcement learning and physics-based simulations to filter implausible artifacts and expand motion repertoire.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If motion capture data is collected to improve motion synthesis quality, then motion data quality is improved, but time consumption and resource requirements increase significantly

Engineering Contradiction:
Improvemotion synthesis qualityVSAvoiddata collection time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent uses video data as a proxy/copy for motion capture data. Video reconstruction data is extracted from regular videos to create motion data representations that approximate mocap data quality without requiring actual motion capture sessions. This copying approach allows the system to bypass time-consuming mocap data collection while still achieving high-quality motion synthesis through the motion model training.

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If motion capture data is collected to improve motion synthesis quality, then motion data quality is improved, but resource requirements increase significantly

Engineering Contradiction:
Improvemotion synthesis qualityVSAvoidresource requirements
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The system copies motion information from readily available video data instead of expending resources on actual motion capture data collection. Video reconstruction data serves as a resource-efficient proxy that captures essential motion characteristics without the high computational and temporal costs of traditional mocap pipelines.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system uses existing video data that is already available in the environment, making the motion data collection self-service rather than requiring external motion capture resources. The video data naturally contains motion information that can be extracted and used for training, eliminating the need for dedicated motion capture equipment and sessions.

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If video reconstruction data is used to expand motion data availability, then quantity of motion data is improved, but physically implausible artifacts are introduced

Engineering Contradiction:
Improvemotion data quantityVSAvoidmotion data accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The motion model acts as an intermediary between video reconstruction data and the final motion synthesis output. It processes the video-derived motion data through learned representations that filter out physically implausible artifacts while preserving the quantity and diversity of motion data. The motion model mediates between the noisy video reconstruction data and the requirements for physically accurate motion generation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system incorporates feedback mechanisms where the motion model is trained iteratively using both video reconstruction data and ground truth motion data. This feedback loop allows the model to learn to correct artifacts in video reconstruction data while maintaining the benefits of having large quantities of motion data from diverse video sources.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250232504A1Machine learning models for generative human motion simulation
Publication Date: 2025.07.17 NVIDIA CORP
  • US20250232504A1 patent drawing
  • US20250232504A1 patent drawing
  • US20250232504A1 patent drawing

AI summary

In various examples, systems and methods are disclosed relating to receive at least one of a text prompt or a kinematic constraint and determine first human motion data using a motion model by applying the at least one of the text prompt or the kinematic constraint to the motion model. The motion model is updated by generating, using the motion model, second human motion data by applying motion capture (mocap) data and video reconstruction data as inputs to the motion model, receiving user feedback information for the second human motion data, and updating the motion model based on the user feedback information. The video reconstruction data is generated by reconstructing human motions from a plurality of videos. Physically implausible artifacts are filtered from the video reconstruction data using a motion imitation controller. The motion imitation controller is updated using at least one of Reinforced Learning (RL) or physics-based character simulations.