Scene Synthesis from Human Motion via Contact Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in capturing and synthesizing realistic human motion in 3D scenes, particularly in creating high-quality large-scale datasets annotated with diverse human motions and varied 3D scenes, which is essential for applications like virtual reality and human-robot interaction, due to the limitations of laboratory settings and costly equipment.
Innovation Solution
A method and system for scene synthesis from human motion, known as SUMMON, that computes 3D human pose trajectories, generates contact labels for unseen objects, estimates contact points, and predicts object placements based on these trajectories, using a combination of modules like human-scene contact prediction and scene synthesis, leveraging temporal cues and contact former models to generate realistic and diverse scenes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional data capture methods using laboratory equipment are used, then measurement precision of human motion and scene reconstruction can be achieved, but device complexity and cost increase significantly
Solution Approach 1:
The patent uses video recordings as copies of real-world human motion instead of direct physical measurement. The system processes video frames to extract pose information, creating a digital representation that replicates the original motion data without requiring complex physical measurement equipment.
Solution Approach 2:
The patent replaces mechanical motion capture systems with a computational approach using video processing and machine learning models. Instead of using mechanical sensors and tracking devices, the system uses neural networks to infer pose from visual data, substituting mechanical measurement with algorithmic processing.
2Adaptability or versatility
If diverse human motions and varied 3D scenes are captured using traditional methods, then dataset quality and variety improve, but loss of time and resources increase due to manual annotation and setup
Solution Approach 1:
The system performs automatic pose estimation and scene reconstruction without requiring manual annotation. The machine learning models process video data autonomously, extracting motion and spatial information directly from the input footage, eliminating the need for time-consuming human labeling efforts.
Solution Approach 2:
The patent pre-trains pose estimation models on large datasets before deployment. This preliminary training allows the system to quickly process new video data without requiring time-consuming manual annotation for each new dataset, enabling rapid adaptation to diverse motions and scenes.
3Productivity
If high-quality large-scale datasets with diverse human motions and varied 3D scenes are created, then productivity of research applications improves, but device complexity and cost of equipment increase
Solution Approach 1:
The patent creates a universal pose estimation system that can process various types of video data for multiple applications including virtual reality, human-robot interaction, and motion analysis. The same core technology serves multiple research purposes, eliminating the need for separate specialized equipment for each application area.
Data Source
AI summary
A method for scene synthesis from human motion is described. The method includes computing three-dimensional (3D) human pose trajectories of human motion in a scene. The method also includes generating contact labels of unseen objects in the scene based on the computing of the 3D human pose trajectories. The method further includes estimating contact points between human body vertices of the 3D human pose trajectories and the contact labels of the unseen objects that are in contact with the human body vertices. The method also includes predicting object placements of the unseen objects in the scene based on the estimated contact points.


