Scene Synthesis from Human Motion via Contact Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in capturing and synthesizing realistic human motion in 3D scenes, particularly in creating high-quality large-scale datasets annotated with diverse human motions and varied 3D scenes, which is essential for applications like virtual reality and human-robot interaction, due to the limitations of laboratory settings and costly equipment.

Innovation Solution

A method and system for scene synthesis from human motion, known as SUMMON, that computes 3D human pose trajectories, generates contact labels for unseen objects, estimates contact points, and predicts object placements based on these trajectories, using a combination of modules like human-scene contact prediction and scene synthesis, leveraging temporal cues and contact former models to generate realistic and diverse scenes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional data capture methods using laboratory equipment are used, then measurement precision of human motion and scene reconstruction can be achieved, but device complexity and cost increase significantly

Engineering Contradiction:
Improveprecision of human motion capture and scene reconstructionVSAvoidcomplexity of data capture equipment and laboratory setup
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses video recordings as copies of real-world human motion instead of direct physical measurement. The system processes video frames to extract pose information, creating a digital representation that replicates the original motion data without requiring complex physical measurement equipment.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces mechanical motion capture systems with a computational approach using video processing and machine learning models. Instead of using mechanical sensors and tracking devices, the system uses neural networks to infer pose from visual data, substituting mechanical measurement with algorithmic processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If diverse human motions and varied 3D scenes are captured using traditional methods, then dataset quality and variety improve, but loss of time and resources increase due to manual annotation and setup

Engineering Contradiction:
Improvediversity of human motions and variety of 3D scenes in datasetVSAvoidtime for data capture, annotation, and processing
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs automatic pose estimation and scene reconstruction without requiring manual annotation. The machine learning models process video data autonomously, extracting motion and spatial information directly from the input footage, eliminating the need for time-consuming human labeling efforts.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent pre-trains pose estimation models on large datasets before deployment. This preliminary training allows the system to quickly process new video data without requiring time-consuming manual annotation for each new dataset, enabling rapid adaptation to diverse motions and scenes.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If high-quality large-scale datasets with diverse human motions and varied 3D scenes are created, then productivity of research applications improves, but device complexity and cost of equipment increase

Engineering Contradiction:
Improveproductivity of research applications in virtual reality and human-robot interactionVSAvoidcomplexity and cost of data capture equipment
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent creates a universal pose estimation system that can process various types of video data for multiple applications including virtual reality, human-robot interaction, and motion analysis. The same core technology serves multiple research purposes, eliminating the need for separate specialized equipment for each application area.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240153101A1Scene synthesis from human motion
Publication Date: 2024.05.09 TOYOTA RESEARCH INSTITUTE INC
  • US20240153101A1 patent drawing
  • US20240153101A1 patent drawing
  • US20240153101A1 patent drawing

AI summary

A method for scene synthesis from human motion is described. The method includes computing three-dimensional (3D) human pose trajectories of human motion in a scene. The method also includes generating contact labels of unseen objects in the scene based on the computing of the 3D human pose trajectories. The method further includes estimating contact points between human body vertices of the 3D human pose trajectories and the contact labels of the unseen objects that are in contact with the human body vertices. The method also includes predicting object placements of the unseen objects in the scene based on the estimated contact points.