Human Action Imitation for Virtual Agents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies fail to effectively enable virtual objects or robots to perform human-like actions in the real world, as they are limited to virtual space interactions and lack the capability to mimic all human actions realistically.

Innovation Solution

An information processing device that generates an environment map, analyzes human actions, and uses machine learning to create an action model, allowing agents to perform actions similar to those of humans by imitating them.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a teacher agent is selected from virtual space agents to learn actions, then the learning process can be implemented, but the agent cannot perform actions similar to those of a human in the real world

Engineering Contradiction:
Improveability to learn actionsVSAvoidrealism of actions
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent uses motion capture technology to record and copy human actions in the real world, then applies this copied motion data to virtual agents. This allows the agent to reproduce human-like movements and actions realistically, resolving the contradiction between being able to learn actions and performing them realistically.

Inventive Principle:
Principle #26Copying

2Ease of manufacture

If virtual space interactions are used for action learning, then the learning process can be simplified, but the range of performable actions is limited

Engineering Contradiction:
Improvesimplicity of learning processVSAvoidrange of actions
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal action learning system that can handle both virtual environment actions and real-world human actions through a unified motion capture and reproduction framework. This multi-functional approach allows the same system to learn from diverse sources (virtual agents and real humans) and apply to various action types, expanding the range of performable actions while maintaining systematic simplicity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If all human actions are attempted to be captured, then complete realism can be achieved, but the complexity of the system increases significantly

Engineering Contradiction:
Improvecompleteness of human actionsVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent employs motion capture technology that records comprehensive human motion data, then selectively applies the necessary portions to the virtual agent based on the specific action context. This approach captures more data than immediately needed (excessive action) but processes and applies only the relevant parts, achieving complete realism without requiring the entire system to handle every possible human action simultaneously, thus managing complexity effectively.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12008702B2Information processing device, information processing method, and program
Publication Date: 2024.06.11 SONY GROUP CORP
  • US12008702B2 patent drawing
  • US12008702B2 patent drawing
  • US12008702B2 patent drawing

AI summary

A configuration that causes an agent such as a character in a virtual world or a robot in the real world to perform actions by imitating actions of a human is to be achieved. An environment map including type and layout information about objects in the real world is generated, actions of a person acting in the real world are analyzed, time/action/environment map correspondence data including the environment map and time-series data of action analysis data is generated, a learning process using the time/action/environment map correspondence data is performed, an action model having the environment map as an input value and a result of action estimation as an output value is generated, and action control data for a character in a virtual world or a robot is generated with the use of the action model. For example, an agent is made to perform an action by imitating an action of a human.