Abstraction Processing Unit for Non-Verbal Action Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in efficiently training deep neural networks for detecting non-verbal actions due to variations in individual differences, environmental conditions, and camera settings, resulting in high imaging costs and difficulty in eliminating user-specific variations.

Innovation Solution

An information processing apparatus that includes an abstraction processing unit to generate abstracted information from a human body model with three-dimensional information, allowing for the expansion of training data using a small amount of input video, and an inference unit to perform label inference using a machine learning model trained on this expanded data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If comprehensive video data is captured to cover all variations in individual differences, environmental information, camera conditions, and obstacles, then the training data comprehensiveness is improved, but the imaging cost becomes enormous

Engineering Contradiction:
Improvetraining data volumeVSAvoidimaging cost
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

The patent creates virtual copies of human bodies through 3D modeling and rendering. Instead of capturing all possible variations through actual video imaging, the system generates synthetic training data by rendering 3D human models from multiple viewpoints and conditions. This copying approach provides comprehensive training data coverage without the enormous cost of capturing every real-world variation.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes parameters of the 3D human models (pose, expression, lighting, background, camera angle) to generate diverse training data. By systematically varying these parameters, the system can create comprehensive training datasets that cover individual differences, environmental conditions, and camera conditions without actually imaging all these scenarios in the real world.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If CG-based synthesis is used to generate training data from a small source video, then the data generation efficiency is improved, but individual differences of users and environmental information cannot be eliminated

Engineering Contradiction:
Improvedata generation efficiencyVSAvoidtraining data representativeness
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the training data generation process into two parts: (1) extracting pose information from a small source video, and (2) generating diverse training data by rendering 3D models with varied parameters. This segmentation allows the system to maintain efficiency while improving representativeness, as the pose extraction from real video combines with synthetic rendering that can eliminate specific user and environmental biases.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary 3D human body model that bridges the source video and the training data. The 3D model serves as a mediator that can be rendered into various conditions while maintaining the pose information extracted from the source video, thereby generating training data that is both efficient to produce and representative of diverse conditions.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of energy

If a deep neural network is trained using a small amount of data, then the imaging cost is reduced, but the model accuracy may be compromised

Engineering Contradiction:
Improveimaging costVSAvoidnon-verbal action detection accuracy
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

By creating virtual copies of human bodies through 3D modeling, the system can generate large volumes of training data without actual imaging. This provides sufficient training data for high-accuracy deep neural networks while avoiding the enormous cost of comprehensive real-world video capture.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system generates diverse training data by systematically varying parameters of the 3D human models (pose, expression, lighting, background). This parameter variation approach allows the model to learn accurate non-verbal action detection across diverse conditions using synthetically generated data, achieving both cost reduction and maintained accuracy.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250200947A1Information processing apparatus, information processing method, and information processing program
Publication Date: 2025.06.19 SONY GROUP CORP
  • US20250200947A1 patent drawing
  • US20250200947A1 patent drawing
  • US20250200947A1 patent drawing

AI summary

An information processing apparatus according to an embodiment includes: an abstraction processing unit (101) that performs abstraction, from a plurality of directions, on a human body model having three-dimensional information and indicating a first pose associated with a first label, generates a plurality of pieces of first abstracted information each having two-dimensional information and corresponding to the plurality of directions by performing the abstraction, and associates the first label with each of the plurality of pieces of first abstracted information. The plurality of pieces of first abstracted information and second abstracted information having two-dimensional information obtained by abstracting a second pose corresponding to the first pose in one domain are used for associating the first label with the second pose in the one domain.