Attention Target Estimation Using 3D Spatial Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods struggle to accurately determine the attention target of communication parties in a space with multiple objects of the same appearance, failing to distinguish which object is being focused on and whether joint attention is occurring.
Innovation Solution
The solution involves using a first-person point-of-view video and line-of-sight information to map objects in a 3D space, calculating the distance between the person's line of sight and each object, and identifying the object with the smallest distance as the attention target, while also considering temporal relevance and interaction among individuals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If prior art methods estimate attention target from place or external appearance, then estimation can be performed using simple criteria, but accurate identification becomes impossible when multiple objects have the same appearance
Solution Approach 1:
The patent transitions from 2D image-based attention estimation to 3D space-based estimation by mapping objects and line-of-sight positions into three-dimensional coordinates. This dimensional upgrade allows the system to distinguish between multiple objects with identical appearances by calculating spatial distances, thereby resolving the ambiguity that plagues appearance-based methods
2Power
If the system uses only external appearance to determine attention target, then processing is computationally simpler, but the system cannot distinguish between multiple objects of the same type
Solution Approach 1:
The patent introduces line-of-sight position information as an intermediary element that mediates between the camera's view and the actual attention target. By calculating the distance between the line-of-sight position and each object's position in 3D space, the system can reliably identify which object the person is actually attending to, even when multiple identical objects are present
Data Source
AI summary
An objective of the present disclosure is to enable estimation of an attention target of communication parties even if there are a plurality of objects having the same appearance in a space. An attention target estimation device 10 according to the present disclosure acquires a first-person point-of-view video IMi captured from a perspective of a person and a line-of-sight position gi of a person when the first-person point-of-view video IMi is captured, identifies positions in a 3D space of objects 31, 32, and 33 extracted from the first-person point-of-view video IMi, and determines an object close to the line-of-sight position gi of the person among the objects 31, 32, and 33 included in the first-person point-of-view video IMi as the attention target of the person.


