AI Scene Gaze Estimation via Neural Network Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies for evaluating viewer engagement with video content fail to accurately estimate interest in specific scenes, leading to inaccurate content recommendations and product promotions, as they rely on metadata associated with the entire content rather than individual scenes.

Innovation Solution

An artificial intelligence information processing device and method that uses neural networks to estimate the degree of gaze and scene information by correlating sensor data with video content, allowing for the identification and characterization of specific scenes a user is interested in, thereby providing more precise metadata for those scenes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If metadata associated with entire content is used for evaluation, then content recommendation can be provided, but accuracy of scene-specific interest estimation deteriorates

Engineering Contradiction:
Improvescene-specific interest estimation accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the content into individual scenes and processes each scene separately. The gaze degree estimation is performed for each scene independently, and scene information is extracted and processed on a per-scene basis. This segmentation enables accurate scene-specific interest estimation while managing complexity through modular processing of discrete scene units rather than analyzing entire content at once.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If sensor information is processed to estimate gaze degree, then user engagement measurement is achieved, but processing time and computational load increase

Engineering Contradiction:
Improvegaze degree estimation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing sensor information and pre-segmenting content into scenes before the actual gaze estimation. Scene information such as timestamps, duration, and content descriptors are extracted in advance. This preliminary preparation reduces the computational burden during real-time gaze degree estimation, thereby reducing processing time while maintaining measurement precision.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If neural network is used for scene information estimation, then estimation accuracy according to artificial intelligence is improved, but device complexity and computational resources required increase

Engineering Contradiction:
Improvescene information estimation accuracyVSAvoidneural network processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The neural network processing is segmented into distinct stages: scene segmentation, feature extraction, and classification. Each stage processes specific aspects of the scene independently. This modular neural network architecture improves estimation accuracy for each scene while managing overall system complexity through organized, stage-wise processing rather than monolithic complex processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12184931B2Artificial intelligence information processing device and artificial intelligence information processing method
Publication Date: 2024.12.31 SONY GROUP CORP
  • US12184931B2 patent drawing
  • US12184931B2 patent drawing
  • US12184931B2 patent drawing

AI summary

An artificial intelligence information processing device that generates information about a scene according to artificial intelligence is provided. The artificial intelligence information processing device includes a gaze degree estimation unit configured to estimate a degree of gaze of a user who is watching content according to artificial intelligence on the basis of sensor information, an acquisition unit configured to acquire a video of a scene at which the user gazes in the content and information about the content on the basis of an estimation result of the gaze degree estimation unit, and a scene information estimation unit configured to estimate information about the scene at which the user gazes according to artificial intelligence on the basis of the video of the scene at which the user gazes and the information about the content.