Gaze Position Detection for Content Recommendation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current information processing apparatuses lack the ability to provide users with personalized content recommendations based on their gaze direction and viewing habits, failing to effectively utilize gaze position data for scene and object identification and metadata analysis.
Innovation Solution
An information processing apparatus that includes a gaze position detection system, which detects and analyzes user gaze on a display screen, generates search information based on focused scenes and objects, and uses this data to recommend content from a storage device, considering priority levels, statistical data from other users, and metadata for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If gaze position detection and analysis functions are added to the information processing apparatus, then content recommendation accuracy and user engagement are improved, but device complexity and processing requirements increase
Solution Approach 1:
The information processing apparatus integrates multiple functions including gaze position detection, scene analysis, object recognition, and content recommendation into a single unified system. The control section performs diverse tasks such as detecting gaze positions, analyzing video content, identifying objects, and generating recommendations, making the apparatus multi-functional and reducing the need for separate dedicated devices for each function.
2Productivity
If real-time gaze position detection and content recommendation are implemented, then user engagement and recommendation relevance are improved, but processing time and computational load increase
Solution Approach 1:
The apparatus performs preliminary analysis of video content by detecting gaze positions and identifying scenes and objects as the user views the content. The control section continuously monitors gaze positions and pre-processes the video data, so when a recommendation is needed, the analysis is already partially complete or can be quickly finalized, reducing the actual recommendation generation time.
Solution Approach 2:
The system implements a feedback loop where the detected gaze position information is continuously fed back to the control section, which adjusts the content analysis and recommendation generation in real-time. This allows the system to focus computational resources on the most relevant parts of the video content that the user is currently viewing, improving efficiency while maintaining real-time responsiveness.
3Reliability
If detailed scene and object analysis based on gaze position is performed, then recommendation accuracy is improved, but information processing complexity and data requirements increase
Solution Approach 1:
The control section extracts only the essential and relevant information from the video content based on the detected gaze position. Instead of processing all video data, the system identifies and extracts key features such as the current scene, prominent objects in the gaze direction, and relevant metadata, discarding unnecessary information and reducing the data processing burden while maintaining recommendation accuracy.
Solution Approach 2:
The apparatus applies different levels of analysis depth to different parts of the video content based on local relevance. Areas or objects that are currently in the user's gaze receive detailed analysis, while other parts receive minimal or no processing. This localized quality approach ensures high accuracy for relevant content while reducing overall processing complexity and data requirements.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
[Object] [Solving Means] An information processing apparatus includes: a gaze position detection section that detects a gaze position of a user with respect to a display screen; a position information acquisition section that acquires position information of the user; and a control section that judges, based on the detected gaze position, a scene and an object in the scene that the user focuses in a content displayed on the display screen, and generates search information from the judged object and the acquired position information.