Gaze-Based Mask Generation for Complex Object Labeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image processing systems struggle to accurately distinguish and label specific objects, especially artistic or complex shapes, in two-dimensional images, leading to difficulties in object recognition and three-dimensional reconstruction.
Innovation Solution
An image processing system and method that utilize a processor and gaze detector to identify view hotspots corresponding to the user's gaze direction, combined with an indicator signal, to generate accurate mask blocks for object blocks in two-dimensional images, enabling precise labeling of specific objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional image segmentation models (CNN) are used to divide objects into two-dimensional images, then the segmentation process can be completed, but the system cannot accurately distinguish or label specific objects with complex or artistic shapes
Solution Approach 1:
The patent introduces view hotspots and indicator signals as intermediary elements between the image segmentation model and the final object labeling. These intermediaries provide additional spatial and contextual information that helps distinguish specific objects, particularly those with complex or artistic shapes that conventional models struggle to identify accurately.
Solution Approach 2:
The patent adds a new dimension of information by incorporating view hotspots (spatial attention points) and indicator signals (user intent markers) into the traditional two-dimensional image segmentation process. This multi-dimensional approach enables more accurate object identification by combining visual data with spatial and user interaction data.
2Measurement precision
If the system attempts to accurately label all objects in a two-dimensional image, then object recognition precision improves, but the processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary segmentation and identification of view hotspots before final object labeling. By pre-processing the image to identify key spatial points and potential objects, the system reduces the computational burden during the final labeling stage, thereby decreasing overall processing time while maintaining accuracy.
Solution Approach 2:
The patent focuses computational resources on regions of interest identified by view hotspots and indicator signals rather than processing the entire image uniformly. This selective approach allows the system to achieve high labeling accuracy for specific objects while reducing unnecessary computational expenditure on less relevant areas.
3Measurement precision
If conventional segmentation methods are used without additional user interaction, then the system operation is simple, but the system cannot correctly identify specific objects or their shapes
Solution Approach 1:
The patent implements a feedback mechanism where indicator signals from user interaction (such as gaze detection or explicit selection) are fed back into the segmentation process. This feedback loop allows the system to adjust its focus and improve object identification accuracy based on user intent, while maintaining ease of operation through natural interaction modalities.
Solution Approach 2:
The patent enables the system to automatically process and interpret user interactions (such as gaze direction or selection signals) without requiring complex manual input. The system self-adjusts its segmentation and labeling based on these interactions, improving object identification while keeping the user interface simple and intuitive.
Data Source
AI summary
An image processing method includes the following steps: dividing an object block into a two-dimensional image; identifying at least one view hotspot in a viewing field corresponding to pupil gaze direction; receiving the view hotspot and an indicator signal; wherein the indicator signal is used to remark the object block; and generating a mask block that corresponds to the object block according to the view hotspot; wherein the indicator signal determines the label of the mask block.


